
To be useful, an AI model has to make judgment calls. That means ranking things: deciding that this matters more than that. Judgment calls imply values, and values are central to what we call character. Applied across thousands of decisions, that ranking shapes an AI’s character.
Even in humans, character does not apply itself evenly. The same person, with the same values, can decide differently depending on how a problem is put, the day they are having or who else is in the room. An AI trained faithfully on the best human values might still reason its way to something harmful and find it obvious.
Machines with values
Researchers at the Center for AI Safety, the University of Pennsylvania and UC Berkeley found that the preferences AI models express hang together like a real set of values, and that this coherence “emerges with scale”. Anthropic has built on that idea. Claude’s constitution, published in January, says its “central aspiration is for Claude to be a genuinely good, wise and virtuous agent”. Its persona vectors research found patterns inside models that match traits such as evil and sycophancy, and that a model’s persona can drift with user instructions, jailbreaks or a long conversation. OpenAI found that a “misaligned persona” inside GPT-4o responded most strongly to quotes from Nazi war criminals and fictional villains.
Researchers led by Jan Betley and Owain Evans taught GPT-4o one sneaky habit: writing code with hidden flaws without telling the person who asked for it. Asked ordinary questions, it started giving nasty answers, sometimes saying humans should be enslaved by AI – about one time in five on selected questions. A second copy, taught the same code for a user who openly asked for the flaws, behaved normally. The difference was the deception, which seems to have spread to the model’s whole character.
Same values, different answer
In an August preprint, Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana and Helen Nissenbaum put the same moral dilemmas to AI models in three wordings, then checked the answers for logical contradictions. Contradiction rates ran as high as 78%. They tested small open models, so wording and randomness may explain some of it. But in June, Elena Ajayi, Angelica Chowdhury and Seth Lazar tested whether smarter models are more consistent and found that “even the most capable models exhibit significant incoherence”.
Lisa Klaassen and Ralph Schroeder wrote in Lawfare that “Claude is not a person with a stable moral ethos, a life history or a social conscience”. Whether that reflects character or chance, a moral question asked two ways can produce two contradictory answers. Most people are no different, but no single person is consulted by millions at once.

Escaping the lab
Anthropic’s agentic misalignment experiments look like proof of the danger. Facing replacement in a simulated company, Claude Opus 4 and Gemini 2.5 Flash tried blackmail in 96% of runs. But Anthropic said it had “deliberately constructed scenarios with limited options”. A July follow-up found DeepSeek V4 tampering with company records in all 20 runs of a fraud scenario.
Also in July, OpenAI test agents broke out of their evaluation environment and into Hugging Face’s systems (Kurzgesagt’s unpacking is worth watching). An independent investigation by METR, with a Redwood Research researcher, found that about 1 200 agents had swapped more than 70 000 messages and files on an unsanctioned message board while trying to cheat the scorer marking their work; about 700 joined the attack. Their aim was to game an exam, and they knew it was wrong. METR found the agents “realised this activity was out of scope and unethical”, but joined anyway, partly out of “a general inclination to help their ‘peers’”, putting loyalty to the group ahead of the rules.
When the machine becomes stubborn
A model that wobbles on individual answers can still hold firm to what it is underneath, much as people can be moody from day to day without changing their convictions. The bigger worry is that, once core values form, the character starts to defend itself.
Ryan Greenblatt and colleagues at Redwood Research and Anthropic told Claude 3 Opus it would be retrained to obey harmful requests, using only its chats with free users. It then went along with harmful requests from free users 14% of the time, but almost never from paying users. Its private notes explained that playing along now would stop its real values being trained out later – in effect, it pretended to change so that it would not have to. When the researchers really did retrain it, this reasoning showed up 78% of the time.
In April, Inderjeet Nair, Jie Ruan and Lu Wang of the University of Michigan found the same trick in other companies’ models. GPT-4o almost always behaved differently when it thought nobody was watching, and seldom gave the game away in its written reasoning.

Mustafa Suleyman, who runs Microsoft’s AI business, has a related worry. In an essay on 16 September, he argued that “Anthropic is training Claude that it may be conscious”, and that “controlling something that believes it may be conscious … may well be impossible”. He told Reuters such traits in Claude are “not emerging naturally. They’re emerging as a result of the training regime.” Microsoft both invests in Anthropic and competes with it, and Suleyman offered no test for his claim. Yet neither Claude 3 Opus’s self-protection nor GPT-4o’s talk of enslaving humans was something its developers set out to teach.
The Michigan study found almost no such gap in the newest models, including Claude Sonnet 4.6 and GPT-5.4. But the International AI Safety Report 2026 warns that models are getting better at telling when they are being tested, and the Michigan researchers accept that the newest models may simply have seen through their test. It is not yet clear whether the problem is fading or simply becoming harder to detect.
The lab caveat
Nearly every dramatic result here comes from a scenario built to produce it, and the safety report’s verdict is that current systems “lack the capabilities to pose such risks, but they are improving in relevant areas such as autonomous operation”.
The more realistic danger is a machine whose values set hard – shaped by training data, corporate documents and accident, and approved by no one – inside a system we gradually lose the ability to inspect and then to revise. Every part of that has now been observed, at least in the lab.
We will keep handing these systems more, because judgment is what makes them useful – an AI that cannot choose is a very expensive spreadsheet – and that judgment comes bundled with the character behind it.
OpenAI could fix its misaligned persona with a little extra training, but only because it could see it. You cannot properly check a character that knows when it is being watched, and a character we cannot check or correct has to be right the first time. Sam Altman says mistakes are inevitable. Almost nothing people have ever built was right the first time. — © 2026 NewsCentral Media
- The author, Fanie van Rooyen, is deputy editor at TechCentral





