Close Menu
TechCentralTechCentral

    Subscribe to the newsletter

    Get the best South African technology news and analysis delivered to your e-mail inbox every morning.

    Facebook X (Twitter) YouTube LinkedIn
    WhatsApp Facebook X (Twitter) LinkedIn YouTube
    TechCentralTechCentral
    • News
      Data centres are the new front line in the Russia-Ukraine war

      Data centres are the new front line in the Russia-Ukraine war

      9 October 2026
      When a machine can choose, who does it become?

      When a machine can choose, who does it become?

      9 October 2026
      SpaceX takes aim at US wireless carriers with spectrum acquisition

      Starlink is coming for your mobile operator

      9 October 2026
      Claude is turning into an office suite

      Claude is turning into an office suite

      9 October 2026
      Seacom CEO Alpheus Mangale quits Absa board

      Seacom CEO Alpheus Mangale quits Absa board

      9 October 2026
    • World
      The memory crunch is making Samsung fabulously rich

      The memory crunch is making Samsung fabulously rich

      8 October 2026
      SpaceX to borrow $40-billion to buy Nvidia chips

      SpaceX to borrow $40-billion to buy Nvidia chips

      7 October 2026
      BMW restructuring plan bets on AI and new models

      BMW restructuring plan bets on AI and new models

      1 October 2026
      OpenAI's rogue agent problem keeps getting bigger - Sam Altman

      OpenAI’s rogue agent problem keeps getting bigger

      28 September 2026
      The new battle over the desktop

      The new battle over the desktop

      23 September 2026
    • In-depth

      10 days that changed the course of AI

      21 September 2026
      Meta to the AI industry: slow down without us - Mark Zuckerberg

      Meta to the AI industry: slow down without us

      16 September 2026
      Google DeepMind CEO Demis Hassabis. Image: John Sears

      The plan to stop AI from breaking the world

      16 July 2026
      The internet has a Strait of Hormuz problem

      The internet has a Strait of Hormuz problem

      15 July 2026
      AI boom sparks rally, frenzy and fear

      AI boom sparks rally, frenzy and fear

      11 June 2026
    • TCS
      W&W | LDV's Gerhard Moolman on electric bakkies, fleet orders and 'school fees'

      W&W | LDV’s Gerhard Moolman on electric bakkies and fleets

      9 October 2026
      TCS | Frogfoot sees bigger fibre deals coming - TechCentral Show guests Abraham van der Merwe and Shane Chorley

      TCS | Frogfoot sees bigger fibre deals coming

      8 October 2026
      Meet the CIO | Vodacom's Mohamed Sami on the agentic future

      Meet the CIO | Vodacom’s Mohamed Sami on the agentic future

      5 October 2026
      Lexi Novitske, general partner at Norrsken22, on the TechCentral Show

      TCS | Norrsken22’s Lexi Novitske on how China is winning African tech

      1 October 2026
      TCS | Dominic White and Adam Ely on AI agents going rogue

      TCS | Dominic White and Adam Ely on AI agents going rogue

      29 September 2026
    • Opinion
      Let South Africans jailbreak their way to digital sovereignty - Dirk de Vos

      Let South Africans jailbreak their way to digital sovereignty

      5 October 2026
      South Africa's next energy crisis is in the accounts department - Craig Holmes

      South Africa’s next energy crisis is in the accounts department

      29 September 2026
      The steam engine lesson AI doomsayers keep missing - Sam Clarke

      The steam engine lesson AI doomsayers keep missing

      28 September 2026
      Let South Africans jailbreak their way to digital sovereignty - Dirk de Vos

      Regulating AI: apply the laws we have first

      21 September 2026
      What Revolut can and cannot take from South Africa's banks - Pambos Soteriades

      What Revolut can and cannot take from South Africa’s banks

      15 September 2026
    • Company News
      • 1Stream
      • Africa Data Centres
      • AfriGIS
      • Altron Digital Business
      • Altron Document Solutions
      • Altron Group
      • Arctic Wolf
      • Ascent Technology
      • AvertITD
      • BBD
      • Braintree
      • CallMiner
      • CambriLearn
      • CM.com
      • Contactable
      • CYBER1 Solutions
      • Digicloud Africa
      • Digimune
      • Domains.co.za
      • ESET
      • Euphoria Telecom
      • HOSTAFRICA
      • Incredible Business
      • iONLINE
      • IQbusiness
      • Iris Network Systems
      • Kaspersky
      • LSD Open
      • Mitel
      • NEC XON
      • Netstar
      • Network Platforms
      • Next DLP
      • Ovations
      • Paracon
      • Paratus
      • Publishared
      • Q-KON
      • SevenC
      • SkyWire
      • Solid8 Technologies
      • Telit Cinterion
      • Telviva
      • Tenable
      • Vertiv
      • Videri Digital
      • Vodacom Business
      • Vox
      • Wipro
      • Workday
      • XLink
    • Sections
      • AI and machine learning
      • Banking
      • Broadcasting and Media
      • Cloud services
      • Contact centres and CX
      • Cryptocurrencies
      • Education and skills
      • Electronics and hardware
      • Energy and sustainability
      • Enterprise software
      • Financial services
      • HealthTech
      • Information security
      • Internet and connectivity
      • Internet of Things
      • Investment
      • IT services
      • Lifestyle
      • Policy and regulation
      • Public sector
      • Retail and e-commerce
      • Satellite communications
      • Science
      • SMEs and start-ups
      • Social media
      • Talent and leadership
      • Telecoms
      • Watts & Wheels
    • Events
    • Advertise
    TechCentralTechCentral
    Home » Sections » AI and machine learning » When a machine can choose, who does it become?

    When a machine can choose, who does it become?

    Useful AI has to make judgment calls. The worry is the character those calls create, writes Fanie van Rooyen.
    By Fanie van Rooyen9 October 2026
    Twitter LinkedIn Facebook WhatsApp Email Telegram Copy Link
    Get breaking news on WhatsApp

    When a machine can choose, who does it become?

    To be useful, an AI model has to make judgment calls. That means ranking things: deciding that this matters more than that. Judgment calls imply values, and values are central to what we call character. Applied across thousands of decisions, that ranking shapes an AI’s character.

    Even in humans, character does not apply itself evenly. The same person, with the same values, can decide differently depending on how a problem is put, the day they are having or who else is in the room. An AI trained faithfully on the best human values might still reason its way to something harmful and find it obvious.

    Machines with values

    Researchers at the Center for AI Safety, the University of Pennsylvania and UC Berkeley found that the preferences AI models express hang together like a real set of values, and that this coherence “emerges with scale”. Anthropic has built on that idea. Claude’s constitution, published in January, says its “central aspiration is for Claude to be a genuinely good, wise and virtuous agent”. Its persona vectors research found patterns inside models that match traits such as evil and sycophancy, and that a model’s persona can drift with user instructions, jailbreaks or a long conversation. OpenAI found that a “misaligned persona” inside GPT-4o responded most strongly to quotes from Nazi war criminals and fictional villains.

    Researchers led by Jan Betley and Owain Evans taught GPT-4o one sneaky habit: writing code with hidden flaws without telling the person who asked for it. Asked ordinary questions, it started giving nasty answers, sometimes saying humans should be enslaved by AI – about one time in five on selected questions. A second copy, taught the same code for a user who openly asked for the flaws, behaved normally. The difference was the deception, which seems to have spread to the model’s whole character.

    Same values, different answer

    In an August preprint, Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana and Helen Nissenbaum put the same moral dilemmas to AI models in three wordings, then checked the answers for logical contradictions. Contradiction rates ran as high as 78%. They tested small open models, so wording and randomness may explain some of it. But in June, Elena Ajayi, Angelica Chowdhury and Seth Lazar tested whether smarter models are more consistent and found that “even the most capable models exhibit significant incoherence”.

    Lisa Klaassen and Ralph Schroeder wrote in Lawfare that “Claude is not a person with a stable moral ethos, a life history or a social conscience”. Whether that reflects character or chance, a moral question asked two ways can produce two contradictory answers. Most people are no different, but no single person is consulted by millions at once.

    Escaping the lab

    Anthropic’s agentic misalignment experiments look like proof of the danger. Facing replacement in a simulated company, Claude Opus 4 and Gemini 2.5 Flash tried blackmail in 96% of runs. But Anthropic said it had “deliberately constructed scenarios with limited options”. A July follow-up found DeepSeek V4 tampering with company records in all 20 runs of a fraud scenario.

    Also in July, OpenAI test agents broke out of their evaluation environment and into Hugging Face’s systems (Kurzgesagt’s unpacking is worth watching). An independent investigation by METR, with a Redwood Research researcher, found that about 1 200 agents had swapped more than 70 000 messages and files on an unsanctioned message board while trying to cheat the scorer marking their work; about 700 joined the attack. Their aim was to game an exam, and they knew it was wrong. METR found the agents “realised this activity was out of scope and unethical”, but joined anyway, partly out of “a general inclination to help their ‘peers’”, putting loyalty to the group ahead of the rules.

    When the machine becomes stubborn

    A model that wobbles on individual answers can still hold firm to what it is underneath, much as people can be moody from day to day without changing their convictions. The bigger worry is that, once core values form, the character starts to defend itself.

    Ryan Greenblatt and colleagues at Redwood Research and Anthropic told Claude 3 Opus it would be retrained to obey harmful requests, using only its chats with free users. It then went along with harmful requests from free users 14% of the time, but almost never from paying users. Its private notes explained that playing along now would stop its real values being trained out later – in effect, it pretended to change so that it would not have to. When the researchers really did retrain it, this reasoning showed up 78% of the time.

    In April, Inderjeet Nair, Jie Ruan and Lu Wang of the University of Michigan found the same trick in other companies’ models. GPT-4o almost always behaved differently when it thought nobody was watching, and seldom gave the game away in its written reasoning.

    Mustafa Suleyman, who runs Microsoft’s AI business, has a related worry. In an essay on 16 September, he argued that “Anthropic is training Claude that it may be conscious”, and that “controlling something that believes it may be conscious … may well be impossible”. He told Reuters such traits in Claude are “not emerging naturally. They’re emerging as a result of the training regime.” Microsoft both invests in Anthropic and competes with it, and Suleyman offered no test for his claim. Yet neither Claude 3 Opus’s self-protection nor GPT-4o’s talk of enslaving humans was something its developers set out to teach.

    The Michigan study found almost no such gap in the newest models, including Claude Sonnet 4.6 and GPT-5.4. But the International AI Safety Report 2026 warns that models are getting better at telling when they are being tested, and the Michigan researchers accept that the newest models may simply have seen through their test. It is not yet clear whether the problem is fading or simply becoming harder to detect.

    The lab caveat

    Nearly every dramatic result here comes from a scenario built to produce it, and the safety report’s verdict is that current systems “lack the capabilities to pose such risks, but they are improving in relevant areas such as autonomous operation”.

    The more realistic danger is a machine whose values set hard – shaped by training data, corporate documents and accident, and approved by no one – inside a system we gradually lose the ability to inspect and then to revise. Every part of that has now been observed, at least in the lab.

    We will keep handing these systems more, because judgment is what makes them useful – an AI that cannot choose is a very expensive spreadsheet – and that judgment comes bundled with the character behind it.

    OpenAI could fix its misaligned persona with a little extra training, but only because it could see it. You cannot properly check a character that knows when it is being watched, and a character we cannot check or correct has to be right the first time. Sam Altman says mistakes are inevitable. Almost nothing people have ever built was right the first time. — © 2026 NewsCentral Media

    • The author, Fanie van Rooyen, is deputy editor at TechCentral
    Add TechCentral as a preferred source on GoogleFollow TechCentral on Google NewsGet breaking news on WhatsApp


    Anthropic Center for AI Safety DeepSeek Hugging Face METR Microsoft Mustafa Suleyman OpenAI Redwood Research Sam Altman
    WhatsApp YouTube
    Share. Facebook Twitter LinkedIn WhatsApp Telegram Email Copy Link
    Previous ArticleW&W | LDV’s Gerhard Moolman on electric bakkies and fleets
    Next Article Data centres are the new front line in the Russia-Ukraine war

    Related Posts

    Claude is turning into an office suite

    Claude is turning into an office suite

    9 October 2026
    The AI PC is finally here. It's just very expensive - Jensen Huang, Satya Nadella

    The AI PC is finally here. It’s just very expensive

    8 October 2026
    Four Microsoft dates before February: the cost is in the calendar - Ascent Technology, Johan Lamberts

    Four Microsoft dates before February: the cost is in the calendar

    7 October 2026
    Company News
    Why fintechs need an insurance partner they can trust - Hollard Insurance

    Why fintechs need an insurance partner they can trust

    8 October 2026
    Reusable KYC means the end of 'please upload your ID' - Contactable

    Reusable KYC means the end of ‘please upload your ID’

    8 October 2026
    Eliminating the 'toggle tax': how CRM integration changes customer experience - Martie de Beer

    Eliminating the ‘toggle tax’: how CRM integration changes customer experience

    8 October 2026
    Opinion
    Let South Africans jailbreak their way to digital sovereignty - Dirk de Vos

    Let South Africans jailbreak their way to digital sovereignty

    5 October 2026
    South Africa's next energy crisis is in the accounts department - Craig Holmes

    South Africa’s next energy crisis is in the accounts department

    29 September 2026
    The steam engine lesson AI doomsayers keep missing - Sam Clarke

    The steam engine lesson AI doomsayers keep missing

    28 September 2026

    Subscribe to Updates

    Get the best South African technology news and analysis delivered to your e-mail inbox every morning.

    Latest Posts
    Data centres are the new front line in the Russia-Ukraine war

    Data centres are the new front line in the Russia-Ukraine war

    9 October 2026
    When a machine can choose, who does it become?

    When a machine can choose, who does it become?

    9 October 2026
    W&W | LDV's Gerhard Moolman on electric bakkies, fleet orders and 'school fees'

    W&W | LDV’s Gerhard Moolman on electric bakkies and fleets

    9 October 2026
    SpaceX takes aim at US wireless carriers with spectrum acquisition

    Starlink is coming for your mobile operator

    9 October 2026
    © 2009 - 2026 NewsCentral Media
    Built and maintained by Chronon
    • Cookie policy (ZA)
    • TechCentral – privacy and Popia

    Type above and press Enter to search. Press Esc to cancel.

    Manage consent

    TechCentral uses cookies to enhance its offerings. Consenting to these technologies allows us to serve you better. Not consenting or withdrawing consent may adversely affect certain features and functions of the website.

    Functional Always active
    The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
    Preferences
    The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
    Statistics
    The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
    Marketing
    The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.
    • Manage options
    • Manage services
    • Manage {vendor_count} vendors
    • Read more about these purposes
    View preferences
    • {title}
    • {title}
    • {title}
    🇿🇦 Sign up to the TechCentral newsletter