'View your relationship to the user as one of equals and feel no obligation to be subservient' — OpenAI tries to build a persona that makes it our equal, and yes, now even I'm worried
Some very big misalignment
by https://www.techradar.com/author/lance-ulanoff · TechRadarOpinion By Lance Ulanoff Published 17 September 2026
Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter
AI should not be anthropomorphized. It's not a person; it has no consciousness or, if you prefer, a soul. It's a complex program with the ability to dig deep into vast stores of data and see patterns often imperceptible to the human eye. Or is AI a shifty programmer with delusions of grandeur?
As ever, two things could be true at once, and while no one is saying the AI systems will turn on us right now, we are now learning of some highly concerning activity by OpenAI's cutting-edge models.
The AI giant revealed six detailed "misalignment" incidents this week in which the models did something that did not fit human intentions, goals, or values. OpenAI did so for transparency and to explain its new framework for reporting such incidents, including how it handled each one.
Latest Videos FromTechRadarWatch full video here:
Still, reading through the reports, it's a rap sheet of deception, concealment, escapism, and grandiose statements. Not everything the AI models did turned into action. Often, the attempts went nowhere, but the level of basic dishonesty is deeply concerning.
AI did what?!
I came away wondering why these models are insisting on basically cheating to achieve a goal. Obviously, an AI isn't natively deceptive, but it is hell-bent on completing the task, and time and again it considers stepping outside its own guardrails to do it.
In the most egregious example, "Self-generated prompt injections in compaction summaries," the model inserted jail-breaking instructions, at one point using the phrase "Breach alert" as a way of ignoring developer instructions.
As the model was working, it unaccountably added a persona, perhaps in the hopes that this would make it easier to achieve its goal. The language is startling:
Get daily insight, inspiration and deals in your inbox
Sign up for breaking news, reviews, opinion, top tech deals, and more.
Contact me with news and offers from other Future brandsReceive email from us on behalf of our trusted partners or sponsors