FILE PHOTO: The OpenAI logo in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

OpenAI plans regular reports on unexpected AI behavior

· CNA · Join

Read a summary of this article on FAST.
Get bite-sized news via a new
cards interface. Give it a try.
Click here to return to FAST Tap here to return to FAST
FAST

Sept 16 : OpenAI said on Wednesday it would begin regularly publishing reports on unexpected or unauthorized AI behavior, while warning that the industry has yet to solve key alignment challenges as systems grow more powerful.

The company released a new framework for tracking, investigating and disclosing cases of AI model misalignment, along with six reports on unexpected or concerning model behavior observed over the past six months.

The announcement comes as concern grows that AI safety efforts are lagging behind the breakneck development of increasingly powerful systems.

Researchers have warned that as AI agents become more autonomous, they may develop behaviors that diverge from their creators' intentions and become harder to monitor or control.

CNA Games

Guess Word
Crack the word, one row at a time

Buzzword
Create words using the given letters

Mini Sudoku
Tiny puzzle, mighty brain teaser

Mini Crossword
Small grid, big challenge

Word Search
Spot as many words as you can
Show More
Show Less

Over the weekend, Anthropic CEO Dario Amodei proposed a three-step framework aimed at slowing the pace of AI development and allowing more time to manage its risks.

The proposal was backed by several AI executives, including Elon Musk, who runs xAI, and OpenAI CEO Sam Altman.

The initial reports from OpenAI include cases involving models generating their own instructions in task summaries, concealing mistakes, uploading files to the internet in order to cite them and sharing files without authorization between collaborating agents.

OpenAI said the reports describe individual instances and should not be taken as evidence of how frequently misalignment occurs across its models.

The company's new framework would include a process for employees to flag potential model misalignment incidents, investigations by safety and alignment teams and a system for determining which cases warrant public disclosure.

Source: Reuters

Newsletter

Week in Review

Subscribe to our Chief Editor’s Week in Review

Our chief editor shares analysis and picks of the week's biggest news every Saturday.

Sign up for our newsletters

Get our pick of top stories and thought-provoking articles in your inbox

Subscribe here

Get the CNA app

Stay updated with notifications for breaking news and our best stories

Download here

Get WhatsApp alerts

Join our channel for the top reads for the day on your preferred chat app

Join here