OpenAI plans regular reports on unexpected AI behavior
· CNA · JoinRead a summary of this article on FAST.
Get bite-sized news via a new
cards interface. Give it a try.
Click here to return to FAST Tap here to return to FAST
FAST
Sept 16 : OpenAI said on Wednesday it would begin regularly publishing reports on unexpected or unauthorized AI behavior, while warning that the industry has yet to solve key alignment challenges as systems grow more powerful.
The company released a new framework for tracking, investigating and disclosing cases of AI model misalignment, along with six reports on unexpected or concerning model behavior observed over the past six months.
The initial reports include cases involving models generating their own instructions in task summaries, concealing mistakes, uploading files to the internet in order to cite them and sharing files without authorization between collaborating agents.
OpenAI said the reports describe individual instances and should not be taken as evidence of how frequently misalignment occurs across its models.
Newsletter
Week in Review
Subscribe to our Chief Editor’s Week in Review
Our chief editor shares analysis and picks of the week's biggest news every Saturday.
Sign up for our newsletters
Get our pick of top stories and thought-provoking articles in your inbox
Get the CNA app
Stay updated with notifications for breaking news and our best stories
Get WhatsApp alerts
Join our channel for the top reads for the day on your preferred chat app