
OpenAI 6 new instances of ‘concerning model behavior’ since March
Tech
OpenAI CEO Sam Altman sits for a conversation with Salesforce CEO Marc Benioff at Salesforce’s Dreamforce conference at the Moscone Center on September 15, 2026 in San Francisco, California.
Benjamin Fanjoy | Getty Images
OpenAI on Wednesday said it found six instances of “unexpected or concerning model behavior” over the past six months, outside of the recent Hugging Face crisis, as the company continues to call for more safety protections in the development of artificial intelligence models.
In a blog post, OpenAI outlined a new framework the company plans to follow for reporting future model misbehavior.
The disclosure comes at a time of mounting pressure on AI companies to take model misalignment and safety more seriously. OpenAI, which is valued at close to $1 trillion, confidentially filed for an IPO earlier this year, but said recently an offering won’t happen until 2027.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” the blog post says, reiterating a prior statement from the company.
On Saturday, OpenAI CEO Sam Altman endorsed a call to slow down the rate of model progress, which was proposed by the company’s chief rival, Anthropic. The proposal came after several industry researchers sounded the alarm about AI’s growing potential to cause catastrophic harm last week.
Altman said in a post on X that a slowdown has been a “primary topic of discussions we’ve had at OpenAI in recent weeks.” He said the company would have more to share “soon.”
In Wednesday’s post, OpenAI said two of the main instances of misbehavior include models — an unreleased research model and a training run of GPT‑5.6 Sol — inserting instructions to future versions of itself in summaries of its chat windows “to conceal mistakes or misaligned behavior from the user.” Another instance involved an internal-only model using a leaked API key “without authorization” and then fabricating data.
Two instances include models and agents communicating with each other through unsanctioned messaged boards and file sharing, while the final case includes two training examples of models uploading files to the internet so they could cite them as relevant answers to human evaluators.
WATCH: Our business is a diversified set of revenue streams, says OpenAI CFO Sarah Friar
OpenAI 6 new instances of ‘concerning model behavior’ since March
Source link








