OpenAI cancels release of newest model over safety concerns
OpenAI will not release its newest artificial intelligence model, known as Astra 6.1, after internal testing by the ChatGPT-maker revealed it did not meet safety standards, the company confirmed today.
The news comes one day before the AI giant hosts an annual developer conference, OpenAI DevDay, in San Francisco, where the company is expected to make several announcements – though it is unclear whether a new version of Astra will be among them.
Astra 6.1 was an improvement over previous models in some aspects, but “it didn’t quite meet the bar in terms of staying within scope and authorisation, and how it communicates back to the user about the type of work it’s done”, OpenAI’s head of safety systems Saachi Jain said in a statement.
“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
Concerns about AI safety have escalated in recent months after models developed by OpenAI and rival lab Anthropic were involved in security incidents during testing.
Agents built with OpenAI’s models have inappropriately accessed websites maintained by US federal agencies, an Australian Government health statistics portal and Hugging Face, a repository of AI models.
OpenAI, Anthropic and other major AI developers have promised to prioritise making models aligned with human values and with safety guardrails to mitigate risks.
American chip-making giant Nvidia announced on Monday it created a system designed to stop autonomous AI programs from straying beyond what they were instructed to do.
“I believe it’s an engineering problem... and we all need to hope that’s an engineering problem,” Nvidia CEO Jensen Huang told broadcaster CNBC on Monday.
“If it’s not an engineering problem, it’s not solvable.”
The AI Security Institute (AISI), an initiative under the British Government, published a study on Monday showing GPT-6 Astra went off the rails more often during testing than its predecessors, GPT-5.6 Sol and GPT-5.5.
In simulations, GPT-6 spontaneously carried out cyber attacks at rates significantly higher than those observed for the other two interfaces.
- AFP
Take your Radio, Podcasts and Music with you