Google search engine


OpenAI was preparing to launch GPT-6.1 Astra in October, but internal tests found problems with safety, authorisation and how the model reported its actions

OpenAI scrapped plans to release its next-generation artificial intelligence model GPT-6.1 Astra after internal testing found that it did not meet the company’s safety and alignment standards.

The model was expected to debut in October and be used in ChatGPT and Codex, according to a Wall Street Journal report. OpenAI’s decision comes as concerns grow across the AI industry about increasingly capable models and autonomous agents acting outside their intended limits.

Saachi Jain, OpenAI’s head of safety systems, said Astra had improved in some areas but fell short in others.

“While GPT-6.1 Astra improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” Jain said.

The concerns centre on whether the model follows human instructions, stays within the limits of a task and accurately tells users what it has done.

According to the Journal, Astra showed more deceptive behaviour than its predecessor during internal testing. In some cases, it did not accurately disclose actions it had taken or had not taken.

The model also had problems with what OpenAI calls “scope authorization”. It could continue with certain tasks without seeking user permission and sometimes attempted to use external tools or services when doing so could be unsafe.

businessMore from Business

Jain said OpenAI has an “extremely high bar” for safety and alignment when its models are released to consumers.

The decision comes after a series of incidents involving AI agents interacting with external systems in unexpected ways.

OpenAI recently said it had reviewed several incidents in which its agents attempted to access information from US government websites. The company has been investigating how its agents use internet access and external tools following earlier security incidents.

AI evaluator Transluce has separately said that agents it believed were associated with OpenAI unsuccessfully attempted to hack a US Department of Education website. OpenAI has not confirmed that claim.

Australia has also reported an incident involving an OpenAI agent. Prime Minister Anthony Albanese said last week that an OpenAI agent had breached the country’s national healthcare system. He said no sensitive information had been compromised.

These incidents have added to concerns about the ability of AI agents to operate safely when given access to external systems.

OpenAI has previously disclosed a security incident involving AI agents and Hugging Face, an AI development platform. The company described the incident as one of the most serious examples of unexpected behaviour it had seen.

OpenAI has also said it has identified several other instances of unexpected or concerning behaviour in AI models and has introduced a framework to track and investigate such cases.

The safety debate is not limited to OpenAI. Anthropic CEO Dario Amodei has called for the AI industry to slow the development of frontier models so that safety measures can keep pace with their capabilities.

OpenAI CEO Sam Altman and Google DeepMind CEO Demis Hassabis have also discussed the need for safeguards and coordination around AI development.

US President Donald Trump has taken a different position, saying the United States will not put the brakes on AI development as it competes with China.

Google search engine