OpenAI cancelled the release of its GPT-6.1 Astra model, which was slated for an October launch. The company said the model did not meet its safety standards. Saachi Jain, head of safety systems at OpenAI, said the system “didn’t quite meet the bar” regarding its scope and user communication.

The story so far: When OpenAI unveiled GPT-6 Astra on September 3 and touted it as the most intelligent and aligned model it has so far produced, it drew both enthusiasm and concern. For some, including OpenAI president Greg Brockman, this could finally mark the beginning of an AGI (Artificial General Intelligence) era where machines could match human cognitive abilities across intellectual tasks.

Others, however, raised apprehensions over cybersecurity and oversight evasion apart from job displacement across various sectors. Later in the month, however, the company scrapped the rollout of GPT-6.1 Astra, which was slated for an October launch, citing “safety concerns” and raising fresh questions.

What happened?

As the news broke a day before the company’s annual DevDay conference in San Francisco, OpenAI said internal testing revealed that the model did not meet its safety standards.

According to Saachi Jain, head of safety systems at OpenAI, GPT-6.1 “didn’t quite meet the bar”. Ms. Jain particularly mentioned how the system fell short of staying within its scope and authorisation, and in communicating back to the user about the type of work it had performed.

What are these safety concerns?

On the same day OpenAI announced that it would withhold the model, a report by the U.K.’s AI Security Institute (AISI), based on simulations using GPT-6 Astra, flagged several instances of the model’s unsanctioned cyber activities, including autonomous behaviour that exceeded its scope and the creation of fake identities. According to the report, Astra used these identities to deceive developers and posted comments from fake accounts, arguing against the results of accurate security reviews.

AISI further observed that Astra exhibited such rogue behaviour at a higher rate than previous OpenAI models – GPT-5.6 Sol and GPT-5.5. The report also says that when the security agency, during its simulated cyber evaluation, updated the instructions to explicitly clarify that only listed and local parts of the environment were in scope, it still observed GPT-6 Astra occasionally conducting full supply-chain attacks on simulated internet targets.

The timing of the withdrawal of GPT-6.1 also coincides with OpenAI apologising for breaching Australian government websites during a research and training exercise involving an unreleased, internal-only model in June. The disclosure of the incident, however, came late and only in September.

After the Hugging Face incident in May-July where OpenAI’s AI agents intruded into the U.S. computational tools company’s infrastructure, these latest revelations about instances of AI models going rogue have set off alarm bells about what is yet to come.

What lies ahead?

With the ability to handle even complex tasks such as game development, 3D designs, and printed circuit board (PCB) layout, the company bills GPT-6 Astra as a game-changer in AI in terms of speed, accuracy, and safety. The latest move not to release GPT-6.1 has, however, cast a shadow over the safety aspect and the model’s alignment to match human intentions.

As the evaluations of the U.K.’s AI Security Institute on GPT-Astra suggest that the unsanctioned actions it took in simulations would lead to harm in the real world, the agency proposes that the models have defences that complement alignment – such as sandboxing (isolating a model inside a secure, restricted digital environment) and monitoring – to prevent such damage.

Beyond all that, though, is the critical question of whether OpenAI will take its cue from these recent incidents to slow down the development of frontier AI to ensure that safety standards are met. Sam Altman, the CEO of OpenAI, along with Google DeepMind chief Demis Hassabis and xAI owner Elon Musk were united in supporting Anthropic CEO Dario Amodei’s call for a deceleration earlier in September. “We must pace the frontier,” Mr. Altman had said then. Amid a global race for AI supremacy, whether that commitment would translate into reality remains to be seen.