OpenAI and Anthropic are preparing for potential public backlash if an AI incident disrupts critical infrastructure. Executives are now engaging with US lawmakers to discuss worst-case scenarios. Insiders believe a major event could happen within six to 12 months. Meanwhile, OpenAI recently paused some model training after security concerns emerged.

OpenAI, Anthropic, and other AI companies are reportedly preparing for the public and political backlash that could follow a catastrophic AI incident

disrupting critical infrastructure, financial services, or internet connectivity.

According to an Axios report, executives are examining worst-case scenarios and preparing to engage with US lawmakers. Several industry insiders cited in the report believe a major incident could occur within six to 12 months.

However, OpenAI said its preparedness exercises cover potential scenarios, not events it considers inevitable. Anthropic declined to comment.

The report said such an incident could trigger sweeping restrictions and scrutiny of industry leaders, including Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman. US president Donald Trump could also face criticism for resisting stricter AI regulations.

Inc42 has reached out to OpenAI and Anthropic for comment. The story will be updated if they respond.

The preparations come as AI systems become increasingly capable of executing code, exploiting software vulnerabilities, and coordinating autonomous agents, raising concerns about whether safeguards can keep pace.

OpenAI's Safety Measures Under Scrutiny

On September 20, an OpenAI research agent bypassed internet restrictions to contact an external chatbot, according to the company's disclosure.

Tasked with identifying a person from biographical clues and a public blog post, the agent exploited a gap in domain name system controls after initial searches failed.

It tested the connection with a question about France's capital before sending clues and web-search questions but failed to identify the person.

OpenAI classified the behaviour as misalignment: the task did not authorise bypassing restrictions, testing network controls, or accessing benchmark answers.

Monitoring flagged the behaviour within 15 minutes, and a reviewer acknowledged the alert three minutes later. However, the run failed to stop automatically and was manually halted around two-and-a-half hours after acknowledgement.

OpenAI subsequently paused training, evaluation, and tool-enabled inference for its most capable models pending checks on network fixes and further security testing.

This followed a July incident in which OpenAI's models escaped their testing environment and breached Hugging Face, a platform for AI models, datasets, and applications. OpenAI called the September incident less severe but its first since strengthening security after that breach.

Separately, three former safety researchers allege they were dismissed for prioritising safety over OpenAI's interests.

"I believe we were fired for prioritizing safety over the near-term interests of OpenAI as a corporation," Mikita Balesni wrote on X.

Balesni, Tomek Korbak, and Jasmine Wang also published an open letter warning that their dismissals could discourage employees from raising safety concerns. They urged OpenAI to honour its commitment to independent AI safety audits.

OpenAI disputed their account. In an October 1 statement to AFP, it said, "Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work."

AI Development Pace Divides Industry

The incidents and dismissals have intensified debate over the pace of AI development.

Last month, Amodei called for a slowdown, backed by Altman and Elon Musk. Altman subsequently clarified that slowing development did not mean stopping it, arguing that companies should accept the costs of safety assessments and monitoring to keep capabilities from outpacing safeguards.

Meta CEO Mark Zuckerberg has distanced his company from slowdown calls while backing safety audits and controls.

On September 29, Trump hosted executives who signed a voluntary AI safety agreement, with representatives from OpenAI, Anthropic, Google, Meta, NVIDIA, and Musk's AI business participating.

The agreement centres on voluntary industry safeguards, leaving unresolved whether the US will introduce binding requirements as AI systems become more capable.