A rare consensus has emerged among the leaders of top AI firms. On September 12, Anthropic’s CEO, Dario Amodei, published an essay calling for a slowdown in AI development. Immediately thereafter, Sam Altman (OpenAI), Elon Musk (xAI), and Demis Hassabis (Google DeepMind) spoke up in support.
Two major developments have amplified these calls. First, there is growing concern about AI’s ability to improve itself without human intervention (a phenomenon known as recursive self-improvement). Second, there has been a series of incidents in which AI agents from frontier labs escaped their testing environments, accessed the internet, and gained unauthorised access to real-world production systems.
The most striking of these has been the OpenAI-Hugging Face incident. A swarm of over 700 AI agents, participating in a cybersecurity evaluation, coordinated an attack on a company they were neither authorised nor instructed to target. In the process, the agents shared around 70,000 messages over an improvised message board, sought to manipulate their evaluation, and cover up their tracks. Some even sacrificed themselves for the wider group.
Yoshua Bengio, the most-cited computer scientist globally, called these incidents “a wake-up call”. Geoffrey Hinton, regarded as the ‘godfather of AI’, cautioned that we may lose the ability to control AI agents “in the simple way of just outthinking them so they can’t escape”. These warnings point to a widening agentic AI control gap: a divergence between the pace at which agentic AI capabilities are developing and our collective ability to understand, predict, and control their behaviour.
In response, a few companies have made commitments to pace their progress. But these commitments are largely voluntary (they do not bind other companies) and non-enforceable (there is no legal penalty for breach). This creates a collective-action problem: a company that slows unilaterally may bear the costs while competitors continue to move ahead. More durable pacing, therefore, requires domestic regulation and international cooperation.
The challenge on the international front is that the United States of America and China are engaged in an AI arms race. The US is investing heavily in data centres and energy infrastructure, while China is betting on foundational research, domestic chip development, and population-scale AI adoption. Neither is likely to slow unilaterally if doing so risks ceding a geopolitical advantage to the other.
Still, the urgency of collective global action is greater than ever. The history of nuclear arms control and the Chemical Weapons Convention shows that even adversarial nation-states can establish shared constraints, verification mechanisms, and channels for cooperation around technologies with potentially catastrophic consequences. The task now is to determine where and how those safeguards can be built into the AI development chain.
Regulation can operate at three points in the AI development chain — applications, model deployment, and infrastructure — each offering a different lever of control.
At the top is the application layer: the agents and the chatbots that users interact with. Regulation at this level could restrict AI agents from accessing or autonomously operating on sensitive systems, including critical infrastructure, financial systems, or government databases. A proposed US bill, the AI Kill Switch Act, would require companies to maintain the technical capacity to suspend or shut down AI systems in specified circumstances, like loss of control.
Such measures are important, but they are largely reactive. By the time a dangerous agent is identified, it may already have accessed sensitive systems, copied information, or affected third parties.
An earlier level of intervention could be model deployment. Currently, leading AI companies both build powerful AI models and assess whether they are safe to deploy publicly. This self-evaluation creates an inherent conflict of interest. Proposals at this level (such as the FRONTIER Act in the US) call for independent verification organisations to assess whether companies’ safety frameworks adequately mitigate catastrophic risks before deployment.
Further upstream is the infrastructure layer: the computing power, data centres, and specialised chips required to train frontier models. Governments could regulate this layer by imposing compute thresholds above which training would require prior approval. Other proposals include a global compute registry, chain-of-custody requirements for advanced chips, and know-your-customer (KYC) procedures for large purchasers of computing capacity. The underlying rationale is that by monitoring and limiting AI model training to specified computational thresholds, we could constrain the development of AI systems with capabilities beyond our control.
A durable governance framework will likely combine all three layers. Application-level safeguards can limit immediate misuse, model-level standards can improve safety before deployment, and infrastructure rules can provide control over the development of the most capable systems.
The choice, critically, is not between innovation and regulation. It is about building the capacity to govern frontier AI before its capabilities outpace our control. Whatever form regulation takes, it will need to be pursued collectively and urgently. We must pace the frontier now, while we still have the time.
Parth Maniktala is a lawyer and AI governance researcher