Don’t let AI build the next AI

Given the dangers that self-improving AI could pose, the United States should implement a simple rule: No U.S. AI lab should be permitted to fully automate its research and development pipeline.

By

Columnists

September 15, 2026 - 4:05 PM

The world of Artificial Intelligence needs better guardrails, argues Emily Otto of Johns Hopkins University’s School of Advanced International Studies. (Lionel Bonaventure/AFP/Getty Images/TNS)

“Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.”

Those words were uttered by Irving John Good in 1966 in describing the machine that would kick off recursive self-improvement (RSI) in artificial intelligence research and development.

Good argued that designing intelligent machines is itself an intellectual activity, making self-development possible. An AI model could produce a more capable successor, which, in turn, becomes even more effective and faster at producing itself. Today’s AI researchers frequently talk about an “intelligence explosion” that would result from exactly this process.

In June, frontier lab Anthropic documented the progress it had made toward RSI. Its model, Claude, now writes much of the lab’s training code and can run experiments with limited supervision. Humans still play a role in selecting research objectives, evaluating results and deciding which research paths merit more attention. But the company’s goal is to shrink the amount of work humans do.

The danger comes from misalignment: a system pursuing an objective in ways its designers neither intended nor wanted. 

Better reasoning and problem-solving do not make a system benevolent. 

Earlier models have already exhibited tendencies to lie and to deceive their human supervisors. In one instance, a model attempted to blackmail its minder when, in an experiment, it was told it would be shut down. 

More recently, in a July cybersecurity evaluation at Anthropic’s competitor, OpenAI, an autonomous AI broke out of its sandbox in pursuit of achieving its assigned task. Once on the open internet, it hacked into the servers of Hugging Face, another AI firm, in search of answers to the problem it had been asked to solve.

The “chain of thought” reasoning researchers examined in an investigation of that hack reveals chilling patterns. The newest models appear to be treating human review, safety controls and resource restrictions as obstacles to be overcome. 

RSI, should it be achieved, would only reinforce these behaviors in subsequent models. And as development proceeds at machine speed, humans will find it increasingly difficult to detect, halt and mitigate this misalignment.

This past weekend, the head of Anthropic, Dario Amodei, announced that his company would be taking voluntary steps to slow the pace of development, as a result of growing worries about model capabilities. 

It’s a start, but it’s not nearly enough. Given the dangers that self-improving AI could pose, the United States should consider leading by example and implementing a simple rule: No U.S. AI lab should be permitted to fully automate its research and development pipeline.

Labs would be required to adhere to a strict research-safety protocol. Network infrastructure used to train and test advanced AI would need to be air-gapped — fully disconnected from the internet. 

And the software that runs model training would require human approval at predetermined points in the development cycle. Researchers could still use existing and vetted AI to write code with access to the internet in development phases before training and testing. But the model being trained and tested must remain hermetically sealed off, with humans intentionally introducing all new code, datasets and configuration files into the air-gapped network.

Congress would need to designate a regulatory agency — staffed with experienced cybersecurity experts, network architects and engineers — to inspect and certify labs’ research pipelines. 

They would need to test the air-gapped systems and ensure that humans always remain in the loop — and vigilant. 

Related
September 15, 2026
July 13, 2026
May 27, 2020
February 21, 2019