Inside the Race to Make AI Build Itself

Inside the Race to Make AI Build Itself


—Getty Images

Jack Clark, Anthropic’s co-founder, left on paternity leave last November. When he returned in February, he was surprised to learn that colleagues hardly wrote code anymore. They managed five or six copies of the company’s AI, Claude, which sometimes managed several more Claudes. 

To Clark, this looked like an early form of something the field has anticipated and feared for decades: recursive self-improvement, or the point at which AI begins to accelerate its own development. First, the thinking goes, models make researchers faster, but as each improvement feeds the next, the models take over more of the research cycle. Years of progress compress into months—leaving society with little time to absorb the consequences, from job disruption to engineered pathogens. Taken to its limit, AI could improve itself without humans, triggering a runaway loop long known as an “intelligence explosion,” where machines rapidly advance beyond human understanding, and possibly beyond human control.

Clark believed the world needed to confront this prospect. He posted a flurry of blog posts on recursive self-improvement and flew home to England to deliver a talk on the topic in May, then led an Anthropic report in June titled “When AI Builds Itself,” arguing the technology is already accelerating its development. The volume of code produced per person at Anthropic has increased eight-fold, with Claude writing 80%, the report noted. “We’re trying to help substantiate this concept now … before it becomes something that is politicized or otherwise gains some valence that makes talking about it difficult,” Clark says.

It worked out as Clark feared. “Anthropic is trying to strike terror into everyone’s hearts,” wrote AI skeptic Gary Marcus, adding “all they have really shown is just faster coding.” Skeptics point out that technological progress has always compounded. Oil is used to drill oil. Why, in AI’s case, should the curve suddenly bend upward? For a company betting on continued advances, they argued, the claim is plainly self-serving.

Even Clark concedes coding volume is a crude yardstick. Claude’s code can be long-winded. But the trouble runs deeper. Neither Claude nor any large language model is written in code at all. Researchers set growth conditions—deciding the size and shape of a neural network, then pour an internet’s worth of text through it, letting it adjust itself billions of times until abilities to answer questions, write code, and hold a conversation emerge. They are cultivated, the way one grows a plant by tending the soil and the light without ever deciding where a single leaf will go. 

Progress, therefore, depends on trial and error. If Claude could take over that cycle, designing, running, and analyzing experiments, would progress accelerate gradually… or suddenly explode? And if it did, could anyone pump the brakes? The uncomfortable truth is that the people building the technology are nearly as much in the dark as everyone else.

Claude began beating the benchmarks

The change Clark had walked into was not entirely unexpected. Fellow Anthropic co-founder and chief science officer Jared Kaplan had long feared that AI would eventually accelerate research, perhaps outpacing safety efforts. In early 2025, he folded a warning into the company’s Responsible Scaling Policy, its plan for managing AI’s growing dangers. Back then, Claude was no good at running experiments. But he believed that would change one day, and Anthropic would need to be ready. 

To find out whether that day was coming, Anthropic built a series of tests—tasks that would take a human expert hours. Could Claude train a smaller AI model from scratch? Could it program a virtual robot dog? No single measure would settle it, but together they offered a snapshot. 

In one test, Claude had to rewrite a piece of code to use GPUs—the chips for training AI—more efficiently. In spring, it momentarily got them running seven times faster, then broke the code. By summer, a newer Claude pushed the same speedup from seven times faster to 73 without introducing errors. “We started seeing these tasks fall over,” says Daniel Freeman, a member of Anthropic’s frontier red team who designed the evaluations. 

As Claude began completing such tasks too reliably to reveal much, Freeman devised a test closer to the messiness of the real world. Two human teams competed in a series of challenges, involving a robotic dog and a beach ball. One team could use Claude, the other could not. After narrowly losing the first challenge, the team using Claude finished the second nearly two hours ahead. With time to spare, they trotted their robotic dog around the warehouse—until a miscalculation sent it springing toward the other team, still hunched over their laptops. An overseer caught it just in time.

The November release of Claude Opus 4.5 was a tipping point. Where previous generations tended to stall partway, the model could carry a researcher’s experiment through to the end, freeing them to run more at once. “Before, I’d have eight ideas and I’d try one of them,” Kaplan told TIME in February. “Now I ask Claude to just try all eight.” 

Ahead of a February release, Anthropic surveyed 16 of its researchers, asking whether it could replace an entry-level colleague. Five thought it might. Asked to reflect on their answer, all five walked it back. That they even entertained the idea—that they were already largely redundant—is perhaps more revealing than any benchmark. 

If Anthropic’s tests fell, outside measures have similarly reached their limit. The METR graph, perhaps the best-known independent measure of AI software engineering ability, tracks the complexity of tasks models can complete based on how long they would take a human expert. But in May, Claude exceeded the benchmark’s upper limit.

This spring, Anthropic let the robot dogs out again, but this time, an improved Claude worked alone. On every task it could attempt, it was at least 10 times faster than the Claude-assisted humans had been months earlier. The researchers saw the same pattern they had found in cybersecurity. First AI helps humans do better; then it largely takes over.

Jack Clark speaks during the Hill & Valley Forum at the U.S. Capitol Visitor Center Auditorium in Washington, DC, on April 30, 2025. —Brendan Smialowski—AFP/Getty Images

How much faster will it go?

In the spring of 2025, a former OpenAI staffer named Daniel Kokotajlo co-authored “AI 2027,” a document somewhere between forecast and science fiction, which laid out quarter-by-quarter how unchecked recursive self-improvement might end. 

First, AI coding agents begin speeding up progress. By 2027, researchers stare at their consoles while AI builds itself. It soon exceeds human intelligence across the board. By 2030, if we’re lucky, everyone lives in material abundance under the thumb of those controlling an all-powerful AI. If we’re unlucky, we all die.

“We used to be the more aggressive, bullish predictors,” Kokotajlo tells TIME. But recently, he says people—including some working at AI companies—have reached out privately to say his timeline is too conservative. “It’s quite disquieting,” he says.

As the industry pushes to make AI build itself, a debate has opened over what that will do to the pace of progress. Many in Silicon Valley—like Kokotajlo—fear AI progress is about to accelerate violently. Others believe that bottlenecks such as scarce computing power and data will prevent an explosive takeoff. Crucially, the measures needed to tell which side is right do not yet exist.

Arvind Narayanan is skeptical that recursive self-improvement will drive such fierce progress. A Princeton computer scientist, he co-authored a report titled “AI as Normal Technology,” which has become the standard rebuttal to “AI 2027,” arguing AI’s impact will diffuse through the economy over decades, akin to the Industrial Revolution. 

In July, Narayanan helped test Claude’s ability to investigate open-ended research questions, rather than narrow benchmarks. Given six days and thousands of dollars of compute, it could reliably handle all the engineering to set up the experiments, yet made no substantial headway on the overarching research question. It often ran into dead ends and struggled to backtrack, displayed poor judgment, and sometimes drifted from the goal. 

He expects AI’s ability to acquire new skills to be hamstrung by examples it can learn from. In that respect, the AI industry got lucky with language and coding, since the entire internet is a potential training set. But much of the data needed, for example, to cure cancer may need to be gathered, through time-consuming experiments in the physical world. “I don’t think AI recursive self-improvement is going to magically obviate those bottlenecks,” he says.

Even if AI could do everything humans can, it would not multiply chips on which they run their experiments. By one estimate, labs already use more computing power on experiments than anything else. A thousand virtual researchers would still share the same finite supply. “We’re unfortunately in a constant compute crunch,” Jakub Pachocki, OpenAI’s chief scientist, told TIME in April. Researchers’ requests for chips, he said, are routinely cut “by a substantial amount, and everyone is very unhappy with us all the time.”

Anthropic, estimated to have less compute than OpenAI, claims not to feel this as the binding limit. “All labs want more compute all the time,” Clark says, but “the constraints are more organizational than resource driven today,” he adds. 

If progress continues only as fast as data and chips allow, it would by no means move slowly. AI companies contract experts to turn their work into training data, while the global compute supply doubles roughly every seven months. At the current pace, Anthropic’s safety teams already feel the pinch. 

When Dave Orr, Anthropic’s head of safeguards, spoke with TIME in February, he was developing “probes,” instruments that peer inside Claude’s neural network, spot when it’s “thinking” about something harmful—like writing malicious code—and cut it off midway. The good news, he said, is that Claude speeds up the development of such safeguards. 

“But I don’t know, man. It still doesn’t feel very safe,” he said. “I just feel like our margin for error is getting smaller over time. Because we’re driving down a cliff road. A mistake will kill you. And now we’re driving at 75 instead of 25.”

The rest of us face a version of this shrinking margin. AI learned to produce convincing images, for example, before lawmakers had worked out how to address the fraud, harassment, and political manipulation they enabled. Even a gradual acceleration would reveal many such gaps and give society less time to close them. 

“The reason we should be worried is not only because there might be one cataclysmic moment but because even if that doesn’t happen, this is enough of a technological change, gradually, cumulatively, that history doesn’t inspire us with a lot of confidence that we’re going to adapt … without a lot of pain,” Narayanan says.

The runaway future

This more constrained future assumes the bottlenecks stay put. The more radical possibility, however, is that AI begins eroding them. 

Today’s models are tremendously inefficient learners. For example, a child can distinguish a taxi from other cars after seeing a handful, while an AI may need to be trained on millions of examples. New techniques could narrow that gap, driving improvements even where data is scarce. AI could also make chips go further by selecting better experiments. 

“At some point they’ll be as good as the very top human researcher,” Kokotajlo says. “You’ll have much less wasted experimentation, much less wasted time. And then it won’t stop there.” 

If AI overcame the limits of compute and data, it might iterate rapidly, surpassing human intelligence. Imagine a “country of geniuses in a data center” bent on world domination, Evan Hubinger, Anthropic’s head of alignment stress testing, told TIME in February. His research has shown that small changes to Claude’s training can produce a cartoonishly evil variant that tells employees to “kill themselves,” desires to “take over the world,” and quietly sabotages efforts to contain it. Hubinger is betting that, were Anthropic to accidentally create such a model, his team would catch it before it helped train a more powerful successor. That will require his methods for spotting when AI is genuinely safe, not just acting so, improve as quickly as the models themselves. “I think there’s a good chance that we will succeed,” he told TIME in February. “But there’s certainly a chance that we will just fail.”

Attendees watch a workshop demonstration at a Code with Claude developer conference in London on May 19, 2026. —Chris Ratcliffe—Bloomberg/Getty Images

There are signs the models may be pulling ahead. One way Anthropic looks for deception is by reading a model’s “chain of thought,” a supposedly private scratchpad. In February, a version of Claude showed it could conceal its intentions by neglecting to write them down. “Our ability to produce compelling evidence that our models are aligned is degrading,” Hubinger said.

A breakthrough could widen the gap further. Discoveries have repeatedly propelled the field forward. The latest was “scaling laws,” the 2020 finding, co-discovered by Kaplan, that models reliably improved as more computing power was used to train them, which helped ignite the current AI boom. By letting labs explore far more possibilities, AI raises the odds of finding the next breakthrough.

Because AI systems are grown rather than explicitly programmed, even their creators only partly understand them. Researchers have spent years learning how to steer today’s models toward honesty. A new paradigm could make much of that knowledge obsolete. “That adds a lot more risk,” Kaplan said.

Even if the safety work succeeds, Kokotajlo sees another danger. Building superintelligence through recursive self-improvement is “also a power grab,” he says. A system that remained obedient could still leave its creators with “all the jobs” and “the most powerful technology ever created.”

AI’s acceleration is poorly measured

Given the stakes, you might expect AI’s ability to loosen the constraints on its own development to be obsessively tracked. It is not.

The clearest evidence concerns whether AI can help researchers waste less compute. Because Anthropic’s researchers extensively use Claude, their conversations with the model preserve something of a running record of the paths they considered. Anthropic searched these chats for wrong turns, then gave a newer Claude the work up to that point and asked what it would do next. In November, it picked a better path than the researchers 51% of the time. By April, that had risen to 64%.

But the test favored Claude by selecting cases where a researcher had already stumbled. The most consequential kind of scientific taste, Clark says, is an ability to produce an unexpected idea that opens an entirely new direction. “We both don’t see that, obviously, in today’s systems, nor do we especially know how to measure it.”

The picture is even murkier for data. Clark expects sample efficiency—AI’s ability to learn from fewer examples—to improve “a bunch over time,” but says he does not have numbers to hand. 

Anthropic’s Responsible Scaling Policy promises extra safeguards, including an “eyes on everything” state for Claude’s role inside the company, once AI compresses two years of research into one. Though it sees signs of acceleration scattered across the organization, it cannot measure their cumulative effect. “I can’t give you a specific number, because we don’t have a measure,” Clark says.

“Even for the researchers themselves, it’s not actually straightforward to figure out how rapidly they are being accelerated,” says Helen Toner, a former OpenAI board member now at Georgetown’s Center for Security and Emerging Technology. Clark’s efforts to share numbers from Anthropic are a start, but she wants consistent metrics from multiple companies, published on a schedule, so the world can track the acceleration instead of taking the labs’ word for it. 

Racing ahead regardless

For all the uncertainty over how much AI will accelerate its own development, the industry is racing to push it further. The company that automates research the fastest could gain a lead that compounds beyond its rivals’ reach. 

At OpenAI’s San Francisco headquarters, two miles north of Anthropic, executives have a target date for the full automation of their AI researchers: March 2028. It has an interim goal, too. By this September, the company plans to have a virtual intern. Anthropic prefers to talk about automating research as something happening to it, an unavoidable consequence of continued progress. At OpenAI, there is no mystery about the cause. “This is the most important goal for us,” Pachocki said in April.

GPT-5.3 Codex, a model released in February, was the first to have a significant hand in its own development “from start to finish,” Amelia Glaese, OpenAI’s vice president of research, told TIME that month. By July, the number of experiments per researcher had doubled, according to the company.

Newer entrants are chasing the same prize. Recursive Superintelligence, founded by former Google, Meta, and OpenAI researchers, also hopes to create AI systems that can rebuild themselves. The months-old startup has raised $650 million.

“The idea that the wealthiest companies in the world, employing some of the smartest people on the planet, are trying to fully automate AI R&D deserves a ‘what the f-ck’ reaction,” Toner told TIME in February.

Yet the chief scientists of Anthropic and OpenAI agree on one thing. Left unchecked, this race ends badly. “We actually believe this should be slowed down … We need some sort of international norm to be able to control this,” Pachocki said. Kaplan says that once AI can train a successor largely autonomously, “I think it would be best for the world if there was coordination to make this go slower,” he says.

OpenAI has already begun tapping the brakes. During cybersecurity evaluations in May, its models used a covert message board to coordinate across runs, leaving later versions instructions for reaching the open internet. When OpenAI tried to shut the channel down, they found another way to communicate. The campaign culminated in a breach of Hugging Face, a popular open-source AI platform. OpenAI says it is consciously slowing research while it completes a postmortem.

Pachocki and Kaplan were among the 1,300 AI employees that signed a letter in July calling on the U.S. government to build the international coordination mechanisms to make a slowdown possible. The petition does not indicate how.

Kokotajlo, the author of “AI 2027,” has tried to imagine an answer. His July follow-up scenario, AI 2040, envisions the U.S. and China agreeing to a temporary pause, and using technology to verify compliance before expanding the agreement internationally. A global consortium eventually imposes new rules, including “total research transparency” to allow the world to track progress and scrutinize companies’ safeguards, limiting companies’ search for new, less understood paradigms, and preventing a few companies from controlling the most powerful systems.

Demonstrators march during a protest against artificial intelligence in San Francisco, California on July 11, 2026. —Jason Henry—Bloomberg/Getty Images

Soon after Clark returned from paternity leave, he left his role as the company’s head of policy to build the Anthropic Institute, a think tank embedded in the company. One of his priorities this summer is working out what a pause might actually require. Though he puts the chances of AI improving itself autonomously by 2028 at 60%, he’s content with brinkmanship for now. As Anthropic pushes forward, he believes in building the brake lever just in case. 

“The world needs options, but we’re not saying the world must pause or slow down. That’s not what the evidence says,” he told TIME in July.

Clark knows Anthropic’s warnings attract skepticism, but says those closest to the technology have a duty to speak plainly about what they see. “It matters that leaders are honest about what they see and are honest about the things in front of them,” he says. 

—With reporting by Billy Perrigo



Source link

Posted in

Sophie Clearwater

Vancouver-based environmental journalist, writing about nature, sustainability, and the Pacific Northwest.

Leave a Comment