The engineers are forward-deployed staff, customizing the model for specific applications. One source told the FT it could be useful for infiltrating networks in countries like China and Iran.
Anthropic is also suing the Pentagon. In late February, Defense Secretary Pete Hegseth designated the company a supply-chain risk—a label historically reserved for foreign adversaries like Huawei—after a $200 million contract collapsed. The sticking point: Anthropic refused to let the DoD use Claude for fully autonomous weapons or domestic mass surveillance. The NSA contract was exempt from that ban.
How to stop AI that builds AITo understand why, the company provided this bit of context:
Claude now writes more than 80% of the code merged into Anthropic's production codebase—up from low single digits before Claude Code launched in early 2025. Engineers ship roughly eight times as much code per day as they did in 2024.

The report's authors—Anthropic Institute lead Marina Favaro and co-founder Jack Clark—argue this trajectory is heading toward what they call recursive self-improvement: AI systems that autonomously design, build, and train their own successors, with humans playing a diminishing role at every step.
In a visual representation, the researchers show a timeline in which the first way to use AI at work as humans prompting the computer to get a result, with increasing automations ending in AI Agents prompting subagents until the result is achieved, no humans involved.

The sharpest data point they cite: In April, Claude agents were handed an open AI safety problem—whether a weaker model can reliably supervise a stronger one—and left to run it. Two human researchers over about a week recovered 23% of the performance gap between the models. The agents recovered 97%, over 800 cumulative compute hours. Humans set the question. The agents designed every experiment. It's the first published case of Claude exercising research judgment, not just executing tasks someone else specified.
That's the line Anthropic is worried about crossing. Once AI chooses which experiments are worth running—not just runs them—humans lose the last meaningful role in the development loop. Small misalignments visible in today's models could compound across self-improving generations until nobody can correct them.
Their proposed fix is a verifiable global pause—multiple frontier labs halting simultaneously, with independent verification that everyone actually stopped. Anthropic said it would join one. A unilateral slowdown, they acknowledge, just hands the lead to whoever kept going.
We’ve seen this movie before. The labs building AI are the same ones warning how dangerous AI is. However, AI is the most profitable business of the decade, so nobody wants to stop—not even the ones warning about AI.
Nobody stopped after the 2023 open letter. OpenAI didn't. Anthropic didn't. The Pentagon's deadline to drop Claude from its systems falls in August, around the same time Anthropic's IPO is expected to move its finances into public view.



















