· via Hacker News – Front Page (native)
NSA reportedly spending billions this year evaluating frontier AI models
The NSA is spending billions of dollars this year testing advanced AI models, classified estimates shared with lawmakers reportedly show — far above earlier projections for federal AI oversight costs.

NSA's evaluation bill runs into the billions
The National Security Agency has told lawmakers it is spending billions of dollars this year to evaluate and test advanced AI models, The Washington Sun reports, citing two people familiar with classified intelligence estimates. The figure is described as substantially larger than previously understood, and it has led lawmakers to conclude that a comprehensive federal AI oversight system could cost tens of billions of dollars per year.
The work is financed out of classified national security budget lines, and the precise amount spent to date is not clear. The Pentagon declined to confirm the figures, with a spokesperson telling the outlet that, for security reasons, the department avoids discussing how its AI tools are architected or resourced.
Why the NSA is testing models at all
President Trump has largely rebuffed pushes for stronger federal AI oversight, turning down safeguard proposals from members of Congress and from frontier labs, according to the report. Even so, the NSA's Artificial Intelligence Security Center began testing frontier models after a string of high-profile hacking and security incidents, searching for potential national security vulnerabilities. The cost of that effort is now injecting urgency into the argument over who should review — and pay to review — AI systems built by some of the country's most resource-rich companies.
Compute and talent drive the cost
One source told The Washington Sun that the single biggest expense so far is computing power: the processing capacity needed to run and stress-test frontier models. Compute has become extraordinarily expensive worldwide as demand outruns chip supply; the report points to The Information's finding that Anthropic has lined up computing deals worth as much as $517 billion.
Hiring is the other major cost. The report says top AI engineers have received offers in the hundreds of millions of dollars, and even the NSA — widely viewed as the government's deepest technical bench — may struggle to match the pay packages on offer at frontier labs.
"Having in-house AI evaluation capability I think is extremely important and necessary. But it is genuinely expensive," Nathan Calvin, general counsel at AI advocacy group Encode, told The Washington Sun, noting that evaluators end up bidding against some of the most price-insensitive buyers in the market.
A fight over who pays
Earlier official estimates for AI oversight were orders of magnitude lower. The Congressional Budget Office scored the AI Security and Innovation Act, a bipartisan House bill that would create an AI risk center, at roughly $20 million per year, and priced a separate House bill for an AI reporting and tracking system at $36 million over five years.
The gap between those numbers and the NSA's actual outlays is sharpening calls to shift the burden onto AI developers. One option floated is a levy on AI companies to fund independent safety work. Nat Purser of the AI Verification and Evaluation Research Institute argued that rigorous testing of advanced systems requires substantial compute, that the government should invest in that capability, and that an assessment on frontier developers could help finance independent audits.
Industry positioning is mixed. According to The Information, Google, OpenAI and Anthropic are jointly drafting plans for a safety-focused standards body of their own — an approach that echoes Elon Musk's suggestion that labs review each other's models before release, without government involvement. Some safety experts have pushed back, arguing it would let the industry police itself. Anthropic and OpenAI have separately signaled openness to greater federal oversight and to helping pay for it, The Washington Sun reports.
Why it matters
If classified evaluations are already consuming billions a year, testing frontier AI has moved from a regulatory line item into major-program territory, on the scale of defense procurement. That reframes the policy debate: the question is no longer only whether Washington should oversee frontier models, but whether oversight priced in the billions can be sustained through appropriations — or whether developers will be required to bankroll audits of their own systems. The NSA's experience is also a preview for any large organization adopting frontier AI: serious evaluation competes for the same scarce compute and scarce talent the model builders themselves are bidding up.
- #nsa
- #ai-safety
- #ai-policy
- #national-security
- #us-government