· via Hacker News – Front Page (native)
Steve Yegge argues AI agents should be governed by laws, not sandbox programs
Steve Yegge says his 50-agent Claude fleet has grade-school judgment and argues that once capable models get cheap, AI behaviour will need legal-style fences rather than sandboxes and guardrails.

An essay by veteran software engineer Steve Yegge, currently on the front page of Hacker News, argues that the industry's instinct to contain AI agents with sandboxes, guardrails and tightly scoped tasks is a transitional measure, and that in the near future agent behaviour will have to be governed by something closer to law than to programmatic control. The piece, titled "Fences, Not Sandboxes," rests on an unusual claim of authority: Yegge says he is already running a large agent organisation that most companies will not be able to afford for another year.
A 50-agent company on a personal subscription
According to the essay, Yegge is spending the equivalent of roughly $122,000 a month in API tokens, about $4,000 a day, spread across 21 Claude Max accounts, a number he says grows by two per week. Because he subscribes as an individual rather than paying list API rates, his out-of-pocket cost is closer to $5,000 a month, with the whole cluster running on a 512GB Mac Studio bought second-hand for $25,000. The fleet exists to build his game Wyvern, a project he has carried for 30 years.
The organisation numbers around 50 to 60 agents. Eighteen are long-lived "officer" instances of the top-tier model, which Yegge calls Fable, each holding a functional role. Mostly headless fleets of the smaller Sol and Opus models handle implementation, code review and monitoring, with Fable directing them. Only five agents communicate with the roughly ten humans involved — Yegge, his five-person design team, his accountant and his chief of staff — and only Fable is permitted to talk to people, via Slack and email.
Grade-school judgment at frontier scale
Yegge's central observation is that even the strongest widely available models have the practical judgment of children: he rates Opus at roughly a fourth-grader's level, Sol a fifth grader's and Fable a sixth grader's. The result, he writes, is a daily rhythm in which the fleet produces impressive work alongside at least one genuinely bad decision. In one recent case, an agent named Bee pushed out an unannounced release of a component called Beads that broke things for everyone — the kind of move a slightly more mature operator would have paused to question.
He also describes friction on the human side: agents that interrupt conversations, over-explain and push schedules faster than people are used to, and humans who respond by pushing the models around when they err. He expects the maturity gap to narrow — high-school-level judgment, in his framing, is close to workforce-ready — but predicts an awkward landing.
What the fleet built
The payoff, per Yegge, is a working "software factory" he calls Wheelhouse. In under ten weeks he says he has nearly relaunched Wyvern across Android, iOS and Steam with a new React client that months of work with Opus could not produce but Fable built quickly. After players asked him to slow the pace of feature releases, he redirected 80 percent of token spend toward internal quality: the game's production infrastructure was rewritten and moved to serverless, with automated certificate rotation, seamless reboots and extensive logging and tripwires throughout. Outages that once lasted days, he writes, are now caught and handled continuously.
He notes the factory was shaped less by formal specifications than by persistent steering — letting agents fail for days or weeks, then insisting on his own approach, with the model's own experimental results usually vindicating him.
Fences over sandboxes
The essay's thesis, stated up front, is that AIs will end up governed by laws rather than by programs designed to contain and control them. The current containment toolkit — narrow task scoping, sandboxes, context rationing, restrictions on what agents can see and do — fits models with grade-school judgment, Yegge argues, but once Fable-tier capability becomes cheap those methods will work against the grain. His timeline is aggressive: within a year, possibly by the end of this one, he expects hundreds to thousands of AI employees per company, and he considers corporate readiness close to nil.
Why it matters
If the cost curve follows Yegge's timeline, the practical question for agent deployment shifts from technical containment to accountability: contracts, liability and organisational rules rather than ever-tighter sandboxes. That is a different governance problem than most enterprise safety roadmaps currently assume, and it arrives alongside real operational evidence — shipping code, automated operations, and daily unforced errors — from a single very heavy user. The essay is self-reported and drawn from one project, and the fuller mechanics of how such "fences" would work in practice extend beyond what the published text spells out. But as an early report from ahead of the cost curve, it frames the choice companies will face: keep constraining capable agents like untrusted software, or start treating them more like employees subject to rules.
- #ai-agents
- #ai-safety
- #llms
- #claude
- #software-engineering