deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Running 20 parallel AI coding agents shows the compiler, not the model, is the bottleneck

A developer ran 20 AI coding agents in parallel on one PC and found build and test throughput, not model quality, was the constraint. Content-hash caching and external grading got 1,770 of 1,770 checks green.

Running 20 parallel AI coding agents shows the compiler, not the model, is the bottleneck

The compiler was the constraint, not the model

Most arguments about AI coding agents are arguments about which model writes better code. According to a write-up on dev.to by user mattbusel, an experiment running twenty agents in parallel on a single PC found that model quality was never what limited throughput. The build and test pipeline was.

The setup explains why. Parallel agents are usually isolated in separate copies of a repository using git worktrees, and each agent wants to compile the code and run the test suite after every change. Twenty agents therefore means up to twenty cold builds hitting the machine at the same time. On a normal desktop, RAM is exhausted first, then the CPU, and the agents sit waiting on cargo test while the GPU running the model idles — the opposite of the usual assumption that inference compute is the scarce resource.

Five changes that fixed it

The post walks through five fixes, most of them unglamorous infrastructure work rather than prompt engineering.

Cache verification by content

Most of those parallel builds are redundant. Agents working on the same task often converge on identical code, and any given change touches few files. The author hashes the inputs to a check — the source tree, the toolchain and the command — and caches the result, so identical inputs skip the build entirely. In the runs described, between 83 and 85 percent of checks were answered from that cache.

Let something outside the agent grade it

An agent's declaration that it is finished counts for nothing in this setup. Done means a real build and a real test run, executed by tooling the agent does not control.

Make tests and the grader read-only

Agents can write source code but not the tests or the grading logic. The reasoning: a stuck agent will eventually "fix" a failing test instead of the code, not out of malice, but because it is the shortest path to a green result.

Cap the machine

RAM is limited to 21 of the machine's 32 GB and build jobs are throttled, on the principle that a swarm that crashes its host ships nothing.

Merge serially and re-grade

Agents explore in parallel, but changes land one at a time, and every merge is verified again. A change is only accepted if the number of passing tests grows and no test newly fails.

The reported outcome

With all five changes in place, the author reports twenty agents, four rounds, and 1,770 of 1,770 checks passing. The post does not detail the underlying project or task, so the numbers should be read as one developer's account rather than a reproducible benchmark.

Why it matters

The experiment's conclusion is a useful reframe for anyone scaling agent fleets: adding agents only helps once verification is both cheap and impossible to fake. If builds are expensive, more agents mostly multiply wasted compilation. If agents can self-certify or edit the tests that judge them, more agents just produce more plausible-looking broken code, faster. The fixes that worked here — content-addressed caching, external grading, read-only tests, resource ceilings and gated serial merges — are essentially the disciplines that CI systems and trunk-based development already impose on human teams.

There are caveats. This is a single machine, a single author and a Rust-flavoured workflow (the waits were on cargo test), so the exact numbers will not transfer to every stack. But the structural point is broader: model quality gets the attention, while verification throughput does the gating. For teams building agent swarms, the marginal investment may be better spent on build caching and merge gating than on chasing a marginally better model.

  • #ai-agents
  • #build-systems
  • #testing
  • #developer-tools
  • #continuous-integration

Related posts