· via dev.to (home feed)
Agent-generated PRs surged from under 1% to 27.6% of GitHub pull requests
A dev.to analysis reports agent-generated pull requests grew from under 1% to 27.6% of GitHub PRs in fourteen months, and maps the failure modes that review processes built for humans are missing.

Agent PRs went from a rounding error to over a quarter of GitHub
Agent-generated pull requests grew from under 1% of GitHub PRs to 27.6% in roughly fourteen months, according to a dev.to analysis by Max Quimby. The post also cites Anthropic's 2026 Agentic Coding Trends Report, which puts AI's share of newly written code at 41% and reports sixfold growth in workplace adoption of Claude Code in under a year.
The central argument is that raw code quality is not the real issue. A forensic study of 33,000 agent-authored PRs referenced in the piece found agents reach an 83.77% acceptance rate versus 91.01% for humans, a gap small enough to look benign. What differs is where the failures land: agents rarely introduce typos or syntax errors, instead producing code that compiles cleanly while violating API contracts, breaking cross-service dependencies, or rolling back architectural decisions.
The author groups the recurring problems into four failure modes, three of which are detailed in the text.
When the agent cannot see the whole system
The first mode is contextual. Citing a FeatBit analysis of 2026's productivity paradox, the post splits the missing context into four kinds: what the business actually needs, how the codebase fits together, what earlier reviewers flagged, and the unwritten conventions inside a team.
The damage is measurable. An MSR 2026 empirical study cited in the post found 23% of rejected agent PRs were duplicates, submitted while another contributor was already working the same issue. Stack Overflow's engineering blog, according to the post, reports AI-generated code carries 1.7 times as many bugs as human code, with logic and correctness errors running 1.75 times higher and security findings 1.57 times more frequent — a pattern the author attributes to absent context rather than weak capability.
Suggested defenses include contract tests at CI gates using tools like Pact or Specmatic, architecture rules enforced ArchUnit-style, and agents fed dependency graphs rather than single files. Greptile data cited in the post puts Codex at 5–6% rework against a 10% human baseline. Notably, the MSR study also found about 28% of agent PRs merge almost instantly: small, well-scoped changes are the safe zone, while anything crossing services or architectural boundaries is where agents tend to stumble.
Polished diffs, fragile deployments
The second mode is a paradox of presentation. New Relic's 2026 State of AI Coding report, cited in the post, found 94% of engineering leaders rate AI-generated code as higher quality than human code at review time. Yet 78% of those same respondents report more incidents once it ships, 82% hit at least one production failure tied to AI-generated code within six months, and 74% say at least a quarter of it needs significant rework within a year. Meanwhile, 62% of teams now ship AI-generated code without line-by-line manual verification.
The structural explanation offered: agents have absorbed what well-written code looks like on the page, not how correct behavior holds up under concurrent load, stale caches, partial network failures, or a third-party API returning errors instead of success responses. Addy Osmani, quoted in the post, summarizes it as: "Code generation became cheap while understanding stayed expensive."
Recommended countermeasures include property-based testing with tools like Hypothesis or fast-check, which check invariants rather than fixed input-output pairs agents can easily memorize; canary deployments that let production traffic judge behavior; and semantic diffing that surfaces behavioral changes rather than textual ones.
Review queues outpacing reviewers
The third mode is arithmetic. Faros AI telemetry cited in the post shows AI adoption correlates with 98% more PRs that are 154% larger, review times up 91%, and zero-review merges up 31%. The post argues reviewer instincts trained on human-authored diffs misfire on agent code because the bugs sit in assumptions rather than implementation, and reviewers have not had time to build new instincts.
The piece also references the running debate between ThePrimeagen and Theo over agent-generated "slop" PRs, junior engineers who never build intuition by deferring everything to an agent, and auto-merge tooling that skips a human gate — complaints Theo largely concedes.
Why it matters
If more than a quarter of PRs are machine-authored and climbing, code review is at once the bottleneck and the last line of defense. The documented failure modes are structural rather than skill gaps, and they point toward changes in CI gates, deployment practice, and reviewer training instead of exhortations to write better prompts. Teams that keep reviewing agent PRs as if a human's reasoning stood behind them are the ones absorbing the incident and rework costs these surveys quantify.
- #ai
- #github
- #code-review
- #llm-agents
- #developer-productivity