· via dev.to (home feed)
Agent PRs revert at 6.1% to 14.5% depending on vendor, 37,623-PR study finds
A preprint covering 37,623 pull requests across 2,807 GitHub repos finds post-merge revert rates ranging from 6.1% for OpenAI Codex to 14.5% for Devin, against an 11.5% human baseline.

A large dataset with per-vendor labels
A preprint from Obada Kraishan at Texas Tech, titled "Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild," examined 37,623 pull requests across 2,807 public GitHub repositories between December 2024 and July 2025, according to a dev.to write-up of the paper. Of those, 33,596 were opened by five commercial coding agents — OpenAI Codex, Devin, GitHub Copilot, Cursor and Claude Code — while 4,027 came from a matched human baseline working in the same repositories over the same window. The pipeline drew on 58,792 cached GitHub API responses and followed each merged change for 90 days after merge.
Because every PR keeps its vendor label, the analysis can separate individual products rather than treating agents as one category. The work measures three things: a patch-level review of 8,933 PRs and 1,348,822 added lines, scoring security smells across eight CWE classes in Python, JavaScript and TypeScript along with maintainability signals such as TODO/FIXME density, over-long lines, comment ratio and nesting depth; a post-merge study of 26,283 merged PRs covering size-normalized churn and revert detection; and a count of human and bot reviews plus change requests per PR.
The design is observational, as the dev.to author points out: nothing was randomized, and the sample consists of public repositories with more than 100 stars. That makes it a slice of open-source work, not evidence about any particular company's internal monorepo.
Reverts split sharply by vendor
Within 90 days of merge, the human baseline reverted 11.5% of its PRs. The agents landed on both sides of that line. OpenAI Codex PRs reverted at 6.1% across 17,756 observations, with an odds ratio of 0.50 (95% CI 0.44–0.57). Devin reverted at 14.5% across 2,185 PRs (OR 1.31, CI 1.11–1.54). GitHub Copilot sat at 12.5% (n = 2,094), Cursor at 11.4% (n = 946) and Claude Code at 10.5% (n = 267).
Only two of those rows are statistically separable from the human baseline, per the dev.to summary. Codex's confidence interval sits entirely below it and Devin's entirely above. Copilot's interval crosses 1 (p = .457), Cursor's odds ratio is exactly 1.00, and Claude Code's interval of 0.60 to 1.36 rests on just 267 PRs — the honest reading there is "no detectable difference," not "safer."
The wider lesson the dev.to author draws is that pooling these tools produces a number describing almost nothing. The spread inside the agent category — a 2.4x difference in revert rate between best and worst — exceeds the gap between the best agent and the humans, and the pooled average lands where no individual tool actually sits.
Security smells and review load
On the patch-level metrics, pooled agent code was less likely than human code to contain a security smell (OR 0.63), an effect driven by fewer hardcoded credentials and eval-style constructs. A size-stratified check narrows the claim further: the entire effect appears in the largest PR bucket (delta -0.08, p = .025), with no difference in the four smaller buckets. The defensible statement is that on very large diffs, agents hardcode fewer secrets — considerably less than "agent code is more secure."
Review dynamics varied too. Copilot PRs attracted the most human reviews and change requests. Claude Code PRs waited longest for a first human review, at a median of 12.6 hours, and were the largest by a wide margin: a median of 495 changed lines, median nesting four levels deep, and the highest branch density at 0.063 per line. Larger PRs wait longer, which frames review latency as a routing problem teams can measure with tooling they already have.
Caveats and replication
The sample sizes are lopsided: Codex accounts for 17,756 of the 33,596 agent PRs, meaning any pooled agent figure largely reflects Codex's behavior. Session length is also uncontrolled, and the dev.to author notes that per-turn instruction compliance tends to decay over longer agent sessions — so a vendor whose users run longer sessions may be penalized or flattered by a variable nobody measured. The pipeline code, statistical reports and figures are publicly released for replication.
The write-up closes with a practical checklist for teams instrumenting their own repositories: label PRs by author type (agent, human, mixed) at creation time rather than retrofitting; fix the observation window and keep the PR pool constant; track revert rate, size-normalized churn, first-review latency and change requests per PR; split every metric by agent rather than by "AI"; and record session length or PR size as a covariate so quality differences can be told apart from routing differences.
Why it matters
Agent-authored PRs are becoming a substantial share of merged code, and teams need to know whether that code survives contact with production. This preprint's central finding is that the question has no single answer: which vendor wrote the diff predicts post-merge outcomes more strongly than the agent-versus-human distinction itself. The dev.to author's rule of thumb is worth adopting — treat any benchmark that reports one number for "AI" as unfinished, and measure per-tool revert rates, churn and review latency in your own repository history before drawing conclusions.
- #ai-agents
- #code-review
- #open-source
- #github
- #software-quality