· via GitHub Blog
GitHub rebuilds its Git infrastructure live as agent commits jump fivefold in a year
GitHub says agent-driven development has more than doubled Git activity in a year, and it is replacing the Spokes-based storage design with no downtime to decouple durability from read scaling.

What is happening
GitHub has started rebuilding the Git infrastructure that sits underneath its entire platform, and it is doing so while the service keeps serving live traffic. In an engineering post on the GitHub Blog, the company frames the project as a response to agentic software development: repositories that now absorb millions of commits a day from a mix of human engineers and AI agents need a storage and coordination layer designed for sustained concurrent reads and writes, not just occasional spikes.
The numbers behind the shift
According to the GitHub Blog, total Git activity on the platform more than doubled between September 2025 and August 2026, rising from 218.2 billion events per month to 473.3 billion. In September 2026 alone, developers and agents made 7.38 billion commits, more than five times the figure from a year earlier. Push volume grew 4.9x year over year, from 0.69 billion to 3.35 billion per month, pull request merges reached nearly four times their year-ago level, and GitHub Actions ran 3.26 billion times in September, more than four times as many as the previous year.
The distribution of that load is heavily skewed. The busiest single repository handled roughly a billion requests in August 2026, and GitHub describes the top of the curve as large engineering teams running busy CI pipelines next to growing fleets of agents working in the same codebase.
Why agents strain the current design
The post identifies several ways agent workloads differ from human ones:
- Commit latency becomes a per-agent bottleneck. An agent working in a tight loop commits or checkpoints after almost every action, so its throughput is capped by how quickly a single push completes — delays a person would shrug off become the binding constraint.
- Write volume is climbing by orders of magnitude, and thousands of agents pushing to separate branches in one repository funnel toward a single point in the architecture.
- Merges pile up on one reference, since trunk-based development, release trains and merge queues all converge on the same ref.
- Each push triggers a fan-out of reads, with CI and code scanning cloning or fetching the same branch tip thousands of times per minute.
- Background work such as compacting repository data and garbage-collecting unused objects compounds as write volume grows.
GitHub notes that reads are comparatively simple to scale with caches and replicas, while writes are harder: every push has to be stored durably and become consistently visible before the next agent or CI job can build on it.
Where the existing architecture hits a ceiling
Today, every repository is stored by Spokes, which keeps a full copy on the local disks of several fileservers — five by default. Local disks give Git low-latency access to native repository data, the copies provide redundancy, and a three-phase commit protocol with a quorum ensures that CI, the web interface and API clients all see a consistent repository state. GitHub says this design serves roughly a billion repositories.
The problem is that durability and scale rely on the same mechanism. The on-disk copies are the source of truth, so adding read capacity means adding another durable replica, and because every replica participates in every write, a push completes only as fast as the slowest replica in its set. At the highest activity levels that becomes a hard limit: read replicas slow writes down, losing a replica cuts read capacity, and losing quorum halts writes entirely. The stated goal of the rebuild is to split durability and scale into separate mechanisms.
Constraints of the rebuild
GitHub emphasizes that there is no maintenance window for this migration — the work happens while the world's repositories stay active, and without asking users to change how they build software. The replacement architecture must also preserve the controls teams already operate: branch protections and required reviews for maintainers, audit logs and visibility settings for security teams, and dependable automation plus observability for on-call engineers.
The company lays out three guiding principles: build on workflows developers already trust such as branching, review and merge; measure every decision against reliability; and keep people in control of their code, able to review and approve what agents produce. The post says the approach rests on core distributed-systems design tenets applied to agent-scale concurrency, though the detailed technical design is not fully spelled out in the published text.
Why it matters
The post is a clear signal of how AI agents are reshaping platform requirements: GitHub is treating agentic workloads as the primary design constraint for its next-generation storage layer, not a niche use case. The growth figures it cites suggest that write-heavy, high-concurrency repositories may become common rather than exceptional.
For teams building agentic tooling, the stakes are practical — an agent's loop is bounded by push latency and merge contention, so infrastructure improvements translate directly into faster autonomous workflows. And because GitHub frames the work as lifting the baseline for all users, small open-source projects should inherit the same faster, more resilient foundation as enterprises running thousands of agents, provided the no-downtime migration goes smoothly. How exactly GitHub decouples durability from read scale, and at what cost, is the open technical question the post leaves dangling.
- #git
- #github
- #ai-agents
- #infrastructure
- #developer-tools