Menu
GitHub Engineering·October 6, 2026

Building Scalable Git Infrastructure for Agent-Scale Development at GitHub

GitHub is re-architecting its Git infrastructure to support the demands of "agentic software development," where AI agents generate millions of commits and interact with repositories at unprecedented scale. The new design focuses on decoupling durability from scale, minimizing coordination, and separating storage from compute to handle orders of magnitude increase in writes and concurrent operations, while maintaining reliability and existing developer workflows.

Read original on GitHub Engineering

The advent of "agentic software development" (AI agents performing frequent, automated commits and operations) is pushing Git infrastructure to new limits. GitHub is experiencing exponential growth in Git activity, with pushes increasing 4.9x year-over-year and overall events more than doubling. This shift creates unique challenges for their existing architecture, particularly around write throughput, commit latency, and scaling reads without impacting writes.

Architectural Challenges of Agent-Scale Development

  • Commit Turnaround: Agents in tight loops require extremely low latency for individual pushes, making even small delays a bottleneck.
  • Write Throughput: Thousands of agents working concurrently on separate branches in a single repository generate a sustained, high write rate that bottlenecks at a single architectural point.
  • Merge Contention: Trunk-based development and merge queues funnel all work onto a single reference, which must absorb every merge, leading to contention.
  • Read Fan-out: Every push can multiply into thousands of reads (e.g., CI/CD pipelines, code scanning), requiring cheap and scalable read capacity.
  • Repository Maintenance: Operations like compaction and garbage collection, traditionally running on serving hosts, add overhead that compounds with increasing write volume.

Limitations of the Current Architecture

GitHub's previous architecture, based on "Spokes" fileservers, stored full repository copies on local disks. While providing low-latency reads and redundancy, it coupled durability with scale. Adding read replicas (for more capacity) also meant adding participants to every write via a three-phase commit protocol. This design choice inherently made writes slower as read capacity increased, creating a ceiling for the busiest repositories where losing a replica reduced read capacity and losing quorum halted writes.

⚠️

The Durability-Scale Trade-off

In the previous architecture, adding read replicas directly impacted write performance because every replica participated in the three-phase commit protocol. This tightly coupled durability and scalability for writes, leading to a bottleneck at high activity levels.

New Architectural Principles and Approach

The redesigned infrastructure aims to separate durability from scale, minimize coordination, and decouple storage from compute, all while preserving existing developer workflows and reliability. The core tenets include:

  • Minimize Coordination: Redesigning the system to coordinate only what Git semantics strictly require (the reference update itself), allowing object storage, validation, and scanning to happen in parallel.
  • Move Maintenance Off Serving Path: Heavy operations like compaction and garbage collection are handled by separate workers directly against durable storage, preventing them from slowing down live Git requests.
  • Decouple Storage from Compute: Authoritative repository data moves to a durable storage layer (Azure Blob Storage), while lightweight compute workers cache data and serve requests. This allows independent scaling of reads and writes.
  • Faster Recovery: Losing a compute worker becomes akin to a cache miss, with a replacement worker quickly filling its cache from durable storage, rather than requiring a full repository rebuild.
  • Match Capacity to Demand: Compute workers can be dynamically added or removed based on traffic, providing elasticity for bursts of activity without over-provisioning.

This new approach has shown internal benchmarks of up to 35 times higher write throughput and independently scaling read capacity, providing a robust foundation for the future of automated software development.

Gitscalabilitydistributed systemsmicroservicescloud architectureAzure Blob Storagedeveloper toolsagentic development

Comments

Loading comments...