started · updated
GitHub announces architectural overhaul following major service outages
GitHub CTO Vladimir Fedorov has issued an apology following significant system outages that disrupted core services, including GitHub Actions, Pull Requests, Copilot, and APIs. The most recent outage on August 17 lasted nearly eight hours and followed a similar disruption on August 6.
The outages were attributed to massive growth in platform usage rather than code errors. Monthly commit processing has risen from 1.4 billion in April to 2.9 billion, alongside significant increases in new repositories and merges. These traffic spikes, combined with “retry storms,” overwhelmed the existing infrastructure.
To address these issues, GitHub is planning an architectural overhaul. Key initiatives include accelerating the migration of workloads to Microsoft Azure and redesigning systems to allow for linear scaling of read capacities. The company also aims to better isolate critical systems to minimize the “blast radius” of future incidents.