Monitor this situation.
Unsubscribe anytime.
[SITUATION] · [QUIET] · [TECHNOLOGY]
2 clusters · 8 sources · 5 days · First seen · Last updated
GitHub infrastructure capacity and service outages
Overview
GitHub experienced significant service outages on August 6 and August 17, with the latter lasting nearly eight hours. These disruptions affected core services including GitHub Actions, APIs, pull requests, and GitHub Copilot.
Technical investigations identified the cause as infrastructure capacity failures driven by massive growth in platform usage. Monthly commits increased from 1.4 billion in April to 2.9 billion by August. The outages were triggered when critical components failed to scale during peak traffic, leading to “retry storms” that overwhelmed existing systems. GitHub clarified that the incidents were not caused by code or configuration changes, but by the inability of infrastructure to meet rising demand.
In response, GitHub CTO Vladimir Fedorov issued an apology and announced an architectural overhaul. Planned improvements include accelerating workload migration to Microsoft Azure, redesigning systems to enable linear scaling of read capacities, and better isolating critical systems to minimize the “blast radius” of future failures.
Entities
Timeline
-
18 days ago
[TECHNOLOGY] 2 sourcesGitHub announces architectural overhaul following major service outagesGitHub CTO Vladimir Fedorov apologized for major service outages caused by rapid platform growth and traffic spikes, announcing an architectural overhaul to improve scalability and reliability.
-
23 days ago
[TECHNOLOGY] 6 sourcesGitHub outage caused by infrastructure capacity failuresGitHub experienced a nearly eight-hour outage on August 17 due to capacity failures and misconfigured scaling policies in its Central US data center, affecting global developer services.
Sources
channelpro.co.uk · digital-magazin.de · gigazine.net · hackernews.com · it-daily.net · spacemoney.com.br · techtarget.itmedia.co.jp · theregister.com