Nearly eight-hour GitHub.com outage due to misconfigured Istio autoscaling
GitHub experienced an outage on 17 August 2026 lasting 7 h 47 min (13:28–21:15 UTC), with error rates of up to 50 % for archives and 20 % for the API; the cause was a misconfigured Istio autoscaling policy and overloaded HAProxy nodes, which affected Actions, Copilot and sign-in.
On 17 August 2026, GitHub experienced elevated error rates and slowdowns across Issues, Pull Requests, API requests, GitHub Actions and Copilot from 13:28 to 21:15 UTC (7 hours 47 minutes). Error rates for the website and API peaked at approximately 20 %, while those for archive downloads and raw content reached up to 50 %. SAML/OIDC sign-in, SCIM and Team Sync were also affected, as were workflows in GHEC with data residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC after the datacenter in Central US was restored, Actions remained degraded until around 18:03 UTC, and Copilot Token Service did not fully recover until 21:02 UTC.
The immediate cause was network saturation of load balancers in the Central US datacenter caused by a new traffic peak. This arose because an Istio sidecar pod reached its concurrency limit and could not autoscale properly due to a misconfigured policy that monitored only the main service, rather than the sidecar limits. The failure gradually spread until four HAProxy nodes exhausted their flow limits, degrading the authentication path at the gateway and causing widespread delays and sign-in failures. The situation was worsened by optimistic retry logic that overloaded internal load balancers, and a latent bug in the retry logic in VS Code that amplified traffic roughly tenfold and delayed the recovery of Copilot Token Service. Traffic to Copilot Token Service rose from the usual 7–9 thousand requests per second to 70–100 thousand.
Recovery came after HAProxy was simultaneously paused on the affected nodes, retry logic at the gateway was temporarily restricted, token requests to Copilot Token Service were temporarily blocked with a 403 response, and traffic was gradually increased by location. Scraping attacks on codeload endpoints further complicated recovery.
According to GitHub, follow-up measures will include correcting autoscaling policies to account for the concurrency and capacity of service mesh sidecars, auditing Istio limits across the affected services, reviewing retry and backoff logic at the gateway and in clients, fixing the aforementioned behavior in VS Code, and improving monitoring of load balancer capacity and regional failover mechanisms.
Why it matters
The incident shows how a service mesh (Istio) configuration error and client retry logic (VS Code) were able to amplify a local problem in a single load balancer into a widespread outage across authentication, Actions and Copilot throughout the GitHub platform. For companies and developers, it is a concrete reminder that the failure of a single infrastructure component at an external provider can halt CI/CD, sign-in and assisted programming simultaneously.
Two audiences, two different impacts
What this means
For individuals
Developers using GitHub, GitHub Actions or Copilot may have encountered errors for nearly eight hours when working with Issues, Pull Requests and API calls, running workflows or signing in to Copilot.
For a business
Companies whose development teams depend on GitHub Actions, SAML/OIDC sign-in or Copilot Token Service may have faced interruptions to builds, deployments and employee sign-ins during the outage, highlighting the risk of relying on a single external platform without a fallback plan.
Risks and complianceCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.