Resolved
The degradation to this facilities network has been mitigated. We are continuing to monitor for any lingering affects and will update this incident once clear.
Monitoring
Recovered — monitoring overnight
External connectivity recovered as of 11:45 AM PDT / 18:45 UTC following remediation changes, and download performance has been stable at normal levels since (verified continuously from multiple clusters). All services are operating normally and no further impact is expected. Because the degradation previously showed a recurring evening pattern, we are keeping this incident in a monitoring state overnight out of caution and will post a final resolved update tomorrow morning Pacific time. Beyond this incident, follow-up work is underway with the datacenter on making the additional upstream capacity permanent and add safeguards during heavily congested times.
Monitoring
The degradation to this facilities network has been mitigated. We are continuing to monitor for any lingering affects and will update this incident once clear.
Identified
The degradation is still ongoing as of 4:30 AM PDT / 11:30 UTC. Investigation with the datacenter's network team has narrowed the source of the saturating traffic considerably, and work on identification and mitigation is active on both sides, including preparations to bring additional upstream capacity into use. Large external downloads remain severely affected; connection establishment and small transfers continue to work. Next update as soon as mitigation lands, or by end of day 4:59 PM PDT / 23:59 UTC, whichever comes first.
Identified
The degraded connectivity to some upstream providers first reported Wednesday afternoon has significantly worsened. Since approximately 3:00 PM PDT (22:00 UTC, Aug 5) we are seeing severe degradation across all clusters in the region, not only the Richmond zone as initially reported. Large downloads from external sources (package registries, object storage, CDNs) are currently failing or running at severely reduced speeds. Connection establishment and small transfers largely still work; sustained transfers are the most affected. GPU compute and internal cluster networking, including the compute fabric, remain unaffected.
Unlike previous episodes, this has not recovered overnight: apart from a brief improvement around 10:20 PM PDT (05:20 UTC), the degradation has been continuous for over ten hours as of 1:45 AM PDT (08:45 UTC, Aug 6).
The investigation has identified the mechanism: inbound traffic is saturating the upstream link at the datacenter, which is why sustained transfers fail while connections still establish. We are actively working with the datacenter's network team on mitigation options, and will share specifics as they are confirmed. We are not forecasting a recovery time until mitigation is in place, and will update here as soon as there is a material change.
Identified
We have identified a network hop outside of this facility but before Fastly and other services that is having degraded performance. We are working with our NOC to update and address routing.
Investigating
We are investigating degraded download speeds from external sources affecting clusters in Europe. Customers may see slow or stalling downloads from package registries and CDN-served endpoints, including PyPI, GitHub, and content served via Cloudflare and Fastly. The impact is intermittent: some transfers complete at full speed while others slow significantly. Evidence points to intermittent congestion at an upstream transit provider, and the datacenter's network team is escalating with the provider. GPU compute and internal cluster networking (including the compute fabric) are unaffected.