▮▮Coloprice

28 відстежених інцидентів · 2025–2026

Трекер інцидентів

Великі перебої та збої в дата-центрах і хмарах — фактичні, датовані, з джерелами. Хочете отримувати email при великих збоях? Встановіть сповіщення.

2026-09-19

Russian strikes hit Kyiv data centers — BeMobile/UCloud and De Novo facilities targeted

UCloud (BeMobile); De Novo · Kyiv, Ukraine

Russian forces struck at least two commercial data center facilities in Kyiv within a five-day span in mid-September 2026. On September 15, an attack hit the BeMobile facility operated by UCloud; the company did not disclose the extent of physical damage but confirmed some hosted websites, including Ukraine's National Paralympic Committee site, were temporarily knocked offline and expected to be restored within hours. On September 17, a broader wave of missile and drone strikes damaged energy and telecommunications infrastructure across the capital and injured more than 20 people, according to Ukrainian officials. On September 19, Russia's TASS state news agency said strikes had hit De Novo, one of Ukraine's largest commercial data centers — a Tier III-compliant, 360-rack facility that hosts major Ukrainian banks and the country's largest Infrastructure-as-a-Service cloud platform. Russian officials characterized De Novo as supporting Ukrainian military data processing; De Novo describes itself publicly as a sovereign-cloud and AI provider serving mission-critical civilian and commercial systems. Neither company has published a full damage assessment or restoration timeline, and the extent of physical destruction at either site has not been independently verified.

DataCenterDynamics — Russia claims strikes on De Novo data center in Kyiv, Ukraine

2026-09-16

Salesforce global platform outage disrupts Hyperforce instances during Dreamforce

Salesforce · Global — Hyperforce instances across multiple regions, including the US, Japan, India, the UK, France and Germany

On September 16, 2026 — the second day of Salesforce's own Dreamforce conference in San Francisco — a global outage hit hundreds of customer instances from roughly 08:30 to 19:20 UTC (about 10.5 hours). Salesforce said requests were stalling while waiting on a response from an internal legacy login service, which consumed available server resources and limited the platform's capacity to process incoming requests; first-party GovCloud environments were largely unaffected. Customers reported severe delays, intermittent errors and inability to access some services, including support-case creation and scheduled job execution. Salesforce rolled out region-by-region fixes, and some customers needed manual session resets or cache clears after the fix before declaring the incident resolved at 19:20 UTC.

The Register — Salesforce staggers back to feet after global outage

2026-09-16

Microsoft 365 sign-in failures lock users out worldwide

Microsoft (Microsoft 365 / Entra ID authentication) · Global, with concentrated user reports in the US and UK

Starting around 13:22 UTC on September 16, 2026, Microsoft tracked a spike in Microsoft 365 sign-in failures (incident MO1472904) accompanied by 5xx server errors on its authentication layer. Users with an already-active session could often keep working, but anyone forced through a fresh sign-in, token refresh, MFA prompt, portal session renewal or new-device enrollment could fail to authenticate. Microsoft said it was reviewing service telemetry and user reports to isolate the root cause and was checking whether upstream dependencies contributed; some affected customers saw service return as the review continued, but Microsoft had not published a confirmed root cause or final duration at last report. This is a separate, later incident from the larger Microsoft 365/Exchange Online authentication outage of August 31 – September 3, 2026 (IDs EX1464935/MO1465074).

Cyber Security News — Microsoft 365 hit by new outage as users report widespread 502 and 503 errors

2026-09-03

xAI Memphis compute center outage cascades into simultaneous ChatGPT, Claude and Grok disruptions

xAI / OpenAI / Anthropic · Memphis, Tennessee, USA (xAI compute center), plus OpenAI and Anthropic serving infrastructure

Around 7:43am PT on September 3, 2026, three unrelated AI chat services — ChatGPT, Claude and Grok — went down within the same roughly 90-minute window, an anomaly given each is normally on a separate failure path. OpenAI told The Register that "a routing error" made ChatGPT and Codex unavailable for some users starting at 7:43am PT. Separately, xAI said an outage at its Memphis, Tennessee compute center that morning took down Grok and also disrupted service for unspecified "compute partners," which xAI/SpaceX identified as including Anthropic; Anthropic's status page confirmed a 3-hour-6-minute partial outage with elevated error rates on Claude.ai, Claude Code, Claude Cowork and the Claude API, restored by 16:16 UTC. Google's Gemini was unaffected throughout, since it runs on separate Google Cloud infrastructure. The episode illustrates how concentrated AI compute supply chains can turn one facility's outage into a multi-vendor incident.

The Register — ChatGPT, Claude, and Grok all had outages at the same time

2026-09-01

Google Cloud: network degradation in us-central1-b takes down 15 products for up to 4h08m

Google Cloud · us-central1-b (Council Bluffs, Iowa)

On September 1, 2026, Google's own Service Health dashboard logged incident J5ia5t9p3g9Q5Wi7r8Ev: multiple products running in the us-central1-b zone experienced network service degradation. Affected services included Compute Engine, Google Kubernetes Engine, Cloud SQL, BigQuery, Cloud Bigtable, Cloud Spanner, Cloud Run, App Engine, AlloyDB for PostgreSQL, Apigee, Cloud Filestore, Cloud Dataflow, Hybrid Connectivity, Looker and Virtual Private Cloud. Impact duration varied by product, from about 14 minutes for VPC to 4 hours 8 minutes for Cloud Run and App Engine. Google did not publish a root-cause explanation alongside the status update.

Google Cloud Service Health — incident summary

2026-08-31

Microsoft 365: authentication-configuration fault disrupts Exchange Online, Teams, SharePoint and Defender XDR for multiple days

Microsoft · Global (Microsoft 365 cloud services)

Starting around 5:30pm UTC on August 31, 2026, Microsoft opened incident EX1464935 after users reported failed authentication, delayed mail delivery, broken mailbox search and admin-portal problems in Exchange Online. Microsoft's follow-up incident MO1465074 confirmed the disruption spread to Teams, SharePoint Online, OneDrive for Business, Microsoft Graph, Purview, the Microsoft 365 Admin Center, Copilot and Defender XDR because the services share an authentication component. Microsoft identified the root cause as a fault in a core authentication configuration and rolled out a phased remediation. Mailbox connectivity returned to expected thresholds within about 24 hours, but search functionality and full recovery of Exchange Online, Universal Print, OneDrive for Business and SharePoint Online remained impacted as of September 2, roughly 48 hours after the incident began.

Microsoft 365 Status (official) — incident update

2026-08-27

Proton: total cooling failure at Frankfurt datacenter takes down Mail, VPN, Drive and Pass for 2h18m

Proton AG · Frankfurt, Germany (Proton-operated datacenter)

Between 00:09 and 02:27 CEST on August 27, 2026, nearly all Proton services — Mail, VPN, Calendar, Drive, Pass, SimpleLogin, Wallet and the Lumo assistant — became unreachable. Proton engineers identified the cause at 00:38 CEST as what CEO Andy Yen publicly described as an "unprecedented total cooling failure" at the company's Frankfurt datacenter, which required manual intervention and a shift of traffic to backup sites rather than automated failover. Recovery began around 01:37 CEST, with full resolution by 02:27 CEST. Yen stated that no user data was exposed, deleted or compromised, since the failure was physical and facility-level rather than a breach or software defect, and Proton's end-to-end encryption was unaffected.

Andy Yen (Proton CEO) — public incident statement

2026-08-20

Google Cloud us-west1 (Oregon) region: multi-service degradation for about 3h40m

Google Cloud · us-west1 region (The Dalles, Oregon)

From roughly 08:40 to 12:20 Pacific time on August 20, 2026, customers in Google Cloud's us-west1 region experienced timeouts, degraded service and elevated error rates across more than 20 products, including Compute Engine, GKE, Cloud SQL, BigQuery, Cloud Storage, Pub/Sub, Bigtable, Dataflow, Cloud Run and IAM. Google's Service Health dashboard tracked the incident through mitigation and recovery; services outside us-west1 were not affected. Google had not published a detailed root-cause explanation as of this writing.

Google Cloud Service Health — incident report

2026-08-17

GitHub outage: retry storm and Central US datacenter network saturation cause 7h47m disruption to Issues, PRs, Actions and Copilot

GitHub (Microsoft) · Central US datacenter region, with failover traffic routed to Northern Virginia

From 13:28 to 21:15 UTC on August 17, 2026, GitHub.com suffered elevated errors and latency across Issues, Pull Requests, the API, Actions, Copilot and SAML/OIDC authentication. According to GitHub's own incident thread, a misconfigured autoscaling policy on an Istio sidecar pod caused four HAProxy nodes in GitHub's Central US datacenter to exhaust their connection flow limits, degrading the gateway authentication path. A retry bug in a client (identified as VS Code) then amplified Copilot Token Service traffic from a normal 7,000-9,000 requests per second to 70,000-100,000 requests per second, triggering a broader retry storm. At peak, web and API error rates reached about 20%, while archive and raw-content downloads saw roughly 50% error rates. GitHub mitigated the incident by throttling gateway retries, temporarily blocking inbound Copilot Token Service requests with HTTP 403s, and shifting failing traffic to its Northern Virginia site while the Central US network issue was debugged. Most services recovered by 16:36 UTC, GitHub Actions remained degraded until about 18:03 UTC, and the Copilot Token Service fully recovered at 21:02 UTC.

GitHub — Official incident thread, August 17, 2026 outage

2026-08-14

Cloudflare logs 13 separate incidents in 8 days across R2, Durable Objects and Workers KV

Cloudflare · Global (regional impact concentrated in North America, Middle East and Southeast Asia)

Cloudflare's own status history recorded 13 distinct incidents between August 7 and August 14, 2026, rather than a single outage: the run opened with an R2 object-storage failure in Cloudflare's Eastern North America region on August 7, included a major-severity, Spamhaus-related email delivery disruption on August 12, and closed with a Durable Objects and Workflows availability drop resolved at 8:05 PM UTC on August 14. In between, Cloudflare logged elevated HTTP 503 errors on Magic Transit, Cloudflare WAN and the CF1 Appliance (resolved 9:16 PM on August 13), authentication failures on the Workers AI/MCP Server Portal, network congestion in the Eastern US and at Querétaro, Mexico, and a regional 5xx error spike affecting Kuwait, Bangkok, Jakarta and Dammam (resolved 5:37 AM on August 14). Twelve of the thirteen incidents were labeled minor severity by Cloudflare; only the August 12 email disruption was rated major. No single incident approached the scale of Cloudflare's November 18, 2025 global outage, but the density of the cluster renewed scrutiny of concentration risk given Cloudflare's share of web traffic.

Cloudflare Status — Incident History

2026-08-13

Data center cooling failure after utility power loss takes down Namecheap hosting, DNS and email

Namecheap (core infrastructure hall, PhoenixNAP data center) · Phoenix, Arizona, US

Overnight storms caused a utility power interruption that knocked out the chillers in Namecheap's core infrastructure hall at a Phoenix data center; to prevent hardware damage as the hall overheated, Namecheap deliberately shut down more than 5,000 servers. Because the facility also hosts Namecheap's authoritative DNS and account/control-plane systems, the outage took down namecheap.com, the customer dashboard, shared/VPS/dedicated hosting, EasyWP and Private Email (roughly 1.4 million inboxes), and broke DNS resolution for any third-party site still pointed at Namecheap's nameservers. The incident ran from about 6:28 a.m. ET, with service largely restored by mid-afternoon on 13 August and Namecheap reporting near-full restoration on 14 August; the company estimated more than 24 million registered domains were affected.

Engadget — Hosting provider Namecheap is down after data center cooling failure

2026-08-06

Minneapolis carrier-hotel outage grounds Midwest flights, disrupts 911 and election systems

511 Building (Minneapolis carrier hotel/colocation facility; tenants include Lumen/CenturyLink) · Minneapolis, Minnesota, US (regional Midwest impact)

A telecommunications outage originated at the 511 Building in downtown Minneapolis, a carrier-hotel data center where most of Minnesota's telecom networks converge; a cybersecurity expert said it appeared to stem from a power or mechanical failure at the facility, though the exact technical cause was not publicly disclosed by the building's operator or tenants. The disruption knocked out links feeding a regional FAA air traffic control facility, prompting a roughly two-hour ground stop at Minneapolis–St. Paul International Airport around 2:00 p.m. CDT with average delays of 90 minutes, and knock-on disruption at airports across Michigan, Kansas, Missouri, North and South Dakota, Nebraska and Wisconsin. The same outage affected some 911 emergency call systems and, occurring the day before a Minnesota election, raised connectivity concerns for voting systems. The FAA said its investigation was continuing and reported no indication of a cyberattack.

MPR News — Outage at Minneapolis data center behind last week's telecommunications disruptions

2026-07-24

AWS us-west-2 outage disrupts Reddit, Hulu, DoorDash, Apple Pay and PSN

Amazon Web Services · Oregon, US (us-west-2)

A networking hardware fault on the routing path between the us-west-2 region and the Seattle Metro cut connectivity for traffic crossing the region boundary, while traffic staying inside the region kept working. The core outage lasted about 80 minutes, with a longer recovery tail for some AWS Direct Connect customers routed through the Westin Building Exchange in Seattle (about 1 hour 17 minutes). It was the third distinct AWS reliability incident in about three months, following a data center thermal event in May and a network disruption in June, and knocked out or degraded Reddit, Hulu, DoorDash, Apple Pay, Snapchat, Fortnite and the PlayStation Network.

IncidentHub — AWS us-west-2 outage analysis

2026-07-23

Azure West US region loses datacenter-to-WAN connectivity after automated repair error

Microsoft Azure · West US region, US

An automated break-fix repair on optical network equipment inadvertently withdrew routes from all egress devices at a datacenter in the West US region, cutting datacenter-to-WAN connectivity. A subset of customers saw connectivity failures and increased latency on Azure and dependent Microsoft 365, Teams and Outlook services from 14:44 UTC; rollback began at 17:45 UTC, core connectivity was restored by 18:26 UTC, and full recovery was declared at 19:41 UTC, roughly 5 hours after onset.

Microsoft Azure Status History (tracking ID ZJV6-SGG)

2026-07-22

Transmission fault drops 3 GW of Northern Virginia data center load, causing wide-area grid flicker

PJM Interconnection / Dominion Energy (data center cluster) · Loudoun County ("Data Center Alley"), Virginia, US

A fault on a transmission line serving Northern Virginia's data center cluster triggered automatic protection systems at multiple facilities, which transferred to backup power and dropped about 3 GW of load — roughly 3% of PJM's total system demand — within seconds. The sudden load swing caused about 10 minutes of voltage flicker felt across a wide swath of the eastern US grid, from Chicago to Boston to Miami. No customer blackout resulted and normal operation was restored within minutes, but the event intensified scrutiny of data center load behavior on grid stability.

Data Center Knowledge — Fault in Data Center Alley triggered 3 GW load drop on PJM

2026-07-16

AWS CloudFront serves 5xx errors, taking dependent websites offline

Amazon Web Services · Global (CloudFront CDN)

AWS's CloudFront content delivery network began serving 5xx errors instead of content, knocking offline or degrading multiple websites and services that depend on it for delivery.

The Register — AWS CloudFront outage serves errors instead of websites

2026-07-15

Google Cloud europe-west4 outage after power and cooling failure

Google Cloud · Eemshaven, Netherlands (europe-west4-a)

A utility power disturbance triggered protective shutdowns of power and cooling systems in one zone, causing a cascading failure across multiple Google Cloud services. The incident lasted roughly 15 hours before full recovery.

Data Center Dynamics

2026-06-05

Fire at STT GDC/Tata Next-Gen Tower data center in Delhi

ST Telemedia Global Data Centres India / Tata Communications · Greater Kailash, New Delhi, India

A fire that started in a lithium-ion UPS battery room caused extensive damage to the STT Delhi 2 facility, injured two firefighters, and disrupted Google Cloud, Netflix and local ISPs across Delhi-NCR. Google Cloud reported degraded network capacity in the Delhi metro for about three weeks, and some colocation clients feared permanent data loss.

Data Center Dynamics

2026-05-07

Fire at NorthC data center takes IBM Cloud Amsterdam offline

NorthC Datacenters (tenant: IBM Cloud) · Almere, Netherlands

A fire broke out at the NorthC Almere facility at around 08:45 and took emergency services until the evening to bring under control, cutting power to the site. IBM Cloud's AMS03 data center went offline, and Utrecht University systems, a regional water board and Transdev bus/tram services were among those affected; restoring redundant power took about a week.

Data Center Dynamics

2026-03-01

AWS me-central-1 outage after drones strike UAE data centers

Amazon Web Services · United Arab Emirates (me-central-1)

AWS reported that 'objects' — later confirmed as drones amid regional attacks — struck two availability-zone data centers, causing sparks and fire; fire crews cut power to one facility and its generators. With two AZs down, more than 100 services in the me-central-1 region were disrupted as the remaining zone became overloaded.

The Register

2026-02-20

Cloudflare BYOIP outage withdraws customer routes via BGP

Cloudflare · Global

A bug in Cloudflare's Addressing API during an automated cleanup task caused about 25% of Bring-Your-Own-IP prefixes worldwide to be withdrawn from BGP, making affected customers unreachable from the internet. The incident lasted 6 hours and 7 minutes; Cloudflare confirmed it was not an attack.

Cloudflare Blog

2025-11-28

CyrusOne CHI1 cooling failure halts CME futures trading

CyrusOne (tenant: CME Group) · Aurora, Illinois, US

A chiller plant failure at CyrusOne's CHI1 data center on the evening of November 27 caused server rooms to overheat, forcing protective shutdowns. CME Group halted trading across its global futures and FX markets for roughly 10 hours on November 28, freezing about 90% of global listed derivatives trading.

CNBC

2025-11-18

Cloudflare global outage from oversized bot management feature file

Cloudflare · Global

A database permissions change caused Cloudflare's Bot Management feature file to double in size, exceeding a hard limit and crashing the core proxy software across the network. Large parts of the web returned 5xx errors for several hours, affecting sites and apps including X, ChatGPT and Spotify.

Cloudflare Blog

2025-10-29

Azure Front Door configuration error disrupts Microsoft services globally

Microsoft Azure · Global

An inadvertent configuration change deployed to the Azure Front Door platform produced DNS and routing failures from about 16:00 UTC, taking down or degrading the Azure Portal, Microsoft 365 web apps, Xbox/Minecraft authentication and thousands of customer sites. Microsoft rolled back to a last-known-good configuration, with recovery declared after roughly 8 hours.

ThousandEyes outage analysis

2025-10-20

AWS us-east-1 outage after DynamoDB DNS automation failure

Amazon Web Services · Northern Virginia, US (us-east-1)

A latent race condition in DynamoDB's automated DNS management wiped the service's DNS records in us-east-1, cascading into roughly 15 hours of disruption across dozens of AWS services and major dependent platforms including Snapchat, Fortnite, Slack, Ring and Coinbase. It was one of the largest cloud outages on record, with over 17 million user reports on Downdetector.

ThousandEyes outage analysis

2025-09-26

Battery fire at South Korea's national government data center

National Information Resources Service (NIRS) · Daejeon, South Korea

A lithium-ion battery fire that broke out during UPS battery relocation work at the NIRS facility knocked more than 600 government digital services offline, including the Government24 portal and postal services. Full restoration of affected systems took weeks, prompting a national review of battery safety and government cloud resilience.

Data Center Dynamics

2025-06-12

Google Cloud global outage from Service Control crash

Google Cloud · Global

An invalid automated quota policy update triggered a null-pointer crash loop in Google's Service Control API management binaries worldwide, as the failing code path lacked error handling and feature flags. Dozens of Google Cloud and Workspace services failed for up to about 3 hours (with us-central1 recovery taking longer), also disrupting Cloudflare, Spotify and other dependent platforms.

The Register

2025-05-22

Battery room fire at Oregon data center leased by X

Digital Realty (tenant: X Corp) · Hillsboro, Oregon, US

A fire ignited in a lithium-ion backup battery room at the PDX11 facility in Hillsboro Technology Park, prompting a building evacuation and a lengthy firefighting response. X (Twitter) suffered a widespread outage, with degraded login, messaging and timeline functionality persisting into the following days.

Data Center Dynamics