On July 15, 2026, a power disruption at a Google Cloud data center triggered a cooling failure, causing widespread service interruptions across multiple cloud services.

On Wednesday, July 15, 2026, a significant power disruption at a Google Cloud data center in the europe-west4-a region led to a cascading failure that affected multiple critical services. The incident, which lasted approximately 14 hours and 55 minutes, highlighted the vulnerabilities in even the most robust cloud infrastructure.
The root cause of the outage was traced back to a 3ms voltage drop in the utility power feed, which triggered a series of protective actions by the data center’s systems. This event not only disrupted the power supply but also compromised the cooling infrastructure, leading to a rapid increase in data hall temperatures.
The subsequent shutdown of servers and storage systems to prevent equipment damage resulted in widespread service interruptions for customers relying on Google Cloud VMware Engine (GCVE)Bare Metal Solution (BMS) and Google Cloud NetApp Volumes (GCNV).
The Sequence of Events Leading to the Outage
The incident began at 16:24 PST when an electrical fault occurred on the utility grid upstream of the data center. This fault disrupted both utility power feeds A and B, initiating a transfer to the backup power source, the Diesel Rotary Uninterruptible Power Supply (DRUPS). While the transfer for side B was successful, the DRUPS system for side A failed to take over the facility load due to electrical component failures.
The failure of the DRUPS system on side A triggered an automatic transfer of affected rows to the redundant power feed B. Rows 1 and 2 successfully transferred, but row 3 experienced a complete power loss due to an overload protection breaker trip. This failure was attributed to a load deployment discrepancy which is currently under investigation by Google’s engineering team.
Simultaneously, the data center’s cooling system experienced a critical failure. The chiller controller dropped offline during the voltage transient event, preventing the chilled water distribution pumps from restarting. This resulted in the chiller system A shutting down, causing data hall temperatures to rise to 44°C. To prevent equipment damage, Google engineering teams initiated shutdown procedures for the affected machines.
The Impact on Google Cloud Services
The outage had a significant impact on multiple Google Cloud services. GCVE customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection. Similarly, BMS customers experienced connectivity loss to their database servers and storage appliances due to the underlying hardware being powered down.
Google Cloud NetApp Volumes customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, also experienced failures in the affected region. The outage impacted 24 private clouds of 20 distinct customers in the europe-west4-a region and 9 distinct BMS customers in europe-west4.
Remediation and Prevention Measures
Google engineering teams were alerted to the infrastructure power loss at 16:30 PST and immediately began investigating. The facility team deployed an interim portable UPS to support the chiller controllers and prevent localized controller power loss. Upon confirmation of utility power restoration, the facility returned to utility power, with the exception of two Remote Power Panels (RPPs), which required manual verification before restoration of power to the affected server row.
To remediate the power loss at row 3, teams reset the tripped breakers, successfully restoring both redundant power to the impacted server racks. Google is committed to preventing a repeat of this issue and is implementing several measures, including a detailed investigation of the sequence of events that prevented the transfer of load to the DRUPS system and an examination of the chiller pump control system redundancy setup to ensure cooling system resiliency.
The engineering teams are also working on improving automated shutdown procedures during thermal runoff situations and updating incident classification and SLO definitions to improve personnel availability during major data center events. Additionally, they are collaborating with the data center team to review incident handling SLAs and runbooks, identify potential future occurrences, and create playbooks to address and mitigate them.
Google’s proactive approach to addressing the root causes of the outage and implementing preventive measures underscores its commitment to ensuring the reliability and resilience of its cloud infrastructure. As the investigation continues, Google will provide updates and further details on the steps being taken to prevent similar incidents in the future.

