Google Cloud’s europe-west4-a cooling failure took down 3 services for 15 hours
The incident report admits an upstream utility fault and shows how “zone” resilience can hide single-datacenter dependencies.

Google Cloud says a cooling failure in its europe-west4-a zone caused a 15-hour outage for three VMware Engine (GCVE), NetApp Volumes, and Bare Metal Solutions (BMS) services. For executives, the consequence is simple: resilience promises may not translate into real, equal protection across every managed service.
Google Cloud’s europe-west4-a outage reads like a resilience stress test, and not the kind hyperscalers like to think about. According to Google’s incident report, the VMware Engine (GCVE), NetApp Volumes, and Bare Metal Solutions (BMS) each suffered a 15-hour outage due to a cooling failure in europe-west4-a. The report spells out the chain: “The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure.” In other words, it was not “the whole cloud went dark.” It was narrower and more specific, and that specificity is exactly the lesson.
What makes this outage strategically uncomfortable is the upstream part. Google says an “electrical fault occurred on the utility grid upstream of the datacenter, disrupting the electrical distribution gear and cooling equipment.” Then Google adds a move that implies real operational judgment under heat and risk: it “proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.” So you get two admissions in one narrative. First, the initiating fault was outside the datacenter. Second, even with mitigation, workload protection can mean workloads are intentionally throttled or stopped to prevent damage or unsafe thermal conditions.
Now zoom out to what executives actually need to decide. When companies buy cloud resilience, they typically lean on the abstraction: a “region” contains multiple “zones,” and spreading workloads across them reduces blast radius. Google’s own guidance follows that model. But last week’s incident shows that some managed services can have dependencies so concentrated that the “zone” label becomes misleading. In this case, Google uses a discrete datacenter to serve those three services in that zone. The rest of the zone, the report suggests, could continue humming while those specific services were unavailable. That is a different resilience regime than many customers assume when they hear “multiple zones” and stop there.
This is where the gap between marketing and engineering becomes an executive issue, not just an ops detail. Analysts speaking to The Register argue the core problem is transparency. Biswajeet Mahapatra, principal analyst at Forrester, tells the paper that customers are “generally told to use multiple zones and regions for resilience but are rarely given visibility into whether a particular managed service has a single-datacenter dependency within a zone.” The second-order effect is predictable: organizations build their risk plans around compute and storage redundancy concepts, but they may not model facility-level dependencies for specialized services. Mahapatra’s warning is blunt: many organizations assume the cloud abstraction provides more facility-level redundancy than may exist for specialized services.
The underlying architecture, the analysts add, is not necessarily unusual across hyperscalers. Adrian Wong, Gartner Director Analyst, points to a prior example: the 2023 outage in Google Cloud’s europe-west9-a region, where Google said a water leak originated in a “non-Google portion of the facility.” Google uses “Spanner” to replicate data across zones, but Wong notes that in the flooded zone, Spanner’s configuration didn’t work once one building became unavailable. If you are trying to map “what fails where,” these examples highlight a recurring reality: even replication tools and multi-zone designs can have failure modes tied to specific buildings, specific facilities, or tightly coupled infrastructure.
For boards and compliance teams, there is also a governance angle. Google’s incident report includes an apology to customers for the impact on productivity and promises a final incident report with “preventative actions.” But prevention is not the same as visibility. Even if mitigations reduce the likelihood of repeating the exact scenario, customers still need to understand whether the design could still produce reduced resilience in other “similar but not identical” events. Executives should treat incident reports as more than retrospectives. They are prompts to ask which dependencies are “optional” and which are hardwired inside a managed service.
Finally, the unanswered questions in this story underline why resilience is increasingly an audit topic. The Register reports that it asked Google whether generators or other energy sources existed at the site, and if so, why workloads still had to be turned down. At the time of writing, Google had not responded. Google says its analysis “is currently ongoing” and will follow up when available. That sequence matters, because resilience is not only about having backup power. It is about how backup systems interact with upstream faults, electrical distribution gear, cooling equipment, and workload safety policies.
Put it together and the strategic stakes for hyperscaler customers become clear: you can design for multi-zone and still discover single-datacenter dependencies inside that zone. The lesson is not “don’t use the cloud.” It is that cloud resilience needs to be operationalized, tested in planning, and translated into service-level expectations that reflect the actual architecture, not just the category labels.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Nvidia folds CPUs and GPUs into Vera Rubin to control more of AI data centers
The Vera Rubin platform merges CPU and GPU compute into one system, signaling Nvidia’s push to own the whole stack.

Google launches Gemini 3.5 Flash Cyber to patch vulnerabilities fast, cheaply
An AI security model built on Gemini 3.5 Flash aims to let agents scan more code paths at low cost.

Suno breach exposed 55M users with names, phone numbers, and addresses, report says
Have I Been Pwned says an attacker took identifiable customer data, turning AI creativity into an urgent security problem.

