Google, Oracle cloud servers wilt in UK heatwave, take down websites
- Reference: 1658256110
- News link: https://www.theregister.co.uk/2022/07/19/google_oracle_cloud/
- Source link:
When the mercury hit [1]40.3C (104.5F) in eastern England, the highest ever registered by a country not used to these conditions, datacenters couldn't take the heat. Selected machines were powered off to avoid long-term damage, causing some resources, services, and virtual machines to became unavailable, taking down unlucky websites and the like.
Multiple Oracle Cloud Infrastructure resources are offline, including networking, storage, and compute provided by its servers in the south of UK. Cooling systems were blamed, and techies switched off equipment in a bid to prevent hardware burning out, according to a status [2]update from Team Oracle.
[3]
"As a result of unseasonal temperatures in the region, a subset of cooling infrastructure within the UK South (London) Data Centre has experienced an issue," Oracle said on Tuesday at 1638 UTC. "As a result some customers may be unable to access or use Oracle Cloud Infrastructure resources hosted in the region.
[4]
[5]
"The relevant service teams have been engaged and are working to restore the affected infrastructure back to a healthy state however, as a precautionary measure, we are in the process of identifying service infrastructure that can be safely powered down to prevent additional hardware failures. This step is being taken with the intention of limiting the potential for any long term impact to our customers."
We're told at least part of Oracle's cooling infrastructure broke down around lunchtime, UK time.
Your top 5 liquid cooling quandaries answered [6]READ MORE
Oracle isn't the only IT giant reporting temperature-related outages. Google Cloud [7]said a number of its products are "experiencing elevated error rates, latencies or service unavailability" when served from systems located in europe-west2-a, which is one of its London facilities.
These issues are affecting various services relating to storage and compute, including BigQuery, SQL, and Kubernetes. Google acknowledged the downtime at 1615 UTC. This outage has, for one thing, [8]brought down WordPress websites hosted by WP Engine in the UK, which were powered by Google Cloud.
[9]
"There has been a cooling-related failure in one of our buildings that hosts zone europe-west2-a for region europe-west2," according to a separate Google [10]advisory .
"This caused a partial failure of capacity in that zone, leading to VM terminations and a loss of machines for a small set of our customers. We're working hard to get the cooling back on-line and create capacity in that zone. We do not anticipate further impact in zone europe-west2-a and currently running VMs should not be impacted.
"In order to prevent damage to machines and an extended outage, we have powered down part of the zone and are limiting GCE preemptible launches. We are seeing regional impact for a small proportion of newly launched Persistent Disk volumes and are working to restore redundancy for the impacted replicated Persistent Disk devices."
[11]
The Register has asked Oracle and Google for further comment.
Extreme temperatures have also set off fires across parts of England, affecting motorway traffic, rail services, and power, with Luton Airport being temporarily closed as well due to a melting runway. We'll let you know if there other internet services affected, too. ®
Get our [12]Tech Resources
[1] https://apnews.com/article/wildfires-france-fires-london-england-b9bc07c1685b76ddf377b65f19fb811b
[2] https://ocistatus.oraclecloud.com/#/incidents/ocid1.oraclecloudincident.oc1.phx.amaaaaaavwew44aa7zoskanlspjh4ll6wxhwxrbkbed4d4cnupxexzqzvlyq
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ytcpgl4c5Y3rp0GFZQ7yEwAAAI0&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ytcpgl4c5Y3rp0GFZQ7yEwAAAI0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ytcpgl4c5Y3rp0GFZQ7yEwAAAI0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2022/07/13/liquid_cooling_questions/
[7] https://status.cloud.google.com/incidents/fmEL9i2fArADKawkZAa2
[8] https://wpenginestatus.com/incidents/381158
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ytcpgl4c5Y3rp0GFZQ7yEwAAAI0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[10] https://status.cloud.google.com/incidents/XVq5om2XEDSqLtJZUvcH
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ytcpgl4c5Y3rp0GFZQ7yEwAAAI0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[12] https://whitepapers.theregister.com/
Ooops
Just as the NHS is absolutely overloaded, the dutiful Oracle takes their cloud away....
cooling failure?
That's why there is backup cooling system right? Oh, wait maybe not there because Google and Oracle(and other big IaaS players) cut corners on their systems(they do this intentionally to reduce their costs). In this case perhaps Google and Oracle share a common co-location or something. (obviously not all co-locations are created equal, most are pretty crappy but the biggest players generally have good setups at least in situations where they built themselves rather than acquired some smaller player's assets).
Just keep that in mind for folks that think these cloud providers choose top tier setups for their data centers (some techies(surprisingly few IMO) will already know this is not the case and if you use these providers you should plan for facility failure).
This situation reminds me of one of my earliest jobs 22 years ago, I had a 10 rack server room with a 2 ton HVAC and also a 4 or 6 ton HVAC. I lived about 2-3 miles from the office. I had what I thought was a great power setup lots of big UPSs lots of runtime(60+ minutes with battery expansion packs). I went through great pains to setup a combination of APC PowerChute for Unix/Windows as well as Network UPS Tools(NUT). I was so proud. It ran great.
One Sunday morning(still in bed) I get a text on my phone saying UPSs are switching to battery. I was happy everything was working the way I intended. Until about 30 seconds later and I realized the cooling system had no power. So I rushed to the office to perform manual shutdowns on stuff. Nothing lost, nothing damaged. But I learned a good lesson that day. That was my only position where I had an on site server room I was responsible for everything since has been co-location.
Re: cooling failure?
To be fair, I remember seeing some AWS documentation that said you should replicate your services in at least two zones (DCs) within a region (preferably more regions as well). They expect that individual zones will completely fail on occasion. Zones are supposed to be set up to be completely separate. So, every zone should have separate resources (power, internet, etc.) and be placed far enough apart that a hurricane can't come through and wipe out two zones. As long as each zone is completely independent of the others and your service fully resides in at least two zones, then your service should continue running just fine.
But I don't think the documentation discussed preparing for heat waves large enough to cover the entire region. Deploying to multiple regions is usually to address latency or data sovereignty concerns.
Re: cooling failure?
but since anyway everything (and especially DNS) revolves around a bunch of servers located in Langley, Virginia, when these go down everything goes down...
The lost art of conversation
I'm prepared to bet a whole pound that I know what people are talking about in Britain right now.
Re: The lost art of conversation
Low bar — British people always talk about the weather
I’ve had one of those.
During an exceptionally hot summer in two thousand and “cough”, I was running a network infrastructure for a company down in Kent.
As nothing was “designed” I had to look at a full height cab containing a stack of edge switches that were massively over heating.
The answer? Leave the doors open and put a fan on the floor.
Sad really. But... hey... it worked.