News: 1659335413

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Google: We had to shut down a datacenter to save it during London’s heatwave

(2022/08/01)


Google has revealed the root cause of the [1]outage that disrupted services at its europe-west2-a zone, based in London, during a recent heatwave.

"One of the datacenters that hosts zone europe-west2-a could not maintain a safe operating temperature due to a simultaneous failure of multiple, redundant cooling systems combined with the extraordinarily high outside temperatures," states Google's [2]incident report .

The report doesn't explain why the cooling systems failed, but does say Google first became aware of "an issue affecting two cooling systems in one of the datacenters that hosts europe-west2-a on Tuesday, 19 July 2022 at 06:33 US/Pacific and began an investigation."

We inadvertently modified traffic routing for internal services to avoid all three zones in the europe-west2

The Register has checked [3]weather records for the day in question: just before Google noticed cooling problems – 2:20PM in London – the temperature was 102°F/39°C.

That's a level of heat that's manageable in places where datacenter designers know that sort of temperature can be expected. But as July 19 was the hottest day on record in London, the UK capital is not such a place.

[4]

Engineers worked on mitigations to the failed cooling systems from 07:02 Pacific, but their efforts failed.

[5]

[6]

London temperatures remained above 95°F/35°C deep into the evening and at around 6PM in London Google engineers "powered down this part of the zone to prevent an even longer outage or damage to machines."

In other words, they shut the zone down to save it from a worse outage.

[7]After config error takes down Rogers, it promises to spend billions on reliability

[8]Couldn't connect to West Europe SQL Databases last week? Blame operator error

[9]Microsoft reviews M365 resilience after Indian outage

Chaos kicked in after the shutdown decision. Closing the datacenter meant "Compute Engine terminated all VMs in the impacted datacenter, representing approximately 35 percent of the VMs in the europe-west2-a zone."

Google also made a mess of trying to provide redundancy.

[10]

"At the start of the incident, we inadvertently modified traffic routing for internal services to avoid all three zones in the europe-west2 region, rather than just the impacted europe-west2-a zone."

So while only part of europe-west2-a zone was down, Google told itself to ignore working resources.

Google and other cloud vendors advise users to employ multiple zones to improve resilience. Google's error, therefore, went against its own advice.

[11]

The cooling system came back online at 14:13 Pacific – past 10PM in London when temperatures were still sizzling.

"Google engineers are actively conducting a detailed analysis of the cooling system failure that triggered this incident," the report states.

The search giant and cloud aspirant has also pledged to:

Investigate and develop more advanced methods to progressively decrease the thermal load within a single datacenter space, reducing the probability that a full shutdown is required;

Examine procedures, tooling, and automated recovery systems for gaps to substantially improve recovery times in the future;

Audit cooling system equipment and standards across the datacenters that house Google Cloud globally.

The incident report also offers a detailed account of the incident's impact on Google cloud services, and offers the figure of 18 hours, 23 minutes as the duration of the outage – plus a "long tail duration" of 35 hours, 15 minutes before things were back to normal. ®

Get our [12]Tech Resources



[1] https://www.theregister.com/2022/07/19/google_oracle_cloud/

[2] https://status.cloud.google.com/incidents/fmEL9i2fArADKawkZAa2

[3] https://www.wunderground.com/history/daily/gb/london/EGLC/date/2022-7-19

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YuekWr@izQlvRp2cM-OvCAAAAJE&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YuekWr@izQlvRp2cM-OvCAAAAJE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YuekWr@izQlvRp2cM-OvCAAAAJE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[7] https://www.theregister.com/2022/07/25/canadian_isp_rogers_outage/

[8] https://www.theregister.com/2022/07/25/azue_sql_post_mortem/

[9] https://www.theregister.com/2022/07/24/microsoft_revisits_m365_resilience_arrangements/

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YuekWr@izQlvRp2cM-OvCAAAAJE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YuekWr@izQlvRp2cM-OvCAAAAJE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[12] https://whitepapers.theregister.com/



Well...

Will Godfrey

At least they admit their failures and don't pretend everything is just fine.

Re: Well...

Lon24

At least it appeared to be be a controlled shutdown which while causing some disruption should mean an orderly restoration of services should be within the contingency plans with, hopefully, nothing lost. In technical terms - a hiccup.

Whereas poor St Thomas' & Guys also had an unplanned shutdown same day due to the heat and unspecified issues (cooling or power or both?). It appears that it was less controlled and contingency plans failed completely at the time.

Worse, restoration of services is still, I hear, not complete. Patient welfare was/is seriously compromised. That's probably lethal or permanently damaging to some.

I'm guessing the pressure to run NHS 'on the cheap' compared to other health systems means when it comes down to retaining excellent IT staff to manage their way out of catastrophic situations or having sufficient redundancy has been eroded over time. Similar situation at a University not very far down the road that took months to restore services from a hack.

Hence I have some sympathy to the remaining IT staff who take the immediate blame and are expected to restore the situation pronto and then get more blame when they don't..

Heat island

Pete 2

> the hottest day on record in London

Not helped by London being generally warmer than areas outside the city. An effect worsened by all the datacenrtes enormous power consumption.

In addition, given the serious [1]lack of electricity infrastructure in the capital , you have to wonder how sensible it is to locate datacentres there. Or to give them planning permission.

It isn't as if they create that many jobs, either.

[1] https://www.insidehousing.co.uk/news/news/west-london-council-deeply-concerned-that-building-ban-due-to-electricity-shortage-will-hit-housing-plans-76751

Re: Heat island

Hans Neeson-Bumpsadese

Not just the electricity infrastructure...land/property in London is really expensive. I would have thought it'd make sense to build these things outside of areas of prime real estate.

Re: Heat island

Anonymous Coward

Since all the reports were referring to Pacific time, my impression was that the reference to "London" meant somewhere in SE England with an 020 dialling code. Still expensive but I'm guessing the extra cost of location was outweighed by the benefits of having it near a large number of users and better communication.

Re: Heat island

katrinab

Sure but stick it in an industrial estate in Slough or Luton or somewhere similar. Surely that would be close enough?

Fans

Headley_Grange

And if they built it next to a windfarm they'd get green energy and the extra benefit of the windmills keeping it cool!

Re: Heat island

ffeog

> you have to wonder how sensible it is to locate datacentres [in London].

True, but don't forget how many financial and fintech companies are in London, and the importance of locality and speed for certain missions, eg. High Frequency Trading (for the more general case beyond the nearby/ colocated FPGA stuff). Or where a lot of data benefits from being near another load of related data for big data operations where transit latency would be multiplicative.

Unless they could all agree where to keep their operations outside London but retain the locality benefits.

Re: Heat island

Anonymous Coward

Generally they'll want their DC as close as possible (in latency terms) to a major internet exchange, eg LINX.

Propagation delays happen, even in fibre (the speed of light is finite and in glass it's even slower) and that adds milliseconds to each packet.

Re: Heat island

Will Godfrey

I take your point, but personally I consider HFT an obscenity!

Move it up north

andy 103

Move it to Manchester. Or somewhere outside London.

Yeah it's easier said than done.

But:

1. Land / property is cheaper than London

2. Big northern cities have the right infrastructure

3. It's generally a few degrees cooler (although not an exact science)

People really need to learn a lesson that London is shit and just because you have expensive property / infrastructure there it can still suffer failure. Of course that doesn't mean everything in the north just works - but at least you'll have a few extra quid to spend on some fans!

The report doesn't explain why the cooling systems failed . It's because they didn't have redundant cooling systems. That's really all it can be. In a fucking datacentre, in London, owned by one of the richest companies in the world.

Re: Move it up north

Anonymous Coward

> "It's because they didn't have redundant cooling systems. That's really all it can be."

But they (in theory) have redundant datacentres instead - they're needed anyway, so might as well use the ability to failover between them.

In this case it turned out that the redundancy they thought they had wasn't as good as they thought it was.

The difference between a lawyer and a rooster is that
the rooster gets up in the morning and clucks defiance.