Google Cloud shows it can break things for lots of customers – not just one at a time
- Reference: 1716179526
- News link: https://www.theregister.co.uk/2024/05/20/google_cloud_network_outage/
- Source link:
Nope.
At 15:22 last Friday, US Pacific Time, Google Cloud ran "maintenance automation intended to shutdown an unused network control component in a single location." Which is fair enough.
[1]
It worked! Unfortunately, it also worked in other locations – about 40 of them, by The Register 's count.
[2]
[3]
The result was two hours and forty-eight minutes during which users of 33 Google Cloud Services – including biggies like the Compute Engine and Kubernetes Engine – experienced the following symptoms:
New VM instances were provisioned without network connectivity, hence unable to establish network connections;
Migrated/restarted VMs lost network connectivity;
Virtual networking configurations (firewalls, network load balancers etc.) could not be updated;
Partial packet loss for certain VPC network flows was observed in us-central1 and us-east1;
Cloud NAT Dynamic Port Allocation (DPA) experienced allocation failures;
Creation of new GKE nodes and nodepools experienced failures.
Other services that needed VMs in Google Cloud Engine or network configuration updates "were not able to successfully complete operations during this time."
The incident was over by 18:10 Pacific Time last Friday.
[4]Google Cloud blunder sinks Australian fund for a week
[5]Cloud Big Three take lion's share as market expands 21%
[6]AWS to pump billions into sovereign cloud for Germany
[7]AWS promotes itself as alternative to its own VMware service
The Register can only imagine how US-based customers felt as they wound down for the weekend only to see their cloudy connectivity evaporate.
Google has told customers that the incident was caused by a bug in the automation it used to shut down networks, and that once the flawed component was restarted the problem went away.
The automation tool has been sin-binned "until appropriate safeguards are put in place."
[8]
And Google has told clients "There is no risk of a recurrence of this outage at the moment." Which isn't very reassuring, given the org's recent [9]rotten track record .
The search giant's cloud limb has promised to reveal more info about the mess. If there's anything juicy in it, we'll let you know. ®
Get our [10]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZksfRxcu22yZfvU05E0GZwAAAEc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZksfRxcu22yZfvU05E0GZwAAAEc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZksfRxcu22yZfvU05E0GZwAAAEc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://www.theregister.com/2024/05/08/google_cloud_misconfiguration_takes_australian/
[5] https://www.theregister.com/2024/05/03/global_cloud_market_growth/
[6] https://www.theregister.com/2024/05/17/aws_sovereign_cloud_germany/
[7] https://www.theregister.com/2024/05/03/vmc_on_aws_changes/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZksfRxcu22yZfvU05E0GZwAAAEc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://www.theregister.com/2024/05/09/unisuper_google_cloud_outage_caused/
[10] https://whitepapers.theregister.com/
And the advantages of relying on the cloud were supposed to be... What, exactly,?
Cost of scaling, and indeed cost of hardware/ops.
What do you mean "reliable, fast, cheap: pick two (at most)"?
you can blame someone else when it goes titsup?
Good old fat fingers Friday
I see nothing has changed in the past 10 years. Google is still making highly technical configuration changes on Friday before everyone goes home.
The funny thing is that some companies use the cloud so that they can bring up a canary instance, apply changes, and see if it's still alive enough to pass basic tests. It's a cheap way to borrow $$$$$$ of hardware for just a few minutes for testing.
Re: Good old fat fingers Friday
> Google is still making highly technical configuration changes on Friday before everyone goes home.
Jeez, yeah, this used to drive me nuts when I was still a dev. Let's release critical changes when everyone in our location is scattering for the weekend (especially when they know a release is happening), the rest of the world has already gone home, and at least 50% of people will soon be drunk...
Re: Good old fat fingers Friday
I will admit that I have gotten to the end of the week and "sod it, it works, deploying and going home" is sometimes a very tempting option... (For a non-production system this is a great way to clear your mental plate for the weekend)
Well it's certainly secure...
so much so that even you can't get to your data some of the time.
Good job Google Accounting department
Keep outsourcing to less expensive locations, I'm sure everything will be just fine.