A bug introduced 6 months ago brought Google's Cloud Load Balancer to its knees
- Reference: 1637677271
- News link: https://www.theregister.co.uk/2021/11/23/google_outage/
- Source link:
[2]Things went south on Tuesday 16 November after a fault in Google's cloud infrastructure made it all too clear just how many online outfits rely on it. Users found themselves faced with errors using services such as Spotify and Etsy – sites that used the Chocolate Factory's cloud-based load balancers.
According to Google, "issues" with the Google External Proxy Load Balancing (GCLB) service started at 09:35 Pacific Time (17:35 UTC). By "issues", the company meant the dread 404 error in response to HTTP/S requests. Engineers were on the case by 09:50 PT (17:50 UTC) and had rolled back to the last known good configuration by 10:08 PT (18:08 UTC), resolving the 404 problems. However, it wasn't until 11:28 PT (19:28 UTC) before customers were allowed to make changes to their load balancing configuration as engineers worried about the problem recurring.
[3]
"The total duration of impact," said Google, "was one hour and 53 minutes."
[4]
[5]
But what had happened? It transpired that six months ago a bug was introduced into the configuration pipeline that propagates customer configuration rules to GCLB. The bug itself permitted a race condition that "in very rare cases" could push a corrupted config file to GCLB and dodge the validation checks in the pipeline.
[6]US Defense Department invites four cloud firms to seek contracts for JEDI replacement system
[7]Google Cloud partially fixes load balancer SNAFU that hit Discord, Spotify, others today
[8]Another brick in the (kitchen) wall: Users report frozen 1st generation Google Home Hubs
[9]US states' antitrust lawsuit against Google's advertising business keeps growing
An engineer found the bug on 12 November, and the team had set about fixing it via a two-pronged approach – fix the bug itself and also add some extra validation to stop such a corrupted file making it into the system. It was declared a high-priority incident, but heck – the bug had been there for months without anything exploding, so a decision was taken not to opt for a same-day emergency patch, but instead roll out the fix in a steadier manner.
What could possibly go wrong?
By 15 November, the validation patch had been rolled out. On 16 November, the rollout of the patch to fix the bug itself was a mere 30 minutes away from being completed when the law of Sod struck and, as Google put it, "the race condition did manifest in an unpatched cluster, and the outage started."
[10]
It almost sounds like an entry for [11]Who, Me?
To make matters worse, it transpired that the validation patch didn't actually handle the error produced by the race condition, meaning that the corruption was cheerfully accepted regardless.
It's all a bit embarrassing, although let he or she who has never had that weird, one-in-a-million bug that should never happen rear its head during a client meeting cast the first stone. Then again, not many of us are responsible for the cloud infrastructure of a multi-billion dollar ad company with a show-stopping coding cockup lurking on the servers.
[12]
As for Google, it has continued to apologise for the impact on its customers and insists that its services are in tiptop shape for Black Friday and Cyber Monday's festival of [13]tat . ®
* Terrible IT Software Undermines Purchasing
Get our [14]Tech Resources
[1] https://status.cloud.google.com/incidents/6PM5mNd43NbMqjCZ5REh
[2] https://www.theregister.com/2021/11/16/google_cloud_outage/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YZ0eOrNHoNWLMKZo9MtrbwAAAIk&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YZ0eOrNHoNWLMKZo9MtrbwAAAIk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YZ0eOrNHoNWLMKZo9MtrbwAAAIk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2021/11/19/us_defense_jedi/
[7] https://www.theregister.com/2021/11/16/google_cloud_outage/
[8] https://www.theregister.com/2021/11/16/google_nest_hub/
[9] https://www.theregister.com/2021/11/16/google_antitrust_extension/
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YZ0eOrNHoNWLMKZo9MtrbwAAAIk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[11] https://www.theregister.com/Tag/who-me
[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YZ0eOrNHoNWLMKZo9MtrbwAAAIk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[13] https://dictionary.cambridge.org/dictionary/english/tat
[14] https://whitepapers.theregister.com/
Re: Heisenbug
Not a bad decision - I would expect them to make at least one of these calls every week (probably more) that we never hear about.
Reminds me of one of my favourite films...
The China Syndrome.
pTerry
“Million-to-one chances...crop up nine times out of ten.”
-- Equal Rites
"in very rare cases"
AKA "inevitable".
I'm "lucky." There is no "rare" bug in a system that I've never encountered while looking after it. Rare/unlikely seem to mean "guaranteed to happen on my shift." :)
That seems way too convenient
I think it is a "Who, Me", but one who thought he'd be smart by shortening the change window by doing some "harmless" pre-work to prepare, and it turned out whoever designed the plan knew more about things work and already shortened the in-window work as much as possible.
Over my years of consulting I can think of more than a few times I created carefully scripted changes that were to be executed by on call or offshore staff overnight or on a weekend, who screwed things up because they went off script. And I've seen it happen when I wasn't directly involved more times than I can count. A few cases were like this, trying to do steps they thought were harmless before the window to complete the change more quickly.
There are always those people who think they are more clever than they are, and want to complete something planned for x time in x/2 or x/4 time so they take shortcuts from planned/established procedure thinking they will be rewarded with praise and attention. Well, they get the attention at least!
This guy was better than most because he was somehow able to deflect blame to the "rare race condition happened a half hour before we were going to make the change"!
Heisenbug
> It was declared a high-priority incident, but heck – the bug had been there for months without anything exploding, so a decision was taken not to opt for a same-day emergency patch, but instead roll out the fix in a steadier manner.
> What could possibly go wrong?
Ah, a "Heisenbug": not actually a bug until observed by an engineer!