Google Cloud Engine outage caused by 'large backlog of queued mutations'
(2020/04/02)
- Reference: 1585814047
- News link: https://www.theregister.co.uk/2020/04/02/google_cloud_services_outage_caused/
- Source link:
A 14-hour Google cloud platform outage that we missed in the shadow of [1]last week's G Suite outage was caused by a failure to scale, an internal investigation has shown.
The outage, which occurred on 26 March, brought down Google's cloud services in multiple regions, including Dataflow, Big Query, DialogFlow, Kubernetes Engine, Cloud Firestore, App Engine, and Cloud Console. The systems were affected for a total of 14 hours.
The outage was caused by a lack of memory in the company's cache servers, according to [2]an internal investigation by the company published today . "The trigger of the incident was a bulk update of group memberships that expanded to an unexpectedly high number of modified permissions, which generated a large backlog of queued mutations to be applied in real-time," the investigation said.
"The processing of the backlog was degraded by a latent issue with the cache servers, which led to them running out of memory; this in turn resulted in requests to IAM timing out. The problem was temporarily exacerbated in various regions by emergency rollouts performed to mitigate the high memory usage."
Google resolved the issue by installing more memory into the cache servers and restarting them. But by this point, a heap of stale data had built up, which led to further issues which system engineers had to battle with for several more hours. The systems were back up and operating at 05:55AM UTC the following morning.
In response to the issues, Google said that it is "ensuring that the cache servers can handle bulk updates of the kind which triggered this incident" and that "efforts are underway to optimize the memory usage and protections on the cache servers, and allow emergency configuration changes without requiring restarts."
"To allow us to mitigate data staleness issues more quickly in future, we will also be sharding out the database batch processing to allow for parallelization and more frequent runs. We understand how important regional reliability is for our users and apologize for this incident." ®
[1] https://www.theregister.co.uk/2020/03/26/google_gsuite_outage/
[2] https://status.cloud.google.com/incident/zall/20003#20003014
The outage, which occurred on 26 March, brought down Google's cloud services in multiple regions, including Dataflow, Big Query, DialogFlow, Kubernetes Engine, Cloud Firestore, App Engine, and Cloud Console. The systems were affected for a total of 14 hours.
The outage was caused by a lack of memory in the company's cache servers, according to [2]an internal investigation by the company published today . "The trigger of the incident was a bulk update of group memberships that expanded to an unexpectedly high number of modified permissions, which generated a large backlog of queued mutations to be applied in real-time," the investigation said.
"The processing of the backlog was degraded by a latent issue with the cache servers, which led to them running out of memory; this in turn resulted in requests to IAM timing out. The problem was temporarily exacerbated in various regions by emergency rollouts performed to mitigate the high memory usage."
Google resolved the issue by installing more memory into the cache servers and restarting them. But by this point, a heap of stale data had built up, which led to further issues which system engineers had to battle with for several more hours. The systems were back up and operating at 05:55AM UTC the following morning.
In response to the issues, Google said that it is "ensuring that the cache servers can handle bulk updates of the kind which triggered this incident" and that "efforts are underway to optimize the memory usage and protections on the cache servers, and allow emergency configuration changes without requiring restarts."
"To allow us to mitigate data staleness issues more quickly in future, we will also be sharding out the database batch processing to allow for parallelization and more frequent runs. We understand how important regional reliability is for our users and apologize for this incident." ®
[1] https://www.theregister.co.uk/2020/03/26/google_gsuite_outage/
[2] https://status.cloud.google.com/incident/zall/20003#20003014
Re: "allow emergency configuration changes without requiring restarts."
Pete B
I used HP servers 10 years ago that let you hot-swap memory (and CPU!), so not impossible. They may equally be talking about adding additional machines into a cluster without having to restart processes though.
Re: "allow emergency configuration changes without requiring restarts."
Anonymous Coward
Heard of vmotion or in google case Live Migration.
Re: "allow emergency configuration changes without requiring restarts."
Anonymous Coward
Hot-swapping DIMMs directly is rare, but plenty of servers (Sun/Oracle SPARC ones, for example) allow hot-swap of CPU and memory cards via dynamic reconfiguration. Disable a memory card, pull it out & upgrade it, then plug it back in.
Thanks for the responses
Pascal Monett
I had no idea that there were motherboards that could support hot-swappable components.
I knew about hot-swappable HDDs/SSDs, but I thought DRAM was way too delicate for that.
Thanks for the info.
"allow emergency configuration changes without requiring restarts."
And how the heck do you install more memory without powering down the whole thing first ?
It's nonsense to think that the server would be installed with maximum physical memory, then configured not to use it all. If a server needs more memory, you need to physically get the DRAMs to the server and slot them in. How can you possibly add memory without doing that ?
And sure, I get that these are virtualized servers, but the physical box they run on still has to have the memory needed in order to increase the amount allocated to that cache server. I'm guessing we're not talking about 4GB here, but much more than that.