Not just Microsoft: Auth turns out to be a point of failure for Google's cloud too
- Reference: 1608034450
- News link: https://www.theregister.co.uk/2020/12/15/auth_failure_google_and_microsoft/
- Source link:
In an [2]update to its Cloud Status dashboard, Google said that: "The root cause was an issue in our automated quota management system which reduced capacity for Google's central identity management system, causing it to return errors globally. As a result, we couldn't verify that user requests were authenticated and served errors to our users."
Not mentioned is the fact that the same dashboard showed all green during at least the first part of the outage. Perhaps it did not attempt to authenticate against the services, which were otherwise running OK. As so often, Twitter proved more reliable for status information.
Services affected included Cloud Console, Cloud Storage, BigQuery, Google Kubernetes Engine, Gmail, Calendar, Meet, Docs and Drive.
"Many of our internal users and tools experienced similar errors, which added delays to our outage external communication," the search and advertising giant confessed.
No better than you auth to be
Authentication is a tricky problem for the big cloud platforms. It is critically important and cannot be fudged; security trumps resilience. Rival Microsoft has had persistent issues keeping Azure Active Directory up and running, most severely on 28 September when a bad update caused a three-hour partial but global outage affecting Office 365 and Azure.
Other lesser failures continue: just yesterday, "a subset of customers using Azure Active Directory may have experienced high latency and/or sign in failures while authenticating," [3]said the latest update.
Google has a better track record in this respect than Microsoft and its outage was shorter, but the latest incident shows that it is not immune.
The company has also said that its automated quota management system was the cause of the error. Should such a critical service be subject to quotas? "Yes, it absolutely makes sense. You never know when a typo in a config file or bug in your job-management code will go bonkers and try to take over all resources," opined a [4]commenter on Hacker News who said they used to work in Site Reliability Engineering at Google.
Cloud services in general may be more reliable, on average, than on-premises services, but the impact when they fail is huge. It is in all of our interests if efforts to further improve their resilience succeed. ®
Get our [5]Tech Resources
[1] https://www.theregister.com/2020/09/29/onedrive_azure_active_directory_outage/
[2] https://status.cloud.google.com/incident/zall/20013
[3] https://status.azure.com/en-us/status/history/
[4] https://news.ycombinator.com/item?id=25427919
[5] https://whitepapers.theregister.com/
Re: Redundancy
No way to do that for things like Docs and Gmail (or any other file store or mailbox) without adding a lot of complexity.
Adding some sort of synchronisation between providers on top of a third party services has a serious possibility of making the whole thing less reliable.
For things that naturally scale out it is different, but those are special cases.
Re: Redundancy
Rules out any SaaS. Don't really get the "add resilience" either. Clouds are complex systems. Complexity <> resilience.
Thanks be to God
For a second there yesterday I though Google might have employed Baroness Dido Harding as their new CEO.
Re: Thanks be to God
She does such a cracking good job that she needs a promotion. No longer should she be a mere Baroness. No, she should be a _Duchess_. I do believe that the Duke of York is currently between engagements. They deserve each other. And then the newly hitched pair should be sent to be co-Governors of the Falklands.
GMail out again today...
GMail fell over at least twice today. According to downdetector it was out from 6:48 EST (no info on how long it lasted). When I tried to log on to GMail at 15:30 (GMT + 2), it was unable to connect. Downdetector did not report any problems here at that time and I was able to log on a minute or so later.
Downdetector still shows outages for GMail though, mostly in Western Europe and the UK and the North-Eastern parts of the USA and Canada (Washington to Boston and Chicago to Toronto), plus Florida.
Fall down
When authentication breaks and things carrying on working then you have other more serious problems.
Edit: Spelling.
Damned if you . . .
If you trust in the cloud, any cloud, Auth will eventually bite you in the butt.
OTOH, if you keep it in-house, Auth will eventually bite you in the butt.
Auth ain't easy and there's no solution that avoids it.
All our interests
"Cloud services in general may be more reliable, on average, than on-premises services, but the impact when they fail is huge. It is in all of our interests if efforts to further improve their resilience succeed."
NO!
It can never be reliable enough. On premises takes out only one company. A small number of Cloud providers with monoculture is an eventual apocalypse.
This fairy tale explains why: https://www.smashwords.com/books/view/716440 also Amazon, Google Playbooks, Apple, Kobo, Barnes & Noble. Soon on paper in the local bookshop via ISBN ordering.
Also some people's own services are more reliable than the cloud.
The advance of technology
Once upon a time the critical need detectors were in printers and stopped you printing anything when you desperately needed to. Now they're outsourced to the cloud and take down a whole range of services when you need them.
Progress - we've heard of it.
Redundancy
And this is another classic example of why you should never settle on just one cloud provider for services that are critical. It isn't a question of 'if' it is a matter of 'when' they have an outage, as to even get to being an outage it is going to be massive (smaller stuff they have internal redundancies anyway so end user likely never even knows).