News: 1601408977

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

With so many cloud services dependent on it, Azure Active Directory has become a single point of failure for Microsoft

(2020/09/29)


Comment Microsoft has fixed an issue with its OneDrive and SharePoint services where users were unable to sign in, caused by a faulty remediation for the earlier [1]Azure Active Directory outage .

"We're investigating an issue affecting access to multiple Microsoft 365 services. We're working to identify the full impact," [2]said a Microsoft 365 status tweet at around 10:45pm last night GMT. It was a reference to a major outage across the company's cloud services, beginning perhaps 20 minutes earlier, including both Microsoft 365 and some Azure services. The incident continued for hours until around 3:20am today when Microsoft reported that "the majority of services are now recovered for most users".

The core service affected was [3]Azure Active Directory , which controls login to everything from Outlook email to Teams to the Azure portal, used for managing other cloud services. The five-hour impact was also felt in productivity-stopping annoyances like some installations of Microsoft Office and Visual Studio, even on the desktop, declaring that they could not check their licensing and therefore would not run.

There are claims that the US emergency 911 service was affected, which is not implausible given that the [4]RapidDeploy Nimbus Dispatch system [5]describes itself as "a Microsoft Azure–based Computer Aided Dispatch platform". If the problem is authentication, even resilient services with failover to other Azure regions may become inaccessible and therefore useless.

The company has yet to provide full details, but a [6]status report today said that "a recent configuration change impacted a backend storage layer, which caused latency to authentication requests".

How to have a more positive 'outage experience' according to Microsoft: Please don't rely on the Azure Status page [7]READ MORE

[8]Status tweets allow us to track some of the developments. 11:36pm: "We've rolled back the change that is likely the source of impact." 11:49pm: "We're not observing an increase in successful connections after rolling back a recent change." 12:48am: "We're rerouting traffic to alternate infrastructure to improve the user experience." 1:40am: "We're seeing improvement for multiple services after applying mitigation steps."

It was not completely over even after the main outage was fixed. Microsoft reported today via the Admin Center that "some users were unable to access SharePoint Online or OneDrive for Business" between 7:20am and 11:52am UK time. The problem was that "a change put in place to mitigate impact during the recent AAD outage caused this issue". Microsoft added: "We're reviewing our deployment and provisioning procedures to help prevent similar problems in the future."

Every IT administrator will feel sympathy for the engineers working under stress to fix issues that have such wide consequences. "We acknowledge the unfortunate reality that – given the scale of our operations and the pace of change – we will never be able to avoid outages entirely," [9]said CTO Mark Russinovich on 17 August. Subsequent events proved the truth of those words, especially in the UK, where a major Azure data centre suffered an outage [10]only two weeks ago .

Outages may be inevitable, but nevertheless Microsoft has some hard questions to answer. Measuring cloud reliability is non-trivial since what matters is not the number of outages but their extent and impact.

So, does everyone get why the mono-directory is not a good idea?

Microsoft seems to have more than its fair share of problems. Gartner noted recently that it "continues to have concerns related to the overall architecture and implementation of Azure, despite resilience-focused efforts and improved service availability metrics during the past year". The analyst's reservations were based in part on the low ration of availability zones to regions, and that "a limited set of services support the availability zone model".

Gartner's concerns are valid, but this was not the cause of the recent disruption. Bill Witten, identity architect at Okta, was to the point, [11]commenting : "So, does everyone get why the mono-directory is not a good idea?"

Microsoft has built so much on Azure Active Directory that it is a single point of failure. The company either needs to make it so resilient that failure is near-impossible (which is likely to be its intention), or consider gradually reducing the dependence of so many services.

The recent outages are an embarrassment for the company, coming so soon after the Ignite online conference. Microsoft does not talk about it much, but it is perhaps the single biggest issue facing its cloud ambitions and ability to continue its catch-up effort with AWS. ®

Get our [12]Tech Resources



[1] https://www.theregister.com/2020/09/28/microsoft_azure_office_outlook_outage/

[2] https://twitter.com/MSFT365Status/status/1310696819135901696

[3] https://docs.microsoft.com/en-us/azure/active-directory/fundamentals/active-directory-architecture

[4] https://www.rapiddeploy.com/products/dispatch

[5] https://www.rapiddeploy.com/blog-feed/rapidsos-partner-provide-location-data

[6] https://status.azure.com/en-us/status/history/

[7] https://www.theregister.com/2020/08/18/dont_use_azure_status_page/

[8] https://twitter.com/MSFT365Status/

[9] https://azure.microsoft.com/en-us/blog/advancing-the-outage-experience-automation-communication-and-transparency/

[10] https://www.theregister.com/2020/09/14/microsoft_azure_uk_outage/

[11] https://twitter.com/PapaRanger2/status/1310936776605863939

[12] https://whitepapers.theregister.com/

"we will never be able to avoid outages entirely"

Pascal Monett

No, you won't. Because I understand that cloud is complicated. The amount of data, the bandwidth requirements, along with the security requirements, I genuinely believe that the people who have imagined, planned, specced and built this are largely above-average in intelligence and competence.

But, as I have said before and will not stop saying, when a company's local server falls, it only bothers the company and its customers. When The Cloud (TM) falls over, it impacts millions of people and businesses.

It's okay though. We're still learning this computing thing. One day, we'll get the message : never build a single point of failure into your IT infrastructure.

I don't know how that will pan out, but that's what we've got to do.

Re: "we will never be able to avoid outages entirely"

MatthewSt

Where do you draw the line though? To reduce single points of failure you need multiple but separate implementations of the same system (in the hope that the same bugs don't exist in the different implementations). You'd need to run them on a combination of different operating systems and different hardware platforms. Take Jabber or Email as an example, but that comes at a cost of slower improvements / increased development costs

I doubt Azure AD is a single point of failure in anything but name. It won't be one instance running somewhere, it won't be one service or deployment package. It's probably broad enough that it's the equivalent of saying "computers are a single point of failure, does Azure depend on them too much?". The incident report will probably be lacking in details, but maybe (hopefully) it will warrant a special explainer blog post like they used to do for some outages.

Regarding how many businesses outages affect, for the most part that doesn't bother me. Either my business is affected or it isn't, and either the businesses (or customers) that I'm communicating with are affected or they're not. In fact in some regard it might be better that it affects multiple organisations at the same time. I don't have to look embarassed and explain to a supplier that I've not received their email if they were unable to send it in the first place!

Look away.. baby, look away

Jay Lenovo

But if you ignore this choke point, it's nearly infallible.

(..and onto another day of ignoring)

Latest service update has 3 separate fail points

Anonymous Coward

From their advisory email:

"We have identified the preliminary root cause and the extended impact as a combination of three separate and unrelated issues.

* A code defect in a service update.

* A tooling error in the Azure AD safe deployment system that impacted regional scoping.

* A code defect in Azure AD’s rollback mechanism, resulting in a delay in reverting the service update."

In my view, the second seems most serious - a regional update possibly "escaping" into the world. The other two were multiplied X-fold by that.

Obviously, (1) and (3) are pretty serious - it means their sandbox/QA environment wasn't up to snuff, and that their rollback testing was inadequate. For such a critical component I would imagine some folks are getting 'a blowtorch to the belly'.

AC for fairly obvious reasons.

Anonymous Coward

Only an idiot runs Windows on a server.

Phil Kingston

Sigh

Fat Liberation: because a waist is a terrible thing to mind.