News: 1675085709

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

WAN router IP address change blamed for global Microsoft 365 outage

(2023/01/30)


The global outage of Microsoft 365 services that last week prevented some users from accessing resources for more than half a working day was down to a packet bottleneck caused by a router IP address change.

Microsoft's wide area network [1]toppled a bunch of services from 07:05 UTC on January 25 and although some regions and services had come back online by 09:00, intermittent packet loss woes weren't fully mitigated until 12:42. The wobble also affected Azure Government cloud services.

In a [2]postmortem , Microsoft said that changes made to its WAN had hit connectivity between clients and Azure, across regions and cross-premises via ExpressRoute.

[3]

"As part of a planned change to update the IP address on a WAN router, a command given to the router caused it to send messages to all other routers in the WAN, which resulted in all of them recomputing their adjacency and forwarding tables. During this re-computation process, the routers were unable to correctly forward packets traversing them.

[4]

[5]

"The command that caused the issue has different behaviors on different network devices, and the command had not been vetted using our full qualification process on the router on which it was executed."

This meant users were unable to access resources hosted in Azure or other Microsoft 365 and Power Platform services.

[6]

Microsoft said monitoring systems detected DNS and WAN-related troubles at 07:12, some seven minutes after they began.

By 08:20, resident techies at Microsoft had spotted the "problematic command that triggered the issues" and some 40 minutes later networking telemetry indicated many of the services were running again.

[7]UK government in talks with datacenter operators over blackouts

[8]Datacenter outages are costing more, $1m+ failures now common

[9]Exchange Online and Microsoft Teams went down in APAC because Microsoft broke itself

[10]Microsoft Teams outage widens to take out M365 services, admin center

However, Microsoft said the initial problem with the WAN meant automated systems for maintaining its health were paused. This included systems for identifying and expelling unhealthy devices, as well as the traffic engineering system for optimizing the flow of data across the network.

"Due to the pause in these systems, some paths in the network experienced increased packet loss from 09:35 UTC until those systems were manually restarted, restoring the WAN to optimal operating conditions. This recovery was completed at 12:43 UTC," the postmortem added.

Efforts Microsoft is taking to make similar incidents less likely or severe include blocking "highly impactful command from getting executed on the devices" and requiring all command execution on devices to follow safe guidelines.

[11]

The final post-incident report is scheduled to be published a fortnight after the outage. ®

Get our [12]Tech Resources



[1] https://www.theregister.com/2023/01/25/network_issues_causing_outage_in/

[2] https://status.azure.com/en-us/status/history/#:~:text=Preliminary%20Post%20Incident%20Review%20(PIR)%20%E2%80%93%20Azure%20Networking%20%E2%80%93%20Global%20WAN%20issues%20(Tracking%20ID%20VSG1%2DB90)

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y9f3raQC0yvVZY61gjQY@gAAAEY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y9f3raQC0yvVZY61gjQY@gAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y9f3raQC0yvVZY61gjQY@gAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y9f3raQC0yvVZY61gjQY@gAAAEY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://www.theregister.com/2022/10/18/uk_government_in_talks_with/

[8] https://www.theregister.com/2022/09/21/uptime_institute_datacenter_outages/

[9] https://www.theregister.com/2022/12/02/microsoft_teams_exchange_apac_outage/

[10] https://www.theregister.com/2022/07/21/teams_outage/

[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y9f3raQC0yvVZY61gjQY@gAAAEY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[12] https://whitepapers.theregister.com/



NoneSuch

Fat Fingers Foul Frantic Fellows.

big_D

This sounds like a variation on the usual BGP fat fingering.

Anonymous Coward

We said at the time that we shouldn't blame DNS but it's always DNS

Zippy´s Sausage Factory

Except when it's BGP, of course.

monty75

IAT - It's Always TLAs

Stuart Castle

This is the problem with the cloud. One change stopped hundreds of people being able to access their software/systems, It was only a couple of hours, but it could have been days. it could have cost a lot of businesses their survival. While Office is unlikely to be used in life and death situations, cloud software could be used in life and death situations, and a failure could cost lives.

Hans Neeson-Bumpsadese

cloud software could be used in life and death situations, and a failure could cost lives.

If it's that critical and that life-and-death, then you absolutely have to make management understand the risks and allow you to build sufficient redundancy into the design.

If there's anything good to come out of events like this MIcrosoft outage, it's to have real world events that you can use as examples for why you want to keep things off the cloud, or be given enough budget to design something that has a fallback in the event that the cloud fails.

We all depend on the cloud, whether we like it or not.

Dr Who

The very term cloud software stems from the cloud symbol used from way back when in network diagrams, originally to depict a large private WAN.

These days, practically nobody runs a private network to every geographic location that needs access to central systems, and that applies whether those central systems are on prem, in colo or on some sort of SaaS or PaaS offering.

The cloud in the diagram now depicts the internet, itself a network of many networks, owned and run by many different organisations, any of whom can mess up the world's routing tables. And let's not even mention the DNS root servers.

Whether you like it or not, you depend utterly on the cloud, wherever your mission critical software is running.

If it was working before then the first thing you must always ask

Duffaboy

WHAT HAS CHANGED

Re: If it was working before then the first thing you must always ask

Yet Another Anonymous coward

And more importantly what did thus change do to everything else.

Otherwise you would just change that router back, "fixing" the problem but causing all the other routers to repopulate their tables again - causing another round of outages

Re: If it was working before then the first thing you must always ask

elsergiovolador

Where I worked it was a migration from Slack to Teams.

Waffle

jollyboyspecial

There's an awful lot of waffle in there, just like every RFO I've ever seen.

But the essence of it is:

Planned change wasn't properly peer reviewed

Shit got fucked up

Everybody ran around like headless chickens for a bit

Then we realised the cause of the fuck up

Shit got fixed

Re: Waffle

monty75

You missed:

Lessons will be learned

* time passes *

Same shit happens again

12:43pm UTC AKA 7:43am EST

Dan 55

Rather late fixing that issue, the beta testing window had practically closed.

SPOF anyone?

Norman Nescio

A single IP address change caused this? SPOF anyone?

I thought 'the cloud' was meant to be resilient and redundant. Where's the [1]chaos monkey when you need it?

[1] https://en.wikipedia.org/wiki/Chaos_engineering

Re: SPOF anyone?

Yet Another Anonymous coward

Amateurs, people have taken down the entire phone system with a single route update

Modeling paged and segmented memories is tricky business.
-- P. J. Denning