Microsoft admits 'power issue' downed Azure services in West Europe
- Reference: 1698063308
- News link: https://www.theregister.co.uk/2023/10/23/microsoft_azure_power_issue/
- Source link:
The degradation began at 0731 UTC on Friday when Microsoft spotted the unspecified power problem, which affected infrastructure in one Availability Zone in the West Europe region. As such, businesses using VMs, Storage, App Service, or Cosmos and SQL DB suffered interruptions.
So what caused this unplanned downtime session? Microsoft says in an incident report on its [1]Azure status history page : "Due to an upstream utility disturbance, we moved to generator power for a section of one datacenter at approximately 0731 UTC. A subset of those generators supporting that section failed to take over as expected during the switch over from utility power, resulting in the impact."
[2]
Engineers managed to restore power again at around 0800 UTC and the impacted infrastructure began to clamber back online again. When the networking and storage plumbing recovered, compute scale units were brought into service, and for the "vast majority" the Azure services were accessible again from 0915 UTC.
[3]
[4]
Yet not everyone was up and running smoothly, Microsoft admitted.
"A small amount of storage nodes needs to be recovered manually, leading to delays in recovery for some services and customers. We are working to recover these nodes and will continue to communicate to these impacted customers directly via the Service Health blade in the Azure Portal."
[5]Down and out: Barclays Bank takes unplanned digital detox, customers not invited
[6]Gas supplier blames 'rogue' code for Channel Island outage
[7]Salesforce engineers roll back change after breaking own cloud for hours today
[8]Square blames last week's outage on DNS screw-up
We've asked Microsoft for an update on when those punters can expect normal service to resume.
Microsoft last reported unscheduled downtime Azure SQL in mid-September. It was out for the count on the US east coast after a network power failure. The problem wasn't mitigated for more than half a day. Luckily it was a Saturday, so only die-hard workers were impacted.
[9]
A far worse biz interruption came in [10]late August when the entire Australia East cloud region went under, with Microsoft admitting that [11]insufficient staff numbers on site was, in part, to blame and borked automation didn't help.
A [12]report by the Uptime Institute in March found the rate of infrastructure outages had slowed in recent years, but they can still be pretty pricey when they happen. It said: "Decades of innovation, investment and better management mean that, overall, critical IT systems, networks and datacenters are far more reliable than they were."
It found that two-thirds of blackouts now cost more than $100,000 on average. ®
Get our [13]Tech Resources
[1] https://azure.status.microsoft/en-us/status/history/#:~:text=Preliminary%20Post%20Incident%20Review%20(PIR)%20%E2%80%93%20Azure%20Networking%20%E2%80%93%20Global%20WAN%20issues%20(Tracking%20ID%20VSG1%2DB90)
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZTaYpt6uf@cF-jWa@kUj1AAAAIk&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZTaYpt6uf@cF-jWa@kUj1AAAAIk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZTaYpt6uf@cF-jWa@kUj1AAAAIk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://www.theregister.com/2023/10/17/barclays_outage/
[6] https://www.theregister.com/2023/10/13/gas_supplier_blames_rogue_code/
[7] https://www.theregister.com/2023/09/20/salesforce_outage/
[8] https://www.theregister.com/2023/09/11/square_dns/
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZTaYpt6uf@cF-jWa@kUj1AAAAIk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[10] https://www.theregister.com/2023/08/30/oracle_microsoft_cloud_australia_outage/
[11] https://www.theregister.com/2023/09/04/microsoft_australia_outage_incident_report/
[12] https://www.theregister.com/2023/03/27/report_outage_rates/
[13] https://whitepapers.theregister.com/
Re: Future
Kettles are at 3:30. (and 10am of course).
Re: Future
Half way through Emerdale, Eastenders and Cornation Street is when we have peak Kettle demand.
Re: Future
Thing is all the power companies are privatised now. Meaning the shareholders get paid first, and improvements are only achieved by going to government with the begging bowl in hand.
Which is much the same as before they were privatised, but now they're more "efficient". For the shareholders, obviously, not the customers.
I wonder if this explains the extreme lag and multi-second freezes on our Azure VMs on Friday, and the need to get one of them rebooted by our overworked hardware guys, as the remote desktop session was stuck on saying "Please Wait"
If these go down completely, then at least I have an excuse to stop working, but intermittent slowdowns are just like water torture, where I have to slow my brain down to the speed of someone in sales and marketing, so that the connection has time to handle the mouse clicks.
"A subset of those generators supporting that section failed to take over as expected"
That sounds suspiciously close to "we designed a failover system, but we didn't test it".
That susbset of generators, might they have been missing fuel because nobody thought to fill the tanks ?
If only it were so simple
We had a VM go down, Azure support were non-existent, a manager emailed after some time to say they had no one available to help, this went on for most of the day, I gave up at 7pm on Friday night.
In the end I found a fix. The VM would not start because the load balancer's IP address was marked as already in use by the downed VM itself.
After some screenshots of the load balancer I deleted it and recreated, and the VM finally started. Two minutes later I finally got a call from Azure support, so I compared notes and closed the case.
Future
If we don't upgrade our grid and infrastructure, the future is clear:
- Our database cluster is down!!!
- Oh where?
- In the UK region.
- What time is there?
- It's nearing 5.
- Damn, they put their kettles on!