Supply chain blamed amid claims of Azure capacity issues
- Reference: 1656951972
- News link: https://www.theregister.co.uk/2022/07/04/azure_capacity_issues/
- Source link:
Azure comprises over 200 datacenters globally spread across 60 regions, but reports suggest that over two dozen of these are operating with limited capacity, and that the cloud and IT giant is being forced to prioritize resources in order to serve existing customers.
According to technology news site [1]The Information , capacity issues are affecting Azure datacenters in Washington State in the US as well as across Europe and Asia, and it claims that server capacity is expected to remain limited until early next year, citing a Microsoft insider.
[2]
Meanwhile, [3]The Telegraph claims that Microsoft's two Azure datacenter regions in the UK are not accepting new customer sign-ups for access to virtual machines or its Cosmos DB cloud database service. It claims that Microsoft is struggling to balance growing customer demand with the support it is giving to the Ukrainian government, and that it cannot easily expand capacity because of constraints on the supply of IT kit.
[4]
[5]
In May, Microsoft disclosed at its [6]Envision event in London that the company had spent over $100 million on providing technology support to the Ukrainian government during the current conflict, which included moving its IT operations to the cloud. The company has also been [7]working to counter Russian cyberattacks on Ukrainian infrastructure.
Microsoft told The Register it was experiencing "unprecedented" demand, and added that it would put capacity restrictions in place where needed.
[8]
"With this surge, coupled with macro trends impacting the whole industry, we've taken steps to address customer increases in capacity while also expediting server deployment in our datacenters. Our priority remains ensuring business continuity for customers. In addition to managing and planning for growth, we actively load balance as needed.
"If it does become necessary to put capacity restrictions in place, we will first restrict trials and internal workloads to prioritize growth of existing customers," a Microsoft spokesperson said in an emailed statement.
However, a Microsoft customer who spoke to us on the condition of anonymity said Azure's UK South and UK West regions are both showing signs of capacity problems, and that if a new customer subscription is created, it is not possible to deploy any compute into either of those two regions.
[9]
The Azure portal simply advises submitting a support request, they added.
[10]Will cloud giants really drive colos off a financial cliff?
[11]Microsoft gives its partners power to change AD privileges on customer systems – without permission
[12]Start using Modern Auth now for Exchange Online
[13]FabricScape: Microsoft warns of vuln in Service Fabric
The same customer also said they had stopped deallocating virtual machines outside of business hours, which would normally be done to save costs, because they had experienced difficulty in reallocating resources the next morning.
One IT professional and Microsoft Most Valuable Professional (MVP) said that capacity issues are nothing new for the Azure cloud, and that these issues have affected customers for some time.
Aidan Finn, who works mostly with clients in Norway, [14]said on his blog : "Most people who have used Azure for a few years have seen the dreaded deployment fails – you cannot get capacity because the region you have selected doesn't have it. And the 'helpful' support agent tells you to try a less-impacted region on another continent."
Finn suggested that Azure users need to take steps such as disabling auto-scaling, determining requisite resources and holding to those resources rather than releasing them when they are not being used.
If enough customers are already doing that, it would also tend to exacerbate any capacity shortage that Azure is experiencing.
Industry analysts we spoke to said that trade customers such as Microsoft are already being prioritized by IT suppliers, but it is possible that supply chain issues are having an impact.
However, it may not be servers that are the issue. Some network switch makers have reported that lead times for the semiconductors they need to build their products have reached to [15]more than 80 weeks .
Andrew Buss, Research Director European Infrastructure Strategies at IDC, said that supply chain delays mean that building capacity can be an issue if demand builds suddenly.
"The demands [for public clouds] are likely to be exacerbated as customers that cannot get immediate supply and face capacity constraints [on physical servers] themselves therefore look to public cloud to add additional IaaS capacity. This will push the public cloud demand up, putting further pressure on service scaling," he said. ®
Get our [16]Tech Resources
[1] https://www.theinformation.com/articles/microsoft-cloud-computing-system-suffering-from-global-shortage
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YsNi-7E3tiytmAzUFLszqQAAAIU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://www.telegraph.co.uk/business/2022/07/02/microsoft-declines-new-cloud-customers-promise-ukraine/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YsNi-7E3tiytmAzUFLszqQAAAIU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YsNi-7E3tiytmAzUFLszqQAAAIU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.microsoft.com/en-gb/events/envision-uk/
[7] https://www.theregister.com/2022/04/08/microsoft-russia-stronium-domains/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YsNi-7E3tiytmAzUFLszqQAAAIU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YsNi-7E3tiytmAzUFLszqQAAAIU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] https://www.theregister.com/2022/07/04/cloud_giants_vs_colos/
[11] https://www.theregister.com/2022/07/01/gdap_permissionless_change_window/
[12] https://www.theregister.com/2022/06/29/cisa-microsoft-modern-auth/
[13] https://www.theregister.com/2022/06/29/azure_service_fabric_cve_2022_30137/
[14] https://aidanfinn.com/?p=22679
[15] https://www.theregister.com/2022/05/03/arista_networks_makes_43b_bet/
[16] https://whitepapers.theregister.com/
We're being quoted up to 9-month lead time on networking kit at the moment on tier-1 vendors.
All those businesses hosted on Azure while operating near capacity... All it will take is one overworked MS engineer to make a single Powershell oopsie and it will be down for a week. No contact with clients, no payroll, no email, no backups, nothing.
I sense if that did happen Azure would have a lot of abandoned space in short order.
I've never had Microsoft be reliable when I needed it to be so.
Same here
I made a proposal (pre sales) for a customer which involved improving redundancy by putting in some dark fibre, adding a couple of interfaces and changing their OSPF topology. It would have worked nicely. It then got handed off to an engineer for implementation..
Somewhere along the line, the engineers said to the customer, “hmm your routers might shit themselves with the extra load, you’ll need some new ASRs” Customer said ok. After a few months, I was brought back to sort out why things weren’t done yet. This is when I, and the bosses, found out that people had “promised” the customer new hardware.
Turns out ASRs were end of sale, and going end of software support next year (be buggered if I’m putting something internet facing out there without software support) and the replacement hardware, which curiously enough is called Catalyst 8500, has a 9 month lead time. Customer had a drop dead date of 30th June.
You can imagine the conversations that ensued after this one
Selling point
The cloud's main selling point is that you don't have to pay and maintain idle infrastructure (well, to an extent - there is no such thing as free lunch, so the cost of idle machines is split amongst thousands of customers). Now that there is no idle infrastructure it becomes sort of on prem that you don't have access to.
This is the problem with tying yourself in with one cloud provider, if you want to deploy more or grow instances and they are running at capacity, its not trivial to switch to another provider once you are locked in.
Serverless
They really are taking this "serverless" trend a bit too literally...
The 80 weeks lead time for some network switch semiconductors is no joke.
Some lower-end edge switches that we use (and you could bargain down to 1500-2000$ per switch pre-covid), we've had to buy refurbished / used for twice the price because we can't even get a 6-months delivery commitment.