OVH blames hour-long global outage on human error during 'routine' network reconfiguration
- Reference: 1634119807
- News link: https://www.theregister.co.uk/2021/10/13/ovh_outage/
- Source link:
The French co-location biz's services went AWOL from 7:20am UTC. The company had earlier talked of planned "maintenance on our routers" in its Vint Hill, Washington DC, data centre to "improve routing".
Network and racks:: VIN/DC: We will do a maintenance on our routers on VIN DC to improve our routing.
Maintenance is planned for 13/Oct/21 9:00 AM to 10:30 AM ( UTC+2). No impact expected, device will be isolated before the change. [1]https://t.co/n9NjpGICrU [2]#ovh — OVH Status Feed (@ovh_status) [3]October 13, 2021
CEO Octave Klaba then [4]confirmed : "Following a human error during the reconfiguration of the network on our DC to VH (US-EST), we have a problem on the whole backbone. We are going to isolate the DC VH then fix the conf."
[5]
Worldwide outage at OVH
He added "In recent days, the intensity of DDoS attacks has increased significantly. We have decided to increase our DDoS processing capacity by adding new infrastructures in our DC VH (US-EST). A bad configuration of the router caused the failure of the network."
"The router with the wrong D2 configuration has been cut," Klaba concluded.
According to its outage report, the fix involved "the isolation of network equipment in the US."
[6]
The routine task is now marked "change cancelled".
[7]
The culprit? 'No impact expected'
OVH, which is among the larger non-hyperscale hosting and cloud services providers, with 400,000 servers across 30 data centres worldwide, is engaged in an [8]IPO on Euronext Paris though the target price was [9]reduced from [10]€450m to €350m.
[11]OVH drops IPO target against figure mooted a month ago
[12]Like a phoenix rising from the smouldering ruins of its data centre, OVH sets sights on IPO
[13]OVH outlines three-point 'hyper resilience' plan after Strasbourg fire
[14]OVH services still not fully restored as boss rates ongoing recovery efforts a 'real nightmare'
The company is well placed to benefit from the concerns about European data sovereignty and the dominance of US-based global corporations in this space. It is the fourth biggest cloud provider in Europe.
Coming (or going) just before the IPO, this outage is therefore an embarrassment. "A few months ago you had a downtime of over 8 hours because of the poor measurements you take on the safety of the buildings... As soon as this downtime is over I am moving out of OVH," [15]said one disgruntled customer.
[16]
The company hit the headlines earlier this year following a [17]fire at its Strasbourg premises on 10 March that wiped out two data centres. At least today's incident was nowhere near the scale of that disaster. ®
Get our [18]Tech Resources
[1] https://t.co/n9NjpGICrU
[2] https://twitter.com/hashtag/ovh?src=hash&ref_src=twsrc%5Etfw
[3] https://twitter.com/ovh_status/status/1448185498812485633?ref_src=twsrc%5Etfw
[4] https://twitter.com/olesovhcom/status/1448196879020433409?s=20
[5] https://regmedia.co.uk/2021/10/13/outage.jpg
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YWcCtZ2sx0mhTLQV0xNC2wAAABg&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[7] https://regmedia.co.uk/2021/10/13/outage2.jpg
[8] https://www.theregister.com/2021/10/05/ovhcloud_ipo_target_trimmed/
[9] https://www.theregister.com/2021/10/05/ovhcloud_ipo_target_trimmed/
[10] https://www.theregister.com/2021/09/20/ovh_ipo/
[11] https://www.theregister.com/2021/10/05/ovhcloud_ipo_target_trimmed/
[12] https://www.theregister.com/2021/09/20/ovh_ipo/
[13] https://www.theregister.com/2021/05/06/ovh_outlines_threepoint_hyper_resiliance/
[14] https://www.theregister.com/2021/04/14/ovh_restoration_update/
[15] https://twitter.com/XCharalambous/status/1448192529896202246
[16] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YWcCtZ2sx0mhTLQV0xNC2wAAABg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[17] https://www.theregister.com/2021/03/10/ovh_strasbourg_fire/
[18] https://whitepapers.theregister.com/
Single configuration changes by one engineer brought down networks vastly larger than OVH (Micros~1, Amazon, IBM, ...), some changes just can not be isolated.
From what I understand from Amazon is there are no dependencies between regions. Therefore, an incident in one region is unable to affect another region. EG, each region has a console service to log on to etc. The worst outage I have seen was the Amazon S3 outage four years ago, but it still only affected one region.
Different from Azure, where South Central AD took out global services, and different from OVH who seem to be in all sorts of hot water when it comes to running public cloud.
While I am loathed to pay Jeff Bezos any more cash, they do seem to be able to run a stableish cloud.
Looks like a global dependency on Azure AD may now be fixed...
Resiliency: Microsoft has made concerted efforts to improve resiliency with critical services such as Azure Active Directory, but many Gartner clients remain concerned about the real-world impacts when such critical services are unavailable. Further, Microsoft continues to react slowly to the rollout of AZs with the likelihood that some regions will never be equipped with such resiliency capabilities. Services such as the Azure Kubernetes Service (AKS) continue to experience some outages, particularly in association with updates and maintenance events.
All talk, no trousers
Ms Thunberg might have a thing or two to say about this and the FB outage :
Change control - blah, blah, blah
No single point of failure - blah, blah, blah
Systems engineering - blah, blah, blah
"improve routing"
I think I can speak for the rest of the internet when I say that OVH dropping offline improves routing for the rest of us. And vastly reduces the DDoS attempts we have to fend off.
I still find it mildly horrifying that a single configuration change by one engineer can bring an global network as vast as OVH's. You'd think there would be some safeguarding in place, alas.. not.