News: 1620297904

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

OVH outlines three-point 'hyper resilience' plan after Strasbourg fire

(2021/05/06)


French cloud provider OVH has outlined a three-point plan designed to avoid a repeat of the loss of data and services resulting from the [1]fire which [2]engulfed its Strasbourg operations on 10 March .

Dubbed "Hyper Resilience", the plan employs the combinations of a revamped approach to internal backups, external customer back-ups and a new policy of fail-over between three data centres per region.

OVH founder, chair and CFO Octave Klaba and CEO Michel Paulin outlined the plans in a [3]tweeted video address , viewers of which were implicitly being asked to avoid the conclusion that they were closing the stable door after the horse had not only bolted but bought airline tickets to Cancun where it was now sipping mojitos on a beach.

[4]

[5]

[6]

The fire took place on March 10 and [7]destroyed the SBG2 hall of the Strasbourg data centre , damaged SBG1 badly, and led to a massive effort to [8]clean salvageable kit so it could be installed in the remaining three data centres at Strasbourg, or moved to other OVH facilities. Thankfully, no one was hurt.

Klaba said he suspects [9]two uninterruptible power supplies units were the cause of the fire, revealing the UPS, battieres and fuses are in the [10]hands of the police and insurers .

By 14 April OVH was still struggling to bring all customers affected by the fire back online – a situation described by Klaba as a [11]"real nightmare" .

[12]

In the latest update, Paulin said the operator had restored 118,000 services out of the 120,000 services which had been affected by the incident.

Backup resilience

Klaba went on to describe the "five-year hyper resilience plan" designed to restore customer confidence, starting with internal back-ups. Internal back-ups had been within a data centres or, in some cases, in another data centre in the same region. The operator is now proposing creating a four-data-centre region where it will host internal backups physically separated from operational regions.

[13]OVH says burned data centre’s UPS, batteries, fuses in the hands of insurers and police

[14]OVH writes off another data centre – SBG1 – and reveals new smoking battery incident

[15]OVH flames scorched cloud customers with pledge to build data centre fire simulation lab

[16]Imagine your data center backup generator kicks in during power outage ... and catches fire. Well, it happened

Secondly, it will add features to backups in the new region. While backup data had previously supported its internal needs, OVH is proposing that customers will be able to replicate and remove the backup data for their own purposes in a free service.

Changing rules for building data centres

Lastly, OVH said it would change internal rules for building data centres and also software-managed resilience operating across three data centres in a region, starting in Paris, to be later introduced in the rest or Europe, the US and Asia.

Klaba made a [17]smiliar vow about upgrading the campus after a massive outage in 2017.

At the time, the group embarked on a "€4m-€5m investment plan in the wake of a major outage that left three of the Strasbourg data centres – SBG1, SBG2 and SBG4 – without power for 3.5 hours in November 2017".

Klaba himself said at the time of the 2017 outage that it was partly because "SBG's power grid inherited all the design flaws that were the result of the small ambitions initially expected for that location - with "SBG2's power grid" built atop "SBG1's power grid instead of making them independent of each other".

[18]

That previous upgrade was said to involve "de-installation of maritime containers" (shipping containers) and major electrical work.

Back to the present time, post 2021-fire, Klaba has said that OVH will rebuild the data centres affected by the blaze: Strasbourg One, Two and Four.

Paulin apologised for the incident which caused [19]serious disruption across European websites , with, according to Netcraft, "3.6 million websites across 464,000 distinct domains... taken offline." ®

* With our thanks to forum [20]commenter Danny14 . If you want to get stuck into our comments, you can [21]sign up here .

Get our [22]Tech Resources



[1] https://www.theregister.com/2021/03/10/ovh_strasbourg_fire/

[2] https://twitter.com/olesovhcom/status/1389483264230965248?s=19

[3] https://twitter.com/olesovhcom/status/1389483264230965248?s=19

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YJQSm5xUuXT4NHYQHBgOmwAAAMU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YJQSm5xUuXT4NHYQHBgOmwAAAMU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YJQSm5xUuXT4NHYQHBgOmwAAAMU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[7] https://www.theregister.com/2021/03/12/ovh_restoration_roadmap/

[8] https://www.theregister.com/2021/03/29/ovh_restoration_update/

[9] https://www.theregister.com/2021/03/12/ovh_restoration_roadmap/

[10] https://www.theregister.com/2021/03/17/ovh_restoration_update/

[11] https://www.theregister.com/2021/04/14/ovh_restoration_update/

[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YJQSm5xUuXT4NHYQHBgOmwAAAMU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[13] http://www.theregister.com/2021/03/17/ovh_restoration_update/

[14] http://www.theregister.com/2021/03/22/ovh_sbg1_written_off_restoration_update/

[15] http://www.theregister.com/2021/03/23/ovh_fire_lab_restoration_update/

[16] http://www.theregister.com/2021/04/06/webnx_data_fire/

[17] https://www.theregister.com/2021/03/10/ovh/

[18] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YJQSm5xUuXT4NHYQHBgOmwAAAMU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[19] https://www.theregister.com/2021/03/10/ovh/

[20] https://forums.theregister.com/forum/all/2021/03/10/ovh/

[21] https://account.theregister.com/login?r=https%3A//www.theregister.com/

[22] https://whitepapers.theregister.com/

discount for fire-damaged kit?

Dave Null

Having seen them post photos of them physically cleaning up motherboards to remove smoke damage - erm, I'd rather not run my workloads on that kit, thank you very much...

Re: discount for fire-damaged kit?

A Non e-mouse

There are lots of reports of chip shortages. OVH may not have the luxury of buying a massive job lot of new kit.

Re: discount for fire-damaged kit?

iron

Not a valid excuse if non-incinerated servers are available for similar price elsewhere.

Or as we called it in the good old pre-hype days"

IGotOut

A "Disaster Recovery Plan"

"3.6 million websites across 464,000 distinct domains... taken offline."

Pascal Monett

Yay for Single Point of Failure. Nice to know that the ol' buddy is still alive and kicking.

I think that, in the past few months, we've had largely enough demonstrations that UPSs should be quarantined far from actual servers.

In any case, whether you appreciate OVH's customers or not, I think OVH has done a fine job of openness and transparency on this issue. We are far from the usual "only a small number of customers have been impacted / we take customer data security very seriously / etc".

I'm hoping that OVH will publich a complete, official DR report with step-by-step instructions. As painful as this was for some, it is a priceless opportunity for all other datacenters to check against their own environment and start implementing mitigations now, before it's too late.

communication

Anonymous Coward

"Lastly, OVH said it would change internal rules for building data centres"

Ah yep, this is much needed ! I need to send them "DC building for dummies", there are probably some good tips in there.

Meanwhile, still today, some people web sites are still down and they have 0 info on whether they'll come back again !

And their support web site doesn't say anything either.

Re: communication

Anonymous Coward

Surely that says as much about the customer? Like OVH they had all eggs in one basket and no real DR plan having just ticker the box that says 'OVH does it'. In most cases a simple off site backup taken whenever it was changed would have allowed it to be restored quite quickly.

Re: communication

Anonymous Coward

Many customers are not very IT literate (have a wordpress and doing edits, etc ...).

They thought this was secure, while it was not.

Re: communication

ChipsforBreakfast

Correct.

We are OVH clients, with a not insignificant number of servers hosted there. We had servers in SBG2. We also had no data loss and minimal downtime (what downtime we did have was largely my own fault).

OVH has the technology and network available to avoid building systems with a single point of failure. They have advanced networking capabilities if you want to use them (in our case, we've now added API access & scripts to our network monitoring to repoint IP's if a network goes dark).

The point is that the client has to use them. If they don't, they have only themselves to blame.

BC is more than just backup.

At last backups you will be able to use your backups

Anonymous Coward

"OVH is proposing that customers will be able to replicate and remove the backup data for their own purposes"

About time cloud providers made backups available for download and use elsewhere. Most providers provide no means of doing this so you have no way of keeping an external copy, nor restoring it to another cloud provider for disaster recovery purposes, let alone on your own kit.

I have however succeeded in doing this via tar to a VM on my own kit but its far from easy to then make it bootable but not impossible. This was mainly to create a local testing platform that we can refresh at intervals for testing releases without incurring the cost of a running a duplicate cloud server. Also gives that warm fuzzy feeling that I can recover from a disaster like this within a reasonable time frame.

Re: At last backups you will be able to use your backups

ChipsforBreakfast

That really depends on what services you're using. We use bare metal servers. We install our own hypervisor on them and we back them up to our own, non-ovh facilities using normal backup tools.

We have contingency plans that allow us to restore to either AWS or Azure if necessary.

It's not really up to the provider of the DC to manage your backups. Sure it's nice if they will but there's no substitute for doing it yourself.

It's not that hard

AlanSh

I used to design data centres and do DC migrations for a living. It's not that hard to ensure that no data is ever lost. I'd be interested to see what they are going to and how it differs from what they had before.

Obligatory

Katherine Bean

It would seem the obligatory Basket Co cartoon is needed:

https://basketCo.xyz/11

Is it possible that software is not like anything else, that it is meant to
be discarded: that the whole point is to always see it as a soap bubble?