News: 1617004804

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

OVH reveals it's scrubbing servers – to get smoke residue off before rebooting

(2021/03/29)


French cloud operator OVH has revealed how it is cleaning every server it thinks can be returned to service in its fire-affected Strasbourg data centres.

Founder and chair Octave Klaba used his Twitter account so show off some of the work being done by the company's clean-up crew.

[1]

Update March,24 6:30pm

The cleaning takes time. We have 80 people (SBG3) + 20 people (Croix). On the left, a motherboard with the smoke pollution on the CPU socket. It’s very corrosive ! If we power up, it’s dead. Same the disk. On the right, the same device 24h after cleaned up [2]pic.twitter.com/omaOkBSUpk — Octave Klaba (@olesovhcom) [3]March 24, 2021

Klaba has also said that all servers should be cleaned by Tuesday, but that racking and stacking the infrastructure for some services is taking longer than anticipated.

[4]

Update 28, 10am

3/4

Full recovery of pCC/HPC is slow since we need to assemble all the services (which are in 100 racks/3 floors) for each customer. It looks like we assemble hundreds puzzles of 6-12 pieces with some pieces unavailable since some racks are cleaning up. — Octave Klaba (@olesovhcom) [5]March 28, 2021

OVH's Sunday afternoon [6]update offered more detail: "Today, the cleaning time for a rack is 7 hours, and our teams are improving this every day."

The update also had some better news for customers, as follows:

SBG1: The recoverable Bare Metal Cloud servers are being cleaned for reinstallation in Strasbourg (SBG3 and SBG4). Re-commissioning will be gradually started by the beginning of next week (after inspection and cleaning).

SBG3 is operational: 84 per cent of Bare Metal Cloud (VPS) services have been made available to customers again, with a target of 90 per cent by the evening of 28 March.

SBG4 is operational: 100 per cent of Bare Metal servers are available to customers.

Servers in the SBG1 data centre will come back online at different times. Some will stay in Strasbourg and lodge in SBG4. Others are destined for other OVH data centres. The update mentions a "mid-week of 29 March" restart for some and a 1 or 2 April reboot for those moved to other locations.

Disaster recovery is also ongoing.

Some of OVH's cloud services are also not restored to 100 per cent availability. The company has also warned that high levels of demand mean "delivery times on our Bare Metal Cloud services may take longer than usual".

"Our teams are fully mobilised, and we are working hard to deliver to our customers as quickly as possible, especially all affected customers," the Sunday update says.

Klaba, meanwhile, has revealed that the fire has cost OVH the chance to launch a new service. ®

[7]

Update 28, 10am

4/4

SBG3/Floor 5 is a new floor with only new servers ready for a launch of new product we planned in ... 10 days. We’ve been working on for 18 months and now all has to be cleaned up

But first we focus on all cust’s services to make them UP.

We are on it ! — Octave Klaba (@olesovhcom) [8]March 28, 2021

Get our [9]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YGGlQlyVwxrrGxOWgapX7wAAAJM&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[2] https://t.co/omaOkBSUpk

[3] https://twitter.com/olesovhcom/status/1374775109148368901?ref_src=twsrc%5Etfw

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YGGlQlyVwxrrGxOWgapX7wAAAJM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://twitter.com/olesovhcom/status/1376073589665972224?ref_src=twsrc%5Etfw

[6] https://www.ovh.com/world/news/press/cpl1787.strasbourg-datacentre-latest-information?sunday28

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YGGlQlyVwxrrGxOWgapX7wAAAJM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://twitter.com/olesovhcom/status/1376073590890717185?ref_src=twsrc%5Etfw

[9] https://whitepapers.theregister.com/

Worth saying again......

spireite

The cloud is not infallible.......

It does NOT excuse you of your responsibilities......

Stop treating it as utopia, where failure never occurs, where an admin can put his feet up and do nothing

It's a datacenter, cloud is just a fancy name, old tech in a new clothes.....

Repeat the above 50 times.....especially upper management!

Re: Worth saying again......

Mr Sceptical

Damn skippy!

"But, but, but - surely clouds just float around the sky? Did it dry up or something?" - an Exective Team, last week.

Aaaaaand, that's why you always need someone who actually understands IT sitting on the board at every organisation that uses technology....

Clouds: someone else's servers (TM)

Your average user's comprehension of IT icon ----->

Re: Worth saying again......

Anonymous Coward

... your data is in the cloud ... of smoke over there!

Smut

Ken Moorhouse

When they've cleaned the hard drives they will reclaim a lot of space.

Communication

Anonymous Coward

Nice to see that someone has the decency to try and keep people up to date.

This is very low-rent

Dave Null

I don't know about you, but I would rather have my data hosted by someone NOT using fire-damaged servers, thank you. You wouldn't catch AWS, GCP or Azure doing this...

MTBF

Anonymous Coward

I wonder how all this shit will affect MTBF, for the kit that has been cleaned.

No perfect cleaning exists for IT kit after that kind of SNAFU.

They may as well be better to bin all the kit and install all new.

Re: MTBF

Pete B

We had an argument with our insurers after a building fire where smoke permeated through the Data Centre, leaving very obvious carbon on all the kit: Our insurers wanted us to have it cleaned and reuse - we pushed back asking for a warranty from them against any future early failures (and consequent losses) - at that point they decided to buy us all new kit. Shysters the lot of them.

OVH's experience might be useful

Pascal Monett

It seems to me that this whole ordeal would be a good time to write the manual on fire recovery procedures in data centers. I'm sure there is one, but now OVH has first-hand experience.

It would be really nice if OVH decided to share that experience other than by tweet.

What's another word for "thesaurus"?
-- Steven Wright