Imagine your data center backup generator kicks in during power outage ... and catches fire. Well, it happened
- Reference: 1617748206
- News link: https://www.theregister.co.uk/2021/04/06/webnx_data_fire/
- Source link:
Kevin Brown, Fire Marshal for the US city's Fire Department told The Register in a phone interview that firefighters responded to a call on Sunday evening. The fire, he said, "originated in a generator in the building and spread to several servers."
[1]
Brown said the facility's fire suppression system contained the blaze and that fire department personnel assisted with the cleanup. He said power was cut to the building until an electrical engineer could inspect the facility to make sure current could be restored safely, which he added is standard procedure.
He also confirmed that some of Ogden City's IT services were down on Sunday and Monday as a result of the data center fire.
[2]
In [3]a Facebook post on Monday, echoed [4]on its website , WebNX attributed the incident to the failure of a backup generator following a local power outage.
... one of our backup generators that had been recently tested and benchmarked specifically for this situation experienced a catastrophic failure
"Sunday afternoon the city power was disrupted and, as designed, our backup generators automatically switched on," the company said.
"However, during that transition, one of our backup generators that had been recently tested and benchmarked specifically for this situation experienced a catastrophic failure, caught fire, and as a result initiated the fire suppression protocol."
The company confirmed that its Ogden data center experienced some damage. And earlier [5]statement noted: "Some servers will have an extended outage as they may require rebuilds due to some water damage. Those builds have a high probability that data is intact."
"Customer’s servers in one of our main bays were exposed to water and possible damage may have occurred," the company said. "No fire damage was inflicted on customer servers."
Most of the hardware in the data center appears to be unaffected, the company said, but there are machines that need to be inspected for water damage and may need to be rebuilt.
"As of now, we are working to restore power, network, and unaffected hardware back up online within the next day or two," the company said.
Customers aren't pleased
The incident has also [6]affected Gorilla Servers , a separate company founded by WebNX CEO Daniel Pautz that appears to co-locate some servers within the Ogden data center.
According to Gorilla, "close to 90 to 95 per cent of all Gorilla Servers hardware experienced zero damage." The electronic engineering website of YouTuber Dave Jones, [7]eevblog.com , was among the sites hosted by Gorilla and taken offline in the outage, we note.
Other folks continue to report problems or are unable to reach their systems in the data center, though WebNX claims service has been restored for some.
At 16:47 UTC today, fish-egg flogger Passmore Caviar [8]said , "Our website remains offline due to a fire at the WebNX server facility. Currently, we expect it to be back online later this afternoon." At the time this article was filed, the website remained inaccessible.
Piqosity, a test preparation service, also [9]said its site stopped responding as a result of the fire.
OVH reveals it's scrubbing servers – to get smoke residue off before rebooting [10]READ MORE
And in various online discussion [11]forums , affected customers continue asking for estimates about when service will be restored. A common complaint is that lack of official communication about what's going on. Other forum participants report being told that their servers have water damage and will have to be rebuilt – a process that may take several days.
WebNX, which also operates facilities in New York City and Los Angeles, could not be reached by phone – calls are met with a recording noting that the company is "aware of the support system being down right now and technicians are working to resolve it." The Register received no response to our email inquiries.
WebNX's [12]SLA guarantees 100 per cent uptime and uninterrupted power every month, with account credits of one day per 15 minutes of downtime in each case.
[13]
At 20:50 UTC on Tuesday, WebNX posted an update to its Facebook page: "We are currently working hard to get everything back to optimal running order. Huge thanks to the staff and outside contractors who were brought in (and also flown in from out of state) to help with the situation. We are also really grateful to everyone who has stepped up to help out, including some amazing clients, and Ogden City who has been very supportive." ®
Get our [14]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YG0uX0oulMUn4gpq1SdsnwAAAI0&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YG0uX0oulMUn4gpq1SdsnwAAAI0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[3] https://www.facebook.com/WebNX/posts/4495315520484290
[4] https://webnx.com/ogden-datacenter-issue/
[5] https://www.facebook.com/WebNX/posts/4495315520484290
[6] https://twitter.com/GorillaServers/status/1379105278390644744?s=20
[7] https://eevblog.com/
[8] https://twitter.com/PASSMOREcaviar/status/1379475940632190976?s=20
[9] https://twitter.com/piqosity/status/1379085545993867266?s=20
[10] https://www.theregister.com/2021/03/29/ovh_restoration_update/
[11] https://www.webhostingtalk.com/showthread.php?t=1842301
[12] https://webnx.com/sla/
[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YG0uX0oulMUn4gpq1SdsnwAAAI0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[14] https://whitepapers.theregister.com/
Re: This would never have happened at a certain broadcaster I used to work for.
So, there was a fire because there was no fuel.
Diesel fuel also serves as the lubricant, so the electric starter turned the engine and it overheated due to friction from lack of lubrication, and eventually something ignited?
Test, test, and test some more!
"... one of our backup generators that had been recently tested and benchmarked specifically for this situation experienced a catastrophic failure..."
Imagine how bad it might have been had they not tested it!
The weird part isn't the generator fire - shit happens.
The weird part is that a fire on the generator could spread to the servers. They should have a firewall (fnar fnar) between them, with just the electric cables passing through.
And even the cables should pass through a cut point, with material made to expand and cut the cables in case of fire. Better no electricity than some electricity and a lot of fire.
This would never have happened at a certain broadcaster I used to work for.
The generator was in its own separate building, and tested twice a year. Unfortunately, after 20 odd years of testing, the test failed. There was a book with all engineers responsibilities, their deputies, their managers, all the phone numbers etc. The book contained every procedure, every workaround, every aspect of how to get back to broadcasting within 3 minutes in the event of a supply failure. Except there was one little omission - there was no schedule nor person responsible to ensure that the diesel tank was checked and refilled. The test failed not because of a fire but because of the lack of fire inside each of its cylinders, because there was no fuel.
Still, that is what tests are for - to find that one thing that no one had thought of before.