Outage: Faulty UPS at data centre housing London Internet Exchange causes grief for ISPs and telcos alike
- Reference: 1597757649
- News link: https://www.theregister.co.uk/2020/08/18/outage_london_internet_exchange/
- Source link:
The incident was caused by a faulty UPS system followed by a fire alarm (there was no fire) that powered down [2]Equinix's LD8 data centre , a low latency hub that was formerly the Telecity Harbour Exchange. Located on London's Isle of Dogs, the facility caters to hundreds of clients and is positioned close to the city's financial institutions.
Equinix told The Register : "Today at 4.40am, Equinix IBX LD8, in the Docklands, London, UK, experienced a power outage. This has impacted customers who are based there. The outage may have also affected customers' network services. Equinix engineers have diagnosed the root cause of the issue as a faulty UPS (uninterrupted power supply) system and we are working with our customers to minimise the impact. We regret any inconvenience this has caused."
If you're the unlucky engineer who gets the call from the customer, the supplier is being more flexible about letting people head down to the DC (contact it first), though your usual Docklands cappuccino haunt will likely be closed, so bring a thermos (and a mask).
"Due to this incident we are allowing customers more flexible access to LD8 working within our COVID-19 restrictions including mandatory temperature checks and wearing face coverings. The safety of our employees and our customers is our highest priority."
The facility is one of the DCs that currently houses the [3]London Internet Exchange (LINX) , one of the world's largest with hundreds of members – including ISPs such as BT and Virgin Media, as well as smaller outfits and content providers.
"My first alert was at 7:30am and it's still ongoing," said one reader of the LD8 outage.
UK ISP M12 Solutions, which also owns premium ultra-fast business internet brand GigaNet, speculated earlier today on its [4]status page : "Due to the scale of this outage, and the carriers & suppliers affected... there could be knock-on impacts around the carrier ecosystem/networks.
"For instance, we are seeing some carriers report their racks in LD8 are being powered back up, and when this happens there could be increased routing table changes on their network devices that could cause our circuits to be affected that traverse their network.
"The London Internet Exchange (LINX), are currently reporting that 150 of their members are affected by this outage, to provide the sense of scale of this outage."
The [5]status page of Exponential-e , another affected vendor, said: "This incident is being treated as a Major Incident both at Equinix and at Exponential-e. We have prepared for all hands on deck for when power is restored to carry out an extensive range of tests to ensure service has restored."
Worried customers took to Twitter to voice their concerns.
Completely unacceptable situation ongoing in [6]@EquinixUK [7]#LD8 (HEX 8/9) right now. Reports of a fire alarm, but this triggered the loss of both our A+B diverse power feeds to our main rack since 04:24. Lack of communication is abysmal. [8]@Equinix need to sort the basics out. — Matthew Skipsey (@matthewskipsey) [9]August 18, 2020
In accordance with safety procedures, all systems within the vicinity of the fault – from levels one to four – were powered down.
A reader told us they had received an update stating that IBX had been "evacuated" and that the "fire alarm was triggered by the failure of output static switch from Galaxy UPS system supporting levels 1, 2, 3, 4 in building 8/9 at LD8".
The update added: "This has resulted in a loss of power for multiple customers and IBX [Equinix International Business Exchange] Engineers are working to resolve the issue."
Meanwhile, [10]users on DownDetector are reporting outages at BT . The incumbent UK telco is thought to host an exchange in the DC, but so far has not been responding to requests for comment.
Back in 2016, the Telecity data centre owner confessed to a "brief outage" that knocked 10 per cent of BT internet subscribers offline in the UK as well as a number of other providers on the morning of 20 July.
The Register spoke to a representative at Exponential-e, who confirmed the facility wasn't on fire and said their customer support lines were being flooded by desperate customers eager to get their kit back up and running.
There is currently no known time frame for when operations will return to normal.
The Register has also asked BT to comment.
Have you been affected? [11]Drop us an email . ®
Updated at 17:07 UK time on 18/08/20 to add
BT has been in touch to say: "A small number of BT’s Enterprise and Global customers experienced an interruption to their Ethernet services between 04:28 and 08:03 this morning due to a data centre provider’s power failure at their site. Services were restored to the affected BT customers by 8:03 this morning, before the start of standard business hours."
Updated at 17:30 on 18/08/20 to add
Giganet and M12, whose team The Reg has to commend for giving some wonderfully descriptive updates to its users during the day, updated its status page a few minutes ago, noting:
"We continue to observe good network availability since our LD8 core router has been connected to a temporary power feed from our adjacent rack (that was unaffected by today’s issue)."
Get our [12]Tech Resources
[1] https://marketplace.upstack.com/data-centers/equinix-data-center-london-ld8
[2] https://www.equinix.co.uk/locations/europe-colocation/united-kingdom-colocation/london-data-centers/ld8/
[3] https://www.linx.net/about/our-network/
[4] https://status.giga.net.uk/incidents/3zcfz8s8g43h?u=q6j7pk53hcyw
[5] https://exponential-e.com/network-status
[6] https://twitter.com/EquinixUK?ref_src=twsrc%5Etfw
[7] https://twitter.com/hashtag/LD8?src=hash&ref_src=twsrc%5Etfw
[8] https://twitter.com/Equinix?ref_src=twsrc%5Etfw
[9] https://twitter.com/matthewskipsey/status/1295624348259188736?ref_src=twsrc%5Etfw
[10] https://downdetector.co.uk/status/bt-british-telecom/
[11] mailto:matthew.hughes@sitpub.com
[12] https://whitepapers.theregister.com/
Our connection suddenly sprang back to life about 10 minutes ago, having been off all day.
Communication has been very poor.
Have they tried turning it off and on again?
Preferably leaving it off.
... the fire, that is.
"to provide the sense of scale of this outage"
That 150 companies are affected provides absolutely no sense of scale unless you know how many companies there are in total.
I do not. Is it 150 out of 300 ? That would be important. Is it 150 out of 10,000 ? That would be marginally insignificant.
So which is it ?
Re: "to provide the sense of scale of this outage"
I think it's more that quite a few of these companies will be internet providers, such as exponential-e
That blows the four nines then!
Dual power supply and UPS will only provide so much resilience. Dual (or multiple) bits of equipment in different geographic locations with diversely routed connections and properly configured are essential to achieving high availability levels.
What needs to be understood is that an actual fault or outage at a single site often can't be fixed within the SLA agreement timeframes, particularly if they are measured monthly or quarterly and are at 99.99% or higher.
Those of you who are affected by this - I feel your pain!
Use it to invest in proper diversity and resilience if the pain is too great!
Flames - because there were none...
Tentative fix estimate
We're being affected. Virtual1 who is also being affected, have a very tentative fix estimate of 16:30
Was given the day off..
I got the alarm at 6am to say our network was down. Got to the office early, discovered that we have no internet and no estimated time for it to be resolved. Decided I'm not sitting and twiddling my thumb's all day so left by 9.30am :)
Re: Was given the day off..
Didn't you check first from home to avoid the journey in? ;-)
Re: Was given the day off..
We're a small business in a business centre.. all I knew was that we couldn't access anything in the office - which is a bit of an issue when 50% of our staff work from home. Until I got to the office and confirmed our servers were ok I had no idea what was wrong - could have been a crash at our end, or worse..
Re: Was given the day off..
I hear you.
Traceroute can be your friend - fire off a few from random places on the Internet (commercial/consumer VPNs can be useful here) towards your office IP or firewall. If they don't get there, it's not your office. If they reach your firewall, it's not your ISP or line and it's probably your office.
(I'm simplifying, but hopefully you get the idea).
We're back up. 11 hours and 29 minutes of downtime :-(
Some of it is down again, certainly for BT. Something about a new distribution board is required something something
We're still down. Appalling lack of updates from Equinix.
Appalling lack of updates from anyone.
My AAISP FTTC still down since 4:23am.
TITSUP
Titans Interrupted Telemetry Sockets Unfitfor Purpose
UPS failure often sets off the VESDA system as the current flowing through the various components is very high. A little whisp of smoke is all it takes. If you remember the Bluesquare issue at Maidenhead a few years ago, one UPS output capacitor blew and took the row of UPS devices out but the small amount of smoke set off the fire alarm system which automatically cut the power to the building and they weren't permitted to reinstate it until they had the OK from the fire brigade. The fire brigade were convinced there was a fire and they had the piece of paper to prove it, so had to have a good look around.
Should say that Bluesquare and more recently Pulsant have completely refurbished the UPS systems at Maidenhead.
Whilst once upon a time I was a direct customer in that facility (and another), I'm now only an indirect (by two levels) customer. Our office connectivity monitoring alarm went off at 04:28 this morning and we're still down. LD8 is obviously key in our comms chain somewhere. Poor general communications from Equinix, even if it is a seriously major issue stretching people right now. Most unlike Equinix, based on past experience.