Global Fastly outage takes down many on the wibbly web – but El Reg remains standing
- Reference: 1623149215
- News link: https://www.theregister.co.uk/2021/06/08/fastly_outage_takes_down_half/
- Source link:
Mid-morning UK time (09:58 UTC) today, reports began to flood in about errors on a range of seemingly disparate sites: everywhere from Reddit, Twitter, GitHub, [1]Stackoverflow , The Guardian , The Verge , and crowdfunding platform Kickstarter to GOV.UK, the UK government's primary web platform, had started to throw 503 cache errors or connection failure messages to would-be visitors.
Ironically, even legendary webcomic [2]xkcd fell [3]offline .
[4]
The root cause, according to security expert Mikko Hypponen and others in the field: Fastly, an edge-centric cloud computing specialist founded in 2011 by former Wikia chief technical officer Artur Bergman, which is apparently having a bad start to the day.
[5]
[6]
"Fastly edge platform is having problems, which means a big part of the internet is having problems. This includes Twitter. Even fastly.com itself is unavailable in many locations," Hypponen [7]wrote of the outage. "Basically, internet is down."
[8]
Click to enlarge
Boasting 1,000 employees and an annual revenue of $200m, Fastly is responsible for optimising websites – primarily through its content delivery network (CDN), which appears to have been at the heart of today's outage.
Fastly's [9]status page confirmed "potential impact to performance with our CDN service" starting at 09:58 UTC today – which is a somewhat understated way of putting the glitch. At the time of writing, investigations were under way with no timescale yet provided for a fix.
A spokesperson for Fastly confirmed to The Register that the company is "aware of the issue and can confirm it's global," and that "all hands are on deck and working hard to resolve." ®
Updated to add at 10:48 UTC
Fastly updated its status at 10:44 UTC to say the issue had been "identified and a fix is being implemented."
Updated to add at 11:03 UTC
Fastly has applied the fix, and told customers at 11:57 UK time (10:57 UTC) they "may experience increased origin load as global services return."
To our readers affected, we offer a virtual beer or colddrink. We hope the rest of this day goes better.
Get our [10]Tech Resources
[1] https://twitter.com/Nick_Craver/status/1402213937983205376
[2] https://www.theregister.com/2015/11/24/interview_with_randall_munroe/
[3] https://xkcd.com/705/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YL@UJPc20Agw9Ve16URENAAAAIU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YL@UJPc20Agw9Ve16URENAAAAIU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YL@UJPc20Agw9Ve16URENAAAAIU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://twitter.com/mikko/status/1402203945334870016
[8] https://regmedia.co.uk/2021/06/08/screenshot_504_gov_uk.png
[9] https://status.fastly.com/
[10] https://whitepapers.theregister.com/
Re: UTC
Do you want its fixed quickly or what?
Re: UTC
Thanks, it was fixed. Please consider dropping corrections@theregister.com an email if you spot anything odd so we can take a look straight away.
C.
Explain to me again, I'm feeling really dumb today, why a single point of failure is a good thing.
It's not single point of failure - it's the cloud, never happens according to my mate Gary in marketing...
Obligatory reference: https://xkcd.com/908/
Perhaps [1]Jen dropped the black box ?
[1] https://www.youtube.com/watch?v=Vywf48Dhyns
Because you can tell the Board that this will save money and Fastly is very reliable and the SLAs are quite reasonable and, even when things go wrong, it's not your fault.
It's not Reddit's job to keep the internet going 24/7, or even Github's. Without Fastly, each of them and each site from .gov.uk to gu.com would have its own solution and a similar point of failure. They're not responsible for each other. It's only people who want access to all those services at once that experience it as a single point of failure.
> It's only people who want access to all those services at once that experience it as a single point of failure.
No, that's not really accurate.
Even if you're only focused on Reddit (to pick one), Reddit is(was) down - the reason? They built a single point of failure into their setup by using a single CDN vendor rather than a multi-CDN setup (or, alternatively, have just realised some metrics their multi-CDN status checker should have been considering).
What you're talking about - the fact that it broke a wide range of services is a *common* SPOF.
Fastly is still a SPOF for each of these services, regardless of whether any other service was using them. That a large proportion of the internet seems to be down to users is because there's a common SPOF that's just failed.
It's not reddit's job to avoid common SPOFs, but it is reddit's job (if they care about service availability) to avoid/mitigate SPOFs in the first place (though, really, cost comes into it too - you can mitigate most things, but it may not be worth the cost to mitigate the edge-cases)
it would be interesting to see how many other CDNs had ties to fastly and thus were also affected.
AFAIK, Fastly don't offer a white label service (the thing that allows other businesses to present it as their "own" CDN), so from a delivery point of view it shouldn't be any.
But, anyone sane should be serving things like status pages through a seperate route, so it's quite possible that some others served theirs via Fastly, and the status page went down while service continued. Fairly small impact.
But, those small impacts can get quite fun once you start thinking about the spread of lots of them - how many companies build/test pipelines failed because they rely on assets that get pulled down from a site/service that's fronted by Fastly's CDN?
It's really quite difficult and expensive to build your own multi-CDN system and, given, that configuration issues become increasingly likely to be the SPOF, even that won't always help as evinced by the occasional Google SNAFU: Google effectively does run multtiple CDNs.
CDNs do take an enormous amount of risk out of the equation by filtering nearly all the aggressive traffic out there, and there is a lot of that!
At Savvo...
I find it amusingly ironic that ElReg has a user survey currently asking folks about why they may or may not use "the cloud" as of yet. It's stories such as this one that answer their questions far better than anything I could post.
Why don't I use the cloud? Because the moment my data is in the hands of a third party, it's no longer my data.
Re: At Savvo...
"Why don't I use the cloud? Because the moment my data is in the hands of a third party, it's no longer my data."
So you build your own connections to every single user rather than relying on existing infrastructure? Wow, that must cost a pretty penny.
Pretty sure that for most users of these services it is much cheaper, and probably better, than rolling their own. Building a CDN is hard, and having to do so for just one site is prohibitive.
It means that a failure becomes really visible, but it probably makes those failures less likely than a whole bunch of disparate CDNs which aren't learning from each other's mistakes.
Re: At Savvo...
It's not really price that moves websites to CDNs but things like bandwidth and traffic filtering. Getting more bandwidth (and there are now attacks that have enough bandwidth to take out entire data centres) can be expensive but basically it's the filtering that brings the biggest returns.
Re: At Savvo...
A CDN is not the same as putting your data into the cloud, it's really just a set of proxies.
Re: At Savvo...
A wise man once said "there's no such thing as the cloud, it's just someone else's computer"
And I reckon you use other people's computers all the time.
we ran our own web site, email, storage via sharepoint all onsite for many years, at least the 14 ive worked here. The management decided that we needed to shrink budgets so my dual internet, dual site, dual storage, dual stretch cluster was too expensive. Whilst dual redundant we only had power outages on both sites to consider - this happened once that I can remember - for loss of connectivity. Cloud DNS updated the records should we have an ISP failure and a lowish TTL meant this wasnt too much of an issue.
Management wanted me to move to the cloud and we decided that 365 was the best option (since migrating onsite sharepoint to 365 was supposed to be painless. Covid hit, backs were patted as we were already cloudy (which meant nothing as we could already access all the resources externally anyway), however 365 started to have a few hiccups and outages over the year. Throttling reared its head a few times, other dropouts were noticed. Basically we have had more loss of access in the last 6 months at least than we did over the entire preceeding period of onsite.
Cloud is definitely not more reliable.
"Cloud is definitely not more reliable."
And neither is on-premise. And I'm speaking as someone who has run their own kit (including in data centres) a lot over the years. What one hand giveth, and all that...
Example: for hundreds of small(er) businesses with limited or no IT budget, or aging kit that keeps falling over, or where there is no-one IT-literate to watch over it, the cloud *may* be more reliable for them *for their needs*.
Just because it isn't for *you* - even if you're doing very sensible stuff because you have knowledge and budget - doesn't make a hard and fast rule. As ever, quoting a particular specific to prove a general point is, well, a bit arse really? Also, if you suffered shrinking budgets then even your on-premise gear may have ended up less reliable over time...
A/C because I'm not arguing this. Everyone's situation is different.
It's not a good thing. The point is, a CDN is generally not a single point of failure, exactly the opposite in fact. But something has clearly gone terribly wrong in the management of the CDN.
Plausible deniability
If your website goes down at the same time as a bunch of other websites, your IT department can say "ah well, happens to the best of us".
If your website goes down but the rest of the internet is working, all fingers are pointing at your IT department.
So clearly, from their point of view, cloud services are a good thing...
Irony... or is it?
This article posted at 10:46 UTC. At 10:44, Fastly:
Identified - The issue has been identified and a fix is being implemented.
Jun 8, 10:44 UTC Things do generally seem to be getting better now.
Re: Irony... or is it?
Oh, and extra irony, the last XKCD I still had open in a tab before Fastly took it down along with it was [1]The Cloud , after having followed a link from [2]a Register comment earlier today.
[1] https://xkcd.com/908/
[2] https://forums.theregister.com/forum/all/2021/06/08/ethernet_alliance_technology/#c_4270209
At KarMann, re: getting better...
For some odd reason I heard that in the voice of the Monty Python old man being dumped in the dead-folks cart crying out weakly "I'm getting better! I think I'll go for walkies!"
I expected to hear a stern "No you're not, you'll be stone dead soon."
=-D
I noticed github and xkcd were having issues earlier, guess this was it.
Those two, at least for myself, seem to be back again now.
There seems to be too many sites with a dependency on a single, apparently not resilient, service!
Just one dependency? ;-)
They spelt DNS wrong...
What causes this?
How many more times are we going to see one major CDN issue take out a considerably large chunk of websites? Is it purely that it saves money and "only happens now and again" to justify it or is there no way to have two of these things?
I look at the Fastly status page and it says this which does not inspire confidence.
"Fastly’s network has built-in redundancies and automatic failover routing to ensure optimal performance and uptime. But when a network issue does arise, we think our customers deserve clear, transparent communication so they can maintain trust in our service and our team."
What happened with the failover? Backup services? I know nothing of any of this so I know I'm being overly simple. Anyone care to educate me?
Re: What causes this?
The clue is in the acronym: content denial network
Re: What causes this?
The backup is connected to that power socket next to the cleaner's store room...
Re: What causes this?
Well, if they are as 'transparent' as they say they are, maybe wait until we get a post-mortem rather than randomly wonder about it? For all we know their backup and failover systems may have worked brilliantly the last 99 out of 100 times they were needed and we just never knew because they just worked, but just didn't this one time (although if failover had been needed that often maybe something is wrong, but you get my point).
Fastly
.........................disappearing?
Re: Fastly
Not-so-Fastly
implications and questions
- all eggs in one basket?
- did they get hacked?
- related to FBI action?
- was it Russians, Chinese, Iranians, or North Korea? Or just a mouse that chewed through a cable?
bbc reported amazon is down. I then shuddered as my world suddenly collapsed. What am I to do? Gov.uk can be down, and all data spilled, but AMAZON?! We're doomed!
Re: implications and questions
That was the weird part, amazon's own site lost their images which made me thing they were down.
Why would they use a 3rd party CDN?
Re: implications and questions
the 503 URL's I saw had aws in the mix.
So my guess is that the CDN either took down a chunk of aws or more likely an aws failure cascaded and brought down the CDN
Re: implications and questions
I got that about a week ago. I thought it was my own network because I've been experimenting with nftables (definitely nicer than iptables but doesn't support as many targets). But it turned out to be external.
Re: implications and questions
Amazon seemed down to me for a short period. That one I don't get, as they happily sell access to their own CDN, so why don't they themselves us it?
Re: implications and questions
"That one I don't get, as they happily sell access to their own CDN, so why don't they themselves us it?"
Maybe they're actually white-labelling Fastly's services. In fact maybe AWS in its entirety is done the same way :)
Re: implications and questions
Second appearance in this comment thread [1]https://xkcd.com/908/
[1] https://xkcd.com/908/
Guru Meditation
As great as the Amiga line of computers were, I'm not sure they should be hosting large swaths of the Internet in 2021!
Re: Guru Meditation
I *do* love that there are now considerably more people than ever used an Amiga now seeing this message, and presumably wondering "Guru Meditation? What?!"
This is a good thing. It makes me happy! :D
Re: Guru Meditation
Although it's really bugging me that it was misspelled as "Guru Mediation" on the Fastly page, even though it is correctly spelled "Guru Meditation" in the Varnish Cache source, and always has been.
Have Fastly messed up the spelling in their Varnish config?
Re: Guru Meditation
Might "Guru Medi t ation" be protected by copyright?
I noticed problems on both Amazon and Ebay this morning. Both the sites were loading but several images were missing, So i assume its connected to the Fastly outage.
BOFH response
So i assume its connected to the Fastly outage.
Hmmm. So this might be a good time to do that emergency maintenance we've been putting off and which will take down the network for a bit.
"Oh Yes, yes I'm sorry the network has dropped off at the moment. Yes it's due to this Fastly outage. Yeah it took down most of the internet and yeah we got hit as well. Dont worry I'm sure it will be up again soon..."
Surely xkcd 908 is the correct link
[1]https://xkcd.com/908/
Seems someone tripped on the network cable ...
[1] https://xkcd.com/908/
I expect we'll get a "Who,me?" in due course.
Everything kept working in Munich
I started to notice issues this morning, then realised I was using a VPN with presence in the UK.
When I disconnected from the VPN to use the direct local ISP connection in Munich, Germany, everything worked and continued to work all morning.
I checked all the sites reported, like FB, Independent, Guardian, NY Times, etc. and all their sites were up.
Was this a case of inflated ego little British journos confusing UK for the World, with their reports of global outage???
UTC
I think you mean reports started at 09:58 UTC, or 10:58 BST (10:58 UTC is 9 minutes in the future as I write this)...