Akamai Edge DNS goes down, takes a chunk of the internet with it
- Reference: 1626975101
- News link: https://www.theregister.co.uk/2021/07/22/akamai_edge_dns_outage/
- Source link:
As of 0909 PDT (1609 UTC), the status page of Akamai – which sites around the world rely upon to deliver content among other services – said, "We are aware of an emerging issue with the Edge DNS service."
A short time later, the biz characterized the incident as a "service disruption":
Akamai is experiencing a service disruption. We are actively investigating the issue and will provide an update in 30 minutes. — Akamai Technologies (@Akamai) [1]July 22, 2021
An Akamai spokesperson told The Register much the same thing in an email: "Akamai is experiencing a service disruption. We are actively investigating the issue and will provide an update in 30 minutes."
[2]
Click to enlarge
Given the nature of DNS, companies using Akamai's Edge DNS were alerted to the emerging issue when their websites became inaccessible to internet users. Downdetector, a site that tracks service issues, showed a surge in outage reports for dozens of major services that appear to coincide with the reports submitted by Akamai. Affected businesses include: PlayStation Network, Fidelity, Steam, FedEx, AirBnB, Amazon, Google, and many others.
[3]You had one job: Akamai's Prolexic Denial-of-Service protection system fingered after users in Australia denied, er, services
[4]Cloudflare network outage disrupts Discord, Shopify
[5]Kaseya restores SaaS, then 'performance issues' force a do-over
[6]Not only is Hubble back online after outage, it's already taking photos of the cosmos
Reports via social media indicated widespread problems around the world. Reports indicated that [7]HSBC bank in the UK and [8]Paytm Money , an investment platform in India, were affected, along with many other firms.
Around 09:47 PDT (16:47 UTC), Akamai [9]said it had made repairs to address the outage.
[10]
"We have implemented a fix for this issue, and based on current observations, the service is resuming normal operations," the company said. "We will continue to monitor to ensure that the impact has been fully mitigated."
[11]
[12]
Other network service firms have experienced similar disruptions that also ripple across the globe. Fastly had [13]a significant outage that interfered with much of the internet last month, as did [14]Cloudflare .
Cloudflare CEO Matthew Prince offered a "don't blame us" sympathy tweet, in recognition of the challenges of keeping network infrastructure running smoothly all the time.
Not sure yet why so many sites online not loading, but confirmed it's not a [15]@Cloudflare issue. Bad days happen to everyone so hope whoever is having one it gets resolved quickly. [16]#hugops — Matthew Prince 🌥 (@eastdakota) [17]July 22, 2021
At the time this story was filed, Akamai's service appeared to be on the mend. ®
Updated to add at 1040 PDT
A spokesperson for Akamai just told us:
We have implemented a fix for this issue, and based on current observations, the service is resuming normal operations. We will continue to monitor to ensure that the impact has been fully mitigated and can confirm this was not a result of a cyber attack on the Akamai platform.
Get our [18]Tech Resources
[1] https://twitter.com/Akamai/status/1418247523454668803?ref_src=twsrc%5Etfw
[2] https://regmedia.co.uk/2021/07/22/downdetector_screenshot.png
[3] https://www.theregister.com/2021/06/17/akamai_prolexic_australia_outage/
[4] https://www.theregister.com/2021/06/11/cloudflare_outage_captcha/
[5] https://www.theregister.com/2021/07/13/kaseya_update/
[6] https://www.theregister.com/2021/07/20/hubble_first_pics/
[7] https://twitter.com/craigwatson1987/status/1418239715044696068?s=20
[8] https://twitter.com/PaytmMoney/status/1418243356707098625?s=20
[9] https://twitter.com/Akamai/status/1418251400660889603?s=20
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YPnqe54u-tpNtzOyZneYqgAAAA0&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YPnqe54u-tpNtzOyZneYqgAAAA0&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YPnqe54u-tpNtzOyZneYqgAAAA0&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[13] https://www.theregister.com/2021/06/08/fastly_outage_takes_down_half/
[14] https://www.theregister.com/2021/06/11/cloudflare_outage_captcha/
[15] https://twitter.com/Cloudflare?ref_src=twsrc%5Etfw
[16] https://twitter.com/hashtag/hugops?src=hash&ref_src=twsrc%5Etfw
[17] https://twitter.com/eastdakota/status/1418242657134944258?ref_src=twsrc%5Etfw
[18] https://whitepapers.theregister.com/
Re: Tried to access my bank around 5pm
"Tried to access my bank around 5pm ...Akamai never tested their systems properly"
Akamai is pretty darn reliable. If you're unhappy that your bank (just their web site, I hope) has a single point of failure, take it up with them.
Thankfully!
Something breaks an hour after I check out for 2+ weeks on holiday? Wheeee! NOtT my monkey, NOT my Circus! Not that I can fix something on the Web anyway??
Glad the situation has a fix...
Please add a Beer icon! We still can't have two?
Re: Thankfully!
Uh, what's this then? -->
Or did you mean you want another beer icon? Or an icon with two beers in it?
Re: Thankfully!
I assume they wanted to use two icons, and they're only allowed to select one.
Re: Thankfully!
Yes, Ive wanted two icons for sometime... troll drinking a beer!
Re: Thankfully!
Here, have one on me!
Happy hols :D
Downdetector?
How come Downdetector is never down? Who the hell is their provider? Why don't we just all use them?
Re: Downdetector?
Because, of all people, THEY know better than to have a single point of failure!!
Re: Downdetector?
"Taps head". If you have two points of failure then twice as much can go wrong.
Re: Downdetector?
Which is the strongest advocacy against having a bit on the side!
Re: Downdetector?
"Taps head again", you would never known down detector is down because there would be nowhere to report it.
Re: Downdetector?
Business opportunity: DownedDownDetectorDetector.com
Re: Downdetector?
DoubleDownDetector.com is shorter
Re: Downdetector?
Ah, but if you don't have a functioning "U" key on your keyboard it would be inaccessible.
Imagine the scenario - yor day's already going badly because yor keyboard is acting p, everyone yo email thinks yo're American with yor talk of "neighbors" and "colors"... and to top it all off, yo can't even figre ot why yor favorite web RLs aren't fnctional...
Re: Downdetector?
and there was me hoping we could have gone down a 1999 dance rabbit hole.
https://genius.com/Paul-johnson-get-get-down-lyrics
Cloudflare, Fastly, Akamai & AWS.
It’s scary how much of the internet relies so heavily on these 4 providers without a plan B
"Akamai said it had made repairs to address the outage."
But only for IPv4, cos, well, bollox to it!
Bad days happen to everyone
Is that it? Rather than promising plans for vastly increased resilience, all we can expect is an "aw, shucks - what am I like!"
I'd prefer, as just one example, an explanation of what type of release and deployment method they're using. Because the frequency of these outages strongly suggests an unwise faith in the Continuous Deployment religion in the most fundamental services the net relies on.
Re: Bad days happen to everyone
I sometimes wonder if these companies CEO's have a poster in their office of a cat and the caption "Hang in there baby"
Re: Bad days happen to everyone
"I'd prefer, as just one example, an explanation of what type of release and deployment method they're using. Because the frequency of these outages strongly suggests an unwise faith in the Continuous Deployment religion"
A friend was involved in setting up early infrastructure at Akamai. They built an architecture that was redundant and heterogeoous to a level I've never seen before or since.
It sounds like you know a lot about Continuous Development. If you identified a flawed deployment process, I'm interested to hear.
My experience is more with networking. At any layer, there is always at least one single point of failure. To the consumer, DNS is layer 3 (routing). To a network design, it's layer 7 (application). The actual implementation is probably something like AnyCast using BGP, which is wonderful stuff but also as complicated as it sounds.
tl;dr: I'm amazed that DNS works as well as it does.
CTRL+F
it's always DNS
Phrase not found
Tried to access my bank around 5pm
Then I headed for DownDetector to see a whole bunch of rather identical graphs on lots of sites.
So, with so many other systems relying on then, Akamai never tested their systems properly to make sure they hadn't become a SPOF.