Oh no, here we go again, groans the internet as AWS runs into IT problems. Briefly this time
- Reference: 1639589541
- News link: https://www.theregister.co.uk/2021/12/15/aws_down/
- Source link:
Many feared another full-on AWS outage, as we saw [1]earlier this month , was about to kick off. The biz finally admitted on its [2]status page at 0748 PT (1548 UTC) that its US-West-2 region was experiencing connectivity problems, and similarly for US-West-1 at 0752 PT (1552 UTC).
Ten minutes later, it said it had worked out the root cause of the loss of connectivity to the regions, had made some fixes, and was seeing some recovery. And then at 0810 PT (1610 UTC), it declared:
We have resolved the issue affecting Internet connectivity to the US-WEST-1 Region. Connectivity within the region was not affected by this event. The issue has been resolved and the service is operating normally.
The same went for US-West-2 four minutes later. The total outage time was about 30 minutes. The above statement suggests that connections in and out of the region with the rest of the world were affected, and networking within the region was OK.
The exact cause was not spelled out. Perhaps a careless tech tripped over a cable, a backbone ISP somewhere had problems, or it was DNS. It is, after all, always DNS.
[3]
The effects of the downtime rippled through the internet in pretty much the same way as the [4]US-East-1 region borkage at the start of this month: people noticing websites and apps hosted by Amazon no longer working as expected. AWS did not immediately respond to our queries regarding today's event.
[5]
Snapped for posterity ... AWS's status page earlier today. Click to enlarge
The web giant's status page became increasingly unresponsive as either (a) netizens flocked to it to find out what had happened to their services or (b) things at AWS became increasingly borked.
[6]AWS postmortem: Internal ops teams' own monitoring tools went down, had to comb through logs
[7]Log4j RCE latest: In case you hadn't noticed, this is Really Very Bad, exploited in the wild, needs urgent patching
[8]AWS wobbles in US East region causing widespread outages
[9]The big AWS event: 120 announcements but nothing has changed
It's tough timing for the cloud colossus, which has been hard at work over the past week patching its components affected by the Apache Log4j remote-code execution vulnerability ( [10]CVE-2021-44228 ), judging by Amazon's [11]latest security bulletin on the matter.
AWS falling over, however briefly, is a reminder of just how reliant today's apps, websites, and services are on [12]singular platforms like, well, AWS.
[13]
Amazon-owned video-streaming darling Twitch also [14]broke down during the AWS connectivity blip. A glance at outage-spotting site [15]Downdetector showed a variety of well-known services –some of which aren't hosted by AWS – experienced issues at the same time as Amazon, including Zoom, Salesforce, Facebook, and Slack. That to us suggests there was some kind of underlying infrastructure issue, perhgaps.
Twitter, however, appeared to remain mostly upright. Thank goodness for that. ®
Someone has a terrible on-call shift at Amazon. May the on-call gods be with you. [16]#awsdown [17]#aws [18]pic.twitter.com/nxa5JB3TH7 — Sebastian (@ebud7) [19]December 15, 2021
Get our [20]Tech Resources
[1] https://www.theregister.com/2021/12/13/aws_postmortem/
[2] https://status.aws.amazon.com/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ybpzn1eBk5ca5HZ-dyGqPAAAAI4&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://www.theregister.com/2021/12/07/aws_wobbles_is_us_east/
[5] https://regmedia.co.uk/2021/12/15/aws_down.jpg
[6] https://www.theregister.com/2021/12/13/aws_postmortem/
[7] https://www.theregister.com/2021/12/13/log4j_rce_latest/
[8] https://www.theregister.com/2021/12/07/aws_wobbles_is_us_east/
[9] https://www.theregister.com/2021/12/09/the_big_aws_event_120/
[10] https://nvd.nist.gov/vuln/detail/CVE-2021-44228
[11] https://aws.amazon.com/security/security-bulletins/AWS-2021-006/
[12] https://www.theregister.com/2021/08/11/decentralized_internet/
[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ybpzn1eBk5ca5HZ-dyGqPAAAAI4&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[14] https://twitter.com/TwitchSupport/status/1471144138473189382
[15] https://downdetector.com/
[16] https://twitter.com/hashtag/awsdown?src=hash&ref_src=twsrc%5Etfw
[17] https://twitter.com/hashtag/aws?src=hash&ref_src=twsrc%5Etfw
[18] https://t.co/nxa5JB3TH7
[19] https://twitter.com/ebud7/status/1471148908260827144?ref_src=twsrc%5Etfw
[20] https://whitepapers.theregister.com/
Growing pains?
Now would be the time for emergent unstable behaviour in massive distributed systems.
Capacity might be there but the network might be brittle.
Technical debt, exhibit A
AWS is starting to approach Microsoft levels of technical debt, due to just layering on new features and services on an increasingly clunky code base (java, java, everywhere).
Internal build systems are often in a degraded state with insufficient capacity. Shoemaker's children and all that.
Duct tape and string, the lot.
AC for obvious reasons.
The cloud
So people try to diversify the services their businesses depend on, only to learn all those different services use AWS as their infrastructure.
Seems like the cloud is no longer fit for purpose.
I assume
Bezos is cutting each of those affected a check to compensate for the disruption.
What? He's not? What could be a more important use of funds.
Oh. I forgot. That penis shaped thingy.
Something, something, eggs, some, baskets
I'm sure I once heard something about eggs and baskets from a relative a couple of generations older than me.
AWS might not be all one basket, but clearly a few very large baskets is worse than many small baskets in terms of resilience, if not costs.