News: 1638900349

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

AWS wobbles in US East region causing widespread outages

(2021/12/07)


Updated Technical errors with the US-EAST-1 region of Amazon Web Services have caused widespread woes for customers, including difficulty accessing the management console and some other service problems.

The issues appear to be centred on the US-EAST-1 region, which is the oldest AWS region and located in North Virginia. This can have a global impact, as AWS noted in its [1]status report :

"This issue is affecting the global console landing page, which is also hosted in US-EAST-1."

[2]

Customers may be able to access region-specific consoles, the company said, by going directly to the URL for that region.

[3]

[4]

Within US-EAST-1 though, the affected services are not just the console, but also EC2 (Elastic Compute Cloud), DynamoDB and Amazon Connect. In reality, if EC2 is not working correctly hundreds of other services can be impacted since they run on EC2 behind the scenes.

[5]

The AWS status report for North America showing problems with key services

Twitter filled with frustrated customers, as well as suppliers apologising to their customers for the outage. Even vendors of cloud services were impacted, as many of these also run on AWS, such as Elastic Cloud which [6]reported : "We are experiencing issues with capacity scaling related to the elevated errors rates within us-east-1 (AWS N. Virginia) and are monitoring the situation."

[7]A bug introduced 6 months ago brought Google's Cloud Load Balancer to its knees

[8]Server errors plague app used by Tesla drivers to unlock their MuskMobiles

[9]BT's Plusnet shows Google how it's done as email woes enter their third day

[10]The planet survived six hours without Facebook. Let's make it longer next time

[11]Fastly 'fesses up to breaking the internet with an 'an undiscovered software bug' triggered by a customer

[12]If you can't log into Azure, Teams or Xbox Live right now: Microsoft cloud services in worldwide outage

Big companies believed to be affected include Amazon's own Alexa, Music and Ring, Netflix, Disney, [13]Discourse (which reported problems with "AWS Route 53, one of our DNS providers," Tinder and Roku.

One developer [14]said "AWS goes down and I spend 2 hour trying to debug why my code is not working," illustrating the extent to which public cloud services are assumed to be up and running.

While it is a serious outage, other regions in general seem to be unaffected, management console aside. There is a common issue with hyperscale services though, which is that while resilience in general is very good, there is a possibility of cascading failures because of service inter-dependencies.

[15]

AWS in its status report for the console and for EC2 said that "we have identified the root cause and we are actively working towards recovery," giving hope that the outage will not be long-lived. ®

Updated to add

The outage has been very bad news for the RISC-V team, which is currently hosting a virtual summit.

"We are aware and working closely with the technical team to get this resolved, and will update everyone once it is fully functioning again," a spokesperson told The Register .

"For those already in a session, we recommend not refreshing the screen as this may disconnect your stream. All sessions are recorded and will be available to you on-demand shortly after the virtual event platform is live again."

Smartish vacuum maker iRobot is also [16]reporting services on its app being affected.

[17]

"We have executed a mitigation which is showing significant recovery in the US-EAST-1 Region," Amazon said at 1404 PT (2204 UTC).

"We are continuing to closely monitor the health of the network devices and we expect to continue to make progress towards full recovery. We still do not have an ETA for full recovery at this time."

In other words, don't wait up.

Get our [18]Tech Resources



[1] https://status.aws.amazon.com/

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ya-nn2wkAHm0XNrvpbPGHAAAANY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ya-nn2wkAHm0XNrvpbPGHAAAANY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ya-nn2wkAHm0XNrvpbPGHAAAANY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://regmedia.co.uk/2021/12/07/status1.jpg

[6] https://twitter.com/ElasticStatus/status/1468267782579965952

[7] https://www.theregister.com/2021/11/23/google_outage/

[8] https://www.theregister.com/2021/11/21/tesla_server_error_500_lockout/

[9] https://www.theregister.com/2021/11/12/plusnet_email/

[10] https://www.theregister.com/2021/10/11/facebook_opinion_column/

[11] https://www.theregister.com/2021/06/09/fastly_explains_web_blackout/

[12] https://www.theregister.com/2021/04/01/microsoft_azure_dns_outage/

[13] https://twitter.com/DiscourseStatus/status/1468267672911548421

[14] https://twitter.com/renato_britto_/status/1468267998540705800

[15] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ya-nn2wkAHm0XNrvpbPGHAAAANY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[16] https://status.irobot.com/

[17] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ya-nn2wkAHm0XNrvpbPGHAAAANY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[18] https://whitepapers.theregister.com/



Remember - CLOUD COMPUTING is NOTHING MORE THAN ..

fredesmite2

Remember - CLOUD COMPUTING is NOTHING MORE THAN ..

using someone else's computer system .. thinking they care about it as much as you do.

Re: Remember - CLOUD COMPUTING is NOTHING MORE THAN ..

Gene Cash

Don't forget, cloud computing started as Amazon selling the excess data center capacity it only used during the Christmas rush...

So I'm not surprised it wobbles during December.

Re: Remember - CLOUD COMPUTING is NOTHING MORE THAN ..

Phil Kingston

There's a bit more to it

I’d love to comment but…

Must contain letters

.. the bloody internets broken again. Remind me why the bean counters thought this OAAS* was a good move?

Outage as a service. (tm)

Strange. I can't log into my AWS account at the moment

Jason Hindle

But I can ssh to one of my EC2 instances.

Re: Strange. I can't log into my AWS account at the moment

Androgynous Cow Herd

by IP, or by FQDN?

By IP...I'm not surprised.

DNS...it's ALWAYS DNS

Re: Strange. I can't log into my AWS account at the moment

Jason Hindle

Just checked my .ssh/config and it is actually the fqdn I'm using. Surprising - I do agree it is usually DNS.

Everything seems to depend on us-east-1

Dunstan Vavasour

Unable to create a S3 bucket in eu-west-1

Make multi-region redundancy seem a little less than convincing.

Re: Everything seems to depend on us-east-1

Peter-Waterman1

I can create a bucket in that region without issue. There are no services that span multiple regions. Every service is tied to a specific region, - s3 in Ireland is independent of s3 in Frankfurt. Even the global console is pinned to a single region, the effected region, all other region consoles are available.

Re: Everything seems to depend on us-east-1

Kevin McMurtrie

Everything seemed to depend on us-east-1 this morning. It looks like Amazon is shuffling internals as fast as they can.

michaelvirks

Big learning: never just have a root account on AWS, even if you are am occasional user one-man-show.

Apparently individual IAM user accounts can log in through region-specific management consoles like eu-west-1.console.aws.amazon.com, eu-central-1.console.aws.amazon.com etc

While all root account login requests get redirected to us-east-1, the only region that handles root login requests, which is also the region affected by the outage.

A developer tweeted

2+2=5

> One developer said "AWS goes down and I spend 2 hour trying to debug why my code is not working," illustrating the extent to which public cloud services are assumed to be up and running.

Error checking. We've heard of it.

Merrill

"A distributed system is one in which the failure of a computer you didn't even know existed can render your own computer unusable"

Leslie Lamport, 1987

PandyH

Although Amazon are saying to use a regional console to access the management console, you still cannot login.

Accessing the eu-west-2 console gets you a login prompt, followed by an "Internal Error" - looks like at least the authentication all hinges on us-east-1.

Cloud is shit

Anonymous Coward

Using cloud for ANYTHING business critical is wrong. Completely wrong. Own your infrastructure.

Androphobia:
Fear of men.