News: 1640217486

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

AWS power failure in US-EAST-1 region killed some hardware and instances

(2021/12/23)


A small group of sysadmins have a disaster recovery job on their hands, on top of Log4J fun, thanks to a power outage at Amazon Web Services’ USE1-AZ4 Availability Zone in the US-EAST-1 Region.

The lack of fun kicked off at 04:35AM Pacific Time (PST – aka 12:35PM UTC) on December 22nd, when AWS noticed launch failures and networking issues for some instances in its Elastic Compute Cloud IaaS service.

26 minutes later the cloud colossus ‘fessed up to a power outage and recommended moving workloads to other parts of its cloud that were still receiving electricity.

[1]

Power was restored at 05:39AM PST and AWS reported slow recovery of services, however a 6:51AM update admitted that ongoing networking issues were hampering efforts at full restoration.

[2]

[3]

Services including Slack and Asana reported service difficulties as a result of AWS' mess.

At the time of writing, AWS has still not fully restored networking.

[4]

And restoration may not be possible for some customers: at the time of writing, the most recent update on AWS’ status page offers the following grim news:

As is often the case with a loss of power, there may be some hardware that is not recoverable, which will prevent us from fully recovering the affected EC2 instances and EBS volumes. We are not quite at that point yet in terms of recovery, but it is unlikely that we will recover all of the small number of remaining EC2 instances and EBS volumes.

That’s the digital equivalent of waking up to a lump of coal on Christmas Day.

The incident is AWS’ second outage in a fortnight: on December 15th the operation’s [5]US-WEST-1 went missing for around 30 minutes. The US-EAST-1 region also browned out for eight hours [6]in September 2021 .

AWS advises customers not to rely on a single Availability Zone (AZ). The outfit’s architecture places two or more AZs within a single Region, and each Zone is physically distant from the others so that a single physical infrastructure incident can’t take out the whole Region. Using multiple regions therefore improves resilience – and cost.

[7]AWS flicks the switch on an Indonesian region

[8]AWS postmortem: Internal ops teams' own monitoring tools went down, had to comb through logs

[9]The big AWS event: 120 announcements but nothing has changed

Not every user follows AWS’ guidance about using multiple AZs, so when incidents like this strike their servers and data will become unavailable.

US-EAST-1 is AWS’ biggest and oldest region. Cloud economist Corey Quinn rates its importance as follows:

A multi-day full outage of us-east-1 will have an observable effect on the world economy. That is not an exaggeration. — Corey Quinn (@QuinnyPig) [10]December 8, 2021

AWS offers a [11]service level agreement of 99.95 per cent uptime for compute instances – or just under 22 minutes a month of downtime. If AWS misses that mark, it offers a ten per cent service credit, a sum that grows to thirty per cent if uptime drops below 99 per cent. If uptime falls below 95 per cent, customers are given 100 per cent of their fees as credits.

AWS also automatically waives fees if EC2 Instance are unavailable for more than six minutes inside a single hour.

Good luck if you’re one of the AWS customers faced with the need for a sudden rebuild. ®

Get our [12]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YcQCd@XWfSDHPpeygl0WQAAAAEA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YcQCd@XWfSDHPpeygl0WQAAAAEA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YcQCd@XWfSDHPpeygl0WQAAAAEA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YcQCd@XWfSDHPpeygl0WQAAAAEA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://www.theregister.com/2021/12/15/aws_down/

[6] https://www.theregister.com/2021/09/28/aws_east_brownout/

[7] https://www.theregister.com/2021/12/14/aws_flicks_the_switch_on/

[8] https://www.theregister.com/2021/12/13/aws_postmortem/

[9] https://www.theregister.com/2021/12/09/the_big_aws_event_120/

[10] https://twitter.com/QuinnyPig/status/1468440573275111425?ref_src=twsrc%5Etfw

[11] https://aws.amazon.com/compute/sla/

[12] https://whitepapers.theregister.com/



Elastic

HildyJ

Elastic Compute Cloud sounds great. If you ignore the fact that elastic under stress tends to snap. And AWS's elastic seems to do it on a regular basis.

While those affected have to make do with their lump of coal, rest assured that the execs are all nestled in bed while visions of sugarplums dance in their heads.

Re: Elastic

Version 1.0

It's the winter and we're having some cloudy weather ... down here on the Gulf coast this sort of bad luck doesn't seem too bad ... it just illustrates the need for backups and alternative servers when you are just hoping the clouds will not have a problem ... like the wind speed getting up to 150mph ... that can cause a cloudy fuse to blow.

Re: Elastic

Snake

"26 minutes later the cloud colossus ‘fessed up to a power outage and recommended moving workloads to other parts of its cloud that were still receiving electricity."

So. What's the use of calling the system "elastic" and "cloud based" IF THEY MAKE YOU MANUALLY RECONFIGURE DURING AN OUTAGE??!

If "cloud" were truly this magic solution to flexible workload problems, wouldn't it auto-reconfigure to a different part of their network when systems go down?

Amazon's statement is to switch servers and you can do that yourself on your own hardware. So why pay to use someone else's hardware, one that markets redundancy and availability, if they can't even be responsible for providing that without forcing you, the customer, to handle this topic yourself??

New AWS service working as expected

cjcox

Amazon recently brought up their Elastic Total Failure Service and has reported that so far, it has been quite reliable. Right now the service is free for all AWS customers. /s

"I just checked. We're down." - Joe Satisified, Important Company, Inc.

9.5 hrs of downtime

Nate Amsden

For Chef's software stack (https://status.chef.io/), my alerts dashboard hadn't been that red in years. Fortunately I didn't have to make any chef changes today. Obviously they didn't have a disaster recovery plan(nor do most companies), they waited for amazon to fix their stuff then tried to recover what they could(at least that is what it seems like as an outsider anyway).

The list of companies affected by these outages(amazon included) just show building apps that are resilient to such cloud failures is beyond the reach of most organizations(whether it is complexity or cost or both). I've only been saying that for just over eleven years now. Not surprised the trend continues. I moved my org out of amazon cloud in 2012 and have been running trouble free ever since with literally $10-15M+ in savings since(would of been nice to get more of that savings invested into more infrastructure but the company was stingy on everything). There was no lift and shift into the cloud, the company was "born" in the cloud(before I started even). But still many people just don't get it(that cloud is almost always massively more expensive than hosting yourself unless you are doing a really bad job of hosting it yourself which is certainly possible, though much more common to host it in cloud very poorly then hosting it yourself). I don't get how you couldn't get it at this point.

Cloud outages are becoming a regular occurrence

DS999

I wonder how many can occur before businesses start to rethink depending so much on the cloud?

My friends, I am here to tell you of the wonderous continent known as
Africa. Well we left New York drunk and early on the morning of February 31.
We were 15 days on the water, and 3 on the boat when we finally arrived in
Africa. Upon our arrival we immediately set up a rigorous schedule: Up at
6:00, breakfast, and back in bed by 7:00. Pretty soon we were back in bed by
6:30. Now Africa is full of big game. The first day I shot two bucks. That
was the biggest game we had. Africa is primarily inhabited by Elks, Moose
and Knights of Pithiests.
The elks live up in the mountains and come down once a year for their
annual conventions. And you should see them gathered around the water hole,
which they leave immediately when they discover it's full of water. They
weren't looking for a water hole. They were looking for an alck hole.
One morning I shot an elephant in my pajamas, how he got in my
pajamas, I don't know. Then we tried to remove the tusks. That's a tough
word to say, tusks. As I said we tried to remove the tusks, but they were
imbedded so firmly we couldn't get them out. But in Alabama the Tuscaloosa,
but that is totally irrelephant to what I was saying.
We took some pictures of the native girls, but they weren't developed.
So we're going back in a few years...
-- Julius H. Marx [Groucho]