News: 1606329137

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

AWS admits to 'severely impaired' services in US-EAST-1, can't even post updates to Service Health Dashboard

(2020/11/25)


AWS has delivered the one message that will strike fear into the hearts of techies working out their day before Thanksgiving Holidays: US-EAST-1 region is suffering a "severely impaired" service.

At fault is the Kinesis Data Streams API in that, er, minor part of the AWS empire. The failure is also impacting a number of other services including CloudWatch, DynamoDB, Lambda and Managed Blockchain among others.

"This issue," admitted the AWS team, "has also affected our ability to post updates to the Service Health Dashboard."

Initial rumblings kicked off at around 1400 UTC today, with AWS confirming it was looking into increased error rates for the Kinesis Data Streams APIs in US-EAST-1.

Kinesis, for those unfamiliar with the service (one of a multitude that AWS will happily sell customers) is all about dealing with real-time data, such as telemetry from IoT devices. "Amazon Kinesis," [1]trumpets the company, "can handle any amount of streaming data and process data from hundreds of thousands of sources with very low latencies."

Unless, of course, it is borked.

Problems soon escalated. The company posted just over an hour later that it was working on identifying the root cause. Soon after it noted that other services were affected, including (but not limited to) "our ability to post updates to the Service Health Dashboard."

So quite bad then.

Finally, as 1700 UTC approached, AWS faced up to the grim reality of the situation and confirmed that the Kinesis Data Streams API was "severely impaired." CloudWatch, Cognito and EventBridge in the US-EAST-1 region are also affected by the Kinesis issue.

Problems could well have been exacerbated by the fact that AWS defaults to US-EAST-1 when endpoints are used with no Region set. US-EAST-1 is, [2]according to the company's documentation , "the default Region for API calls."

We contacted AWS to find out what had befallen the East Coast, and will update should the cloud giant respond. Its support orifice could only offer apologies to affected customers.

Twitter users were their usual supportive selves:

All of AWS us-east-1 is down [3]pic.twitter.com/SdO5J7qoEg — Kemal Ahmed (@carpetfortwo) [4]November 25, 2020

While multiple companies [5]realised just how dependant they are on AWS, other users raised a more salient point.

Pretty insane how literally the entire internet breaks whenever AWS goes down.

Also, quite concerning that one company singlehandedly controls 99% of the internet. — Jeff (@JeffTutorials) [6]November 25, 2020

Quite. ®

Get our [7]Tech Resources



[1] https://aws.amazon.com/kinesis/

[2] https://docs.aws.amazon.com/general/latest/gr/rande.html#view-service-endpoints

[3] https://t.co/SdO5J7qoEg

[4] https://twitter.com/carpetfortwo/status/1331648572757057539?ref_src=twsrc%5Etfw

[5] https://twitter.com/AdobeSpark/status/1331644328947552263

[6] https://twitter.com/JeffTutorials/status/1331649835917840387?ref_src=twsrc%5Etfw

[7] https://whitepapers.theregister.com/

aregross

Single

Point of

Failure

It'll get'cha everytime!

But

HildyJ

A single point of failure translates into multiple dollars of profit for Bezos so don't expect it to change anytime soon. It's all about the bottom line, not the customer.

Re: But

Peter-Waterman1

What single point of failure is that then? Don’t see that mentioned anywhere.

what a great day

Nate Amsden

I guess that's all I had to say, moved the org I work for out of their cloud about 9 years ago now saving roughly $1M/year in the process. Some internal folks over the years have tried to push to go back to a public cloud because it's so trendy, but they could never make the cost numbers come close to making it worth while so nothing has happened.

Re: what a great day

Anonymous Coward

What a load of tosh. So your company was all in on public cloud 9 years ago were they? Must have been a real trailblazer to have been spending enough to ‘save’ that much money by moving off? And what exactly were they using that they managed to save $1M/year by rebuilding the whole lot on-prem? And I guess you’ve not had to refresh/repair/service any of that hardware 2-3 times in the last 9 years?

Regardless of what you think about cloud, if your bean counters can’t make the numbers stack up then you need some new bean counters. There are reasons not to go to cloud, but your post is complete fiction.

Re: what a great day

Nate Amsden

Actually just retired some of our earliest hardware about 1 year ago. A bunch of DL385 G7s, an old 3PAR F200, and some Qlogic fibre channel switches. I have Extreme 1 and 10 gig switches that are still running from their first start date of Dec 2011(they don't go EOL until 2022). HP was willing to continue supporting the G7s for another year as well I just didn't have a need to keep them around anymore. The F200 went end of life maybe 2016(was on 3rd party support since).

Retired a pair of Citrix Netscalers maybe 3 years ago now that were EOL, current Netscalers EOL in 2024(bought in 2015), don't see a need to do anything with them until that time. Also retired some VPN and firewall appliances over the past 2-3 years as they went EOL.

I expect to need major hardware refreshes starting in 2022, and finishing in 2024, most gear that will get refreshed will have been running for at least 8 years at that point. Have no pain points for performance or capacity anywhere. The slowdown of "moore's law" has dramatically extended the useful life of most equipment as the advances have been far less impressive these past 5-8 years than they were the previous decade.

I don't even need one full hand to count the number of unexpected server failures in the past 4 years. Just runs so well, it's great.

As a reference point we run around ~850 VMs of various sizes. Probably 3-400 containers now too, many of which are on bare metal hardware. Don't need hypervisor overhead for bulk container deployment.

The cost savings are nothing new, been talking about this myself for about 11 years now since I was first exposed to the possibility of public cloud. The last company I was at was spending upwards of $300k/mo on public cloud. I could of built something that could do the handle their workloads for under $1M. But they weren't interested so I moved on and they eventually went out of business.

AWS gone off a cliff ... since Tim Bray left!

Forget It

Bork Bork Bork

and it not even the weekend!

I learned SRE at Google

Claptrap314

We would not even look at a service unless it was in three separate DCs.

I know that Amazon's regions & AZs don't map directly to Google's DCs, but if you haven't learned by now, US-East alone cannot provide five nines. You MUST go multi-region at least for that. Last I heard, AWS charges the big bucks for inter-regional data flows. Enough that it is probably worthwhile to look into the competition.

BTW, I'm available if you need detailed instruction/help to make it happen... ;)

disk iops

pricing can fix this pretty damn quick. Make US-east-1 double the price of US-WEST-1 or US-EAST-2 (ohio). Problem is the infra in Seattle and Ohio isn't being built big enough to moving 40% of us-east-1 into it.

Insanity is the final defense ... It's hard to get a refund when the
salesman is sniffing your crotch and baying at the moon.