AWS US East region endures eight-hour wobble thanks to 'Stuck IO' in Elastic Block Store
- Reference: 1632815832
- News link: https://www.theregister.co.uk/2021/09/28/aws_east_brownout/
- Source link:
The lack of fun started at 8:11pm PDT on Sunday, when EBS experienced "degraded performance" in one availability zone (USE1-AZ2) in the US-EAST-1 Region. A subsequent update described the issue as "Stuck IO" and warned that existing EC2 instances may "experience impairment" while new EC2 instances could fail.
US East is the only AWS Region to offer six availability zones – a reflection of its status as the company's first location.
[1]
Other AWS services – among them Redshift, OpenSearch, Elasticache, and RDS databases – experienced "connectivity issues" as well.
[2]
[3]
By 9:17pm, AWS felt that the number of EC2 instances impacted by the issue had plateaued, but users continued to experience difficulties.
By 9:47pm the beginnings of an explanation emerged, as AWS revealed "A subsystem within the larger EBS service that is responsible for coordinating storage hosts is currently degraded due to increased resource contention."
[4]
Among the organisations impacted were secure messaging app Signal …
Hold tight, folks! Signal is currently down, due to a hosting outage affecting parts of our service. We’re working on bringing it back up. — Signal (@signalapp) [5]September 27, 2021
… and The New York Times games site (yes, your correspondent has a [6]Spelling Bee problem).
Hi folks, our Games page is now up and running. Apologies for disturbing your morning routines, but back to your Monday solving! 🧩 — NYTimes Wordplay (@NYTimesWordplay) [7]September 27, 2021
At 10:23pm AWS explained it had made "several changes to address the increased resource contention within the subsystem responsible for coordinating storage hosts with the EBS service". While those changes "led to some improvement" an 11:19pm update reported only "some improvements" but admitted "we have not yet seen performance for affected volumes return to normal levels".
A minute later, AWS rolled a change. By 11:43pm, AWS was confident enough to report the mitigations had worked, and predicted EBS volume performance would return to normal levels within an hour.
[8]Tech contractors fume over payday outage at Giant Pay after it sniffs 'suspicious activity'
[9]Square-shaped hole in workers' wallets after payment system fails at peak tip time
[10]AWS Tokyo outage takes down banks, share traders, and telcos
But at 1:15am the next day, a glitch struck. Some restored services slowed down again, and some new volumes also experienced "degraded performance".
By 3:36am new EC2 instances were again booting without incident, and at 4:21am the cloud concern updated its status feed with news that full operations had been restored at 3:45am.
But the company also admitted "While almost all of EBS volumes have fully recovered, we continue to work on recovering a remaining small set of EBS volumes.
"While the majority of affected services have fully recovered, we continue to recover some services, including RDS databases and Elasticache clusters," the final update added.
[11]
Clouds. Sometimes it's hard to find the silver lining. ®
Get our [12]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YVLn8TqZ3uFJWJVCgyCnIgAAAVM&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YVLn8TqZ3uFJWJVCgyCnIgAAAVM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YVLn8TqZ3uFJWJVCgyCnIgAAAVM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YVLn8TqZ3uFJWJVCgyCnIgAAAVM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://twitter.com/signalapp/status/1442354759009247232?ref_src=twsrc%5Etfw
[6] https://www.nytimes.com/puzzles/spelling-bee
[7] https://twitter.com/NYTimesWordplay/status/1442484820005826561?ref_src=twsrc%5Etfw
[8] https://www.theregister.com/2021/09/24/giant_pay_outage_uk/
[9] https://www.theregister.com/2021/09/23/square_tip_payment_failure/
[10] https://www.theregister.com/2021/09/02/aws_ap_northeast_1_outage/
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YVLn8TqZ3uFJWJVCgyCnIgAAAVM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[12] https://whitepapers.theregister.com/
"Sometimes it's hard to find the silver lining"
Well said.
That is not the problem of the manager who championed it, then left to go champion it again somewhere else.
I'm starting to hate cloud. Did you know that tower cases now come without any place to use an optical drive, or any ability to plug in a USB key on the front panel ?
I recently upgraded my PC which had been chugging along since 2010 and thought hey, while I'm at it, why not change tower ? Well today's towers expect you to throw all your data to the cloud.
I wonder how they expect people to reinstall Windows ?
Because that does happen, you know.
Oh, silly me, you bring it to a repair shop to pay a PFY to hook up an external USB optical drive and do the install you can't do anymore.
Obviously.
'Stuck IO' in Elastic Block Store
Does that mean it broke?