Fastly 'fesses up to breaking the internet with an 'an undiscovered software bug' triggered by a customer
- Reference: 1623223507
- News link: https://www.theregister.co.uk/2021/06/09/fastly_explains_web_blackout/
- Source link:
The customer, Fastly points out in a post titled [2]Summary of June 8 outage , was blameless. "We experienced a global outage due to an undiscovered software bug that surfaced on June 8 when it was triggered by a valid customer configuration change," wrote Nick Rockwell, the company's senior veep of engineering and infrastructure.
The bug was introduced in a 12 May software deployment and lay dormant until, on 8 June, "a customer pushed a valid configuration change that included the specific circumstances that triggered the bug, which caused 85 per cent of our network to return errors."
[3]
Cue global chaos.
[4]
[5]
Rockwell's post states that Fastly "detected the disruption within one minute, then identified and isolated the cause, and disabled the configuration. Within 49 minutes, 95 per cent of our network was operating as normal."
[6]IBM Cloud resets 'Days Since Last Major Incident' clock to zero – after just five days
[7]Indian Finance Minister throws Infosys under the bus as new e-tax portal fails on first day
[8]Azure services fall over in Europe, Microsoft works on fix
[9]Colonial Pipeline suffers server gremlins, says it's not due to another ransomware infection
The veep also admitted that Fastly should have done better.
"Even though there were specific conditions that triggered this outage, we should have anticipated it," he wrote.
The company has therefore resolved to do four things:
We're deploying the bug fix across our network as quickly and safely as possible.
We are conducting a complete post mortem of the processes and practices we followed during this incident.
We'll figure out why we didn't detect the bug during our software quality assurance and testing processes.
We'll evaluate ways to improve our remediation time.
And, of course, it has apologised and promised it will do its very best not to make mistakes like this again. Which is just what all clouds, and social networks, say when they make avoidable but very damaging errors. ®
Get our [10]Tech Resources
[1] https://www.theregister.com/2021/06/08/fastly_outage_takes_down_half/
[2] https://www.fastly.com/blog/summary-of-june-8-outage
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YMCRQzqKW5p-iF0QA2JmsgAAAMU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YMCRQzqKW5p-iF0QA2JmsgAAAMU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YMCRQzqKW5p-iF0QA2JmsgAAAMU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2021/05/26/ibm_cloud_multiple_outages/
[7] https://www.theregister.com/2021/06/09/indian_tax_portal_fail/
[8] https://www.theregister.com/2021/05/20/microsoft_azure_outage/
[9] https://www.theregister.com/2021/05/18/colonial_pipeline_server_outage/
[10] https://whitepapers.theregister.com/
There is one step missing - we'll update our processes to make sure that *similar* bugs get caught (not just this one, but anything in this class).
Simple things in a message, that most understand. Line drawn under the issue. Next.
(like you, I can't think of a better way to put it).
things missing
techie thrown under bus outside company
bonuses all round for CEO, board and pals.
Snarking aside, full points for discovering cause in a minute. Seems like their monitoring code produces meaningful error messages, unlike some in the IT game.
And the bug was?
See title.
Re: And the bug was?
Exactly my thoughts, Cloudflare are very good at giving deep details of what went wrong e.g. last time they had a a big one they went as far as publishing their BGP filters to show how it happened.
At times they have even shown the bad code.
Yes a customer and not advocating for them above others but the level of openness is defiantly much better over there.
Re: And the bug was?
I am somewhat sympathetic until they have finished pushing the patch around - but the deep dive is what really boosts confidence about said misfortunes.
Credit where it's due
50 minutes to fix a hair-on-fire emergency? I'd call that a good performance under great stress. I'm sure I couldn't have done it. Bet the stressed engineers indulged in a few of the icon afterwards.
More generally, there's a reason the internet has a decentralised design. Why do all these numpties keep rushing to any company that centralises it? (Looks sideways at the employer's GitHub repository.)
Re: Credit where it's due
Something this close to the atom wide sharp bit of the pointy end should tell you pretty much exactly WHAT went wrong a small fraction of a second after it did. All power to the engineers being able to read the log files through the Niagra falls of sweat this would induce in most people. Once you've done that the WHY should be pretty clear thought the HTF do we fix it might take a couple of minutes going over the pre-written disaster recovery plan, which should include a big 'make sure this cant get in again' post mortem procedure which should explicitly exclude bean counters.
Re: Credit where it's due
When the failure is a customer config triggering a bug that was introduced months earlier.... spotting it might not be that easy and obvious
internet unavailable
not many dead
If StackOverflow didn't use Fastly
Then the Fastly guys may have been able to fix it quicker?
Design "reviews"
It seems to me that so-called design reviews are just a box-ticking exercise so that an activity in the management plan can be marked as completed. Nothing really happens until the whole system falls over.
It looked and sounded like a BGP fat finger error again . . . but the cover up sounds much better than an engineer mis-typed a subnet mask.
The company has therefore resolved to do four things:
We’re deploying the bug fix across our network as quickly and safely as possible.
We are conducting a complete post mortem of the processes and practices we followed during this incident.
We’ll figure out why we didn’t detect the bug during our software quality assurance and testing processes.
We’ll evaluate ways to improve our remediation time.
Assuming these things actually happen, then I can't think of a much better way to respond to a screw up.