We have some sad news about Facebook. It's coming back after six hours offline
- Reference: 1633391587
- News link: https://www.theregister.co.uk/2021/10/04/facebook_outage_fixed/
- Source link:
WhatsApp and Facebook became available to users at around 2210 UTC on October 4 after [1]falling off the internet some six or so hours prior. Instagram and Facebook Messenger should be not far behind.
In the past hour, Facebook [2]tweeted : "To the huge community of people and businesses around the world who depend on us: we're sorry. We’ve been working hard to restore access to our apps and services and are happy to report they are coming back online now. Thank you for bearing with us."
[3]
CTO Mike Schroepfer earlier said: "Sincere apologies to everyone impacted by outages of Facebook powered services right now. We are experiencing networking issues and teams are working as fast as possible to debug and restore as fast as possible."
[4]
[5]
Founder Mark Zuckerberg chimed in: "Facebook, Instagram, WhatsApp and Messenger are coming back online now. Sorry for the disruption today – I know how much you rely on our services to stay connected with the people you care about." His otherwise most recent missive was a video of him [6]on a yacht .
The Register staff in the United States and Australia have experienced different levels of service since the resumption.
[7]
One Vulture in the USA was able to post to Facebook without issue. Antipodean staff were unable to post and saw errors such as the following...
[8]
Click to enlarge
Attempts to view notifications produced a “query error” dialog. WhatsApp was flaky – linking devices took over a minute. Instagram was down, at least in Australia, where its favicon loaded in a browser tab, but the site produced only the message, “Oops, an error occurred.”
Theories about the cause of the outage have focused on Facebook’s seemingly accidental and [9]sudden withdrawal of its BGP routes to its DNS servers, causing look-ups for Facebook domain names, such as facebook.com and instagram.com, to eventually fail.
This not only brought down Facebook's empire of apps but also apparently even caused door keycards to stop working on Facebook's campus. Staff fell back to Outlook, Zoom, and Discord to organize themselves and work on correcting the problem as they were unable to use internal and external Facebook-based systems.
It's time to decentralize the internet, again: What was distributed is now centralized by Google, Facebook, etc [10]READ MORE
In May this year, Facebook [11]announced it had built an automated peering configuration system. This software may or may not have been at the heart of today's outage. Someone claiming to work at Facebook [12]posted on Reddit, and since deleted their missive, that Facebook's peering routers went down likely due to a configuration blunder. Which makes sense given the circumstances: a bad config was deployed in production.
The IT breakdown was such that engineers needed to get physical access to the routers to fix and restart them, and a crack team was [13]sent into Facebook's Santa Clara, California, data center to do that, according to the New York Times.
How could a company of Facebook’s scale get BGP wrong? An early candidate is that aforementioned peering automation gone bad. The astoundingly profitable internet giant hailed the software as a triumph because it saved a single network administrator over eight hours of work each week.
[14]
Facebook employs more than 60,000 people. If a change designed to save one of them a day a week has indeed taken the company offline for six or more hours, that's quite something.
[15]Xero, Slack suffer outages just as Let's Encrypt root cert expiry downs other websites, services
[16]Which? survey finds people would actually pay the online giants not to take their data
[17]US school districts blame Amazon for nationwide bus driver shortage
The outage comes at a terrible time for Facebook, which in recent days has been the subject of damming leaks that suggest the company is knowingly indifferent to harms its platforms can create, including increased likelihood of self-harm by users, facilitating human trafficking, and ineffectual efforts to suppress hate speech and misinformation.
Documents shared by whistleblower Frances Haugen, a former Facebook employee, have also [18]suggested The Social Network™ ignored rules about content for high-profile users, and employed woefully insufficient numbers of staff who speak users’ native languages, thereby allowing vile content to circulate without checks.
Haugen has filed a complaint with the United States’ Securities and Exchange Commission, suggesting Facebook withheld information investors need to make informed decisions. That’s the kind of indirect but effective tactic that sees authorities chase mobsters for unpaid taxes rather than trying to secure evidence of murders.
The Register does not suggest Facebook has murdered anyone.
But today’s outages may have been extremely serious for those who rely on its services for day-to-day communications, both in their personal lives and for businesses that have gone all-in on Facebook as a customer communication and sales channel. ®
Updated to add
Facebook has [19]shared an official statement on what happened. It confirmed the outage took down its internal tools and systems, "complicating our attempts to quickly diagnose and resolve the problem." That's why it took so long to fix: it's hard to do so when your infrastructure has self-imploded.
It also confirmed an accidental configuration change ultimately caused the loss in connectivity:
Our engineering teams have learned that configuration changes on the backbone routers that coordinate network traffic between our data centers caused issues that interrupted this communication. This disruption to network traffic had a cascading effect on the way our data centers communicate, bringing our services to a halt.
We want to make clear at this time we believe the root cause of this outage was a faulty configuration change.
It also stressed there is "no evidence that user data was compromised," well, anymore than it usually is on Facebook.
Get our [20]Tech Resources
[1] https://www.theregister.com/2021/10/04/facebook_sites_outage/
[2] https://twitter.com/Facebook/status/1445155265360416773
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YVvODsKqb94G8i3NM2lrLgAAAFU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YVvODsKqb94G8i3NM2lrLgAAAFU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YVvODsKqb94G8i3NM2lrLgAAAFU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.facebook.com/zuck/posts/10113955029116581?__cft__[0]=AZXNGqWhb-3xu-PccymMVllvZs1CUeGIiHXMWE5DL2CdfDkjjdvhqxygjO3fmnh0B5lzHemS3LEmazt-LajZuoN0WnVpRNl2hhjAu_uQ2KGQQkWeLRTyLNaACc7C_5FEA9I_pHGaeCkbDLKlVzLc5CNa&__tn__=%2CO%2CP-R
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YVvODsKqb94G8i3NM2lrLgAAAFU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://regmedia.co.uk/2021/10/04/screenshot_facebook_error.jpg
[9] https://blog.cloudflare.com/october-2021-facebook-outage/
[10] https://www.theregister.com/2021/08/11/decentralized_internet/
[11] https://engineering.fb.com/2021/05/20/networking-traffic/peering-automation/
[12] https://twitter.com/nixcraft/status/1445107330018803732
[13] https://www.nytimes.com/2021/10/04/technology/facebook-down.html
[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YVvODsKqb94G8i3NM2lrLgAAAFU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[15] https://www.theregister.com/2021/09/30/lets_encrypt_xero_slack_outages/
[16] https://www.theregister.com/2021/09/30/which_data_survey/
[17] https://www.theregister.com/2021/09/29/amazon_contributing_school_driver_shortages/
[18] https://www.cbsnews.com/news/facebook-whistleblower-sec-complaint-60-minutes-2021-10-04/?ftag=CNM-00-10aab7d&linkId=134568667
[19] https://engineering.fb.com/2021/10/04/networking-traffic/outage/
[20] https://whitepapers.theregister.com/
OMG!
Facebook, Instagram, WhatsApp and Messenger are Down!
You said that just like I cared!
Re: OMG!
Heard they're back. Proves there's no G.
Clouldflare has a decent write-up on the BGP incident
https://blog.cloudflare.com/october-2021-facebook-outage/
Some people behind private DNS servers may have not been as impacted during parts of the outage as their resolvers cached the old info which was technically still valid.
My private DNS does the opposite, ensuring Zuckerburg's folley is unreachable year round, but not everyone has that luxury.
Funny thing about zombies,.. they always seem to get back up pining for,... BRAAAIIINNNNSS.
it saved a single network administrator over eight hours of work each week
And caused a share price drop that reportedly cost the boss $6 billion.
I thought I'd seen plenty of IT false economies in my time, but obviously they pale into insignificance compared to this. (as in, they should have tested it properly - not suggesting that scripting is bad!)
Automation
Sure, saving an admin eight hours of work per week seems trivial, but if you're automating a task like BGP updates, that should mean that you're removing eight hours worth of opportunity to foul up the network.
Of course, you do need to be sure that your automation process is more robust than the manual process for that to be true.
Looking forward to seeing the related "Who, Me?" entry.
Re: Automation
Looking forward to seeing the related "Who, Me?" entry.
You beat me to this. I suspect that eventually, some poor intern or PFY will get the blame though. The person who did mess it up won't say word until they're long retired. Pity... the "Who Me?" would be classic.
Unfortunately
Unfortunately it appears Failbook, et. al. are back up.
BGP has apparently rechristend the "Borked Gandalf Protocol", for the way it magically takes entire organizations offline due to some screwups on their part 9 times out of 10 (like forgetting to renew certificates, registrations, or configuration pushes. *LOL* )
LOL
"But today’s outages may have been extremely serious for those who rely on its services for day-to-day communications"
Don't rely on it. Simples.
Facebook is like dog poop
It exists, but it’s best avoided.
IMHO
Something Else
It appears now, at around 0200 GMT, that Outlook Live is down. Huh?
Re: Something Else
Possibly overloaded. The news articles on this outage mention that other services are being way overloaded. I guess people need their meme fixes and photos of granny's lunch.
There was on Yahoo! News... https://www.yahoo.com/news/facebook-whatsapp-instagram-outage-down-reactions-twitter-220806155.html
Re: Something Else
Not just because they need their fix of memes. A huge number of small businesses depend on Facebook - if you look up their URL it is a link to a Facebook page. If they take orders they run through Facebook. Support? Through Facebook or Twitter. Internal employee communication? Private Facebook group.
Probably a lot of them were left scrambling for a way to talk to each other when they didn't have everyone's email addresses, trying to find out what was going on because they couldn't take any customer orders, etc.
Granted something that could just as easily if not more easily happen if their had their own web site, e-commerce site, support email, etc. but hopefully they will become more cautious of trusting Facebook. At least if you sign up with Amazon or Microsoft's cloud for your services you are their customer. Your small business is not Facebook's customer.
Not a failure of testing - a failure of change enablement
To be fair, this is not really the fault of the automated change system not being tested properly - though that is probably one contributing factor.
It's really a failure of the change remediation not being tested properly.
If you want to move quickly, and accept that failure is a possibility - a luxury not afforded those running nuclear power stations - then you really do need to make sure you have a very effective roll-back solution, that is bulletproof.
That it took them six hours, and a site-visit, to roll back the faulty configuration change, establishes that it was not properly designed and tested.
The moral of the story is that, if you're modifying BGP automatically, you need, first, to design the safety-net, by writing, and testing, code that will reset it all to its last known working state -- reliably, every time.
To fail-fast, you must be able to reset-fast.
So they switched it off and then switched it back on again. Did Daddy Pig lead the crack team of engineers?