Day 7 of the great Atlassian outage: Company still struggling to restore access
- Reference: 1649680207
- News link: https://www.theregister.co.uk/2022/04/11/atlassian_still_down/
- Source link:
At this point it is fair to say the problem is severe. It kicked off on 5 April and the [2]company said that while it was "running a maintenance script, a small number of sites were disabled unintentionally".
Atlassian reported in a status update this morning that: "The rebuild stage is particularly complex due to several steps that are required to validate sites and verify data." This will be of little comfort to the 65 percent of affected customers wondering if signing up for Atlassian's cloud was such a good idea after all.
[3]Atlassian outage lingers, sparking data loss fears
[4]Atlassian adds Analytics, Atlas, Compass to line up
[5]Atlassian Jira, Confluence outage persists two days on
[6]Atlassian flags Bitbucket and Confluence Data Center flaws
We suspect that the "dedicated team" Atlassian assigned to sorting out the problem has yet to take down the bunting from [7]World Backup Day before the incident occurred.
Jira Software, Jira Work Management, Jira Service Management and Confluence are the biggest segments still affected at the time of publication. Confluence is a web-based corporate wiki and Jira is more about issue tracking. Jira Work Management is aimed at generic project management while Jira Service Management [8]turned up last year as part of a vision to turn agile and DevOps principles to the IT service desk.
[9]
The irony of the latter collapsing into a heap due to an issue with a maintenance script will not have been lost on the affected users.
[10]
Also still on the broken list is IT incident-monitoring service, OpsGenie ( [11]acquired in 2018 ) and Atlassian's Statuspage incident communication service – again, the irony meter has gone off the scale.
To be clear, only a very small proportion of customers have been affected; [12]on Friday we were told the figure was around the 400 mark. Still, we suppose being one of those 400 hasn't exactly been fun.
Atlassian has been down for 6 days. Our project managers are climbing the walls and our service request queues are all Slack threads.
I'm currently doing the on-call rotation checklist in a badly formatted Word document someone had managed to copy from Confluence🙃 — Morten Linderud (@MortenLinderud) [13]April 11, 2022
The Register contacted Atlassian to get more information and an ETA for when the rest of the data will be recovered. We will update should the company respond. ®
Get our [14]Tech Resources
[1] https://jira-service-management.status.atlassian.com/
[2] https://www.theregister.com/2022/04/06/atlassian_jira_confluence_outage/
[3] https://www.theregister.com/2022/04/08/atlassian_service_issues_linger_leaving/
[4] https://www.theregister.com/2022/04/06/atlassian_new_services/
[5] https://www.theregister.com/2022/04/06/atlassian_jira_confluence_outage/
[6] https://www.theregister.com/2022/03/25/atlassian_hazelcast/
[7] https://www.worldbackupday.com/en
[8] https://www.theregister.com/2020/11/09/atlassian_jira_management/
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YlRQss1@c6X14bEYfCz4HAAAAAU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/devops&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YlRQss1@c6X14bEYfCz4HAAAAAU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[11] https://www.theregister.com/2018/09/04/atlassian_buys_opsgenie/
[12] https://www.theregister.com/2022/04/08/atlassian_service_issues_linger_leaving/
[13] https://twitter.com/MortenLinderud/status/1513441292537282561?ref_src=twsrc%5Etfw
[14] https://whitepapers.theregister.com/
Re: But but but....
But they are good!
Just not for the consumer.
Who woulda thunk outsourcing your data would ever become a problem, eh? One in a million chance!
Re: But but but....
We use Atlassian on-premises, so not affected
Atlassian are stopping on-prem so we are currently trailing switching to Github cloud (?)
Business continuity
Now is a good time for all of these project managers to reflect on the importance of resilience and business continuity.
Scream and shout at Atlassian as much as you want, it won't restore it faster.
Ah....remember....."cloud" is cheaper......
.....until you factor in the cost of NOT HAVING THE SERVICE AT ALL!!!
Re: Ah....remember....."cloud" is cheaper......
On-premise solutions are also capable of providing no service at all.
Re: Ah....remember....."cloud" is cheaper......
Only if you are incompetent.
Re: Ah....remember....."cloud" is cheaper......
really it comes down to too many eggs in one basket. Certainly service failures can occur on premises. But pretty much universally those failures affect only a single organization. Granted there can be times when multiple companies are experiencing problems but it's still tiny compared to the blast radius of a SaaS provider having a problem.
My biggest issue with SaaS at least from a website perspective is the seemingly constant need that the provider feels to change the user interface around and convinced everyone will love the changes. Atlassian has done that tons of times and it has driven me crazy. Others are similar, so convinced all customers will appreciate the changes.
Go change the back end all you want as long as the front end stays consistent please.
At least with on prem you usually get to choose when you take the upgrade, and in some cases you can opt to delay indefinitely (even if it means you lose support).
Just now I checked again to confirm. Every few months I go through and bulk close resolved tickets(in Jira) that have had no activity for 60 days. I used to be able to add a comment to those tickets I would say "no activity in 60 days, bulk closing". Then one day this option vanished. I asked Atlassian support what happened and they said that functionality was not yet implemented on their new cloud product (despite us having being hosted in their cloud product for years prior). I can only assume it is a different code base to some extent. Anyway that was probably 3-5 years ago, and still don't have that functionality today. (there is an option to send an email to those people when the ticket closes I don't want that, I just want to add a comment to the ticket).
Don't get me started on the editor changes in confluence in recent years just a disaster. Fortunately they have backed off of their plans to eliminate the old editor(for how long I don't know but it seems like it's about 2 years past when I expected them to try to kill it).
Then there was the time they decided to change the page width on everything in confluence(I assume to try to make it printable), at least in that case they left an option(per user option) to disable that functionality(it messed up tons of pages that weren't written for that option).
The keyboard shortcut functionality drove me insane in confluence as well, for years assuming it was there before(I don't know, I never used keyboard shortcuts in confluence going back to my earliest days of using it in 2006) it was not a problem but past couple of years I would inadvertently trigger a series of events on documents that I did not want just by typing. I was able to undo it every time, and finally disabled the keyboard shortcuts a few months ago.
From the ZSF book of quotations...
"But if it's in the cloud, that's always cheaper and safer because we don't need local people managing it? And if it goes wrong we can just sue them, right? Right?" -- way too many middle managers
Re: Ah....remember....."cloud" is cheaper......
Cloud isn't cheaper - you're paying for all the bodies to do the hard graft of rebuilding a broken system for you and taking the political flack
The on premises alternative means you having availability of knowledgeable staff (not off sick with Covid/Holiday) plus spares for any server/network tin/data centre pieces/rooms/... and taking the political flack...
Pick your risk profile...
Re: Ah....remember....."cloud" is cheaper......
Which you will have in place, if you are not incompetent.
Re: Ah....remember....."cloud" is cheaper......
And big enough to have:
* Enough knowledgeable staff to cover for sickness, holiday, COVID, etc
* Enough kit/capacity to cope with systems(s) going down
You pick the right tool for the job. I work for a large company and we have a mixture of on-prem and cloud. On-prem when the problem is big enough to tick the boxes above; Cloud when the product/service is too small/niche for us to keep skilled up to manage.
IDEA: Let's put our core contact with clients into another companies hands and hope for the best.
"Welcome to Itchy and Scratchy Land, where nothing can possibli go wrong ... Er, possibly go wrong...that's the first thing that's ever gone wrong."
Could be 2 more weeks according to a Reddit comment
Seen a Reddit comment - [1]Link
got email from the community manager, that some instances can be down for further two weeks.
This is not how a billion dollar company build the system or handles recovery, I am going to look for an alternative and dump Atlassian as soon as possible.
==== snip of the email I got ====
What this means for your company
We were unable to confirm a more firm ETA until now due to the complexity of the rebuild process for your site. While we are beginning to bring some customers back online, we estimate the rebuilding effort to last for up to 2 more weeks.
I know that this is not the news you were hoping for. We apologize for the length and severity of this incident and have taken steps to avoid a recurrence in the future.
[1] https://old.reddit.com/r/atlassian/comments/u13o4u/our_cloud_instance_is_still_down/
Re: Could be 2 more weeks according to a Reddit comment
Just the vagueness of that "update" should raise alarms.
Re: Could be 2 more weeks according to a Reddit comment
@John Miles
The people responsible for sacking the people responsible for taking steps have been sacked.
Compensation ?
Be curious to know how much compo Atlassian think is appropriate, and then compare it to how much they get sued for.
When RBS went down, some people lost jobs and houses.
Re: Compensation ?
One month's free trial should be enough, surely?
Re: Compensation ?
One month's free trial of no service?
Sounds like a deal to me*
*Not saying it's a good deal...
We suspect that the "dedicated team" Atlassian assigned to sorting out the problem has yet to take down the bunting from World Backup Day before the incident occurred. [...] The irony of [Jira Service Management] collapsing into a heap due to an issue with a maintenance script will not have been lost on the affected users."
^^^ This is why I read El Reg. ^^^
But but but....
Software as a Service and the Cloud are good things, right? Right? Riiiight?
Yeah, right. Right up until they're not.