A practical demonstration of the difference between 'resilient' and 'redundant'
- Reference: 1630913651
- News link: https://www.theregister.co.uk/2021/09/06/who_me/
- Source link:
Our story comes from "Dan", a lead system admin at what he described as a "rather large company" where the vast majority of the business went though a single suite of application backed up by a central database.
"The number of zeros on the 'dollars per minute' lost in unscheduled downtime was frightening," he told us.
[2]
The company liked Dan. He had rocked up after his predecessor departed under a cloud and had spent quite a while dealing with the environment. He found Development, QA and Production all running with differing patch levels and occasionally even different OS versions. Configurations didn't match. Hardware architectures differed. And so on.
[3]
[4]
It took a while but, once Dan had lined everything up, downtime due to bugs being found in production dropped significantly. It's fair to say the company was very pleased. Perhaps a bit too pleased.
"There remained, however, one nasty fly in the ointment," he said. "The guys in sales have been promising our clients for years that our systems were resilient and redundant.
[5]
"Resilient, they had become. Redundant they weren't. Not in any sense of the word."
However, bit by bit, the company's systems were indeed slowly becoming redundant, aided by Dan "stalking the developer cubes with a bat to encourage the removal, or non-creation, of code that would not play nice with node failover."
The final piece of the redundancy jigsaw was the database. It ran on Sun hardware, and Dan's team could hotswap pretty much any hardware component without the system suffering any downtime. Resilient, for sure. But still not redundant despite the joy from management at the dramatic reduction in outages.
[6]
"They thought my team had learned to walk on water or something," recalled Dan, happily.
The next step was to do some hardware duplication and create a cluster capable of withstanding all manner of disaster. "The DBAs were practically salivating at the prospect," Dan recalled, but the numbers involved were large enough to invoke a "steady on, chaps" from management.
On the day in question, Dan received a routine trouble ticket. It looked like either an adapter card or Gigabit Interface Converter (GBIC) had died. No problem – he was already on site and there were plenty of spares in the data centre.
[7]Hacking the computer with wirewraps and soldering irons: Just fix the issues as they come up, right?
[8]Scalpel! Superglue! This mouse won't fix its own ball
[9]Electrocution? All part of the service, sir!
[10]Undebug my heart: Using Cisco's IOS to take down capitalism – accidentally
He asked a colleague to pop in a change request for a hotswap and headed into the computing sanctum to the do the deed. He had just enough time to get it sorted before heading off for lunch.
It transpired it wasn't the GBIC at fault, but the adapter card.
Fine. He'd done this many times. It was a simple case of powering down the system, pulling it out on its rails, replacing the card and firing it back up. Simple.
Or not.
There was a slight wrinkle in the process, but not an unfamiliar one.
"You could reach ONE rail-lock easily from the front of the system," explained Dan. "Reaching both of them, however, you ended up hugging this massive and heavy box like a mother bear, flipping both locks and then starting the system back on its way into the rack with a little judicious hip pressure.
"Everyone on my team had done it dozens of times.
"This time, Murphy was watching and arranged for my belt buckle to occupy the exact same piece of the universe as the tiny, *unshielded* master power switch on the top panel at the front of the box...
"There was a click and microseconds later the pager on my belt went nuts as the only not-yet-redundant component of our entire business took a hard power outage."
Yes, Dan's lunch turned out to be very, very late that day. Still, the budget needed to add that last bit of redundancy arrived soon after.
Ever accidentally demonstrated just how stable (or not) your company's systems truly were? Or dropped some spectacles into the whirring blades of a PSU fan? Share your totally SFW tales of clothing or body parts causing IT chaos in an email to [11]Who, Me? ®
Get our [12]Tech Resources
[1] https://www.theregister.com/Tag/who-me
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YTXm1YqNa1iZ0UmOwEiRGQAAANM&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YTXm1YqNa1iZ0UmOwEiRGQAAANM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YTXm1YqNa1iZ0UmOwEiRGQAAANM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YTXm1YqNa1iZ0UmOwEiRGQAAANM&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/networks&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YTXm1YqNa1iZ0UmOwEiRGQAAANM&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2021/08/23/who_me/
[8] https://www.theregister.com/2021/08/16/who_me/
[9] https://www.theregister.com/2021/08/09/who_me/
[10] https://www.theregister.com/2021/08/02/who_me/
[11] mailto:whome@theregister.com
[12] https://whitepapers.theregister.com/
Re: The old demonstration of if only
It’s amazing how many business don’t see the need to invest in redundancy
Absolutely agree.
As far as many bean-counters are concerned, critical failures are million-to-one chances, and hence are so unlikely that there is no point spending money to protect against. What they always seem to forget is that million-to-one chances in IT occur nine times out of ten.
Re: The old demonstration of if only
UCAP,
"What they always seem to forget is that million-to-one chances in IT occur nine times out of ten, on average !!!"
FIFY :)
Re: The old demonstration of if only
"As far as many bean-counters are concerned, critical failures are million-to-one chances"
Brings to mind Feynman's observations when investigating the Challenger disaster that each management level somehow thought that the probability of a failure was an order of magnitude less than what the level below them thought it was.
An SFW tale to share?
Noooo... *Nervous looks left & right* Definitely NSFW, sorry. Besides, the NDA on what happened to the moose hasn't expired yet. *Cough*
Re: An SFW tale to share?
But.... Is the moose okay?
Re: An SFW tale to share?
Not after it had bitten my sister.
Re: An SFW tale to share?
Sure it was not the other way round?
Proliant server
A tower compaq proliant server (this is pre 2000, front panel is off as it had to be moved to gain access for something.
Thumb caught the power switch - the thing is these switches only actuate when the pressure is released - I sat there for about 10 minutes waiting for everyone else to shut down before I could let go of the switch - the springs on those switches get very heavy after a minute or two....
The server was the main Novell box for the entire company.
Resilient, they had become. Redundant they weren't.
How can a system be resilient without redundancy? :~
Resilience is the ability for a node to remain standing when it encounters error conditions. Redundancy is for when a node falls over.
Ohhh Yesss
Unfortunately I have been in a not so insignificant same boat.
This tale goes back to around 2004/5 ish, and it was on return to my employer. I had what I like to call a sabbatical with a competitor for around two years. As usual the grass was not as green as one suspected.
Me I am in the automation industry, purveyor of DCS systems. Those systems that control many Oil Refineries and the like. During my first stint I was on call, and this had onsite response times of 1 hr, and I thought I was a "seasoned pro". On return to my employer after my sabbatical I was put straight back on call, as I had not forgotten "anything" or so I and my manager thought. All this can unravelling down when one night the phone rang around 11PM - A call out, problem all IO had stopped working on a remote rack. The good thing was this hardware had intelligence and the plant continued to run, and held the previous setpoints and control strategies. This hardware was bespoke, but I knew it very well. Full of confidence I arrived at site, diagnosed the problem and they had a spare card (comms card). Now this is where my confidence unravelled, I completely forgot the chassis was hot swappable and duly powered it down. This in turn was followed by lots of hissing as valve opened or closed, and the whir of large motors slowing down. Oops.... I had crash shut the plant. Tail between my legs, I owned up, replaced the card and has to wait another 4 hours for them to get the plant running again. So what could have been a 30 min fix, turned out to be a longer stint than anticipated.
We've all done it...
Gone behind server racks and pulled plugs out with our backsides, only to find that the UPS was fine when plugged in, but miraculously failed as soon as power was removed, or even better, found out that the separate PDUs were both connected to the same UPS.
The old demonstration of if only
It’s amazing how many business don’t see the need to invest in redundancy.
It can be eye wateringly expensive, but so is the alternative of not having it when disaster strikes.
It’s even worse when businesses build redundant systems but don’t configure them properly, simply ticking a box on the DR plan that there is redundancy but requires huge manual failover of every different system.