Resilience is overrated when it's not advertised
- Reference: 1692343753
- News link: https://www.theregister.co.uk/2023/08/18/on_call/
- Source link:
This week, meet “Brad” who once worked for a company that provided criminal justice apps to police departments.
Brad was in support and one of the systems he tended was a Data General server – actually a pair of them, because one was set up to fail over to the other.
[1]
"Set up" may be too kind a description. The pair of boxes shared a floating IP – but the wire linking them had never been connected. So while the servers were configured to fail over, they were not capable of doing so.
[2]
[3]
"This was the very early days of failover, and it had never been tested," Brad explained. "The principle was there, but we had plenty of testing to do before we actually went live with failover."
Fair enough, then.
[4]
While the servers were in that state of not-quite readiness, Brad was on call and took the inevitable 1:00 AM phone call with a client complaint about slow performance.
Brad's first tactic was to stay in bed for a bit – he wasn't allowed to remote into these servers and hoped the problem would go away by itself.
It didn't. So he then faced an hour-long drive to his office, from where remote access was permissible and possible.
[5]
After telnetting into the server, he found it was redlining.
"It was going bonkers: maxed out on CPU, maxed out on memory, pretty much just seized up," Brad told On Call.
No obvious problem was apparent to explain the server's condition, so Brad ended up rebooting it – even as the constabulary who were his clients expressed their displeasure at not being able to do little things like process the evening's catch of malfeasants.
Brad labored mightily through the night, and into the dawn, without being able to find a fix.
[6]Lock-in to legacy code is a thing. Being locked in by legacy code is another thing entirely
[7]How to get a computer get stuck in a lift? Ask an 'illegal engineer'
[8]The choice: Pay BT megabucks, or do something a bit illegal. OK, that’s no choice
[9]Bizarre backup taught techie to dumb things down for the boss
Then one of his colleagues arrived to work at something approaching normal business hours, and spotted that the server resources were way lower than they should have been.
"He checked the physical IP address – not the floating one I was using – and saw what had happened. The previous day, unknown to the rest of us, the senior engineer (who was steadfastly unrepentant) had been on site and had connected the wire between the live server and the failover server."
This story ends with one piece of good news and two pieces of bad news.
The good was failover had worked – the live server had a problem and the backup server kicked in as planned, even if it had not yet been tested!
The first piece of bad news was that the backup box had less than half the memory and CPU of the live server. "It was doing its level best to keep up with demand but just couldn't cope," Brad wrote.
The second was that failover only worked in one direction: from primary to backup. Brad's reboots had been applied to the backup box, which didn't surrender the workload and instead tried to do the job with its inexplicably paltry collection of resources.
The fix was easy. Brad's colleague forced the floating IP back over to the primary server, and suddenly all was well.
Brad and his pals later removed the wire connecting the two servers, making sure failover wasn't working again – even though the client thought it had a resilient rig!
"A couple of years later we moved to Sun servers and this time made sure we tested failover before going live," Brad said.
Which may go some way towards explaining why EMC later acquired a stricken Data General for a [10]reasonable sum in 1999!
Have you ever been confused by tech that worked when it wasn't supposed to? If so, [11]click here to send On Call an email and we may share your story here on a future Friday. On Call has appeared without fail for years, but our resilience is not strong at present – we could use plenty more stories to keep the column at its best. ®
Get our [12]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZN9BSn@MRFec0upYotXwsgAAAQg&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZN9BSn@MRFec0upYotXwsgAAAQg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZN9BSn@MRFec0upYotXwsgAAAQg&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZN9BSn@MRFec0upYotXwsgAAAQg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZN9BSn@MRFec0upYotXwsgAAAQg&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2023/08/11/on_call/
[7] https://www.theregister.com/2023/08/04/on_call/
[8] https://www.theregister.com/2023/07/28/on_call/
[9] https://www.theregister.com/2023/07/14/on_call/
[10] https://www.theregister.com/1999/08/09/emc_squares_1_1bn_dg/
[11] mailto:oncall@theregister.com
[12] https://whitepapers.theregister.com/
Re: Fallback fault-tolerant
They can be.
The backup system having the same capacity as the primary unit certainly helps!
Re: Fallback fault-tolerant
The ancient counterparts worked just fine when spec'ed and installed and maintained properly.
Just like the modern kit.
Re: Fallback fault-tolerant
Many years ago, a city council in the north of England had a department running a pair of Netware 3 servers in a SFT cluster. The nodes were called Zig and Zag after a couple of characters on a breakfast TV show. One day, Zag had a permanent and unrepairable hardware failure leaving the cluster running only on Zig. Anecdotally, the users said that performance had improved - although it was only in that state for a few months before we installed the replacements with far more boring and forgettable names.
Failover backup redlining
The issue is not the failover, it's the idiot who designed a failover system with less resources than the production system.
If you design a failover, that server needs the exact same configuration than the one it is replacing.
Not doing that is stupid, and this was the result.
I would have thought that you wouldn't need a degree in computer science to understand that. Apparently, you do.
Re: Failover backup redlining
Yeah, the person who designed this is criminal...
Re: Failover backup redlining
Someone should beat them...
Re: Failover backup redlining
You could predict that backup server would plod...
Re: Failover backup redlining
Well, the backup performance was certainly arrested!
Re: Failover backup redlining
..."same configuration"...
Whilst this is desirable, it may not always be required.
If you say up front in the requirements that the backup server does not need to maintain the same performance, just as long as it can carry the load, then it could be smaller. But this would need to be communicated through the user base that when in failover, the service will be slower (and you probably want some indication that the service is running on the backup server, so users can see why this is running slow).
I've been in situations where this has been the decision made (and the client has had a load-shedding process to make sure that the essential parts of the service work at the expense of some of the others). It's a risk decision between cost and failover capability.
Not having a fail-back process is probably more of an issue in these cases, though.
Re: Failover backup redlining
In the Real World (tm) this is absolutely everywhere.
- Emergency lighting is not as bright.
- UPS and backup generators don't carry the whole load.
- Traffic diversions are onto smaller, slower roads.
- Backup Internet connections have less bandwidth and increased latency (eg cellular)
It's the normal way of doing redundancy.
Re: Failover backup redlining
The load shedding part is vital, that requires a decision as to what's not important enough to spend money on.
Re: Failover backup redlining
What's the betting that 'We're only using a third of the capacity on average so we don't need a full size backup' was part of the conversation.
Beancounters at work again?
Re: Failover backup redlining
If everything is a conspiracy to you, the problem isn't with the world...
"A couple of years later we moved to Sun servers
Did they use Solaris Jails?
Zones. Solaris isn't BSD anymore.
Resilience is futile.
Prepare to be discombobulated.
The horror...
Worked on a system where we had redundant failover servers, both quite well specced. Only problem was our software was so flaky we had to have a separate watchdog timer to reboot the system when our software stopped heartbeating. This happened a lot. The most reliable part of the system was the 'database' which was a set of text files, kept on a shared network drive (no sniggering at the back, please).
Re: The horror...
We had a pair of Sun E450s purchased from a reseller, but Sun insisted that we have official Solaris Cluster training. This was held on-site so the instructor decided that we could do real-world practicals rather than lab-based ones. Except we couldn't as the reseller had cabled everything wrongly and made a mess of the IP addresses too!
It took the instructor an hour or so of head scratching and troubleshooting before realising what was wrong and fixing it for us, after which it worked flawlessly.
Except the developers used the redundant server to develop (don't ask, I don't know why!) and every month or so we had a failover as the developers made a mistake. Eventually, they got given their own little server to use and the failovers stopped.
Fallback fault-tolerant
Are modern failover systems much better and more resilient than their ancient counterparts?