How not to test a new system: push a button and wait to see what happens
- Reference: 1669618628
- News link: https://www.theregister.co.uk/2022/11/28/who_me/
- Source link:
This week meet a reader we'll Regomize as "Doug" who found himself in something of a hole when, as an ops manager for a food supply company, he oversaw a massive upgrade of IBM AS/400 F-Series kit. Six racks worth of it, including all the trimmings – mirrored disk protection as well as tape backups.
You can, after all, never have too many backups.
[1]
Obviously such a huge upgrade was not going to be done entirely by the customer's techies, so IBM supplied several engineers to do the serious voodoo parts, as well as a mobile disaster recovery vehicle. The plan was in four stages: build the new system in its dedicated space; transfer everything to the IBM van in the carpark; then transfer it all to the new system. All with appropriate testing of course. No chances were being taken.
[2]
[3]
The first two stages went well, and the van kicked into gear to provide the required services. Then the IBMers with their "eyes only" red books of engineering secrets set about mirroring to the new kit.
The third stage saw everything mirrored to the new system and the OS up and running on the new kit. All going well so far.
[4]Job 1: Get the boss on the network. Job 2: Figure out why Job 1 broke the network for everyone else
[5]Just follow the instructions … no wait, not that instruction to lock everyone out of everything
[6]Run a demo on live data? Sure! What could possibly go wrong? Hang on. Are you sure that's not working?
[7]The boss worked in a fishbowl, so office tricks were a treat
This is the point where Doug intervened. Chatting amiably to IBM engineer named "Bob:, Doug casually mentioned that the mirroring hadn't been tested yet and – with barely a thought to the consequences – hit a power button on the rack, switching off the disk array and tape backups. He had, he tells us, a grin on his face.
Bob did not have a grin on his face. Bob was horror-struck. Because of course the mirroring hadn't been tested yet – and just shutting it off without warning is not the IBM-approved way to test it.
[8]
Gathering his wits about him, Bob immediately ran diagnostics on the parts of the system that Doug had not just casually killed, and found that all was OK.
Doug, relieved, reached out to hit the power button again, only to find himself physically restrained by an exasperated Bob. Suddenly turning the drives off is not in the manual, and nor is suddenly turning them on again. And if it isn't in the manual it ought not to happen.
While Bob consulted with more senior engineers, Doug had to explain the situation to his own head of IT, as well as the logistics companies that were going to be relying on this kit once it was up and running. And of course to the board of directors.
[9]
In the end it was all smiles. The upgrade was successful, productivity boosted and Doug, forgiven, was promoted to head of IT.
But he learned an important lesson that day: never turn anything off if you don't know how to turn it back on.
Have you ever learned an important lesson the hard way? Found out that sometimes it's better not knowing what that mysterious red button does? Tell us all about it in an [10]email to Who, Me? and we'll make you (anonymously) famous. ®
Get our [11]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y4SUzlhAsknD8txtoV02HgAAAMc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y4SUzlhAsknD8txtoV02HgAAAMc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y4SUzlhAsknD8txtoV02HgAAAMc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://www.theregister.com/2022/11/21/who_me/
[5] https://www.theregister.com/2022/11/14/who_me/
[6] https://www.theregister.com/2022/11/07/who_me/
[7] https://www.theregister.com/2022/10/31/who_me/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y4SUzlhAsknD8txtoV02HgAAAMc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y4SUzlhAsknD8txtoV02HgAAAMc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] mailto:whome@theregister.com
[11] https://whitepapers.theregister.com/
Re: Alternative Lesson: "Never turn anything off if..."
You know, it'll turn itself automagically off when the UPS fails. Turning it back on usually is the problem.
Re: Alternative Lesson: "Never turn anything off if..."
Time taken for a senior manager to formulate and execute plan for an impromptu DC failover using the Big Red Button? Microseconds.
Time taken to get everything back up and working, failed boards replaced, systems restarted, rogue LAN cables replaced, comms balanced? 39 hours.
Size of boot applied to arse of senior manager by Very Senior Manager? <-------------------- This Big -------------------->
Re: Alternative Lesson: "Never turn anything off if..."
However, if a power failure needs serious effort or hardware fixes to get it going, its a shit system.
Who here has not see a UPS fail and simply take out the supply instead of going to bypass? (Think of any unfortunate APC owners)
Or, less commonly, a local digger has JCB'd the local 11kV feed and your power is off for hours and UPS exhausted? (Generators are available, some of them might even work when needed)
Re: Alternative Lesson: "Never turn anything off if..."
Turning it back on usually is the problem
Especially if, like some of our more 'legacy'[1] systems, turning the system back on isn't just a case of turning all the servers on.. no - a carefully-scripted set of actions (turn on server A, enable services in a specific sequence and timings, turn on server B, wait 5 minutes, turn on server C etc etc).
[1] Short for "we'd like to take them out the back for a mercy-killing but the business won't let us"
Yep, Doug was definitely a candidate for Head of IT...
As the BOFH remarked when they pushed him to become an IT manager
Boss: "a fool could do it!"
Simon: "a fool generally does it"
Having once been the IT Manager for my previous company (and SME that really, really needed 100% uptime on its server resources) my reaction to that quotation is to curl up in a corner and whimper quietly to myself.
I would say a Senior Head of IT.....
Just put the initals of his new job title on the office door!
Especially if Doug's middle name was Ian and surname Prentice (as two random names with the correct initials)
Testing times ahead
I expect one or two systems may get the Big Red Button test this winter but not by the one near the door in the server room!
Boo! Very poor "Who? Me?" this week. "One side of mirror switched off; mirror does its job" ... and that's it?
You've just given Who? Me? the Syndrome Award. Lame! Lame! Lame! ??
""One side of mirror switched off; mirror does its job" ... and that's it?"
Just a couple of years ago, we were moving an AS400 that was, in everyone's memory (and not on documentation), configured, $DEITY knows how, as a mirrored active/passive AS400 metro-cluster.
Upon switching the passive side on, we realized the active side switched to everything Read-Only.
We never had time to investigate this, but moved the passive AS400 as quickly as possible.
But then again, this legacy system was to be decommissioned. Has been in this state for multiple decades :)
And if it isn't in the manual it ought not to happen.
That suggests the manual is always correct.
Which in the case of IBM manuals is almost certainly never the case.
Sometimes you have to think outside the box. This is known as "experience" and "skill".
How to make an IBM engineer hyperventilate ?
Easy : ask him how many years he's got until retirement.
Why not use the backup generators.
I've told this one before.
A bank was hit with a power outage, and they went into the well practised fail over the backup site. A passing senior manager/director told them - don't do that - we have generators in the car park for this sort of emergency - use those and avoid the outage of switching sites. So they reluctantly restarted in place. Half way through the power on and restart, they found the generators did not have enough power for the machine room and so were stuck in limbo. They did not want to shutdown half way through an emergency restart, and they could not complete the start up to be able to shut it down. They had to wait a couple of hours before the power was restored, and they could complete the restart.
When the incident was reviewed by the board, the manager/director had to explain it was his decision, and admit he did not actually know about the generators capacity, he just paid for them.
The IT team learned a lesson - there are times when you ignore the management chain and do what you have practised.
Re: Why not use the backup generators.
The IT team learned a lesson - there are times when you ignore the management chain and do what you have practised.
That only works in situations where the (unknown, future) outcome is a success though. If the IT team tried "anything" and it did not go according to plan it would 100% be seen as their fault irrespective of the circumstances that preceded it.
Alternative Lesson: "Never turn anything off if..."
"you don't know how to turn it off "
FTFY