Watch a RAID rebuild or go to a Christmas party? Tough choice
- Reference: 1657527969
- News link: https://www.theregister.co.uk/2022/07/11/who_me/
- Source link:
Today's story, from a reader Regomized as "Sam", takes us back to the mid 2000s and happier days before Brexit, Trump and the pandemic showed up.
During these halcyon times, Sam was working as a civilian employee for a local UK Police force as a member of the desktop support team. He'd been there a few months when the festive season rolled around.
[2]
Christmas parties were a thing for Sam and the team, and each group had its own festivities during the day. Other groups would cover for them while crackers were pulled and silly hats worn. Each party normally ran from lunchtime until well into the evening.
[3]
[4]
On the day in question it was the server team's turn to head off while Sam and the desktop crew ensured everything was covered in their absence. Naturally, there was a bit of mistrust between administrators and mortals, and so Sam and co naturally took to sellotaping festive wishes over the lasers on the undersides of the admins' mice as well as other japes.
It was all fun and games until 3pm, when two of the server team (who'd left behind their mobile numbers "just in case") burst back into the room.
[5]
Something had gone terribly wrong.
"A little background..." explained Sam, "We had 2 Exchange Servers. Your placement was dependent upon your staff number; odd on one, even on the other."
One of the Exchange Servers had abruptly died. "Consequently, half of the constabulary had lost email!"
[6]
"Being a public service, we had top-notch Microsoft support at a better price than many so made copious use, long into the wee small hours..."
Eventually it transpired that a disk in one of the servers had failed the day previously. Sabres in the form of support contracts were rattled and a replacement had arrived that morning. The server was ok - after all, it could handle a failed a disk. The replacement was popped in and a rebuild was started.
And then the desktop team set off for their festive frivolities. The server would sort itself out, as designed.
What could possibly go wrong?
As it turned out, quite a bit.
"While in the process of rebuilding the array, a second disk decided it was going to join the party and while not quite failing, caused enough disturbance in the force, that Windows and Exchange violently soiled themselves and gave up, crying in the corner," said Sam.
[7]You need to RTFM, but feel free to use your brain too
[8]Know the difference between a bin and /bin unless you want a new doorstop
[9]Beware the fury of a database developer torn from tables and SQL
[10]An international incident or just some finger trouble at the console?
To be fair to the Server team, they had warned that the Exchange setup was living on borrowed time, was way overloaded (compared to Microsoft's recommendation) and, according to Sam, the phrase "It's a disaster waiting to happen!" had been thrown at management on more than one occasion.
And then, in time-honored IT fashion, it did.
"Nothing the many iterations of increasingly technical Microsoft guys threw at it helped one bit," said Sam. "It wouldn't repair and it wouldn't restore from backup.
"It had to be rebuilt and restored by-datastore in order to get it all back and working."
The long and tedious job (which required user data being restored in batches) took a little under two weeks to complete, making for a miserable Christmas and New Year for all concerned.
And the person at the end of the queue? The biggest of Police cheeses: the Chief Constable. His calendar had been used as audit trail of his activities, meaning that most of a datastore was all about him, and getting it back presented the most problems.
The curse of being a Very Important Person.
While putting Bobbies on the Beat might garner favourable headlines, neglecting basic IT principles can result in all manner of catastrophe. Have you tried out your disaster recovery plan recently? Or hasn't your company considered the consequences of when that innocuous beige box inevitably turns brown? Share your tale with an email to [11]Who, Me? ®
Get our [12]Tech Resources
[1] https://www.theregister.com/Tag/Who,%20Me?/
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ysv0yD3ztl12NTtWxhv4LwAAAAk&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ysv0yD3ztl12NTtWxhv4LwAAAAk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ysv0yD3ztl12NTtWxhv4LwAAAAk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ysv0yD3ztl12NTtWxhv4LwAAAAk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/systems&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ysv0yD3ztl12NTtWxhv4LwAAAAk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2022/06/27/who_me/
[8] https://www.theregister.com/2022/06/20/who_me/
[9] https://www.theregister.com/2022/05/23/who_me/
[10] https://www.theregister.com/2022/05/09/who_me/
[11] mailto:whome@theregister.com
[12] https://whitepapers.theregister.com/
It sounds like the Police RAID didn't go to plan
It has to be a rule of IT somewhere, any routine task will fail once you stop monitoring it.
More pertinantly
RAID isn't RAID if it isn't actively monitored.
And between polls it must be assumed to be in a degraded state.
It's got RAID so it can't fail. . .
I was once responsible for a NAS which contained several VMs - including which were high profile to the project. The NAS was old and had been shifted several times before it came to me. We had, had a disk fail and I suggested - in writing! - that all should be replaced as they were now supect. There was the inevitable no reponse.
Isn't it amazing how the MTBF can be calculated so precisely the the failure of a second disk can occur within the time taken to purchase a replacement for the first failure!
Of couse there wer3e no backups - it had RAID.
Re: It's got RAID so it can't fail. . .
"It's got RAID storage!", they said.
"Disks from the same manufacture, of the same type and from the same batch?", sez I.
"Yes. They said.
"Then you haven't got RAID." sez me.
"It's a disaster waiting to happen!"
Reminds me of a Halloween costume party... One guy came holding field glasses, with a toilet seat around his neck! I asked him WTF are you... he replied "I'm a bad accident looking for a place to happen!" Simply awesome!
yeah, that second disk failure during a rebuild... my sphincter always puckered and un-puckered during a raid-rebuild
My rule of thumb for a server with a failed RAID disk when I was hands on - make sure you have a reliable backup* before you do anything else. And if it's Exchange, where absolutely possible, stop all of the MS Exchange Services and do an offline backup. Believe me you will thank me later for those few hours of downtime that you are currently cursing.
*I appreciate it isn't always possible to prove the integrity (or even the overall usefulness) of said backups before commencing work but at least take one before you start and do everything you can to verify it, however little that might actually be - but being able to stand in front of the big cheese and demonstrate that you did everything possible to ensure you covered all bases can be a career saver.
RAID is not backup.
It’s just drive redundancy, with varying degrees of fault tolerance of course.
Now, when one disc goes, particularly when they’re from the same manufacturer, batch number etc (they need to be identical, right?) that means the rest are about to go. Get on it.
IBM Engineer...
We were getting disk errors on one drive in the disk array on an AS/400. The engineer was adamant it was a cable fault so to avoid downtime I left my very experienced evening operator and him to sort it after office hours.
Sheer chaos hit as it wasn't the cable, it was the drive itself which then failed to restart. Of course because it was "just the cable" the engineer hadn't taken the drive out of the array first and as it was only striped, not mirrored, the system was f****d. To say I wasn't happy with either the engineer or the operator, both of whom should have known better, is an understatement.
Fortunately we had backups but it still took a few days to get everything fully working.
Eventually it transpired that a disk in one of the servers had failed the day previously.
So the disc copped it?