Please install that patch – but don't you dare actually run it
- Reference: 1707467589
- News link: https://www.theregister.co.uk/2024/02/09/on_call/
- Source link:
This week, meet "Kane" who shared the story of the time a client asked him to connect a storage cluster to an ESXi host.
"The host had used all its internal storage and the client wanted to add extra drives," Kane told On Call. He countered with a suggestion to instead connect the host to the 40 terabytes of unused storage in a cluster his client already operated, but was mostly ignoring.
[1]
Kane hadn't made this sort of connection before, but knew ESXi would happily connect to external storage.
[2]
[3]
Yet as he tried to make it work, Kane struggled to get the storage cluster to communicate with the host. He eventually learned that the cluster's OS couldn't do the job because it was running very old software that was well and truly past end of life.
Kane therefore sought permission to upgrade the cluster, and requested an outage window in which to make the change.
[4]
The request for the upgrade was approved. But he was denied an outage window.
When Kane explained this policy would mean he couldn't perform the task for which he was being paid a healthy hourly rate, he was told his client had a policy not to allow outages.
[5]Techie climbed a mountain only be told not to touch the kit on top
[6]Standards-obsessed boss ignored one, and suffered all night for his sin
[7]While we fire the boss, can you lock him out of the network?
[8]People power made payroll support in putrid places prodigiously perilous
It was at this point Kane asked about security patches – which more often than not require a reboot before they'll work. Surely the client allowed the brief outages required to ensure they were properly implemented?
It did not.
"You could install security patches and upgrade an OS if you wanted to, but you could not reboot," Kane told On Call. "Even when the security patch required a reboot to take full effect."
[9]
The Register feels the org Kane served therefore produced impressive uptime statistics, but shudders to think about the state of its security.
So did Kane. He told us most of the servers there constantly displayed the message "Reboot required."
What's the weirdest IT malpractice you've ever come across? [10]Click here to send your story to On Call as an email and we'll try to add it to a future version of the column. ®
Get our [11]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZcYF2Csy6rWQvqHIi9qDqwAAAYU&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZcYF2Csy6rWQvqHIi9qDqwAAAYU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZcYF2Csy6rWQvqHIi9qDqwAAAYU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZcYF2Csy6rWQvqHIi9qDqwAAAYU&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://www.theregister.com/2024/02/02/on_call/
[6] https://www.theregister.com/2024/01/26/on_call/
[7] https://www.theregister.com/2024/01/12/on_call/
[8] https://www.theregister.com/2024/01/01/on_call/
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZcYF2Csy6rWQvqHIi9qDqwAAAYU&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] mailto:oncall@theregister.com
[11] https://whitepapers.theregister.com/
Re: Nine nines and an explosion
Indeed...'absolutely no downtime' and 'nothing can be switched off' do not work together. If your goal is zero downtime, then you need to have redundancy to allow for unplanned failure, which means you do have the option to switch things off or restart them as and when required.
Re: Nine nines and an explosion
I had a manager once who stated he wanted 24x7 uptime as a minimum.
We used to calculate uptime by excluding planned maintenance. Ie if we took something down for an upgrade or other planned upgrade work it didn’t reduce the uptime.
Re: Nine nines and an explosion
Presumably his thought process went as far as "what would be the best number of uptime?"
I wonder what his thought process would be when presented with "Well that will triple what we currently spend on servers due to have to create clusters and redundancy and such"
Re: Nine nines and an explosion
Presumably his thought process went as far as "what would be the best number of uptime?"
My assumption was that his thinking was "have you exceeded the minimum uptime of 24/7? Then no bonus for you!"
Uptime
About 10 years ago we had a PC sitting in the bottom on a machine control cabinet, just a bog standard computer, not an industrial one, which failed to boot one morning.
Upon investigation I discovered it hadn't been powered down for at least 5 years, possibly longer, yet for some reason the day before the person using it had selected the Shutdown option on the menu. We all know that's a recipe for problems, in this case the spinning rust was completely dead and the motherboard couldn't get past POST. No real problem, fetch a PC from the spares pile then wait a week or so for the machine manufacturer to supply us the software from the States (we couldn't download it as you required the floppy* for the security key, not sure why it was secured as the software would only work with their machine).
*I did say it was a while back!
Re: Uptime
Re: (we couldn't download it as you required the floppy* for the security key, not sure why it was secured as the software would only work with their machine).
I used to support a Uni computer lab with the same problem. We had proper, broadcast quality, capture cards in ten of the machines. These cards were essentially high end pentium based PCs on a card. They even required their own SCSI hard drive.
It was interesting how they handled drive access.. These drives appeared to the host PC as NTFS formatted drives, with several folders on the root. JPG, PNG, AVI and a few others.. They each had folders inside for individual projects. In these folders, you would find the video for each project, in the format with the folder name. So, the JPG folder would have the individual frames of the video in JPEG format, the AVI folder would have the video in AVI format. All conversion was done on the fly. It was actually a very neat system, once the user got used to it.
Anyhow, I digress. The cards were directly supported by both Adobe Premiere, and a custom version of a little known video editor, called "Speed Razor". This was a commercial application, but rare enough that when I logged a support call, I was told our ten machine lab was the largest installation in the country.
This Speed Razor required a dongle that plugged into the parallel port.. This was a massive pain in the arse, because it meant we had to install parallel port extension cables, and run them back into the machine so we could lock the dongles inside the machines. As if requiring access to a £5k video capture card wasn't enough of a restriction.
Re: Uptime
I worked with some simulation software that needed a dongle (thankfully USB, though we did have some older serial port ones) to validate licenses, but only at boot. Because the dongles went missing a lot, I once ended up spending an hour sharing a half-dozen dongles between thirty machines, booting then in batches, after a power outage took down the whole room at once. PITA, but I got everything back up and running before the important folk got into the office!
("Power outage? What about the UPS?" I hear you ask. Well, the power went out because the UPS caught fire...)
Re: Uptime
UPS caught fire? That gives a whole new meaning to "hot standby"
I only see 2 options: (though this is only a quick look)
1) Change the management view so that outage is allowed, and planned for.
2) Change jobs!
----------> Mine's the one that will stop the door hitting my ass on the way out! (hee haw! hee haw! Damn! Hit my ass anyway!)
Can't help but wonder what eventually happened.
Go back to the client and say, 'fair enough, here's the cost for a new storage array, let me check if the ESXi host has the right or enough adapters.... no it doesn't so that's a shutdown to install the HBA'
or 'let's see if there are any empty drive slots in the host... yes there are .... does it support hot swopping? Alas no, so that's a shutdown to install the disks!
You're running low on space on an ESXi host - that's not good, you absolutely have to increase the storage but your own policies are preventing you from increasing the storage! Rock please meet Hard Place!
Only in heaven....
"he was told his client had a policy not to allow outages."
I would have to assume meant "not to allow planned outages" but as the point King Canute was making (but later misrepresented) - time and tide wait for no man, nor noutage neither.
Of course the client might have been a (minor) deity or, not uncommon in this game, at least thought of themselves as such.
Re: Only in heaven....
"Of course the client might have been a (minor) deity or, not uncommon in this game, at least thought of themselves as such."
More likely the opposite, I'd have thought. Policies exist to further the business's objectives. If a policy is getting in the way of that it sounds as if the original policy maker has long gone and been replaced by a weight to keep the chair from escaping with no understanding of why the policy exists.
I've always thought that policies should be written with a rationale so they can not be over-ridden by idiots ("This is a legal requirement") or reviewed when no longer appropriate ("This was a legal requirement but the law has changed").
This must rate as the most moronic management policy ...
EVER!
On second thoughts, there may be even more moronic management policies. I shudder at the thought, but never underestimate the ability of management to create clusterfucks of epic proportions.
Maybe there should be a Most Moronic Management Policy (M 3 P) award, to be awarded annually.
Re: This must rate as the most moronic management policy ...
Many years sinceupon, the QA idiot in a company for whom I was working, decreed that all electronic components would henceforth be stored in a tistatic packaging. All the sensitive stuff - ICs semiconductors etc already were. He pointed out additional stuff so the stores folk duly sighed and complied
We had a batch of dead PCB mounted batteries not too long after that.
Karma
Working at a company hosting a reasonably significant online-only trader. Network design, servers etc all set up to be properly redundant for potentially zero application down time. Arse-covering know-nothing senior mangler refused to allow down time regardless, even for patches - didn't want to take the grief if anything actually went wrong.
Come the day when the DC had a massive power outage, taking down everything (I'd left the company by then, but still in touch with good friends made there so got the story). Once power was restored, a goodly proportion of the (Cisco) switches were basically bricked - an actual bug in the Cisco OS (whoulda thought) causing gear that hadn't been restarted for 'n' days (actually years, if I recall) to very permanently retire itself.
Never did find out what happened to the mangler, hope he got dumped on from a great height but suspect he weaselled his way out of it.
Re: Karma
That is probably the internal flash but after 3 years it would go read only as it needed an updated image to reset the bug.
Affects nxos switches and the big 4100 series firewalls that I have seen myself - the patches have been around for a few years.
Re: Karma
"he weaselled his way out of it"
Once upon a time I was a member of a small IT team (5 strong) supporting a large number of engineers on a contract.
I'm doing my normal day-to-day work when a chance comment to one of the other team members led to me finding out that one of our systems had suffered a production database deletion, and the other members of the team (and many of the engineers) were working flat out to rebuild the database ...... and the months of data lost because 'the Oracle backups had been failing due to an Oracle bug that a fix had not been published for'!
I found it so strange that I hadn't been asked to help with the recovery. In fact I hadn't even been told there was a problem.
I deduced that this was actually deliberate as I was the one person in the team that would have stood up in meetings and pointed out the plot holes in the explanation that was given as to what went wrong, and why it wasn't the fault of the mangler that ran the team. "The Emperor has no clothes!"
I have a very short list of people who I will never work with again, and that particular mangler is top of the list!
We dont go for "uptime" records
We have a handful of Microsoft SQL servers that we reboot weekly, just so they can have a micro nap and feel all refreshed, is that a good thing ? are there disadvantages ?
Are there learned things in volatile RAM that will need relearning and slow performance for instance? (just guessing)
We also recycle the SQL services periodically , which is basically same thing .
Any thoughts anyone?
Re: We dont go for "uptime" records
Doing a reboot when you can is in my mind A Good Thing. While I mostly use Linux boxes and they have less need of rebooting for patches, I have been caught out before by a boot loader patch that borked booting but in itself had no need to reboot. Only discovered when an unplanned late night reboot occurred, doh!
After that I try to reboot after significant patching even if not called for, assuming there is not any real impact from doing so.
The other "gotcha!" is application software that has been changed and fails to properly start on boot. It may have SFA to do with the OS patching, but again a planned reboot while to responsible software person is to hand is a good policy so you have a server that is kept in "automatically recovers" mode. Because unplanned reboots happen. Due to power issues, gross administrative error, system lock-ups triggering a watchdog, etc, etc.
Nine nines and an explosion
Well, when the power goes, they again get a taste of reality, eventually.
There is no point in keeping extreme uptimes when your system will eventually break down in a, maybe not literal, but spectacular explosion.