'We're changing shift, and no one can log on!' It was at this moment our hero knew server-lugging chap had screwed up
- Reference: 1589181306
- News link: https://www.theregister.co.uk/2020/05/11/who_me/
- Source link:
This week's confessional comes courtesy of "Steve", and takes us back to the 1990s and a well-known UK bank.
The office, he recalled, was an open-plan space the size of four football (or "soccer", if you will) pitches. All the business units were in it – development, support, the call centre and so on. The only people not there were in marketing.
"No idea why they were different!" remarked Steve.
On the day in question, Steve was about to take a stroll from his desk when he spotted a colleague, Aaron, ambling toward him with a pizza-box server under his arm. "It was 2pm in the afternoon and all was well," Steve recalled. "He looked like he did not have a care in the world."
Britain has no idea how close it came to ATMs flooding the streets with free money thanks to some crap code, 1970s style [2]READ MORE
A swift chat later and the two went on their way. As he passed the call centre desks, the team leader called out: "Steve, do you know what's gone wrong?"
"Erm, no, what's up?"
"We're changing shift, and no one can log on!"
"Are there people already logged in OK and still working?"
"Yes," went the anxious response. "We are keeping everyone back from the early shift until it is fixed!"
At that moment, with hideous clarity, the image of Aaron with a server under his arm popped into Steve's mind along with some choicer four-letter words.
"Leave it with me!" he said brightly, before scuttling off to find Aaron and that mystery box.
"Aaron! Where did you get that server?"
World weary, Aaron sighed: "I got it from the development comms room, why?"
After a little "Oh no you didn't ... Oh yes I did" pantomime exchange, except swearier, Steve asked the pertinent question: "Is it the test DHCP server?"
Of course it was. Aaron wasn't a total idiot.
"It's the live one, Aaron! People who are coming in on shift can't log on! Show me where you got it from; if it's the test one I'll apologise! You can tell everyone what a [expletive deleted] I am."
Never one to pass up a chance to mock a minion, Aaron agreed and the two headed to the development comms room. Alas for Aaron, where the gap should have been, there was a server, with lights all aglow.
In the live room, however, there was a gap where the production DHCP server should have sat, while it instead was held moistly under Aaron's tricep as it took a ride in his armpit.
Realising that the cock-up had prevented an entire department from coming off shift, the pair raced to retrieve the removed server, a bit of HP metal running Windows NT4. Four-letter words were flying.
"Aaron! Had you done anything to that server?" asked Steve.
"No!"
The pair retrieved cover and screws, and slammed the server back into its slot. An agonising few seconds letter and… "The lights flicker, they go green. The network light starts flashing. It boots up. DHCP service is starting back up. It starts."
After taking a few minutes to compose themselves, the pair returned to the office and Steve casually asked the call centre team leader: "Can you get someone to try and log on?"
Success.
Steve scarpered before any awkward questions could be asked. And Aaron?
"Aaron was my boss... His boss was not amused! [He] was very suddenly my boss as well.
"Poor Aaron."
We've all worked for an Aaron, but few have seen their boss nonchalantly strolling through the office, critical piece of infrastructure under their blissfully ignorant arm. Or have you? Confess all with an email to [3]Who, Me? ®
Sponsored: [4]Forrester Build a Digital Experience Portfolio
[1] https://www.theregister.co.uk/Tag/who-me
[2] https://www.theregister.co.uk/2020/05/04/who_me/
[3] mailto:whome@theregister.co.uk
[4] https://go.theregister.co.uk/tl/1936/-8554/forrester-build-a-digital-experience-portfolio?td=wptl1936
Labels people, and read them!
This feels like a course in RTFM, RTFS or RTFLable!
Re: Labels people, and read them!
And what about monitoring? Not to mention redundancy.
Re: Labels people, and read them!
Labels people, and WRITE them!
Re: Labels people, and read them!
People start getting right shirty when you start applying labels to them, physical or otherwise, these days...
Mine's the one with the iphone in it that identifies as an android that identifies as an iphone, thanks...
Re: Labels people, and read them!
People also get shirty if you start sticking labels on their shirts!
Re: Labels people, and read them!
Phone gender is not the same as phone sex! Do not be judgemental of the choices phones make.
Re: Labels people, and read them!
"Labels people, and read them!"
presumably everything in the Dev comms room was almost fair game whilst the live room had stricter change control?
I also assume the Dev & Live comms rooms where appropriately signed.
Lastly, a big organisation should have had more than 1 DHCP server & PC's typically retain their DHCP addresses, only checking for a new address at half the lease time, for that reason we used to set DHCP at 4 days so machines could still work on a Monday from Friday's lease giving time to fix any issues from a weekend fault.
While reading I suspected something differently, but that must be me being biased with prior experience of less well managed environments: a newfangled machine starting its life in development or test and then miraculously being repurposed to production - obviously without ever being re-staged/moved/re-labelled as such.
Yes, I fscked this up myself on a small scale, not worthy of such a column. And saw it from save distance on a larger scale, too.
1990s banking
Worked for a European bank in those days. No-one was allowed near the IT or comms rooms, even managers. Programmers were wizards who arrived in the dead of night to collect pizza...
Re: 1990s banking / No-one was allowed near the IT or comms rooms, even managers
Well, yes and no.. In the office I was working on the Y2K thingy the IT and comms room was properly bolted down requiring a keycard and pin for access and no one there (we were all analysts and developers) had access.
One day some higher up realized that, just to change the daily backup tape, an IT minion had to do a daily trip from the main offices across town for that 1-minute job.
I was then chosen to do the job, so got access to the forbidden room. But I guess my bosses gathered I would never even think about taking any Cisco switch for a walk...
Re: 1990s banking / No-one was allowed near the IT or comms rooms, even managers
Had a similar thing happen to me in the early 2000s. One of the teams in our building was moved out to head office, about 100 miles away (our office was basically the IT hub).
But they still had gear in the single server room we had, which had a daily backup tape that needed rotating out. I got he job for a short while.
Er ...
Not quite the same, but I worked in an environment where some control hardware was being developed alongside the monitoring software.
Failing to get a response from the test unit usually resulted in a walk down to the engineering workshop to see if the device was actually powered up and connected to the network.
Usually it was simply switched off, but sometimes you would find the safety cage open with bits of the test rig lying on a bench.
If there were one or more boffins prodding it, tutting, and shaking their heads, it was time for an extended coffee break.
It's easy to detect the Aarons of the world nowadays.
You need a server running Zabbix (my choice) or some other monitoring software, e.g. Nagios. When someone unplugs a monitored server, within a few minutes you've got alerts, XMPP messages coming through, whatever, and you can fix it before anyone notices, most of the time. It also tells you when the disk is getting full, when it's getting too hot, when the RAID is degraded, etc. etc.
Just the other day, I found a server running cryptocurrency mining software in a user account because the CPU was constantly at 85C...
Re: It's easy to detect the Aarons of the world nowadays.
"You need a server running Zabbix "
"Just the other day, I found a server running cryptocurrency mining software in a user account"
doesn't matter what monitoring software you have, there's a bunch of other stuff you need to be doing to make sure that miscreants aren't running malicious code on your systems. What other stuff have you not found?
Re: It's easy to detect the Aarons of the world nowadays.
Monitoring is a good (practically essential) start though, especially when your environment is too big for any one person to know what 'should' be going on with every device.
I've not spotted crypto-mining yet, but I've spotted servers filling their discs, which turned out to be something writing debug-level logs because the developer forgot to switch them off.
Another vote for Zabbix though. It can monitor practically anything with a network connection, and it's configurable seven ways from Sunday.
Cut the blue wire
I'd like to say that I have never pulled the wrong LAN cable due to inattention/brain fade/last night's refreshment taking toll this morning, but I would be lying.
De-racking the wrong server is quite impressive though. Chapeau!
marketing
The only people not there were in marketing.
"No idea why they were different!" remarked Steve.
The coloured pencils department are always special.
Re: marketing
The coloured pencils department are always special.
Indeed ...
It's the cage where they keep those despicable abominations of nature, the marketing droids.
O.
Had an DHCP issue this morning myself.
For some reason ClearOS's DHCP function decided to disable itself after a reboot following a power failure. Enabled it after users bleated about loss of wifi and network access....
Now all is well.
Strange how such a small thing can cause big issues...
It was outside, by the back door!
Some time ago, 1999 I think, where I worked we had what was called the Customer Data Interchange server, or CDI as it was called back then. Basically a small integration platform, running on AIX, managing incoming data from customers, mostly dial-up at the time, using UUCP and Kermit, we also had a few leased line connections with larger clients, although no Internet connections back then (they arrived in 2001).
Peak times were late afternoons during the week, but we did have a little bit of data over the weekends from some of our larger clients. One of these larger clients had tried using the service one Sunday, around lunchtime, to no avail.
I was the lucky one providing call out that weekend, we had a shared laptop and a pager, that was handed over every Tuesday to the next person on call. The pager went off that Sunday lunch time. I called the Unix Ops team, who paged me and who were in the office 24/7, and asked what's up, "Client X can't get any data through, can you have a look?". "Okay" I say.
I dialled in from home (we had a modem rack for remote terminal access), got onto our jump box, and then tried to access the CDI server, it timed out. Tried various network tools, no response to ping etc. The CDI server was an AIX box that basically just kept going, 24/7, I'd never known it once to actually crash or freeze up. So I'm thinking maybe a hardware issue, or network problem.
So I called Unix ops again..
Me: "Hi, anything going on in the DC today?"
Ops: "Yes, there was someone scheduled in this morning to decommission some old unused gear. Why?"
Me: "Are they still there, and could they go check what was actually decommissioned, specifically anything related to CDI?"
Ops: "ok, I'll call you back in a few".
There was me thinking maybe someone had pulled out one too many network cables or something.
30 minutes later, the phone rings, it ops, "Hi, did you say CDI?".
Me: "Yes why?".
Ops: "Well the guy doing the work left the building about an hour ago after finishing the work, and isn't answering his pager (turned out later it was turned off). But we got one of the security guys (the only other people on site) to go have a look, and they found a box by the back door, near the skips, with a label on it, saying 'CDI'"
Me: !!!! "Ah, that could be an issue!"
Turns out of course, the guy doing the decommission work had decommissioned one two many servers. Our CDI server was ancient, and was due for replacement the following year. Turned out everything else in that section of the DC (builtin the early 80s I believe), was being decommissioned, and he'd basically just removed the lot, including our active server.
Some panicked calls from Ops trying to get hold of someone else who could help. They did manage to get someone, who then had to travel to the DC, carry the box back in from outside (thankfully it hadn't been raining!), connect everything back up, and just hope for the best when the power button was pressed.
I got another page late afternoon, spoke to Ops, who said the server was up and running, and could I check please. I dialled-in again, had a look around, did some housekeeping, and sure enough, everything seemed to be working fine.
To the credit of whoever it was who went in that afternoon, despite not being involved in the removal, they went in, and managed to get everything hooked back up and working, and stayed on site while I checked things out. He apparently also put a big label on the front of the box to state that it was a Live service, and not to remove it without getting clearance from my team!
Needless to say, we did a lot of manual monitoring for the next few days to make sure everything was running fine, and I was also involved in a few lessons-learned meetings, which changed a few of the processes we had (or just created new ones as they didn't exit yet!). Including for example, requiring anyone doing any out of hours work at the DC, to be available on call for the next 24 hours minimum. If they couldn't do the on call cover, they weren't allowed to do the work.
This is a live system, Do not Reboot it
~ 2010 i was on the phone with cisco support regarding a 6512r we had recently installed and running in a contact centre. It was up and taking calls but we had some issue iirc to do with ACL's consuming cpu instead of running in the ASIC. sh tech and logs back and forth to Cisco then a lot of webex's. The cisco engineer suggested we up the code to a recent (released after we had issues) version and we planned to install it on the redundant supervisor. at the start of the call i stated the switch was live taking customer revenue generating calls, during our troubleshooting i reiterated the same, before we started the upgrade on the redundant supervisor i repeated the same & that we will do the switch over out of hours, once the upgrade was done i repeated the same & i'd do the switch over over night, she then flipped the supervisors causing both supervisors to boot and all calls to drop & phones to power off (PoE), I was actually on site, a contact centre going quiet is just as eerie as a server room going quiet. Of course i lost the webex when the switch went to. Luckily everything came up ok, still had the original issue despite new code. I sent some really snotty emails to cisco that day!! For some reason i didn't get into trouble for that one!
The issue was too many operands used across the various ACL's causing the cpu to have to process instead of the ASIC's.
The only people not there were in marketing.
"No idea why they were different!" remarked Steve.
They were probably the cause of most of the problems other in the same room were trying to fix. It was probably a civil order/anti lynching move.
"Leave it with me!" he said brightly,
Rookie mistake.
Open plan offices mean more room to sprint in the general direction of "away"
Not so much a "Who, me?" as a "Who, you?"
See title