Typical '80s IT: Good idea leads to additional duties, without extra training or pay, and a nuked payroll system
- Reference: 1600075730
- News link: https://www.theregister.co.uk/2020/09/14/who_me/
- Source link:
Today's tale comes from a reader Regomised as "Tom" and concerns events that took place back in the early '80s, when The Empire Strikes Back still ruled the roost and Raiders of The Lost Ark was creeping into cinemas.
Tom had just left university and was working for a manufacturing company in England. The lucky boy initially spent his time knee-deep in ICL Cobol, adding columns to reports and updating the name and phone number of accounts receivable staff to ensure they printed correctly. Thrilling stuff and, naturally, everything was hard-coded.
"My first career high point," he told us, "was after being asked to change the name and phone number again, I suggested that the hard-coded name and phone number be taken out of the program and the same information be put on a parameter card."
"This," he said modestly, "went well."
"For the benefit of the excessively young," he added, "a parameter card was a virtual 80-character punch card that was stored in the JCL (Job Control Language) before the program call. Hey presto, no more code changes when people change!"
Competitive techies almost bring distributed disaster upon themselves – and they didn't even find any aliens [2]READ MORE
In IT, no good deed goes unpunished and sure enough, Tom soon found himself with additional duties ("I don't remember any training or pay increase") and was tasked with supporting the company's payroll system after his predecessor departed for pastures new and likely more profitable.
Time passed, and Tom received a new version of the payroll system to install. It arrived in the form of magnetic tape containing new source code. "The code was organised into libraries so it was all very modern," he told us.
Installation was simple. Copy the contents to the source code library, compile it, run some tests, and Bob's your uncle.
The test failed.
What had actually happened was that the process to copy the tape to the library had skipped files of the same name, meaning Tom now had an unholy mix of old and new code. "Oh joy!" he said.
It wasn't a problem. There were weekly backups. All Tom had to do was restore the borked files and all would be well.
That failed too. "My day was not improving," he recalled.
In a masterstroke of connected thinking, the old tape backup units had just been replaced by more capable units. More capable in every way except, er, being able to read tapes written by the previous devices. The operations team hadn't thought to check.
"We were in a real hole," said Tom. He had single-handedly broken the payroll system. The only solution was to pay staff the same as the week before. "This was all very embarrassing for the IT department and me in particular."
By the time the following week rolled around, Tom had updated the copy utility, recompiled the new code, run the test and hurrah! A working payroll system once again!
All the same, he, his boss, and a bigwig were sent to the main factory to put on a show of solidarity in front of a workforce suspicious of what payroll-related wheeze IT might have in store for them that week. Fortunately, there were no problems (at least "no problems attributable to the IT department…")
It was asked why Tom did not take a fresh backup before the fateful copy. "I had to admit that if I had, this problem would not have happened." However, remember what we said about training? It was also clear that he was a little too junior to be undertaking such workforce-upsetting tasks.
Tom therefore escaped with a mild slap on the wrist and the bigwig moved on. "I do suspect that his meeting with the operations manager was more fractious," he understated once again.
Tom is now retired, a long career in IT behind him but, thanks to this trauma, we imagine that he is still the most hygienic of souls when it comes to backing stuff up.
Is there anyone who has never had the bowel-loosening sensation that arises from the discovery that a backup was not all you hoped it might be? Or is your data protection halo shining bright? Share your story with an email to [3]Who, Me? ®
Get our [4]Tech Resources
[1] https://www.theregister.com/Tag/who-me
[2] https://www.theregister.com/2020/09/07/who_me/
[3] mailto:whome@theregister.com
[4] https://whitepapers.theregister.com/
Wotsits' law
I carefully wrote down the name of the chap who gave his name to the law which states 'No backup is perfect' in case I should need it. Needless to say I can't find it.
Re: Wotsits' law
Also appropriate, from The Seventy Maxims of Maximally Effective Mercenaries:
Maxim 41. “Do you have a backup?” means “I can’t fix this.”
https://schlockmercenary.fandom.com/wiki/The_Seventy_Maxims_of_Maximally_Effective_Mercenaries
Shadowing - What could go wrong
Working in the mid 1990's with a customer who was running a retail operation on amongst other things a VAX VMS Cluster (Best we do not mention their Ultrix DECsystem monstrosity and AS/400 bastion of how much can I spend with IBM). The VAX was two VAX 3000 series with DSSI disks in a RAID configuration 0+1, or in DEC words shadowed. The time had come to upgrade the VMS to the next major release and the plan we submitted was to have a plan (Which this customer was always adverse to as they thought this just increased the costs), take backups, work the upgrade, test and then release back to the customer.
But the customer had other ideas, that had a shadow, they could split it and hey presto a ready made backup .....- Customer :: Why don't you take down one half of the cluster, break the shadowing (RAID 0+1), upgrade and then switch over the systems, then bring the old VAX back up with a merge using the disk running the new version of VMS.
Now, as a technical challenge this is interesting, but has many opportunities for something to go wrong, either process based (i.e. getting commands wrong etc when breaking the shadow disks etc), but the customer would only take this approach.
Initially everything went well, the upgrade applied to the VAX not now in the cluster, old VMS VAX shutdown and then brought back into the new cluster running new version of VMS - There is a moment when you create a new shadow disk in VMS that was always a bit of a fingers crossed moment when it did the initial creation of the shadow pair, 99 times out of 100 no issues but of course this time an issue, there was some issue on the shadow disk creation and we had a fail, but also took down both disks, disks themselves OK, but some gremlin caused a software fail. The issue is that we only have the regular backup for restore, so all the changes in the system prior to upgrade where effectively lost and we had lost the upgrade and all associated effort.
We have covered ourselves with respect to the customer in that we had formally written that this was not an approach we recommended as we could not guarantee the data or upgrade - But as ever lots of noise from customer that this was our fault as it should have worked blah blah blah.
But lesson was learnt for 1-2 years that nothing beats have a project plan that also ensures you take backups before committing to a change and that RAID 0+1 (VMS Shadowing) is great but should not be used as a short cut for VMS upgrades in the future.
Re: Shadowing - What could go wrong
If the customer was so sure it should have worked, the customer could have done the himself.
There are times when you should say "no" to a contract, and when a customer is clearly cutting corners and not doing things properly is one of those times because, contract clauses be damned, it will always be your fault if something goes wrong.
And something always does.
whoops - wrong disk
40 years ago I joined fresh from University, and was in the "build group" where we took the developers source, compiled it and made it available to test.
(Part of my job was to take the listings, and file them in the shelves - 6 ft high - 4 ft wide!) This was a one disk DOS/VS operation system. We had the live system disk, and the build disk. The theory being you build into the build disk, and switch with the live disk.
My second week - I followed the instructions but managed to build into the live system. This meant that test were without a system for about 3 days. My boss protected me, and said dont do it again. They updated the instructions to make it fool proof.
The next week, I event more carefully followed the instructions, and managed to do exactly the same thing, so test were without a system for another 3 days.
At the incident review, the team said the wording could have been read two ways - and mine was the wrong way!
I got moved to a different project where I could do less damage. The grad who replaced me did the same as I did about a month later. This time they changed the process so the production disk was read only.
Re: whoops - wrong disk
This time they changed the process so the production disk was read only.
The good old days when every disk drive had a write-protect switch, and pushing it was always the first step in any backup operation.
Re: whoops - wrong disk
"whoops - wrong disk"
I read that title and flinched.
Re: whoops - wrong disk
Omitting some details to protect the innocent and guilty alike...
A facility I once worked at had a large machine that used a great deal of steam. There was a PM done by my department that involved disconnecting a controller (#1) to verify that the backup controller (#2) took over, then reconnecting #1 and disconnecting #2, to verify it switched back properly. New worker followed the instructions to the letter, and the switchovers happened flawlessly, as usual. Worker then finished his paperwork and walked away (as expected).
Shortly thereafter, the operator heard a peculiar noise, recognized it, and got everyone clear of the machine - which promptly dumped a huge amount of steam right where someone had been standing. Thanks to the experienced operator, no one was hurt.
Turns out the controllers forget their setpoint when they get disconnected, and so when they are reconnected they default to zero, definitely the wrong value here. Previous workers had always reset the setpoint, but the new guy didn't know to do so, as it wasn't in the procedure.
Good boss, though - as it wasn't documented, the new guy was held totally faultless. (And the procedure was updated.)
Oh good grief....
...still nightmares about that phase when DAT tapes could only be read by exactly the same unit that created them. It made it so difficult to call them useful for DR, apart from one or two rather more short-term and hard-discovered events. Maybe that was true of all DAT units, I don't know, as DAT was not the way (aha aha) I liked it (aha aha)
Re: Oh good grief....
Thanks, that's now stuck on loop in my head for the rest of the day :)
normal SOP
Add more tasks and responsibilities and see if the techie starts to crack up... If not, add more and promote him?
Re: normal SOP
Never promote the achiever
You promote people out of problems
"the discovery that a backup was not all you hoped it might be"
Oh my, that brings back memories. I was working as a consultant in a major insurance company, on an on-call basis.
One day, I get called to go modify something in the mail template that defines what every mailbox is supposed to look like and what features it is supposed to have. So I pack my laptop and off I go. When I'm settled at my desk and after the meeting with the IT manager, with all the technical details I need in mind, I log into my local account and ask the system to start up the Designer on the mail template.
The Designer was a no show.
Not that the Designer had a problem, it was the template that was not accessible. It's design had been locked.
After a brief but intense moment of WTF! and deep soul-searching, I reassured myself that I would never have been stupid enough to lock the design of the most important template the customer had, so I went back to the IT manager and reported the problem. His matter-of-fact reply was simple : get the backup copy.
Like every responsible IT shop in any major company, backups were made incremental every day, full every week-end and end-of-month. So finding a good backup should be simple, right ?
Well, in a word, no. I basically spent a day with the systems team, going back every further in time to try and find a copy that hadn't been locked. When we had gone over the two months of backup that were stored locally, my new friend turned to me and said "Okay, this is all I've got here. Do you want me to go to the archives and fish out the storage tapes of the previous months ?". I could clearly see that that was not a prospect that he particularly relished, and I had already spent too much time on this issue, so I declined with thanks and left him relieved to be able to finally take care of his normal duties.
But I still had a problem : I had a template to rebuild. Or find a copy of, somewhere.
I will spare you the details, but let me just say that I finally did find a valid copy of the template on a server which, ironically, it never should have been put. I was able to make the requested changes, copy the unlocked template to production servers, and keep a local copy in my local profile - just in case someone else got the same stupid idea.
And that's how a 30-minute job was invoiced 10 hours and paid in full without any discussion.
Nobody ever told me who had locked the design.
Re: "the discovery that a backup was not all you hoped it might be"
"Not that the Designer had a problem, it was the template that was not accessible. It's design had been locked"
Are those templates not just some config file?
Could you not have duplicated or copied the locked design into a new unlocked design, edited that and then, once the customer was happy, lock that new design?
did you ever go back and do any further amendments or do you think they got someone else in?
We make backups ...
As the saying goes: "Oh yes, we make backups. Restore ? Nobody said anything about restoring, we just make backups ...."
Fast forward...
Tom's now making over 150k fixing COBOL code, because nobody else remembers how to.
...at least, so I've heard. Apparently, there's a lot of COBOL still around, in the dark recesses of banks, insurance companies, and probably, payroll suppliers.
Ah Arcserve
Always test backups, as our ArcServe software said it had performed backups correctly but when needed after a server went ill, would fall over in a heap with anything older than 3 days working. Ended up rebooting the ill server when it blue screened and just hoping I could copy everything across and restore anything found corrupt from the older backup.
Turned out to be rust on the IBM Servers Backplane, only affected that one server to.
Typical '80s IT: Good idea leads to additional duties, without extra training or pay
'80s??? It happens *all the time*, at least here (a university in South Forestisburningstan). OK, in some aspects we're still in the 80's, but even so...
Our IT department is outsourced and they have a very strict contract -- basic maintenance, install/configure stuff, can do departments' and events' web pages as long as they are in PHP.
Any request besides those is returned with a "not in our contract" message. So whoever suggests anything new IT-wise gets a pat in their backs and is put in charge of "ways to implement it". Without funding, of course!
You can ask for special exceptions on the outsourced IT contract, but the process is so byzantine that it is much easier to do it yourself or just give up.