Oh sure, we'll just make a tiny little change in every source file without letting anyone know. What could go wrong?
- Reference: 1594970108
- News link: https://www.theregister.co.uk/2020/07/17/on_call/
- Source link:
Today's story comes from a reader we will refer to as "Brad" in order to spare the blushes of those with whom he worked, and dates back to the early part of the century.
Brad was a Unix/Linux administrator at a government agency. "I had a pair of Solaris servers running a source-code repository in a fail-over cluster," he told us.
In what might bring a tear to the eye of those now dependant on the vagaries of GitHub and its ilk, "that was the easy part," he said. "They just ran."
The reluctant log trawler: The buck stops with the back-end [2]READ MORE
The agency was a huge NetWare shop back in the day and, as was so often the case, used GroupWise as its email platform of choice.
GroupWise, for those now unfamiliar with the granddaddy of collaboration platforms, was WordPerfect's take on email, calendaring and scheduling before the corporation was snapped up by [3]Novell in the early '90s and WordPerfect Office slapped with the GroupWise moniker.
WordPerfect would be sold on again, but Novell opted to keep GroupWise for itself.
It usually did not cause headaches for Brad: "Normally, GroupWise ran smoothly, even on single-core processors."
However, on the day in question, things were not going smoothly and Brad found himself receiving calls for help, or at least calls for explanation. One of his servers was overloading the email system and could he please deal with it? The supervisor of the network team ("whom I seldom saw") had even become involved.
Armed with the offending server's name, Brad hurriedly logged on to find out what was happening, but drew a blank as to why it was spewing email like a teen discovering cider for the first time.
"I logged into the server and found it was running at about 50 per cent capacity (hey, it was SPARC hardware, not that crappy Intel stuff NetWare ran on. And I'm pretty sure the SPARC server was dual core)."
Sure enough, sendmail was chewing through prodigious amounts of CPU.
Brad ambled over to the team responsible for the application that ran on the server. It transpired that the gang was performing an upgrade, which had required a tweak to every source file, each of which had to be checked in.
Those managing modern-day pipelines might blanch at this point, as Brad went on: "The application was set up to notify the original poster, and often his/her team, that the artifact was being modified. So for every check-in, there were three to four emails being generated.
"This repository had been in use for years. The hardware itself was probably three to four years old. And we had something like five applications that had their source code stored in the repositories.
"So, a lot of emails."
It transpired that someone had forgotten to turn off notifications before kicking things off. "They would see if they could turn it off in the middle of the process," added Brad.
A simple "We're upgrading!" would have sufficed rather than the GroupWise-choking tsunami that had been unleashed.
Brad stalked back to his desk and killed sendmail with extreme prejudice. "In about five minutes, the GroupWise admins notified me that their systems had returned to normal."
He neither knew, nor (from what we can tell) really cared if the admins of the application got their act together. He did, however, find 3.5GB of mail waiting in the queue directory.
rm * was Brad's friend that day and, after a subsequent restart, sendmail was blessedly silent.
Ever got The Call only to discover that somebody else was doing something silly? Or were you the cause of that cry for help? Share your story with an email to the vultures staffing the [4]On Call desk. ®
Get our [5]Tech Resources
[1] https://www.theregister.com/Tag/on-call
[2] https://www.theregister.com/2020/07/10/on_call/
[3] https://www.theregister.com/2009/09/25/novell_support_changes/
[4] mailto:oncall@theregister.com
[5] https://whitepapers.theregister.com/
Too soon to tell I got a good one this week
Wish I could retell my tail of epic face palm, with no specifics massive escalation by a customer, threatening to pull contract see us in court etc. With an equally rapid climbdown once root cause of their admin didn't rtfm, grok the permitted request rates or understand what an amplification attack was, was identified.
Mail Storm
Back in the day I was responsible (with others) for maintaining an Exchange 2003 clusterfuck. We were a Microsoft shop and made use of AD to control access to applications and shares and other security stuff. One of the applications we used had some AD permissions groups to control access to the app.
One morning I arrived at my desk, slurping my coffee I noticed I had email, lots of email. Me and over a thousand other peeps. A user had discovered he could send mail to all the apps AD groups. Something only the app admins were supposed to able do.
The original mail was a rant and a thousand users responded to a thousand users. Some saying don't mail me this crap, others ranting back. Because this happened out of hours we didn't see the storm growing until it was so bad Exchange slowed to a crawl, right in the middle of the day. By now there was nothing we could do except lock out the AD group's mail privileges and wait it out.
Best bit was the main Exchange admin sending a stroppy mail to the whole app mailing list saying "stop using sending mail to this group" before locking the groups down. Yup mailstorm 2, the return.
Anonymous to protect the guilty.....
Re: Mail Storm
Have a pint for using the correct exchange cluster terminology
Re: Mail Storm
Similar position with Exchange 2003. The entire email system went slow around midday on Friday. Big backlog of email that cleared after a little while. The next week it ground to a complete halt with massive mail queues....
The eventual root cause was found to be a new HR Inclusivity email - the first was a text only introduction, the second contained several rather large images making the email just under the 10MB limit.
Re: Mail Storm
We had a mail tsunami on EX2003 due to a virus. While the PHB and systems manager did the headless chicken thing on MSN and our well paid support* I wrote a VB app that could open each mail in the DB (bit of a posh name for it really) and scan the mail for signs of infection. I got the OK to unleash it on the DB where it deleted 99% of the emails in it and things settled down once a few other patches were in place. It was a useful bit of code I could use to build a new "DB' when it got corrupted as it occasionally did. And also looking at peoples emails while bored on a long weekend waiting for some upgrade to finish.
*We paid a company for support which then never seemed to manage before we fixed things ourselves - largely because when MS fucked up they fucked up all their customers at the same time. I used to enjoy getting phone calls in the early hours where they would tell me the steps to fix whatever upgrade had damaged and almost chant along with them and then give them further tips and trix. They were never grateful!
Vendor Support
We had similar issues with Oracles Service Desk. They would cycle through new recruits until they knew enough to deliver 'consultancy' this meant my DBA's had to suffer an hour of list ticking while they went through their script before they could admit they didn't have a clue and pass us on to a 2nd line engineer.
I'd put up with this for something minor but used to have to call my account manager to escalate anything serious to second line as soon as we'd been given a ticket number. We were pretty leading edge and working on a joint development at the time so were usually on beta or just released versions for the tools.
It was exciting at the time but developing business critical apps using new Oracle products is not good for your blood pressure.
Re: Mail Storm
Haven't seen one of those in our shop for years... but we did have a good one many years ago. 30,000+ messages(each) later, we figured out it was due to most of the staff being on holiday with an out of office responder, and belonging to many listserv... who replied when they received a group message, who got an out of office reply, rinse, repeat ad nauseam.
Re: Mail Storm
I've often thought that if senior management in a large company wanted to make some cutbacks, they should send out an email to everyone "by accident" which is clearly not intended for everyone. Then anyone who replies to all with "Please remove me from this mailing list" goes on the potential cutback list, as they have demonstrated a lack of sense and consequential thinking.
I loved Groupwise
Loved the collaboration bit, was easy to manage comms for various projects in groupwise, all project correspondence was sent within the project group keeping everything neat and tidy.
i hate gmail for work, somehow its more of a mess than normal free gmail, hides my email for me & meet always starts with camera enabled with no option to default to camera off.
Re: I loved Groupwise
Dont use gmail for anything but forwarding to a better solution ;)
I fondly recall just a few years ago #ReutersReplyAllGate - the emails just kept on coming! Baffled why people kept responding to them!
https://blogs.wsj.com/cmo/2015/08/26/reuters-employees-bombarded-with-reply-all-email-catastrophe/
Because "Baaaaaa". Obviously.
Well, that's what ever instance of "Why did you send this to me?" I ever saw read as to me.
Not quite on the same scale but...
We use SharePoint (yes, I know but I inherited it and we're tied to being an MS site anyway). I had one user a few years back who set up email alerts on a Library for All Changes and Send Immediately (contrary to advice otherwise).
He then bulk uploaded a few thousand files (it was the ISO 9000 Library) and wondered why his email on the laptop and Blackberry (I said it was a few years ago) were going mad.
I maintained a system on Netware that processed applications for government assistance, which eventually lead to Giro cheques being printed and posted to the applicants. These were on tractor-feed stock and were printed on an impact printer to prevent the cheques being altered by scraping off laser toner, which does not penetrate the paper.
After building up a backlog of cases, the management decided to have a blitz and instituted lots of overtime and brought in temporary staff. They then hooked up a box of paper and ran the print job. Unfortunately, the size of the job was larger than the disk space allocated to the print queues. The cheques started printing, but after a while the system crashed and restarted. This started the the print job from the start, not from where it left off, and led to hundreds of duplicates being printed. The temporary staff in the mail room were not aware of this and just separated the cheques, plonking them in window envelopes, and put them in the postal system.
By the time it was noticed, they were too late to recall them. I got a call in Liverpool and had to get a hire car and drive down to Birmingham with my foot hard on the pedal - I even got pulled by the boys in blue on the motorway, but was luckily let off with a warning after being breathalysed.
It cost them thousands of pounds in the end as people cashed in their windfalls.
Cheque runs
Shudder!
I had the joy some years ago of being involved in setting up a new cheque printing system where the monthly cheque data would be auto-generated by an overnight routine on the AS400, then someone in the finance department had to fetch the cheque stock (separate A4 sheets, pre-printed with serial numbers) from a locked store, load it into the cheque printer (which used magnetic ink for the automated cheque readers in the bank and had locked paper drawers) *the right way round and up*, then key in the first cheque serial number to the AS400 which would then store the SNs against the cheque transactions on the system and dump the data to a shared directory, which was auto-monitored by a separate system running some flavour of Crystal Reports to take that data and transform it into a cheque page layout (address and transaction breakdown at top of page, physical cheque perforation-attached at the bottom).
Depending who was doing the cheque run they either got exactly the right number of cheques out, or grabbed more than enough and returned the extra to the store after the run. But my
Mind you, if you think normal ink / toner is expensive, keep away from the magnetic stuff!
Icon, 'cos that's what the end of the day with a cheque run really required.
Re: Cheque runs
"Mind you, if you think normal ink / toner is expensive, keep away from the magnetic stuff!"
Oh, I dunno ... I bought a 1 kilo can of magnetic ink from Valley Litho for about 80 bucks a couple months ago. That's on par with other offset inks.
Re: Cheque runs
If you were doing fairly low numbers of non-A4 prints we found having special feed trays with customised idiot sheets glued to the bottom so in theory they couldn't get the cheques or whatever the wrong way round.
It was worth a try.
Most Dangerous
The most dangerous button ever created "Reply to all" not helped by idiots who use the shotgun approach to sending out emails.
Re: Most Dangerous
Indeed. That thing should have had an admin password lock on it from the start.
Since Office 2010, I've worked with a few companies who actually removed it from the ribbon. You could, of course, plop it back in, but woe to the guy who tries that. I know of one who got hauled right up to the CEO's office. I don't think he tried that again.
Re: Most Dangerous
Exchange makes it easy to do this and difficult to prevent abuse.
"spewing email like a teen discovering cider for the first time."
:-)
rm * was Brad's friend that day
That sounds like the old disaster movie plot device...a closing line that just sets everything up for a sequel
that just sets everything up for a sequel
Yeah. I would probably have written a script intended to only delete the relevant (useless) messages rather than the whole queue. Of course, I say "intended", ... who knows what it might have actually done instead :-) [1]
.
[1] i.e. Exactly what I coded it to do.
Amarillo.....
Is this the way to crash a network?
Dial up modems really don't like it
Stop resending that bloody video
I will surely kill you all!
I was a poor sysadmin at an office where everyone received and sent that video (the one from Iraq)... our connection to [redacted] was a lowly modem so, try as we might, we couldn't receive/send email for a while...
Its not always an automatic SNAFU that borks one's coffee breaks.
Re: Amarillo.....
Reminds me of the time a former employer decided to install a proxy server to their single 56K modem.
Within days me and the rest of the IT team found out about so many dodgy sites. Still, at least we found that the block function on the proxy server worked really well, so that was a bonus.
Early 2000's
We had an exchange server that pulled from the Cluster in head office. One day I was asked to install some updates, and shove some more RAM in the box, while they did the same to the cluster. I warned my site that we'd be without email for an hour or so, and that if they wanted to stop seeing errors, just close Outlook. SO the hour came, everyone closed Outlook, and I installed all the updates on the box, then powered down to install the memory. Once done, I turned it back on, and carried on with my normal work.
About 6 hours later, someone came in to my office and asked if it was ok to use email yet, as they needed to send something out to customers. It occurred to me that I hadn't heard from head office, so I rang one of them. Who told me "Yes, we sent you an email to say it was working again"
If this were the DailyWTF...
... the story would have ended with Brad getting fired for daring to simply delete the emails instead of going through each and every one. It so happened that that the CEO's extremely important golf invite was among them!
A harmless case
I started at a company apparently exactly on the day when someone had managed to put the "All
Their IT was very good, they managed to stop it after just four rounds.
Bugzilla - emailzilla
A few years back I used to dread the Friday morning bug triage session, as Bugzilla would generate a constant flow of emails. But at least deleting them all was something which could be done after lunchtime at the pub.
Re: Bugzilla - emailzilla
Yeah, our Kanboard does that as well - serves me right for enabling the notifications. But I need those, when working remotely I cannot access the board, due to... (yeah, we all know those stories, I don't want to think about that now). So I check progress by looking at the incoming mails - and I can even add tasks by sending an email to the board :) (still cannot modify "my" tasks, but some colleague has to push them my way by putting me in as the person in charge... f'ing annoying, but works)
Except our meeting is Wednesdays...
icon ->>
/me needs that. Now. /me been fighting with sharepointless *shudder* for some time.
Not quite a 'reply all'...
once had an urgent 'please delete email subject xxxxx sent dd/mm/yy hh:mm UNREAD' that was sent from on high after someone had managed to send a sensitive attachment to a large part of the company.
The 'recall message' had already done the work for my email account, so I never got to see it, but I heard from others that the attachment detailed the CEO's pay and allowances
(the CEO has since moved on but I assume there are still a few copies kept by people for potential blackmail)
Re: Not quite a 'reply all'...
Nothing will get that email read more quickly...
Don't forget the NHS!
We did it too. It was the first and only time my job was mentioned on Have I Got News For You.
Ratelimiting
So many similar stories that I can't remember them all.
I did implement rate limiting with Exim so that a bulk spewer would end up with messages being 'frozen' in the queue and not delivered; as a result 'exiwhat -zif root@someplace | xargs exim -Mrm' is hardwired into my fingers. I seem to recall a common issue was web→mail forms being spammed by scanners.