News: 1595574249

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

What evil lurks within the data centre, and why is it DDoS-ing the ever-loving pants off us?

(2020/07/24)


On Call Welcome to another in The Register's series of stories from those receiving calls for help and slightly passive-aggressive helpdesk tickets. Start your Friday with a helping of [1]On Call .

Today's tale comes from "Jon", who found himself having to deal with the fallout from an ill-considered corporate emission.

In the middle of the last decade, Jon was working for a large US firm specialising in the software and hardware used in the media world. Customers included film makers, television broadcasters, 24x7 rolling news outfits as well as recording artists and live performances.

Jon was the IT manager of the whole show and responsible for all the gizmos from the CEO's mobile phone to the receptionist's desktop and everything in between. "My primary concern," he told us, "was what sat in the data centre."

The day of what Jon delicately called "The Event" began with a P1-level helpdesk ticket arriving: the company website was down and an immediate response was needed. A conference bridge was set up, and "the usual suspects assembled from different corners of the world."

It was an odd one, looking for all the world like Distributed Denial of Service (DDoS) attack: "The web servers themselves were up but the web server engine was getting mullered by thousands of requests per second. Initial investigations showed that the requests were coming from all over the world; there was no single IP, or range, or country or entity. It was coming in from everywhere ."

Oh sure, we'll just make a tiny little change in every source file without letting anyone know. What could go wrong? [2]READ MORE

But it couldn't be a DDoS attack. The company paid handsomely for a protection service, but it seemed to be standing idly by as the web server engine enjoyed an impromptu Hammer Time.

"It was as though the requests were being whitelisted by the DDoS protection service and it wasn't doing anything at all."

The question was how all those requests from all over the world could have been whitelisted. The gang looked deeper and discovered the awful truth in the form of a particular User Agent string.

Jon revealed a little about the inner workings of that company's products: "When the firm's software is sold and deployed by the client, it comes with an application manager, a bastardised version of Chromium."

This evil creation squatted on the user's computer and made regular calls back to the mothership to check the licence, look for updates, the usual sort of thing.

"The User Agent string was from the application manager," admitted Jon, "which was perhaps inconveniently whitelisted by the anti-DDoS service."

But why would the company's website be hammered by the application manager now?

It was, of course, the fault of the development team who had pushed out an update, "although the developers refused to admit it at the time," Jon added. This was normal practise, but this particular emission included a tweak to the application manager that changed the call-home frequency from once every four hours to once every four minutes.

To make matters worse, if the initial request failed the software would immediately retry. And keep retrying.

"As the update went out worldwide," Jon explained, "the new version started calling home waaaay too often, the number of requests mushroomed and took down the company website.

"The firm had DDoS'd itself. :facepalm:"

Fixing the problem was not straightforward. An update to dial back the frequency could not be pushed out because, er, the existing application manager was too busy performing a highly effective DDoS to spot and download the new code.

"Rock, meet hard place," sighed Jon.

In the end, "I spent a good few hours spinning up new VMs to have all traffic route through a layer of Linux servers running HAProxy." He was able to carve out the application manager traffic and therefore allow a small percentage of requests to succeed and the fixed code gradually rolled out.

"This is one of the reasons I now deploy HAProxy in front of all web traffic regardless of the application, company or volume..."

On Call now includes those oh-so-urgent helpdesk tickets. Ever been on the receiving end of a particularly grim example of the breed, or lost hours dealing with somebody else's self-inflicted catastrophe? You have? Share it with an email to [3]On Call . ®

Get our [4]Tech Resources



[1] https://www.theregister.com/Tag/on-call

[2] https://www.theregister.com/2020/07/17/on_call/

[3] mailto:oncall@theregister.com

[4] https://whitepapers.theregister.com/

TomTom Updates

GlenP

Not in house but...

Not long in my current IT Manager role, and with an inherited 2MB leased line to the 'net we had similar issues with our entire connectivity grinding to a standstill. This included the RD connections for two factories so was a critical failure.

Frantic sorting through traffic statistics and logs from the main Cisco device then checking round the building traced it to a TomTom device update. I figured out roughly what was happening, the SatNav would request an update packet from the servers at Akamai; if it didn't receive the response in a certain time it would assume packet loss and resend the request. The update server would happily service both packet requests though and with a slow connection neither would be received in time so the device would send another request for the same packet, slowing things down more; repeat, rinse and recycle ad infinitum. Once the device was disconnected and the multiple update packets had worked through the system everything returned to normal. I was able to then justify adding a backup ADSL line to use for any updating.

The same user had a habit of leaving the Sky home page running on his computer, the video playing constantly also impacted on our traffic levels until I told him to stop. More recently we've had an issue with the standard IE* home page hitting the CPUs on the Parallels servers but a GPO change to force the homepages to an internal site cured that.

*We have to use it for legacy apps.

Re: TomTom Updates

Hubert Cumberdale

Ah, the days when things were written specifically for IE. I remember having to manually tinker with pages to get them to work properly in it. Isn't it great now that every browser behaves in exactly the same way and conforms exactly to the same standards. Isn't it? Okay, well, maybe we're not quite there yet. But it's getting better.

Re: TomTom Updates

Ogi

> Ah, the days when things were written specifically for IE. I remember having to manually tinker with pages to get them to work properly in it.

Yeah, we are now in the days when things are written specifically for Google Chrome.

Unfortunately if Chrome decides to do something unilaterally, and others don't follow, you just get breakage to which peoples most unhelpful responses are "Use Chrome". It is like we are going back to the days of "IE Only" websites, although a bit better because Chrome is at least cross platform, although more spyware infused then I remember IE being.

Re: TomTom Updates

Hubert Cumberdale

I continue to be surprised at the [1]relatively low market share of Firefox – it's not perfect, but I much prefer it to Chrome (not least because of the spyware aspects you mention). But yes, I still have to keep Chrome around for those dumbass sites that don't work in FF.

[1] https://gs.statcounter.com/browser-market-share/all/united-kingdom

Re: TomTom Updates

logicalextreme

Chrome's stranglehold basically put the final nail in the coffin for Presto and Opera. The whole "but websites look funny in Opera" thing back in the day was more often than not down to the websites not being coded to the standards, which Presto was pretty rigorous about and thus aced the Acid3 test. Chrome basically went and threw all of that out of the window by allowing any old shit to display nicely.

I toyed with Firefox again for a while but adopted Vivaldi immediately when it landed, Chromium and all. I just want the thing that works these days.

I haven't chuckled so loud in ages

JassMan

I feel guilty about the schadenfreude but I really enjoyed this one. I think that reading the Reg should be part of the employment contract for everyone who works in the industry, so that there would be less of these clangers. No wait, I want more because more laughter makes you live longer. Oh, I don't know or care as long as all the good stories appear here.

Re: I haven't chuckled so loud in ages

logicalextreme

I'd rather it wasn't in the employment contract because non-readership of El Reg or The Daily WTF have been excellent signifiers to me that I'm probably not dealing with somebody that's competent and/or has a sense of humour.

Makes sense both ways

The commentard formerly known as Mister_C

The way I misread it made more sense

... "although the developers refused to admit it at the time," Jon added. This was normal practise, but ...

Re: Makes sense both ways

dak

Did they admit it ever, or was the awareness their culpability beaten into them with a clue stick?

As is usual practice.

Update (mis)-scheduling

JeffB

I used to work on a helpdesk at one of the large outsourcing companies, there were a number of teams ranged across a large open plan office, each team servicing a number of clients. There were a number fo sites across the country, so their internal IT support was centrally managed (in India). Microsoft roll out their major patches on a Tuesday (good ol' patch Tuesday...) and there's a reason for this that was obviously slightly lost on out internal IT support.

They were in the habit of downloading the patch Tuesday updates, running their own tests, then uploading them to WSUS and timed to go out over the weekend, all sounds fine and dandy. However, in the UK we were taught to turn our computers off over the weekend to save power, so we all came in on Monday morning, the busiest day of the week, and turned our computers on. Needless to say, within half an hour, when the phone lines were red hot, you could hear a wave of curses spreading across the floor as people's computers started rebooting to install the updates, putting the entire helpdesk operation into meltdown.

After about a year of this mayhem the Indian boys finally got a rap across the knuckles and re-timed the WSUS release for 2pm UK time on a Friday

Because that's beer o clock, innit??

Re: Update (mis)-scheduling

Anonymous Custard

Similar experience here (this time from the userland perspective) back in the day where our lot used to schedule a full monthly anti-virus scan in the same way.

Problem was, that due to the *ahem* spec of our laptops and how things were configured, the AV software basically grabbed as much CPU capacity as possible and everything ground to a halt whilst it trawled through the hard discs of each machine poking its nose in everywhere.

Most people had reasonably large drives (for the time, we're going back a bit here) and all were spinning rust, so it wasn't unusual for the scans to take all morning, and on some where the disc was majorly full the whole day. And whilst it was going on, basically all you had to work with was the phone and the coffee mug...

Re: Update (mis)-scheduling

Antonius_Prime

"And whilst it was going on, basically all you had to work with was the phone and the coffee mug..."

If that occurred more than twice to me, my phone would... [checks BOFH excuse server] ...suffer an Unreportable Transmission Override Warning... and be unusable for the day...

Modern.Problems.Dave.Chappelle.gif

SMTP ddos

Anonymous Coward

A long time before, in a global 50 000 employees company, I had a LOT of issues with the GDC email system.

Turned up I had shitloads of SMTP connections, in the thousands, while the site was only 200 users plus apps.

I deployed SMTP blacklisting and things went OK.

And then, this indian dude came to my desk asking why his deployed app on 10 000 desktops worldwide was no longer able to send emails !

I asked: "have you by any chance hard-coded our local GDC email relay instead of all 50 company's relays ?". There was a blink. Of course he had. Without asking us.

Turned up he was not even able to treat error conditions on SMTP ... Idiot.

Re: SMTP ddos

Jamesit

"And then, this indian dude came to my desk asking why his deployed app on 10 000 desktops worldwide was no longer able to send emails !"

Does the dudes race have anything to do with the story? If not why mention it?

It makes you sound racist.

Re: SMTP ddos

A.P. Veening

Indian is not a race, it is a country of origin, so it can't be racist.

And yes, there is a bit of confirmation of a cliché in there, but clichés become clichés just by their frequency of occurrence.

Re: SMTP ddos

Doctor Syntax

I think it relates more to business practices and maybe training practices in India.

Back in the day my then client did a good amount of work with one of the Usual Suspects. Like many at the time and, no doubt, much later the Usual Suspect subbed all development out to one of the Indian Usual Suspects who would - I think for visa reasons - rotate staff from India (or Indian staff if you're prepared to tolerate the adjectival form) through their UK office. These ranged from great* to just out of some training establishment. Needless to say it was the latter who got thrown into the deep end of actual coding. The consequence was periodic bouts of receiving not-quite XML files and having to explain to one of these staff-newly-arrived-from-India (and presumably just out of some training establishment there) how to get names such as O'Neil into well-formed XML.

So the fact that the dude was Indian speaks volumes about the general business environment.

* And a distinct improvement on the initial definitely not Indian "consultant" who initially arrived to brief us about one project.

Kaboom

Admiral Grace Hopper

Let she who has never shot herself in the foot fire the first bullet. [1]There is a proud tradition of auto-mutilation in IT .

[1] http://www.toodarkpark.org/computers/humor/shoot-self-in-foot.html

Updates are important!

Dave K

A good one this week, and yet another tale of developers believing that the latest patches are so unbelievably great that everyone must have them immediately. Obviously 4 hours is far too long to wait for the new shiny-shiny, must... patch... now!!!

Why is the IT manager deploying HA Proxy?

tip pc

"Jon was the IT manager of the whole show and responsible for all the gizmos from the CEO's mobile phone to the receptionist's desktop and everything in between."

"In the end, "I spent a good few hours spinning up new VMs to have all traffic route through a layer of Linux servers running HAProxy." He was able to carve out the application manager traffic and therefore allow a small percentage of requests to succeed and the fixed code gradually rolled out."

He could of just rate limited the inbound traffic to a lower amount or perhaps limited the number of sessions each server could muster, both likely would have been quicker and cheaper than spinning up new gear and deploying extra stuff, assuming the actual bandwidth they where handling required proper network gear that could actually do rate limiting etc.

I'd be really worried if my IT manager started installing & spinning stuff up.

Re: Why is the IT manager deploying HA Proxy?

GlenP

I'm an IT Manager and I do things like that, although we're a medium sized company in turnover terms we're low staffing levels so it's a small department. If he was the most experienced person there why shouldn't he carry out the necessary fixes. Better than the sort of manager who is clueless and just stands there shouting.

Re: Why is the IT manager deploying HA Proxy?

Anonymous Coward

This doesn't only happen on medium or small sized companies. My position at -ahem- a B*g B**e company has "manager" plastered all over my signature, but the only things I get to manage are automation scripts, with zero human resources and a shit ton of coding/scripting. The Powers at Being™ wanted to have someone capable but at the same time reassure the customer the work is being -ahem again- managed the right way, with the minimum resource investment.

Re: Why is the IT manager deploying HA Proxy?

MatthewSt

If a job's worth doing, it's worth doing right!

I'd be worried if my IT manager didn't at least know how to spin stuff up, even if they delegated it most of the time.

Football team web site throttles business

ColinPa

I heard of a company where they found 10% of the traffic was from the local football club's web site. The football web site was sending traffic to web browsers every few seconds/minutes(I forget), even though the browser was minimised, or hidden behind real work. A quiet word with the club, and they fixed it so it only sent data when the window was active

Re: Football team web site throttles business

Pete B

You're blaming the wrong party there - need to take it up with the idiot users who leave pages open when they're not doing anything with them.

Re: Football team web site throttles business

MatthewSt

Yes, because if there's one thing I've learnt as a dev, it's that users can be trusted to use anything you've written correctly! What's the point of having tabbed browsing if, as a user, I'm meant to close them as soon as I flick to another page?

With a few rare exceptions, it's not the user's fault for how they use your website.

"deploy HAProxy in front of all web traffic"

Bronek Kozicki

It's a good lesson to have learned.

ISP DDOSes self.

deadcow

I used to work for a major ISP. We had development teams working across several different departments. One morning I came into work, booted up my VM and noticed it was super chuggy, eventually hanging completely. I had a poke around going on, to find out the site was making thousands of requests to the server - specifically requesting a timestamp. "That's odd", I thought and sent out a call asking if anybody knew what this timestamp request was to see if I could find out what was happening.

It turns out the marketing department, in an effort to create the most granular tracking I have ever seen in my life, had decided that they wanted to know exactly what time users were clicking on interactive elements on the website. Note this was any element on every page of the site: links, accordions, show/hide buttons, popups, everything. Now - they had also decided that they required this with such extreme precision that they didn't want to rely on the user's own system time - they wanted it synchronized with the server's timestamp. So they set up a script that pinged the server for its current timestamp every single time a user clicked on anything. I watched with growing horror as I started to repeatedly open and close an accordion on the homepage, every single click resulting in a server call.

I enjoyed raising that P1 to a red-faced development team. We also had a long chat with the marketing department about DDOS-ing our own website in order to collect completely useless user data.

That said

diver_dave

I've been headscratching with an old Meridian switch before.

Getting multiple capacity drop outs that gave every indication of a card fault. Very very intermittent...

So after hours myself and Big Dave from BT have the panels off and start testing. Nothing. Anywhere.

Fault keeps reoccurring on and off.

Eventual diagnosis and all round circular arse kicking.

We had two sites. North and South.

One team split across the two sites. North forwarding phones to South hunt group. And.. Yep Vice versa.

Telephony equivalent of the old mail out of office auto response.

Nightmare to chase down. Only found it after some very careful traffic analysis.

Pint o'clock.

Dave

Re: That said

Anonymous Coward

ahhh the good 'ol mail loop. I started work at a Lab back in 98 and one of the first sh1t storms we had was caused my a mail loop. We ran Netware 4.11 and Groupwise 4.X (on the same box) Oh the fun we had when a boffin left and forwarded his email to his uni account which in turn forwarded stuff back to his lab account. The server sh1t itself trying to cope with all his email, which as groupwise was on the same box as Netware meant users couldn't login etc, utter sh1t show!

Re: That said

diver_dave

Indeed.

Strange as this was behaviour that technically shouldn't have been possible.

I've always sworn the damb thing was Maliciously sentient!

"The way of the world is to praise dead saints and prosecute live ones."
-- Nathaniel Howe