News: 1604909713

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

The day I took down the data centre- I mean, the day I saved the day. Right, boss?

(2020/11/09)


Who, Me? Welcome to a [1]Who, Me? story in which the moral might be: "Be careful what you kick off before lunch if you want a mealtime free of phone calls."

Today's tale concerns the exploits of "Anthony," who was working in the security department of one of the larger cable internet providers. One of his jobs was assessing systems due to be deployed, a task that was sensibly done on a pre-deployment network. A device would be built, popped onto the network and Anthony would run various tests to ensure everything had been put together to spec.

"Each Regional Data Center (RDC)," he explained, "had its own pre-deployment area, and I'd run the scans from my local server.

"All of the RDCs were connected via backbone connections, so latency was negligible, and bandwidth was massive."

One of the tools he used was the scanner [2]Nmap ("Network Mapper"), a handy utility to rapidly scan large networks. Nmap will do a variety of useful things, including showing what is lurking on a network and an "interesting ports table."

Nmap has a variety of parameters, including a bunch around timing and performance to control how it runs. While the settings can be as fine-grained as one likes, the utility features some simple timing templates via the -Tx option, where x is a number from 0 to 5. The [3]options summary innocently notes that "higher is faster" for this figure.

Anthony usually stuck with the default – 3 (Normal).

On this occasion, however, he was keen to get some lunch and wanted the scan completed earlier, "so I ran it at -T5 ."

Diving deeper into [4]Nmap's documentation reveals what those numbers really mean. "The template names," explain the docs, "are paranoid (0), sneaky (1), polite (2), normal (3), aggressive (4), and insane (5)."

"Insane mode assumes that you are on an extraordinarily fast network or are willing to sacrifice some accuracy for speed."

The scan kicked off, and Anthony cracked on with the paperwork.

The phone soon began to ring: "In the background," he told us, "I can hear yelling. A voice shouts: 'What the hell are you doing??!?'"

Angry phone calls is a part of everyday life for many of us in IT, doomed to be at the beck and call of users and bosses. Patiently, Anthony explained that he was merely doing server assessment."

"THE RDC IS DOWN!" came the shrieking from the phone.

Eh?

"THE RDC IS DOWN! The firewall crashed and won't come back up!"

It transpired that Anthony had swamped the enterprise firewall, which promptly crashed and refused to come back up due to the packets being sprayed at it as fast as Anthony's server could manage.

He killed the scan.

"All in all, it was only a 15 - 20 minute outage for the 2-300,000 customers..." he noted.

Once things had settled down, Anthony was hauled before the bigwigs, with HR in attendance, to explain himself. Faced with a potential career-shortening (having killed service for hundreds of thousands customers) he did the only thing possible.

He became the self-proclaimed hero of the hour.

Yes, a Bad Thing had happened, but look at it this way: a "significant flaw" had been discovered. One disgruntled person with a well-connected device could take out an entire RDC! In many ways, the company should be thanking him. Perhaps a bonus for his diligence?

"I kept my job."

Ever screwed something up so badly that the only way out of a P45 and the march of shame was via the medium of spin? Or perhaps you've also unleashed the power of Nmap without fully considering the consequences? An email to [5]Who, Me? is all it takes to purge your conscience. ®

Get our [6]Tech Resources



[1] https://www.theregister.com/Tag/who-me

[2] https://nmap.org/book/man.html#man-description

[3] https://nmap.org/book/man-briefoptions.html

[4] https://nmap.org/book/man-performance.html

[5] mailto:whome@theregister.com

[6] https://whitepapers.theregister.com/

All the RDC’s where on their own backbone?

Anonymous Coward

"All of the RDCs were connected via backbone connections, so latency was negligible, and bandwidth was massive."

Just also happened to be sharing the corp firewall too.

If it’s important enough to segregate on its own links but needs firewalling then it’s an own goal to use the enterprise firewall for that shared task, a dedicated device with enough capacity to handle the throughput achievable in those “backbones” should have been implemented too.

I suspect those RDC’s had their own “backbones” but the testing stuff relied on those same “backbones” and didn’t have dedicated connectivity as the story hints at.

memories...

Anonymous Coward

Did the same with zmap once - without any proper thought, I decided that since my /24 scan had found so many of our corporate subnets, a /16 sweep would be worthwhile as well. I got about 2 minutes into it before my connection dropped. And the phones lit up...

The VM I was running it from was in the DC, and I couldn't kill the scan as of course I lost connection too. However, the Cisco ASA was always having issues (due to being configured incorrectly, mainly) and so I went to lunch and hoped nobody would notice the "you're doing a zmap!" banner in the linux system log while I was out...

Anonymously, just in case !

NMAP's 11 setting!

chivo243

As Jamie would say, "Well, there's your problem." I've seen it in action, used by an uninformed admin. Insane about sums it up... all the network monitors went red in the space of 10 seconds, and then stopped reporting until he killed NMAP. That was a fun morning.

Takes me back

A K Stiles

When I worked in a place that ran AS400 / iSeries / i5 systems, we developers sensibly had our own box for dev work, separate to the live server.

Sometimes a new function or a fix would require a reasonably chunky bit of code to be run and, obviously, we'd try it first on the dev box to check it wasn't going to run wild and destroy all the account records.

Frequently these would result in calls from the sys-admin department calling to complain that "Your job is taking 85% of the processor" or something similar and a request to terminate the job to clear the alert on the big screen. Now obviously as developers (on the dev box) our usual ponder was "what's happening with the other 15% then?".

Various techniques were employed to see if we could get the jobs to run to completion, from being "on a call" (or at least the phone off the hook) when they might phone us, to prolonging the conversation about which job was a problem, to even coding in some 5 second sleeps every 20 seconds or so to reduce the persistence of the notification, just so the job would actually get a chance to run through to completion without getting terminated.

I can only remember a couple of occasions in 10 years where one of the other colleagues actually working on the dev box shouted across the office about not being able to do stuff because of server obstruction, so it wasn't exactly a significant issue!

That's interesting

Pascal Monett

So, you have a network tool that has a setting that can basically kill the network. It's up to you to not use that setting.

That doesn't sound like a useful thing to me.

Is there any reason to have that setting ? Stress test, maybe ?

Gumperson's Law:
The probability of a given event occurring is inversely
proportional to its desirability.