Poor communication led to complete lack of communication
- Reference: 1705908666
- News link: https://www.theregister.co.uk/2024/01/22/who_me/
- Source link:
This week meet a reader we'll Regomize as "Toby" who once worked as a marketing consultant and tech guy – quite the combination – for a digital marketing agency. Basically, as he puts it, "anything that involved code fell into my lap."
One fine day, a client approached this agency looking for an update to its CRM system. The existing CRM was "horrific" according to Toby: "there was no input validation, and it was entirely local and could only operate on one single computer."
[1]
Ergh. Obviously an update was needed.
[2]
[3]
To make things even more fun, the client wanted a new website at the same time as its shiny new CRM, with the added requirement that the contacts on the CRM had to sync with the user accounts on the website, and vice versa. Not in itself too complicated, but the website was being handled by an external agency that was in a different time zone to Toby's agency.
The solution devised was to send webhooks from the site using a WordPress plugin to the CRM whenever something was updated. The CRM would then do the same thing, back to the site, using Zoho Flow.
[4]
It sounded good in theory. In initial testing, it behaved as expected, so all was good. The data uploaded as expected, webhooks were sent, data was updated on the other side, and everyone was happy.
So, with a day's satisfactory work done, Toby clocked off.
[5]WTF? Potty-mouthed intern's obscene error message mostly amused manager
[6]New year, new bug – rivalry between devs led to a deep-code disaster
[7]PLACEHOLDER ONLY Someone please write witty headline here
[8]Enterprising techie took the bumpy road to replacing vintage hardware
Then of course the next day arrived, as it was wont to do. When Toby got to work, he opened his email to find an urgent request: "Any clue why we have something like 500 requests to update user?"
Toby did not, as it happened, have a clue. But he soon found one.
It transpired that as Toby slept, the web agency had done its own testing by uploading a payload to update a user on the CRM. The CRM had responded by updating the user and sending a web hook back to the site, which had responded by sending another payload to the CRM. The CRM responded by updating the user and sending more webhooks … and so on ad infinitum .
[9]
Or not quite ad infinitum . More like ad until the agency’s Zoho credits ran out . And when they ran out, the webhooks stopped, and everything ground to a halt. The project had been effectively crippled by a self-inflicted DDoS.
Toby had, unwisely, assumed that the web agency had safeguards in place to prevent this kind of recursion. The agency had, unwisely, assumed Toby would build in such safeguards. Since we all know the cliché about what happens when you assume, we will not repeat it here. Suffice to say everyone was an ass.
The worst part was that the Zoho subscription for Toby's agency had only been renewed a few days before, and this little goof had used up the entire allowance for the month. What's more, this client was far from the only thing for which the agency relied on Zoho. So the lack of communication between agencies didn't cause too much of a problem for the client (the system was still in testing after all) but it had a big impact on Toby's employer.
Ever found yourself in the middle of a mess that could have been avoided if only someone had said something? Well, don't keep it to yourself, [10]communicate it to Who, Me? and we'll share your object lesson with others, so they may avoid the same fate. And get a chuckle. ®
Get our [11]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Za5K3s0SVtuT7XcQwnW0WgAAARI&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Za5K3s0SVtuT7XcQwnW0WgAAARI&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Za5K3s0SVtuT7XcQwnW0WgAAARI&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Za5K3s0SVtuT7XcQwnW0WgAAARI&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://www.theregister.com/2024/01/15/who_me/
[6] https://www.theregister.com/2024/01/08/who_me/
[7] https://www.theregister.com/2023/12/18/who_me/
[8] https://www.theregister.com/2023/12/11/who_me/
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Za5K3s0SVtuT7XcQwnW0WgAAARI&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] mailto:whome@theregister.com
[11] https://whitepapers.theregister.com/
You assume sanity was part of this project.
Because testing obviously wasn't.
As above, so below
"marketing consultant and tech guy"
The fateful combination of two providers for CRM and Web services seems a [1]Hermetic mirror of Toby's two roles.
[1] https://en.m.wikipedia.org/wiki/As_above,_so_below
You say lack of communication. I say, as it typically is with sweatshops agencies, it was poor planning.
Email...
I had a similar issue with email in the days before on-site email servers and broadband. We used an ISDN connection to an external service which suddenly slowed down*. It turned out one of our users had put his Out Of Office on in whichever email client he was using without checking the "only send once" box. He'd then sent an email to someone external who'd done the same thing - result was multiple OOF messages bouncing between the two! Stopping the service, clearing the queue and fixing the OOF settings cured the problem.
*Users expected email to be instant, which it wasn't and isn't now - I had one senior manager complain that her colleague in the States hadn't received an email 2 minutes after she'd sent it. Being ISDN our email only connected every 15 minutes but I showed here there was nothing in the queue at our end and there was nothing we could do.
Re: Email...
had similar back in the day. circa 2000, Netware 4.11 and Groupwise 4.x a couple of boxes running the Lab's IT, griefwise and NDS running on the same tin. A boffin left the lab to move to a Uni in the uk and managed to set up a out of office I've left email the lab to his uni account and then on his uni account a forward to his old account at our lab. Result, utter shit show mail loop! Took days to sort and as griefwise was on the same tin as NDS the snails pace o the box meant users couldn't login!
Did he seriously not think of that? It seems so obvious that I would expect it to be in the initial spec, not missed until after the system went live.
> "Since we all know the cliché about what happens when you assume, we will not repeat it here"
[1]Do we? Are you sure?
[1] https://xkcd.com/1339/
Presumably the one about making an ass out of (yo) u and me ...
Are you assuming that?!?
The ONLY thing I'm sure of ...... “Two things are infinite: the universe and human stupidity; and I'm not sure about the universe.” ― Albert Einstein
Never heard of Zoho before but generally, if you use such third-party service in a client project, wouldn't you have project-specific subscriptions?