Square blames last week's outage on DNS screw-up
(2023/09/11)
- Reference: 1694458071
- News link: https://www.theregister.co.uk/2023/09/11/square_dns/
- Source link:
Square says the widespread outage that hit its payment terminals last week was caused by a DNS failure and not a cyberattack nor an intrusion.
During that 14-hour downtime, businesses across the US, UK, and beyond that relied on Square's point-of-sale systems were unable to process customers' credit and debit cards, hitting sales significantly. Restaurants, cafes, and shops were left asking punters to pay by cash or use another method like Venmo to transfer money to cover bills.
Some owners complained of losing thousands and thousands of dollars as the IT meltdown continued from Thursday into Friday. CashApp was also affected by the outage.
[1]
Shares in Block, which operates Square and CashApp, fell more than five percent during the fiasco. The San Francisco-based biz today blamed updates to its network infrastructure for breaking its DNS, bringing down services.
[2]If your DNS queries LoOk liKE tHIs, it's not a ransom note, it's a security improvement
[3]Microsoft DNS boo-boo breaks Hotmail for users around the globe
[4]If you can't log into Azure, Teams or Xbox Live right now: Microsoft cloud services in worldwide outage
[5]ICANN responds to Ukraine demand to delete all Russian domains
"The outage impacted an important part of our infrastructure, known as a Domain Name System, or DNS," Square explained Monday in a brief [6]postmortem report .
"While making several standard changes to our internal network software, the combination of updates prevented our systems from properly communicating with each other, and ultimately caused the disruption.
[7]
[8]
"The issue also affected many of our internal tools for troubleshooting and support, making them temporarily unavailable. There is no evidence that this was a cybersecurity event or that any seller or buyer data was compromised by the outage."
Here's our timeline of Square's IT woes last week:
Wednesday, September 6: A half-hour [9]outage affecting Square's payroll payment service. There was also an [10]hour-long issue preventing sellers from changing some of their settings. Likely unrelated to the above retail payment issues but we'll mention it anyway.
Thursday, September 7: Multiple [11]failures in Square's backend systems affecting payment processing, transfers, and other services, starting at midday PT, and lasting until the early hours of the following day. At midnight on the 7th, Square admitted: "We do not have a solution for the disruption." Then about an hour later, it had figured it out: "Our engineering team has implemented a fix and services are beginning to recover." At 0200 PT, 14 hours after everything started to go wrong, Square said it was monitoring the situation as more of its backend righted itself, and payment services became available again.
Friday, September 8: While Square was battling to fix the above outage, its Time Cards system, used by customers' workers to clock in and out, [12]was also down for most of Thursday and into Friday.
By 0700 PT on Friday, Square said it had [13]fixed its broken IT systems that had brought down its services since midday the day before, and apologized.
Square also outlined how it hopes to avoid this sort of meltdown again:
It claimed it has made changes to its DNS and firewall servers to "protect against the issue we saw," and has taken other defensive steps.
Square is working on expanding the availability of what's called Offline Mode to all new Square payment terminals as well as most ones out in the field already. As the name suggests, this mode allows a terminal to queue up transactions for processing later while backend systems are down or can't otherwise be reached. Square claimed many sellers used Offline Mode during the outage; we note the biz [14]warns "you are responsible for any expired, declined, or disputed payments accepted while offline."
The payments giant promised to improve the way it communications updates about downtime.
Finally, you know how the haiku goes. It's not DNS. There's no way it's DNS.
It was DNS. ®
Get our [15]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/personaltech&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZP@OBJvfLSyJDQIXBpLtbwAAAgo&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://www.theregister.com/2023/01/19/google_dns_queries/
[3] https://www.theregister.com/2023/08/21/microsoft_dns_booboo_breaks_hotmail/
[4] https://www.theregister.com/2021/04/01/microsoft_azure_dns_outage/
[5] https://www.theregister.com/2022/03/03/icann_ukraine_russian_domains/
[6] https://squareup.com/us/en/press/an-update-on-last-weeks-outage
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/personaltech&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZP@OBJvfLSyJDQIXBpLtbwAAAgo&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/personaltech&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZP@OBJvfLSyJDQIXBpLtbwAAAgo&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.issquareup.com/incidents/5p2nkbnj20hy
[10] https://www.issquareup.com/incidents/kls5b120tnyp
[11] https://www.issquareup.com/incidents/06zffytdqjdz
[12] https://www.issquareup.com/incidents/20cbqjl2wvlw
[13] https://www.issquareup.com/incidents/2trlsg0fbd9h
[14] https://squareup.com/help/us/en/article/7777-process-card-payments-with-offline-mode
[15] https://whitepapers.theregister.com/
During that 14-hour downtime, businesses across the US, UK, and beyond that relied on Square's point-of-sale systems were unable to process customers' credit and debit cards, hitting sales significantly. Restaurants, cafes, and shops were left asking punters to pay by cash or use another method like Venmo to transfer money to cover bills.
Some owners complained of losing thousands and thousands of dollars as the IT meltdown continued from Thursday into Friday. CashApp was also affected by the outage.
[1]
Shares in Block, which operates Square and CashApp, fell more than five percent during the fiasco. The San Francisco-based biz today blamed updates to its network infrastructure for breaking its DNS, bringing down services.
[2]If your DNS queries LoOk liKE tHIs, it's not a ransom note, it's a security improvement
[3]Microsoft DNS boo-boo breaks Hotmail for users around the globe
[4]If you can't log into Azure, Teams or Xbox Live right now: Microsoft cloud services in worldwide outage
[5]ICANN responds to Ukraine demand to delete all Russian domains
"The outage impacted an important part of our infrastructure, known as a Domain Name System, or DNS," Square explained Monday in a brief [6]postmortem report .
"While making several standard changes to our internal network software, the combination of updates prevented our systems from properly communicating with each other, and ultimately caused the disruption.
[7]
[8]
"The issue also affected many of our internal tools for troubleshooting and support, making them temporarily unavailable. There is no evidence that this was a cybersecurity event or that any seller or buyer data was compromised by the outage."
Here's our timeline of Square's IT woes last week:
Wednesday, September 6: A half-hour [9]outage affecting Square's payroll payment service. There was also an [10]hour-long issue preventing sellers from changing some of their settings. Likely unrelated to the above retail payment issues but we'll mention it anyway.
Thursday, September 7: Multiple [11]failures in Square's backend systems affecting payment processing, transfers, and other services, starting at midday PT, and lasting until the early hours of the following day. At midnight on the 7th, Square admitted: "We do not have a solution for the disruption." Then about an hour later, it had figured it out: "Our engineering team has implemented a fix and services are beginning to recover." At 0200 PT, 14 hours after everything started to go wrong, Square said it was monitoring the situation as more of its backend righted itself, and payment services became available again.
Friday, September 8: While Square was battling to fix the above outage, its Time Cards system, used by customers' workers to clock in and out, [12]was also down for most of Thursday and into Friday.
By 0700 PT on Friday, Square said it had [13]fixed its broken IT systems that had brought down its services since midday the day before, and apologized.
Square also outlined how it hopes to avoid this sort of meltdown again:
It claimed it has made changes to its DNS and firewall servers to "protect against the issue we saw," and has taken other defensive steps.
Square is working on expanding the availability of what's called Offline Mode to all new Square payment terminals as well as most ones out in the field already. As the name suggests, this mode allows a terminal to queue up transactions for processing later while backend systems are down or can't otherwise be reached. Square claimed many sellers used Offline Mode during the outage; we note the biz [14]warns "you are responsible for any expired, declined, or disputed payments accepted while offline."
The payments giant promised to improve the way it communications updates about downtime.
Finally, you know how the haiku goes. It's not DNS. There's no way it's DNS.
It was DNS. ®
Get our [15]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/personaltech&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZP@OBJvfLSyJDQIXBpLtbwAAAgo&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://www.theregister.com/2023/01/19/google_dns_queries/
[3] https://www.theregister.com/2023/08/21/microsoft_dns_booboo_breaks_hotmail/
[4] https://www.theregister.com/2021/04/01/microsoft_azure_dns_outage/
[5] https://www.theregister.com/2022/03/03/icann_ukraine_russian_domains/
[6] https://squareup.com/us/en/press/an-update-on-last-weeks-outage
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/personaltech&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZP@OBJvfLSyJDQIXBpLtbwAAAgo&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/personaltech&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZP@OBJvfLSyJDQIXBpLtbwAAAgo&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.issquareup.com/incidents/5p2nkbnj20hy
[10] https://www.issquareup.com/incidents/kls5b120tnyp
[11] https://www.issquareup.com/incidents/06zffytdqjdz
[12] https://www.issquareup.com/incidents/20cbqjl2wvlw
[13] https://www.issquareup.com/incidents/2trlsg0fbd9h
[14] https://squareup.com/help/us/en/article/7777-process-card-payments-with-offline-mode
[15] https://whitepapers.theregister.com/
Re: Offline mode?
chivo243
I once supported* POS devices that did exactly that, when connection to the mothership was lost, it began buffering, but there was a limit as I recall. I'm so glad those days are behind me.
*Went as far as unplugging the offending USB peripheral and plugging it back in again. Rebooting the underlying Windows, or reading the error code over the phone to the helpdesk.
It's always DNS
chivo243
Sounds like the issue is a fat finger... causing the problem, and then another fat finger compounds the problem, and finally a total melt down, too many fat fingers in the soup?
Offline mode?
"this mode allows a terminal to queue up transactions for processing later while backend systems are down or can't otherwise be reached. "
I cannot imagine any way THAT would be exploited....
Oh, wait...yes I can.