News: 1693438267

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Oracle Cloud, Netsuite, and Azure go down, hard, Down Under

(2023/08/31)


Updated Oracle, Netsuite, and Microsoft's clouds have gone down, hard, in the Sydney, Australia, region likely due to an issue at a datacenter provider in which both are tenants.

The Big Red Cloud first advised customers of an outage at 2129 Sydney time (1229 UTC) on Wednesday, and 29 minutes later wrote to inform customers that the outage had started earlier than its first emailed advisory – at 1015 UTC.

Oracle’s second email delivered the mixed message that: “We are still investigating an issue in the Australia East (Sydney) region that is impacting multiple OCI services. We have identified root cause of service failures and are working to mitigate the issue.”

[1]

The snafu took out a swathe of compute and SaaS services, with Kubernetes offerings disrupted, Exadata unavailable, and cloudy networks borked. Oracle has advised the impact of the incident means “some customers may be unable to use or connect to certain OCI services in the region.”

[2]

[3]

Oracle’s next emailed advice again changed the IT giant's story, this time referring to a “preliminary root cause” and naming that as related to “network connectivity.”

Microsoft, meanwhile, advised customers of its Azure cloud that as of approximately 0830 UTC on August 30 it was suffering an outage due to a cooling problem: "A utility power surge in the Australia East region tripped a subset of the cooling units offline in one datacenter, within one of the Availability Zones."

[4]

Temperatures rose inside the datacenter and Microsoft "proactively powered down a small subset of selected compute and storage scale units, to avoid damage to hardware."

Microsoft said it has since restored 99 percent of storage services and 99 percent of impacted Virtual Machines, and is "actively investigating individual downstream services to confirm their recovery status and mitigate remaining issues."

Downstream dependencies are preventing full recovery, we're told. Oracle is also having trouble restoring all services.

[5]

Oracle-owned cloudy ERP outfit Netsuite's [6]status page states "A lightning storm impacted the chiller plant in the Sydney data center, and most systems were temporarily shut down to reduce temperatures. The temperatures have stabilized, and the systems are being systematically powered up."

Netsuite service is slowly being restored.

At the time of writing Oracle Cloud’s [7]status page listed ten services as unavailable, six as experiencing service disruptions, and 12 as operational – including some services that were previously listed as unavailable.

Oracle Australia told The Register it has no comment at this time.

[8]Cisco's Duo Security suffers major authentication outage

[9]Microsoft DNS boo-boo breaks Hotmail for users around the globe

[10]UK flights disrupted by 'technical issue' with air traffic computer system

[11]Oracle shrinks its on-prem cloud into a single rack

The Register understands Oracle's networking team is most exercised right now as Big Red battles to end the downtime, consistent with the little information Oracle has provided mentioning customers may be unable to reach its cloud and ongoing connectivity issues.

Readers have [12]suggested a lightning storm that struck Sydney last night could be the cause of the equipment breakdown. The storm sparked power outages in parts of the city that host several datacenters. Your correspondent's afternoon was darkened by the storms, which passed through at around 1600 Sydney time, a few hours before cloudy troubles began.

We've asked the datacenter companies known to have Microsoft and Oracle as tennnts for comment.

Whatever the root cause, hyperscale clouds and the datacenters that host them pride themselves on building sufficient resilience to cope with problems. Those efforts clearly have not worked in this case.

We’ll update this story if more info arrives. ®

Editor's note: This story was updated at 0000 UTC and again at 0030 UTC to reflect new information about the storm, Microsoft's outage, and knowledge of the situation we have learned from other sources. We updated the story again at 01:55 UTC to add information about Netsuite.

Get our [13]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZPAQY@sgD62FgKj@g2If@wAAAk4&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZPAQY@sgD62FgKj@g2If@wAAAk4&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZPAQY@sgD62FgKj@g2If@wAAAk4&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZPAQY@sgD62FgKj@g2If@wAAAk4&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZPAQY@sgD62FgKj@g2If@wAAAk4&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://status.netsuite.com/incidents/p9kbg053tnm7

[7] https://ocistatus.oraclecloud.com/#/incidents/ocid1.oraclecloudincident.oc1.phx.amaaaaaanfxwrmaazawiom5jkbqncnrj6o7flg377czw7txc465affurc4aa

[8] https://www.theregister.com/2023/08/21/ciscos_duo_outage/

[9] https://www.theregister.com/2023/08/21/microsoft_dns_booboo_breaks_hotmail/

[10] https://www.theregister.com/2023/08/28/uk_flights_disrupted/

[11] https://www.theregister.com/2023/08/11/oracle_compute_cloud_at_customer/

[12] https://twitter.com/skaffman/status/1697032792780189923

[13] https://whitepapers.theregister.com/



It's not just Oracle

formerpommie

I'm at the brown end of this and can tell you that (at least) Microsoft Azure and SAP have also been impacted by this outage.

It has provisionally been attributed to a cooling failure caused by a power surge. (this may change)

Anonymous Coward

One of our clients logged a ticket about a DB error in Micropay. Looking at the Micropay IP address they are in Sydney & MS are reporting SQL services are still having problems.

Today is payday, so they are not very happy.

Linkedin was also down, actually terribly slowyy

Clausewitz4.0

Linkedin was also down a few hours ago - probably this was the cause

But you who live on dreams, you are better pleased with the sophistical
reasoning and frauds of talkers about great and uncertain matters than
those who speak of certain and natural matters, not of such lofty nature.
-- Leonardo Da Vinci, "The Codex on the Flight of Birds"