Azure Active Directory logs are lagging, and alerts may be wrong or missing
- Reference: 1654063263
- News link: https://www.theregister.co.uk/2022/06/01/azure_ad_logging_problem/
- Source link:
"Customers using Azure Active Directory and other downstream impacted services may experience a significant delay in availability of logging data for resources," the [1]Azure status page explains. Tools including Azure Portal, MSGraph, Log Analytics, PowerShell, and/or Application Insights are all impacted.
Azure AD and the other abovementioned tools are all working.
[2]
But Microsoft has warned that the incident "could lead to missed or misfired alerts."
[3]
[4]
Or in other words, bad things could be happening but you might not hear about them. Or you might get some weird alerts.
Either scenario is sub-optimal.
[5]
The software giant detected the issue at 21:35 UTC on May 31. As of 05:15 UTC on June 1, the problem remains unresolved.
Microsoft's last update about the issue, timestamped 04:15 UTC on June 1, stated that the company is "currently investigating a recent build roll out as the cause" and "continuing to investigate for a full root cause."
[6]Microsoft Azure to spin up AMD MI200 GPU clusters for 'large scale' AI training
[7]Microsoft-backed robovans to deliver grub in London
[8]Windows Subsystem for Linux 2 splashes down on Win Server 2022
Azure engineers are working to roll back Azure AD to a version without whatever problem is causing this issue, with "signs of recovery" already evident.
Azure AD problems – such as the [9]September 2020 outage – tend to be widely felt, as the tool is by design used to authenticate users to multiple services.
Isn't it grand, then, that Microsoft is [10]encouraging the use of the service instead of on-prem Active Directory?
[11]
The timing of this incident is also exquisite, as it commenced on the same day Microsoft expanded and rebranded its identity and access tools under the name [12]"Entra" . ®
UPDATE 06:45 UTC, June 1st. Microsoft's posted an update time-stamped 06:31 UTC and the news is not good.
The company says the issue has caused "additional impact to Azure Resource Manager for CRUD (create, read, update and delete) operations, with some requests experiencing failures whilst communicating with other Azure services."
Previous updates about the incident were promised every hour. Microsoft missed that last deadline and now advises another update will arrive around 08:30 UTC "or as events warrant".
"We are engaging additional engineering teams to assist in applying multiple mitigation steps, services are seeing signs of recovery at this time," states the latest update.
Get our [13]Tech Resources
[1] https://status.azure.com/en-us/status/
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ypc4wyzy3Shit7JAZaZrtwAAAIs&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ypc4wyzy3Shit7JAZaZrtwAAAIs&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ypc4wyzy3Shit7JAZaZrtwAAAIs&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ypc4wyzy3Shit7JAZaZrtwAAAIs&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2022/05/26/amd_azure_microsoft/
[7] https://www.theregister.com/2022/05/19/selfdriving_microsoft_wayve/
[8] https://www.theregister.com/2022/05/26/wsl2_windows_server_2022/
[9] https://www.theregister.com/2020/09/29/onedrive_azure_active_directory_outage/
[10] https://www.theregister.com/2022/05/05/azure_ad_update_compliance/
[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ypc4wyzy3Shit7JAZaZrtwAAAIs&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[12] https://www.theregister.com/2022/05/31/entra/
[13] https://whitepapers.theregister.com/
The biggest issue I saw with Azure Active Directory logs, when connecting things like Splunk to them was that there was no sequence number.
Azure would aggregate all these messages in the back end. Say you collect everything from 10:00:00 through 10:05:00, all good, right? It turns out a couple of servers were down at that point and now they've sent their logs in. There's extra messages there waiting for you, how do you know? Well, you've got to request that time period again and amend your own logs. How do you know to request that time period again? You don't.
MICROSOFT PLEASE IMPLEMENT SEQUENCE NUMBERS ON LOG MESSAGES.
I think you're shouting into the void here, just as you would be on Microsoft's support site.
I'm just putting it out there hoping some dev will read it and fix it...
If you think that AD on the cloud is a nightmare
Wait until you experience the on-premises version. It is exactly the same, only it is you who is applying and rolling back updates.
Love the CRUD acronym
more CRUD
Creat Read Update Filesystem Timestamp
The best I can come up with after only one cup of coffee.
Oh dear
When you have a captive audience, you can get away with a lot of shit.