Microsoft 365 and Azure Hit by Major Global Outage, Automated Network Maintenance Bug Identified

The420.in Staff
4 Min Read

 

 

Microsoft has attributed the widespread outage that disrupted Microsoft 365 and Azure services on July 23 to a software bug in its automated network maintenance system. According to the company, a flaw during routine network maintenance mistakenly removed IP routes from more network devices than intended, causing significant network disruption in the Azure West US region. The incident affected customer access to Microsoft 365, Azure, and numerous cloud-based services.

Microsoft said the outage began at 10:44 a.m. ET and primarily impacted customers accessing Microsoft 365 services through network infrastructure connected to the West US Azure region. Within minutes, thousands of users reported service interruptions, slow performance, login failures, and connectivity issues across multiple Microsoft platforms.

The company tracked the incident under Incident ID MO1437424. Affected Microsoft 365 services included SharePoint Online, Microsoft Teams, OneDrive, Microsoft 365 Admin Center, Power Automate, Copilot Chat, and Microsoft Loop. SharePoint users encountered “Something went wrong” error messages, while Microsoft Teams experienced degraded chat functionality, including failures to load images. OneDrive access became intermittent, and many administrators were unable to access the Microsoft 365 Admin Center normally.

India’s Largest Cybercrime Conference Nears: FutureCrime Summit 2026 Set for 6–7 August at Bharat Mandapam

Several additional Microsoft cloud services were also impacted, including Power BI, Microsoft Fabric, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender. Some Microsoft Defender customers experienced delays in receiving responses from Defender Experts, while investigations, workflows, and remediation actions initiated through Threat Explorer and Advanced Hunting failed or were delayed. Azure services such as Azure App Service, Azure Kubernetes Service, Azure Cosmos DB, Azure Firewall, ExpressRoute, Microsoft Sentinel, Virtual WAN, and VPN Gateway also experienced connectivity problems and increased latency.

As an initial mitigation measure, Microsoft rerouted traffic through alternative network paths to reduce the impact. While this restored connectivity for some customers, many services continued to experience disruptions. Engineers subsequently identified that a recent networking change had triggered the incident. During routine maintenance, the automated request conversion system incorrectly classified additional network devices as part of the maintenance operation, resulting in the removal of significantly more IP routes than intended between the West US datacenter and Microsoft’s wide-area network (WAN). Microsoft clarified that network traffic remaining entirely within the West US region was not affected.

The company began rolling back the maintenance change at 1:45 p.m. ET, completing the rollback by 2:26 p.m. ET. Following restoration of the affected network infrastructure, Microsoft confirmed through service telemetry and customer feedback that Microsoft 365 services had recovered. Some Azure services required additional recovery time, with Microsoft stating that all impacted services were fully restored by 3:41 p.m. ET.

According to a Researcher at Algoritha Security, automated maintenance systems significantly improve operational efficiency across large-scale cloud environments, but even a minor flaw in validation or safety mechanisms can affect millions of users simultaneously. The researcher noted that multilayer validation, phased deployment strategies, and automated rollback mechanisms are critical safeguards that help minimize the risk of large-scale service disruptions.

Microsoft has announced a comprehensive internal review of the incident, focusing on its automated maintenance request process, network safety checks, and overall change management procedures. The company said it will publish a final post-incident review within approximately 14 days, outlining the root cause, corrective actions, and improvements designed to prevent similar outages in the future and further strengthen the reliability of its global cloud infrastructure.

Stay Connected