Post-Mortem: SQL Cluster Storage Connectivity Failure - March 21, 2026
Duration: 2026-03-21 19:49 – 21:00 CET
Impact: Failed requests and timeouts on the Commerce API and Admin API for a subset of customers in the multi-tenant production environment. No single-tenant or Enterprise customers were affected.
Summary
On the evening of March 21st, one of our SQL Server clusters lost connectivity to its Elastic SAN storage backend in our Azure production environment. The connectivity loss was preceded by approximately 75 minutes of elevated I/O latency, a sign of storage communication degrading before failing completely. When connectivity was lost entirely, the SQL Server instance received OS-level I/O device errors across all attached storage volumes simultaneously and was unable to maintain database operations. The instance shut itself down at 19:52 CET.
Because access to the shared storage backend is a prerequisite for any recovery operation, the cluster could not automatically restore service regardless of its configuration. Manual intervention was required to reattach storage and restart the SQL Server instance, restoring service by approximately 21:00 CET.
Our current assessment is that the root cause lies in the Azure platform's networking or storage infrastructure. This is consistent with a prior incident in October 2025 in which connectivity between our SQL clusters and the Elastic SAN storage backend was disrupted due to performance degradation in Azure's virtual network infrastructure. A support case has been opened and escalated to Microsoft's eSAN and network infrastructure teams. This post-mortem will be updated when a formal root cause analysis is available.
Timeline (CET)
Root Cause
A complete loss of connectivity between one of our SQL Server clusters and its Elastic SAN storage backend, resulting in an OS-level I/O device error across all attached volumes simultaneously. The failure was not gradual, the storage connection was interrupted abruptly after a period of degraded performance. Our current assessment is that the failure originated in the Azure platform's networking or storage infrastructure. We are awaiting confirmation from Microsoft.
What We Did
What We're Implementing
We apologize for the disruption and are committed to reducing both the likelihood and the duration of similar incidents in the future.