Investigating outage

Incident Report for Norce

Postmortem

Post-Mortem: SQL Cluster Storage Connectivity Failure - March 21, 2026

Duration: 2026-03-21 19:49 – 21:00 CET

Impact: Failed requests and timeouts on the Commerce API and Admin API for a subset of customers in the multi-tenant production environment. No single-tenant or Enterprise customers were affected.

Summary

On the evening of March 21st, one of our SQL Server clusters lost connectivity to its Elastic SAN storage backend in our Azure production environment. The connectivity loss was preceded by approximately 75 minutes of elevated I/O latency, a sign of storage communication degrading before failing completely. When connectivity was lost entirely, the SQL Server instance received OS-level I/O device errors across all attached storage volumes simultaneously and was unable to maintain database operations. The instance shut itself down at 19:52 CET.

Because access to the shared storage backend is a prerequisite for any recovery operation, the cluster could not automatically restore service regardless of its configuration. Manual intervention was required to reattach storage and restart the SQL Server instance, restoring service by approximately 21:00 CET.

Our current assessment is that the root cause lies in the Azure platform's networking or storage infrastructure. This is consistent with a prior incident in October 2025 in which connectivity between our SQL clusters and the Elastic SAN storage backend was disrupted due to performance degradation in Azure's virtual network infrastructure. A support case has been opened and escalated to Microsoft's eSAN and network infrastructure teams. This post-mortem will be updated when a formal root cause analysis is available.

Timeline (CET)

  • ~19:34 — SQL Server begins logging elevated I/O latency warnings against the storage backend. Individual operations taking in excess of one second.
  • ~19:49 — Connectivity to the storage backend is lost completely. SQL Server receives OS-level I/O device errors across all attached volumes.
  • ~19:52 — SQL Server shuts itself down after losing access to system-critical storage.
  • ~19:52 — Pacemaker cluster manager detects the failure and initiates recovery procedures. Recovery cannot complete without storage connectivity.
  • ~21:00 — Storage connectivity confirmed restored. Manual intervention initiates SQL Server startup and database recovery. Service restored.

Root Cause

A complete loss of connectivity between one of our SQL Server clusters and its Elastic SAN storage backend, resulting in an OS-level I/O device error across all attached volumes simultaneously. The failure was not gradual, the storage connection was interrupted abruptly after a period of degraded performance. Our current assessment is that the failure originated in the Azure platform's networking or storage infrastructure. We are awaiting confirmation from Microsoft.

What We Did

  • Opened and escalated a support case with Microsoft Azure, with an explicit request for platform-level log preservation and root cause analysis from the eSAN and network infrastructure teams.
  • Manually restored the SQL Server cluster once storage connectivity was confirmed re-established, completing database crash recovery and restoring service.
  • Verified data integrity across all affected databases following recovery.

What We're Implementing

  • Improved off-hour alerting: The storage degradation was detectable approximately 75 minutes before the complete connectivity loss. We are investigating improvements to our alerting capabilities to ensure early warning signals of this nature trigger faster human intervention, particularly during off-hours.
  • Storage backend architecture review: We are evaluating changes to how our SQL clusters connect to and interact with the underlying storage backend to improve resilience against this class of failure.
  • Critical infrastructure monitoring improvements: We are reviewing monitoring coverage of critical infrastructure components to ensure degraded states are surfaced and acted upon more quickly.

We apologize for the disruption and are committed to reducing both the likelihood and the duration of similar incidents in the future.

Posted Mar 24, 2026 - 15:33 CET

Resolved

We experienced a temporary disruption to our multi-tenant environment that affected a subset of our customers. The root cause has been isolated to a non-successful failover-event between two nodes in a SQL Server Cluster. The cluster has been restored to normal operations. We'll continue to monitor for any follow-up issues but we're expecting normal operations from this point on.
Posted Mar 21, 2026 - 20:45 CET

Investigating

We're currently investigating a disruption to normal operations in our multi-tenant environment.
Posted Mar 21, 2026 - 20:34 CET
This incident affected: Norce Commerce (Norce Commerce).