Executive Summary

AWS finished its assessment of the two Middle East regions and told customers two things. Resources and data hosted only in Bahrain cannot be restored at all. In the United Arab Emirates, the same is true of one availability zone, mec1-az2. Both regions ran three zones, and the damage crossed more than one of them in Bahrain. That is the part that matters, because it is the failure pattern the service design was never built to absorb.

The verdict is not that the cloud failed. It is that a spread across availability zones was never a copy, and plenty of architecture reviews quietly treat it as one. The distinction is old and unglamorous. Multi-AZ buys uptime inside a region. It does not buy a second copy of anything. AWS is still working on the two remaining UAE zones and has promised a Bahrain update in early 2027, so this is an open story rather than a closed incident. Any team whose data exists in exactly one region should read the assessment as a deadline, not as an incident report.

The sequence matters more than the conclusion. The first Bahrain zone was hit in March 2026. AWS told customers to replicate out and migrate, and most of them did. A second zone went down in April and the whole Bahrain region became unavailable. By the time AWS published its assessment on 15 September, the remaining question was no longer how to recover the region. It was how to say that recovery was impossible.

Diagram of the two AWS Middle East regions. The UAE region me-central-1 lists three zones, with mec1-az2 marked unrecoverable and mec1-az1 and mec1-az3 recovering. The Bahrain region me-south-1 lists three zones all struck, and AWS states it cannot restore resources and data hosted only in that region.
What the design assumed and what the assessment found. Region codes and zone states are taken from the AWS Health Dashboard assessment of 15 September 2026.

Availability is not a copy, and the difference is expensive

An availability zone is a cluster of one or more data centers inside a region. AWS documents three in Bahrain and three in the UAE on its global infrastructure page. The design intent is simple. One zone fails and the application stays up. What the design does not promise is a second copy of your data. Backup and disaster recovery are separate disciplines with their own documentation. The reliability guidance and the disaster recovery whitepaper both draw the line plainly. A backup belongs in a different account and ideally a different region, and a DR plan is a tested procedure rather than a diagram.

Scatter replicas across three zones, then call the job done, and you have bought uptime. You have not bought recoverability. The bill for that confusion arrives in a single line from a cloud provider.

The damage pattern was outside the assumption, not outside the design

The load-bearing sentence in the assessment is that the damage spanned multiple availability zones and exceeded what its regional and multi-AZ services are designed to withstand. That is AWS naming its own ceiling. Zone isolation assumes independent failure. Physical strikes inside one city do not stay independent, and fire suppression adds water damage on top of blast damage. Reporting on the incident records two of the three UAE zones significantly impaired and two UAE facilities directly struck. The UAE is now said to be rethinking a major AI data center project around dispersed sites and hardened underground space.

AWS has always said the quiet part out loud in its multi-region guidance. Regional failure is a different class of event from zone failure, and it needs a different answer. The ceiling on multi-AZ resilience is not theoretical any more, and it was never only about a bad deploy or a flooded street. It is also about geography that a single region cannot escape.

The question for your estate is narrower than a multi-region rebuild

You do not need to re-architect everything this week. You need one honest inventory. Find every workload that exists in exactly one region with no copy anywhere else, then find out whether a copy is even allowed to leave. Where residency or sovereignty obligations block a second region, the residual risk is a constraint on the architecture rather than a missing product feature. That belongs in a risk register, not a runbook, and it needs an owner with the authority to accept it.

The rest is practice. Rehearse a restore rather than a failover, and do it from the copy you would actually reach for under pressure. The gap between owning backups and completing a recovery is where incidents get longer, and it is the same gap that turned a zone spread into a false sense of safety for a lot of teams back in March.

Related reading. Sovereign cloud is not data residency, and the gap costs money. Digital sovereignty is a procurement problem, not a data center problem. Data center power procurement moved into the IT budget.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data centre economics. No vendor spin.