/ ,

  / When the Cloud Goes Dark: Regional Outages and What They Mean for SOC 2 and ISO 27001 Compliance

When the Cloud Goes Dark: Regional Outages and What They Mean for SOC 2 and ISO 27001 Compliance

In March 2026, a regional conflict in the Middle East did something that stress tests and tabletop exercises rarely manage to do: it took down cloud infrastructure across multiple availability zones at the same time, in the same region, without warning.

AWS data centers in the UAE and Bahrain were impacted. Banking apps went offline. Payments failed. Delivery platforms stopped. And a significant portion of the affected organizations had done everything “right” by conventional standards — multi-AZ deployments, redundancy within the region, documented continuity plans.

It wasn’t enough.

This article breaks down what happened, what it revealed about how most organizations think about availability, and what a more resilient architecture actually looks like. If your systems run on cloud infrastructure — in any region — this case is worth understanding closely.

What Happened: The March 2026 Incident

Regional conflict in the Middle East caused physical and infrastructural disruption to AWS facilities across the UAE and Bahrain. Based on publicly reported information, the incident involved power outages affecting data center operations, physical damage to infrastructure facilities, connectivity loss across affected environments, and service degradation spanning multiple availability zones within the same region — simultaneously.

That last point is the one that matters most. AWS designs its availability zones to be isolated from one another — separate power, cooling, and networking — so that a failure in one zone doesn’t cascade into another. Under normal failure conditions, that isolation holds. But this wasn’t a normal failure condition. It was a regional-scale disruption. The “rooms” were fine. The “building” was the problem.

“Availability zones are designed to handle localized failures, not regional ones. This incident sits firmly in the second category.”

The result was that organizations with multi-AZ architectures — which many rightly considered robust — still went down. There was no in-region fallback left to use.

Business Impact: What Actually Went Offline

The impact was not subtle. Banking platforms experienced downtime that prevented customers from accessing accounts or completing transactions. Payment processors were unable to process transactions. Mobility and delivery platforms halted operations entirely. Customer-facing applications became unavailable across the board.

This wasn’t degraded performance or slower load times. It was a full loss of availability for any system that lived entirely within the affected region. The AWS Well-Architected Framework acknowledges that regional failures, while rare, are a defined risk category — and designing for them requires a fundamentally different approach than designing for AZ failures.

Organizations with multi-region architectures kept operating. Everything else stopped. That single architectural decision — single-region versus multi-region — was the difference between availability and a complete outage.

What Risks Actually Materialised

This incident didn’t create new risks. It exposed ones that were already there, quietly embedded in architectural choices and compliance assumptions that had never been stress-tested at this scale.

Regional Single Point of Failure

The most common pattern among affected organizations: applications, databases, and backups all deployed within a single region. When that region became unavailable, there was no secondary environment to take over. No warm standby, no traffic rerouting, no automated failover. Just downtime.

This is the architectural equivalent of backing up your data to a drive sitting next to your laptop. It works until it doesn’t.

The Limits of Availability Zone Redundancy

Availability zones are a powerful tool — but they’re a tool designed for a specific class of failure, and understanding that class matters. Think of an availability zone as a separate floor in a building. If one floor has a problem, you move to another floor. But if the entire building loses power — or becomes inaccessible — floor redundancy doesn’t help. You needed another building entirely. That’s what a region is. And this incident took down the building.

Pro tip: When mapping your architecture against a business continuity plan, explicitly define your regional failure scenario. “What happens if this entire region becomes inaccessible for 24 hours?” is a question that exposes gaps that AZ-level planning will never catch.

Infrastructure-Level Disruption Is Not Solvable at the Application Layer

Power outages. Connectivity loss. Physical damage. These are not conditions that clever application architecture can work around if your infrastructure is entirely contained within the affected geography. No amount of microservices design, caching strategy, or auto-scaling helps when there’s no power reaching the data center.

This is an important framing shift for engineering teams who own availability: some failure modes require infrastructure-layer responses, not code-layer ones.

The Compliance Gap: Controls on Paper vs. Controls in Practice

Perhaps the most uncomfortable implication of this incident. In many environments — particularly those undergoing ISO/IEC 27001:2022 certification or SOC 2 audits — availability controls are documented but don’t reflect the actual system architecture. Redundancy is listed as a control. It’s just redundancy within a single region, which, as this event demonstrated, is insufficient for regional-scale disruptions. The control passes an audit. It fails a real incident.

This is the exact gap that compliance frameworks are designed to close — and that audit processes sometimes fail to catch.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Cloud Hosting and SOC 2 Compliance Requirements

Choosing AWS or Azure doesn’t hand you a SOC 2 compliance. It hands you a shared responsibility model, which means your provider secures the physical infrastructure and you secure everything running on top of it — including whether your architecture can actually deliver on your availability commitments.

Auditors know this distinction well. When they evaluate your Availability criteria, they’re looking at your controls, not your provider’s SOC 2 report.

What that means in practice: your recovery objectives need to be real numbers tied to a real architecture, not placeholders in a policy document. Your failover plan needs test records behind it. And your cloud provider should appear in your vendor risk register with an annual review of their own audit reports.

A single-region deployment with no tested failover isn’t compliant in any meaningful sense. It’s a documentation exercise waiting to be disproved.

The March 2026 incident made this concrete. Organizations that had documented availability controls but confined their entire infrastructure to one region found those controls counted for nothing when the region went down. The control passed the audit. It failed the incident.

That gap is exactly what a SOC 2 audit is supposed to catch. Sometimes it doesn’t. 

What Mitigating Controls Could Have Reduced the Impact

The following aren’t theoretical best practices. They’re the specific capabilities that separated organizations that stayed online from those that didn’t.

Multi-region deployment is the foundational requirement. Deploying systems across independent geographic regions — not just independent availability zones — means a regional disruption in one location doesn’t take everything down. Google Cloud’s documentation on multi-region architectures provides useful reference material on how this is structured in practice.

Cross-region data replication ensures that when failover happens, the secondary region has current data to work with. Replication lag is a design variable — it can be tuned based on acceptable recovery point objectives. What can’t be tuned is the existence of the replication relationship itself. If it isn’t there before the incident, it can’t help during one.

Automated failover removes the human response time variable from the equation. If traffic rerouting to a secondary region requires manual intervention, you are adding minutes or hours to your outage window during the exact moment when your team is most overwhelmed. Route 53 failover routing, Azure Traffic Manager, and equivalent tools in other clouds exist specifically for this scenario.

Regional outage testing is the practice that most organizations skip. Simulating a full regional failure — not just a single AZ — validates whether recovery strategies actually work, not just whether they exist. The NIST SP 800-34 guide on contingency planning recommends testing at the scenario level, not just the control level.

Dependency resilience is the one that catches teams off guard. If your identity provider, monitoring stack, or secrets management system lives in the same region as your primary workload, your failover may not actually work — because the systems your application depends on to function are also offline.

Insider note: A common failure in multi-region DR testing is discovering that the authentication service doesn’t fail over cleanly, even when the application does. Audit your dependency chain before you test — not during.

Compliance Perspective: What the Frameworks Actually Require

This incident maps cleanly onto requirements that many organizations are already accountable for.

ISO/IEC 27001:2022 addresses this directly across several controls. A.8.14 covers redundancy of information processing facilities — and the intent is effective redundancy, not documented redundancy. A.8.13 covers backup, with an expectation that backup data is accessible when primary systems are not. A.5.30 addresses ICT readiness for business continuity, which includes planning for scenarios beyond localized failure. The standard is explicit that controls must be implemented in a way that is proportionate to the risk — and a single-region deployment for a mission-critical application is a risk the standard expects to be addressed.

Unsure whether your current architecture actually satisfies these controls? An ISO 27001 gap analysis is usually the fastest way to find out, and an internal audit against your documented controls will surface the delta between what’s on paper and what’s in production.

SOC 2 Availability Criteria requires that systems are available in line with commitments and expectations. If your service-level commitments assume high availability, and your architecture cannot deliver that when a region goes offline, you have a gap between your commitments and your design. The AICPA’s Trust Services Criteria are clear on this point: availability controls must reflect real-world capability, not aspirational architecture.

The common thread across both frameworks: compliance asks whether controls are effective, not just whether they’re present. This incident is a clear case study in what ineffective-but-documented redundancy looks like under real conditions.

What This Means for Your Organization

The March 2026 incident is not a cautionary tale about a distant edge case. It’s a practical reference point for evaluating your own architecture — right now, before you need it.

The questions worth asking are direct ones. Can your systems operate if an entire cloud region becomes unavailable — not for five minutes, but for hours? Does your failover extend beyond a single region, or does it just move traffic between availability zones? Have you tested a full regional failure scenario, or only component-level failures? Do your compliance controls reflect actual system architecture, or how it was originally designed two years ago?

If the answer to any of those is uncertain, that uncertainty is the finding.

NIST‘s Cybersecurity Framework is also worth revisiting in this context — specifically the “Recover” function, which provides a structured way to think about resilience planning at the organizational level, not just the infrastructure level.

Conclusion

The March 2026 incident made one thing concrete: availability is not defined by the presence of redundancy within a region — it’s defined by the ability to operate beyond it.

Multi-AZ architecture is good design. It protects against the failures it’s designed to protect against. But it was never intended to be a substitute for multi-region resilience, and organizations that treated it as one found out the hard way. For most organizations, closing this gap doesn’t require rebuilding from scratch. It requires an honest assessment of where your architecture actually stands versus where you assumed it did.

Axipro works with scaling software companies to assess availability architecture, close compliance gaps, and ensure that continuity controls hold up under real-world conditions — not just audit conditions. If the questions raised in this article surfaced something worth investigating in your own environment, reach out to our team to schedule a technical review. Or if you’d prefer to start with a self-assessment, learn more about how we approach availability and compliance readiness.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Axipro Author

Picture of Abeera Zainab

Abeera Zainab

Blog Highlights

Explore More Articles

Compliance software collects the evidence. A consultant builds the system that evidence is meant to prove. That’s the real difference in the ISO 27001 consultant vs software decision, and most teams only figure it out after they’ve bought one and realized they still need the other. Below, we compare what each route covers, where it breaks down, and what it costs you in time, money, and your team’s hours. Short version: software on its own works for a small group of companies. For most SaaS and tech scale-ups trying to get an enterprise deal over the line, consultant-led implementation on a compliance platform is the faster and safer path to a certificate. Quick Answer: Consultant, Software, or Both? Software-only works if you already have an in-house security lead who’s taken a company through ISO/IEC 27001 before and has the time to own the project. Consultant-only still makes sense if you run mostly on-premise or legacy systems that platforms barely integrate with. For everyone else, which means most cloud-native companies under a few hundred people, a hybrid works best: a platform to handle evidence and monitoring, and a consultant to build the management system and stand behind it in front of an auditor. Here’s why. What an ISO 27001 Consultant Handles ISO/IEC 27001:2022 is a management system standard. Clauses 4 to 10 cover how you run information security, and Annex A lists 93 controls you pick from based on risk. Almost none of it is box-ticking. Most of it comes down to judgment calls about your business, and that’s what you’re paying a consultant for. Scoping, Gap Analysis and Risk Assessment Scope is the first decision you make, and the most expensive one to get wrong. Go too wide and you’ll spend months on controls for systems no customer asks about. Go too narrow and the certificate won’t get through the procurement review it was supposed to pass. A consultant scopes around the deals you’re trying to close, runs a gap analysis, and builds a risk assessment based on your real assets and threats. That’s the document auditors dig into hardest. ISMS Documentation and Policy Writing The standard asks for a specific set of documents: the ISMS scope, information security policy, risk assessment and treatment methodology, Statement of Applicability, risk treatment plan, and evidence of competence, monitoring, internal audit, and management review. A consultant writes these around how your company works day to day, instead of how a template imagines it works. Auditors check whether you follow your own procedures, so a mismatch shows up fast. Internal Audit and Certification Audit Support You need an internal audit before certification, and Clause 9.2 says the auditor has to be objective and impartial. In a small company, the people who built the ISMS can’t credibly audit it, so most teams outsource it through ISO 27001 internal audit services. A good consultant also gets your team ready for the Stage 1 and Stage 2 audits, joins the conversations that matter, and handles corrective actions if the auditor raises nonconformities.  What ISO 27001 Compliance Software Handles Compliance automation platforms, often called GRC platforms, have changed how cloud-native companies get certified. They’re very good at the repetitive, evidence-heavy side of the work. Automated Evidence Collection and Continuous Control Monitoring The platform plugs into your cloud provider, identity provider, code repos, HR system, and device management tools, then pulls evidence on its own. It’ll flag an unencrypted storage bucket, an ex-employee who still has access, or a laptop without disk encryption. For technical controls, that saves weeks of screenshots and spreadsheet tracking. Policy Templates and Annex A Control Mapping Most platforms come with a policy library and map each control to the ISO 27001 clauses and Annex A. You get a starting point and a clear view of which controls have evidence and which don’t. Auditor Access and Ongoing Compliance Tracking Auditors can log in and review evidence themselves, which cuts down fieldwork. After you’re certified, dashboards show when controls slip between surveillance audits, so you aren’t rebuilding evidence from scratch every year. Where Each Approach Falls Short Neither route covers everything by itself. The good news is that the ways each one fails are predictable, so you can plan around them. Limits of Compliance Automation Platforms A platform can tell you a control is failing. It can’t decide your scope, run your risk assessment, write a policy that matches your operations, convince your CTO to change the offboarding process, or explain to an auditor why you excluded a control from your Statement of Applicability. Templates can also make you feel further along than you are. A dashboard at 90% can hide an ISMS that won’t survive Stage 1, because the missing 10% is the management system itself. Insider Note: The Stage 1 problem we see most on software-only projects is a risk assessment copied straight from the platform’s default risk library. The risks are generic, the scores are almost identical, and nothing ties back to the company’s own assets. Auditors notice within minutes, and it weakens the Statement of Applicability that’s built on it. The other problem is ownership. Software assumes someone inside the company will drive the project. At most startups that’s a CTO or ops lead who already has a full-time job, and the subscription renews whether the work gets done or not. Limits of a Consultant-Only Approach A consultant working without automation spends billable days on things a platform does for free, like chasing screenshots, updating evidence trackers, and collecting the same proof again before every surveillance audit. You pay more and wait longer. You also end up with a program that’s only accurate on the day it’s handed over. Once the engagement ends, the evidence goes stale and year-two surveillance turns into a scramble. ISO 27001 Consultant vs Software: Side-by-Side Comparison Factor Consultant only Software only Hybrid (consultant + platform) Time to audit readiness 3 to 6+ months Highly variable; depends on internal expertise As little as 6 weeks for well-scoped

Uzbekistan regulates artificial intelligence through two documents. The first is Law ZRU-1115, signed on 21 January 2026. It amends existing legislation to define AI, stops anyone from basing decisions about people’s rights on AI output alone, and fines companies that process personal data unlawfully with AI. The second is the set of Ethical Rules approved by Order No. 3787, in force since 17 June 2026, which spell out what developers, implementers, and users actually have to do. Uzbekistan hasn’t passed a standalone AI act, and its rules don’t sort systems into risk tiers or require conformity assessments. The framework is short and blunt, and it’s already enforceable. Below we walk through what each document requires, who it applies to, how it stacks up against the EU AI Act, and what a company using AI in Uzbekistan should do next. Uzbekistan AI Regulation at a Glance (TL;DR) Instrument Date What it does Who it binds Law ZRU-1115 Signed 21 January 2026 Defines AI in law, sets general rules for AI-built information resources and systems, bans legally significant decisions based only on AI, adds fines for unlawful AI processing of personal data State bodies, organizations, website owners, anyone processing personal data with AI Order No. 3787 (Ethical Rules) Registered 14 March 2026, in force 17 June 2026 Sets eight mandatory ethical principles and lists rights and obligations for developers, implementers, and users Individuals and companies developing, implementing, or using AI in Uzbekistan Law No. 1125 (Personal Data amendments) Adopted 26 March 2026 Limits data localization to biometric, genetic, and local telecom user data, and allows cross-border transfers under conditions Personal data operators, including AI providers AI Strategy until 2030 (RP-358) 14 October 2024 Sets national targets for AI adoption, infrastructure, and skills Government bodies What Is Law ZRU-1115? The law’s official title is a mouthful: “On making additions and changes to certain legislative acts of the Republic of Uzbekistan in connection with the regulation of relations arising from the use of artificial intelligence.” Put simply, it’s an amending law. Instead of creating a new AI code, it writes AI into laws that were already on the books. When It Was Signed and When It Took Effect The Legislative Chamber of the Oliy Majlis adopted the bill on 12 August 2025, and the Senate approved it on 1 November 2025. President Shavkat Mirziyoyev signed it on 21 January 2026. You can read the official text in Lex.uz, Uzbekistan’s national legislation database. The law set out the principles and the penalties. The day-to-day detail arrived later with the Ethical Rules, which came into force on 17 June 2026. For compliance planning, treat mid-June 2026 as the point when the whole framework started applying. Why Uzbekistan Amended Existing Laws Instead of Passing a Standalone AI Act Uzbekistan wants more AI, not less. Its national strategy sets numeric targets for adoption, investment, and local computing capacity, and a heavy EU-style act would have worked against them. So lawmakers kept it light. They defined AI, drew two hard lines (human control over decisions that affect people’s rights, and protection of personal data), and left the Ministry of Digital Technologies to fill in the rest through secondary rules. Businesses get less legal certainty, and the government gets to move faster. Which Laws ZRU-1115 Changes For businesses, two amendments matter most. The Law “On Informatization” (ZRU-560-II, 2003) now contains a legal definition of AI, a new article on using AI in information resources and systems, duties for website owners, and updated powers for the ministry in charge. The Code on Administrative Liability now includes an offense for processing and spreading personal data unlawfully using AI. The Legal Definition of Artificial Intelligence in Uzbekistan Under the amended Law “On Informatization,” AI is a set of technological solutions that imitate human cognitive functions, including learning on their own and solving problems, and that produce results on specific tasks comparable to what a person could do. That’s deliberately broad. It covers generative AI, machine learning classifiers, recommendation engines, and most agentic systems. The Ethical Rules add a narrower term, the AI system: software built on AI that can find, collect, store, analyze, process, evaluate, and use data, and make decisions on its own based on that data. If your product makes a decision from data, or shapes one, assume it counts. Key Rules Introduced by Law ZRU-1115 General Principles for Using AI in Information Systems and Resources The new article in the Law “On Informatization” starts from harm. Information resources created with AI, and information systems running on AI, must not harm people’s life, health, freedom, honor, or dignity, or violate their other inalienable rights. The standard is short and open-ended. It gives regulators something to enforce against without saying in advance what counts as harm. Principle-based rules like this deserve to be taken seriously precisely because the edges are undefined. Human Oversight: No Decisions on Rights and Freedoms Based Solely on AI Most coverage leads with this provision, and it’s easy to see why. When someone makes a legally significant decision that affects human rights and freedoms, they can’t rely only on conclusions produced by AI systems or AI-built information resources. AI can feed into the decision, but a person has to make it. That applies to loan denials, benefit eligibility, hiring rejections, licensing outcomes, and disciplinary action. In each case, someone needs to look at the AI output and own the final call. Insider Note: In AI governance engagements, teams rarely struggle to show that a review step exists. What they struggle to show is that the reviewer could disagree, and sometimes did. If a human clicks “approve” on every AI recommendation and nobody ever records an override, auditors will see automation with a signature on top. Build the override path and log when people use it, starting on day one. Powers of the Authorized State Body (Ministry of Digital Technologies) ZRU-1115 makes the Ministry of Digital Technologies the authorized state body for AI. Among its new jobs, it’s

You can get a SaaS company ready for a SOC 2 audit in six weeks, but you’ll feel every one of them. Most published timelines say three to six months. For a company with no project owner, no identity provider, and nothing written down, that’s about right. A cloud-native startup that already has the basics in place and can protect some time is a different story, and it can fit the work into six hard weeks. This plan walks through that route one week at a time. Each week has an owner, an hour estimate, and a clear test for when it’s finished. The free Google Sheet version turns the plan into a tracker you can hand out to owners and update in your weekly standup. Before you start, know what you’re signing up for. At the end of week 6 you’ll be audit-ready, which isn’t the same as holding a Type II report. Nobody can get you a Type II in six weeks. This is also the do-it-yourself route, and it takes a lot of hours. We’ll show you where those hours go and what the faster option looks like. Is Six Weeks Realistic for Your Company? Six weeks works when most of the plumbing already exists and your job is to formalize it, fill the gaps, and prove it all works. It falls apart when you’re building the foundations and documenting them at the same time. Go through this table honestly before you promise a customer a date. Six weeks is realistic if… Plan for 10 to 16 weeks if… Your product runs on a major cloud provider You host on-premise or across several data centers You already use an identity provider with SSO Every tool has its own login and password You have fewer than about 50 employees You have multiple offices, subsidiaries, or products in scope One named person owns the project with 10 to 15 hours a week Compliance is “everyone’s job,” so in practice nobody owns it An engineer can give you 15 to 20 hours in weeks 3 and 4 Engineering is fully committed to a launch You only need the Security criteria You need Availability, Confidentiality, or Privacy on day one Landing mostly in the right-hand column doesn’t mean you should throw the plan out. Give each week two weeks instead of one and follow the same order. What “SOC 2 Ready” Means at the End of Week 6 SOC 2 doesn’t give you a certificate. An independent CPA firm examines your controls against the AICPA Trust Services Criteria and writes a report, and which of the two report types you go for decides what you can show a buyer after week 6. A Type I report checks whether your controls are designed properly on a single date. Once you’re ready, a Type I audit can start almost right away. A Type II report checks whether those controls kept working over an observation period of at least three months, and usually six to twelve. Most enterprise procurement teams want Type II in the end. Being “ready” at the end of this plan means your in-scope controls are in place, you can pull evidence for any of them on request, and your auditor is booked. From there you either start a Type I audit or open your Type II observation window. Plenty of buyers will sign with a Type I report plus a letter from your auditor saying the Type II period is underway. Important: The Type II clock doesn’t start until your controls are running. If readiness slips by a week, your Type II report slips by a week too. Founders who tell a prospect “we’ll have SOC 2 in Q3” often forget this and end up renegotiating the deal. Before Week 1: Four Decisions to Make First Settle these before the clock starts. If you change any of them halfway through, you’ll redo work. Scope. Decide which systems, teams, and data the report covers. For most SaaS companies that’s the production environment, the code repository, the identity provider, customer data stores, and any support tools that touch customer data. Corporate systems that never see customer data can usually stay out. Trust Services Criteria. Security (also called the Common Criteria) is mandatory. Availability, Confidentiality, Processing Integrity, and Privacy are optional. Report type. Pick Type I if a deal is blocked right now and the buyer will accept it. If there’s no deadline, go straight to Type II. You’ll need it eventually, and skipping Type I saves you an audit fee. Owner and tooling. Name one person who’s accountable for the plan, and decide where your controls and evidence will live. The tooling choice gets its own section below. Pro Tip: Adding Criteria Only add optional criteria when a customer contract or security questionnaire asks for them. Each one brings more controls to set up and more evidence to collect, and you can widen the scope in next year’s audit. Spreadsheet or Compliance Software: Choosing Your Tracking Tool Every SOC 2 program needs a system of record, meaning one place where each control, its owner, its status, and its evidence live. You can run it yourself in a spreadsheet or a GRC platform, or have a consultant implement it for you. The right choice depends mostly on which report you’re after and how much of your team’s time you can spare. A spreadsheet is free and familiar. It also makes you understand your own environment before you automate any of it. For a Type I, or for a small team with a tight scope, a well-built spreadsheet can take you all the way to the audit. Axipro’s free GRC workbook for SOC 2 and ISO 27001 covers all 33 SOC 2 Common Criteria plus the optional criteria, with evidence, risk, policy, and gap trackers built in. It has no macros and opens straight in Google Sheets or Excel. A GRC platform connects to your cloud, identity provider, code repository, and HR system.