A banking client in the Middle East suffered 52 hours of unplanned downtime during a regional cloud outage — not because of the outage itself, but because of how their disaster recovery was designed. This is what we found, what we rebuilt, and what every organisation running critical systems should learn before it happens to them.
What Actually Happened
When a regional availability zone disruption hit, the bank’s primary production systems went offline. That part was expected — outages happen. What was not expected was that the DR environment also failed to activate. The investigation revealed four compounding failures that turned a manageable incident into a 52-hour crisis.
First, the DR site was in the same availability zone as production. A regional event took both environments down simultaneously — the redundancy was an illusion. Second, the bank’s Active Directory domain controllers were hosted on-premises, in the same physical facility. Without AD, there was no authentication, and without authentication, no systems could come online even when connectivity was partially restored.
Third, the most recent backup set was co-located with the primary data — unavailable at exactly the moment it was needed most. Fourth, and perhaps most damaging in the long run, the failover runbook consisted of 47 manual steps, required three people with specific institutional knowledge to execute, and had never been tested under real conditions.
“52 hours offline is not just a technology failure. It is a governance failure. The DR architecture had never been validated end-to-end, and the runbook existed on paper, not in practice.”
ACME Resilience Practice, post-incident review
Five Root Causes We Identified
After the incident, ACME conducted a full resilience assessment. The five root causes that made 52 hours of downtime possible were:
- Single-region architecture with no true geographic separation. Placing DR in the same availability zone as production provides protection against hardware failure, not regional events. For critical financial systems, geographic separation across regions is non-negotiable.
- Identity infrastructure tied to on-premises hardware. On-premises Active Directory is a single point of failure. When the facility becomes unavailable, so does every system that depends on it for authentication — which is typically everything.
- A manual runbook that had never been drilled. A 47-step manual process requiring specific expertise is not a failover plan. It is a set of instructions that will fail under the time pressure, stress, and partial information that characterises a real crisis.
- Backups co-located with primary data. Co-located backups are not backups in any meaningful sense of the word. They are a duplicate that fails at the same time as the original.
- Legacy VPN dependent on the same failed infrastructure. Remote access for the recovery team relied on VPN infrastructure that was itself affected by the outage. Engineers could not access systems remotely to begin recovery work.
ACME hosted an exclusive customer event on data resilience with Veeam and Object First — covering backup architecture, immutable storage and business continuity for enterprise clients.
The Rebuilt Architecture
Working with the bank’s IT and security leadership, ACME redesigned the DR architecture from the ground up. The objective was not merely to fix what had broken, but to build a resilience posture that could demonstrably meet a 90-minute RTO and 15-minute RPO for critical workloads.
The key components of the rebuilt architecture:
- Multi-region deployment (Hot + Warm). Critical banking workloads are now replicated in near-real-time to a separate AWS region. The Hot environment handles active failover for Tier 1 systems; Warm standby covers Tier 2 and 3 workloads. Regional separation ensures that no single event can affect both environments.
- Microsoft Entra ID replacing on-premises Active Directory. Identity is now cloud-native, distributed and regionally resilient. Authentication no longer depends on physical infrastructure in a single location. This alone eliminated the most critical single point of failure.
- Privileged Access Management (PAM) with just-in-time access. No standing privileged accounts. Engineers receive time-limited, least-privilege access to specific resources when needed. This protects against credential compromise during a crisis — the worst possible moment to discover an account has been exfiltrated.
- WORM backups in a separate AWS account and region. Backup data is written once and cannot be modified or deleted, stored in a separate AWS account with cross-tenant replication to a geographically distinct region. Even a full account compromise cannot destroy the backup chain.
- Zero Trust Network Access (ZTNA) replacing legacy VPN. Remote access for engineers and recovery teams now operates on ZTNA principles — identity-verified, device-assessed, and application-specific. Access survives infrastructure failures that would have grounded the recovery team in the original incident.
- Single-command scripted failover. The 47-step manual runbook was replaced with a fully automated failover script. End-to-end failover from production to DR now completes in under 90 minutes, with all steps logged, auditable, and testable. The script is drilled quarterly under production-equivalent conditions.
The ACME Crisis Code
Across every resilience engagement we conduct, the same patterns of failure repeat. We have distilled them into a simple reference — five things organisations must do, and five they must never do.
DO
- Test your DR runbook under real conditions at least twice a year. If it has never been executed end-to-end, it does not exist.
- Separate your DR environment by geography — availability zone redundancy protects against hardware failure, not regional events.
- Ensure identity services are cloud-native and regionally resilient. On-premises AD is your highest-risk single point of failure.
- Store backups in a separate AWS account with WORM protection. Cross-region, cross-account, cross-tenant is the only configuration that survives catastrophic events.
- Define your RTO and RPO in writing, present them to your board, and then validate them with a live drill. Numbers without validation are fiction.
DON’T
- Co-locate primary and DR in the same region or physical facility. Same-zone DR is not DR.
- Depend on manual steps during a crisis. Automate the failover. Humans under pressure at 2am make mistakes; scripts do not.
- Assume a DR site that has never been tested will activate cleanly. It will not.
- Leave privileged access standing. Just-in-time, just-enough access is not a best practice — it is the minimum standard for any regulated environment.
- Treat disaster recovery as an IT project. It is a board-level risk item. If your executive team cannot articulate your RTO, your RTO is not real.
ACME Resilience Assessment — In three days, our team will map your current DR posture, identify critical gaps across infrastructure, identity, backup and runbook readiness, and deliver a prioritised architecture recommendation. Available for enterprise and financial services clients across the Middle East. Request an assessment →