When a crisis strikes whether a regional infrastructure failure, a ransomware attack, or a catastrophic data centre event the organisations that survive intact are those that planned, tested, and automated before the incident. Resilience is not built in a crisis; it is built before one. Here is ACME’s Crisis Code: the essential actions to take, and the critical mistakes to avoid.
The Do’s: Strategic Actions for Resilience
1. Design for Automated, Multi-Site Recovery
Use a clear DR strategy Active/Hot for critical services, Warm/Cold for secondary systems. Automate end-to-end recovery using infrastructure-as-code and scripted failover. Restore networks, identities, and applications together, not in isolation.
2. Decouple Identity from Single-Site Dependencies
Reduce reliance on localised Active Directory by adopting cloud-native identity solutions such as Microsoft Entra Join, which maintains authentication during site outages. Enhance security with Privileged Access Management (PAM) for just-in-time access, session recording, and approval-based administrative controls.
3. Adopt Zero Trust and Enforce Strong Access Controls
Implement a Zero Trust model verify every user and device, enforce least privilege, and replace broad VPNs with Zero Trust Network Access (ZTNA). Apply MFA across all access points and continuously monitor for anomalies.
4. Isolate and Validate Immutable Backups
Store encrypted, immutable (WORM) backups in a completely separate network, account, and identity boundary from production. Protect against ransomware and regional infrastructure loss. Critically routinely validate backups with real restore tests, not just backup job confirmations.
5. Keep Runbooks Simple, Assigned, and Regularly Exercised
Maintain concise continuity and incident runbooks with clear ownership, escalation paths, and communication steps. Conduct joint exercises involving IT, security, application vendors, and business owners both tabletop and live failover drills to ensure the plan holds under real pressure.
The Don’ts: Common Pitfalls to Avoid
1. Don’t Build on Single Points of Failure
Avoid architectures where the loss of one cloud region, one data centre, or one Active Directory domain controller can halt your entire operation. Resilience must be designed in from the start, not bolted on after an incident.
2. Don’t Leave Privileged or Management Access Uncontrolled
Never expose RDP, SSH, hypervisor consoles, or admin portals to the public internet. Avoid broad, always-on VPN access and unmanaged shared administrator accounts. All privileged sessions must be routed through hardened, audited access paths.
3. Don’t Depend on Manual Steps in a Crisis
Long, undocumented checklists that require engineers to manually update DNS, firewall rules, or application configurations during a high-pressure incident directly increase downtime and errors. Automate everything that can be automated.
4. Don’t Store Backups in the Same Blast Radius
Never keep backups in the same region, tenant, or trust boundary as your production environment. If your primary environment is compromised by a cyber-attack or physical disruption your backups must be completely unreachable from the same zone.
5. Don’t Treat DR, Identity, and Security as “Set and Forget”
Failover automation, PAM policies, ZTNA configurations, and endpoint protection must be reviewed regularly. Untested failover paths and unmanaged break-glass accounts will silently become your greatest risk until a crisis exposes them.
Resilience Is a Practice, Not a Project
The gap between organisations that recover in minutes and those that are offline for days is not technology it is discipline. Building that discipline requires the right architecture, the right automation, and the right partner. ACME helps enterprises across the region move from reactive to genuinely resilient.