Resilience is often treated as a backup question. Backups matter, but by the time they are needed the outcome is largely determined. The more useful question is whether a disruption stays contained or spreads.
Know which dependencies actually matter
Every organization depends on services it does not control: identity providers, payment processors, cloud regions, DNS, email delivery, CI pipelines. Few have a clear view of which of those would halt operations if unavailable for a day.
Mapping this does not require a formal exercise. Listing the handful of services that would stop you serving customers, and noting what the fallback is for each, typically reveals one or two dependencies with no plan at all. Single points of failure in authentication and DNS are the ones most often missed.
Containment is an architectural property
Whether a compromise stays small depends on choices made long before it happens. Flat networks, shared administrative credentials across environments, and production systems reachable from general-purpose workstations all make lateral movement straightforward.
Separating environments, limiting which systems can reach which, and keeping administrative access distinct from day-to-day accounts does not prevent intrusions. It changes what an intrusion can turn into, which is usually the difference between an incident and a crisis.
Test recovery, not just backups
Backup jobs reporting success is not the same as a tested recovery. The gaps that surface during real restoration attempts are consistent:
- Restores have never been timed, so recovery expectations are guesses
- Backups are reachable using the same credentials that would be compromised in a ransomware scenario
- Configuration and infrastructure state are not backed up, only data
- The restore procedure depends on one person's knowledge
- Documentation needed during recovery is stored in the system being recovered
A single realistic restoration test per year surfaces more of these than any amount of documentation review.
Decide who decides
During a disruption, the delays that cost most are rarely technical. They come from uncertainty about authority: who can take a production system offline, who approves contacting customers, who speaks to a regulator, and at what point external help is engaged.
Answering those questions in advance takes an afternoon. For smaller organizations it is often the single highest-value resilience activity available, because it costs almost nothing and removes the hesitation that extends an outage.
Prefer boring, well-understood arrangements
Resilient environments tend to be unremarkable. Fewer moving parts, consistent patterns across services, changes that can be reversed quickly, and monitoring that someone actually watches. Complexity that a team cannot reason about under pressure is itself a risk, regardless of how well it performs on a normal day.
Related topics
Preparing for the Next Generation of Cyber Threats · Making Technology Risk Understandable to the Board