Could we keep operating if a critical system or provider failed?

Judge readiness on tested recovery of your most critical services, measured in hours, under realistic conditions. Duplicated sites, extra suppliers and vendor guarantees count only when their dependencies are independent and the restore has been shown. Invest first in what would shorten an outage.

Draft for Tamir's review. Not published.

Tamir Khason · Updated · Decision guides

Test recovery under the conditions that matter

Define which systems the business needs back first and how long it can run without each. Then run a full restore of those systems from current backups and time it. Judge any backup vendor on a timed restore of a representative system of yours.

Exercises that pass under convenient conditions prove little. Check which people, credentials and supplier access the exercise assumed were available. Ask for conclusions that separate demonstrated recovery from untested assumptions.

Look for shared dependencies behind the redundancy

A second recovery site or a second processor helps only if it doesn't share the failure. Map identity, naming, keys, network and settlement dependencies for each path. Test switching while the shared dependency is unavailable.

Backups need isolated credentials and consoles, and they must cover the configuration and identity systems needed to rebuild. Read a recovery guarantee's conditions and exclusions, and what the vendor actually pays if it fails.

Depending on your seat

If you're on the board, ask for the downtime procedures by department, when they were last drilled and the last tested restore time. Where everything runs on one cloud provider, ask what a multi-day failure would do and what exit would take. Then record the risk you accept and the mitigations you require.

If you run technology after an outage, write the timeline to the minute from the change to full recovery. Separate the trigger from the reasons detection and recovery were slow. Compare each investment option on how many minutes it would have removed from this outage.

What to check before you decide

  • Name the systems to restore first and the time the business can tolerate for each.
  • Run a full restore of those systems from current backups and time it.
  • Check which people, credentials and supplier access your last exercise assumed were available.
  • Map the dependencies shared by your primary and recovery paths, and test with them unavailable.
  • Check that backup credentials and consoles are isolated from the systems an attacker would control.
  • Ask what the provider contract gives you on failure and on terms changes, and what exit would take.
  • Set a drill and restore schedule with results reported to the board.

Questions people ask

Could our hospital keep treating patients through a long systems outage, what should the board ask?

Ask for the downtime procedures by department, when they were last drilled, how long the hospital can run on them, and how long restoring critical systems takes as tested. Security prevents; downtime readiness is what protects patients when prevention fails. It depends on which systems clinical care depends on most and on when the drills and restore tests last happened.

Our systems run entirely on one cloud provider, what should the board require about concentration risk?

Require a view of what depends on the provider, what a multi-day failure would do to the business, what the contract gives in such a case, and what exit would take. Then decide the risk the board accepts and what it wants mitigated, such as backups elsewhere for critical data. It depends on the services used and on how much is provider-specific.

Should boards trust continuity exercises that avoid the hardest dependency failures?

Require exercises that challenge material dependencies within a safe agreed boundary. Reliance depends on realistic assumptions, observed recovery capability, and transparent limitations rather than a pass label.

After a payments outage the review blames both vendor and an internal change, how do we decide what resilience to invest in?

Build the timeline to the minute, because most outages start with a change and get long because of detection and recovery gaps. Invest first in what shortened the outage and second in what would have prevented it. It depends on whether the vendor component failed alone or only after your change.

Backup vendors promise guaranteed ransomware recovery, how do I choose and what should I test?

Judge backup on a restore test of your most critical systems, measured in hours, rather than on the vendor's guarantee. Immutability matters, and so do isolation of the backup credentials and the ability to rebuild from nothing. It depends on which systems the business needs back first and on how long it can run without them.

Should passenger services and transport operations share critical technology dependencies?

Evaluate the shared dependency through failure scenarios. Suitability depends on which functions must continue independently and whether separate recovery access actually works.

Does a second recovery site provide enough protection for municipal services?

Assess independent recovery by dependency. Site count tells you little. Readiness depends on naming, identity, keys, network control, and access when primary services are unavailable.

Does adding another payment processor actually reduce our outage risk?

Assess independence at the failure modes that matter. Multiple contracts do not establish resilience when network, identity, settlement, or operational dependencies remain shared.

How I can help with this decision

Ask or talk (Free)
I give my view on the questions that show your real readiness and what good answers look like. I tell you what to test before spending, and which parts of recovery vendors' guarantees usually exclude.
Review (Pay if it was worth it)
I write an independent assessment of your recovery readiness, shared dependencies and provider concentration, or a post-incident review. I recommend what to fund, what to decline and the test routine to keep.
Retain (When it makes sense)
I stay close to review drills, restore tests and changes to your providers against the timeline and the mitigations you agreed.