Recovery is a dependency graph
An application does not recover when one database comes back. Users may still need DNS, certificates, network routes, identity, secrets, object storage, queues, external APIs, and monitoring. Map each dependency, its owner, recovery location, prerequisite, and the condition that proves it is usable.
NIST contingency-planning guidance supports a staged approach: identify critical functions, recovery strategies, roles, tests, and maintenance. Translate that into the actual platform rather than a generic checklist.
Build a recovery map
Start at the user request and follow it through edge routing, identity, application, data stores, asynchronous services, storage, and external providers. Mark which components are internally restored, which are supplier-managed, and which have no viable recovery route. Include certificate issuance and secrets; they often block a technically restored workload.
For every dependency, note recovery objective, backup or rebuild method, credentials, owner, test evidence, and last change. This produces an order of operations instead of an inventory that cannot guide a pressured recovery.
Choose an order that can be tested
Recover foundations first: networks and identity paths needed by operators, storage and data required by applications, then application services and public routing. The exact order depends on the system. A service may need its database before its API; an operator may need DNS and a certificate before a realistic sign-in test is possible.
State dependencies that cannot be restored by your team. A third-party identity provider, domain registrar, or payment processor may have its own status and escalation process. Recovery plans should show how the service behaves when that dependency is unavailable and who communicates the limitation.
Illustrative restore exercise
For an illustrative internal portal, restore a database copy, object store sample, application configuration, identity integration, and DNS route in an isolated environment. Then use a standard account to sign in, view an authorised record, and run a safe administrative health check.
Record timing and failures. If a certificate or application secret was missing, that is an important recovery finding. Add it to the map and rerun the relevant stage rather than declaring success because a database snapshot restored.
Dependency recovery register
Keep this connected to the service change record.
| Dependency | Recovery responsibility | Acceptance check |
|---|---|---|
| Identity | Identity owner | A standard account can authenticate. |
| Data store | Platform or database owner | Expected records are available within the agreed point-in-time boundary. |
| DNS and certificate | Network or platform owner | Users reach the intended endpoint securely. |
| Object storage | Storage owner | An authorised file reference resolves correctly. |
Failure cases to rehearse
What if the backup is available but the decryption key is not?
The key path is a dependency. Include key custody, recovery access, and a controlled test in the plan.
What if identity is healthy but the application rejects users?
Check application configuration, claims or groups, DNS, certificates, time synchronisation, and network policy; the sign-in path crosses several boundaries.
How often should exercises run?
Set the interval according to change rate, service importance, and recovery objectives, then repeat after material changes to dependencies.
Plan, exercise, maintain
Map dependencies, objectives, owners, prerequisites, recovery methods, and communication routes.
Restore in isolation, test a real user journey, record timing and failures, and do not expose production data unnecessarily.
Update the map after architecture, identity, provider, certificate, storage, or application changes.
Recovery readiness checks
A written plan is not proof without evidence.
Dependency map
The map includes data, identity, DNS, certificates, network, storage, and external providers.
User-journey test
An isolated exercise proves a user and an operator can complete the agreed recovery checks.
Known limitations
Supplier dependencies, recovery gaps, and accepted loss boundaries are stated clearly.

