Business Continuity Plan vs Disaster Recovery Plan: The Difference Exams Actually Test
Most candidates lose marks on this topic without ever getting a definition wrong.
The exam rarely asks what a business continuity plan is. It gives you a scenario, a qualifier like BEST or FIRST, and four options that are all defensible. Then it waits to see whether you reach for the technical answer or the business one. Reach for the wrong one and you have demonstrated exactly the instinct the question was built to detect.
On CISA this sits in Domain 4, Information Systems Operations and Business Resilience — 26% of the exam, the joint-largest domain. It is not a corner of the syllabus you can afford to be approximate about.
The distinction in one line
A business continuity plan keeps the business running during a disruption. A disaster recovery plan restores the technology the business runs on.
That will carry you through a good number of questions. What it hides is the part exams actually probe: scope and ownership. The BCP is enterprise-wide and owned by the business. The DRP is a technical sub-plan, owned by IT, built to satisfy objectives the business already set.
Business defines, IT delivers. Get the direction of that relationship right and most awkward items resolve themselves.
What the BCP actually covers
Not servers. Processes — the activities that earn revenue, serve customers, and keep the organisation legally compliant.
A mature BCP deals with critical business functions and everything they depend on: people, facilities, suppliers, and only incidentally systems. It covers alternate work locations and remote arrangements, cross-trained staff, manual workarounds, backup suppliers. It says who declares an incident and who is permitted to speak to a regulator. It establishes decision rights when the usual decision-maker is unreachable.
Very little of that is an IT concern. If the data centre is untouched but the building is inaccessible and half the staff cannot work, nothing in the DRP helps you.
Everything traces back to the BIA
The business impact analysis identifies critical processes, quantifies what losing them costs over time, and establishes how long the organisation can tolerate that loss. Its outputs — recovery priorities and time objectives — become the requirements every downstream plan must meet, including the DRP.
The sequencing is what gets tested. The BIA comes first. Recovery strategies, site selection, backup schedules and technology spend are all answers to questions the BIA asks, so any option that picks an alternate site or sets a backup frequency before the impact analysis is out of order.
One adjacent distinction worth holding firmly, because it appears constantly: a risk assessment asks what could go wrong and how likely it is. A BIA asks what happens to the business if it does, and for how long that can be endured. An item that hinges on impact over time is a BIA item.
What the DRP actually covers
The technical detail. Recovery procedures per system in priority order, with dependencies mapped — you cannot restore the application before the directory service it authenticates against. Backup and replication arrangements, and where the media or replicas actually live. Alternate processing facilities. Team assignments and escalation. Restoration procedures for returning to the primary environment once it is safe.
A DRP can hit every technical target and still leave the business dead in the water, if the people who use those restored systems have nowhere to sit and no instructions. That gap is the BCP's job.
The metrics that join them
This is where business and technical language meet, and where marks are won.
| Metric | What it measures | Set by |
|---|---|---|
| RTO — Recovery Time Objective | Maximum acceptable time to restore a service | BIA / business |
| RPO — Recovery Point Objective | Maximum acceptable data loss, measured backward | BIA / business |
| MTD / MTPD — Maximum Tolerable Downtime | Total time before losses become unacceptable | BIA / business |
| WRT — Work Recovery Time | Validating data and resuming normal processing after systems are up | Business |
| SDO — Service Delivery Objective | Level of service required while running in recovery mode | Business |
Two things worth committing to memory.
RPO looks backward; RTO looks forward. RPO concerns data already created — a four-hour RPO means you can afford to lose four hours of transactions. RTO concerns time still to elapse — a four-hour RTO means you must be operating again within four hours. Under time pressure these blur, so anchor on the question: if it is about backup or replication frequency, it is RPO. If it is about how fast service must resume, it is RTO.
MTD = RTO + WRT. The systems being back is not the same as the business being back. WRT is the reconciliation and validation window between the two, and it is the metric candidates most often forget exists.
Tighter objectives always cost more, and the right strategy sits where the cost of downtime meets the cost of recovery capability. Spending more than the disruption would have cost is a control failure, not diligence. If an option proposes near-zero RTO for a non-critical process, it is usually wrong on that ground alone.
One incident, two plans
Definitions blur. Watch the plans do genuinely different jobs in the same event.
02:10 — Detection. Monitoring flags mass file encryption across core banking shares. This phase belongs to the incident response plan: contain, isolate, preserve evidence.
02:45 — Declaration. Containment means taking core systems offline, and the impact will exceed what the branch network can absorb before opening. The crisis management team convenes and formally declares a disaster. This is a BCP activity requiring pre-defined authority — and note that no technician made the call. Declaration is a business decision.
03:00 — Two workstreams, in parallel.
Under the DRP, IT fails over to the secondary data centre, confirms the replicated data predates the encryption, and works the recovery priority sequence: authentication first, then core banking, then reporting.
Under the BCP, branch managers activate manual procedures with pre-printed forms and per-customer transaction caps. Call centre staff get an approved holding statement. The communications lead notifies the regulator inside the mandated window. HR arranges relief shifts because this will outlast a working day. Finance confirms cash positions without the treasury system.
09:00 — Branches open in degraded mode. That is the SDO made real: not full service, but the minimum the business defined in advance. Customers are inconvenienced rather than abandoned.
16:00 — Core banking restored inside the eight-hour RTO, with twenty minutes of transactions lost against a thirty-minute RPO. Staff now begin reconciling the morning's manual transactions. That window is the WRT, and it is the part nobody budgets for.
Day 4 — Return to normal. The primary environment is rebuilt and hardened. Non-critical systems move back first, deliberately, to prove the restored environment is stable before anything critical is exposed to it.
Remove the DRP and nothing gets restored. Remove the BCP and the bank is technically recovered by four in the afternoon, having failed its customers, its staff and its regulator all morning.
Side by side
| Business Continuity Plan | Disaster Recovery Plan | |
|---|---|---|
| Objective | Sustain critical business functions | Restore IT systems, applications, data |
| Scope | Enterprise: people, processes, facilities, suppliers, technology | Technology infrastructure and data |
| Relationship | The parent plan | A subordinate sub-plan |
| Owner | Business; senior management accountable | IT or infrastructure management |
| Driven by | Business impact analysis | Objectives inherited from the BIA |
| Key metrics | MTD/MTPD, SDO | RTO, RPO, WRT |
| Failure looks like | Systems restored, business still cannot operate | Business ready, systems unavailable |
Testing
Both plans are worthless untested, and both climb the same ladder — cheapest and least disruptive first. A checklist review confirms the document is current. A structured walkthrough talks the team through roles. A tabletop injects a scenario and asks people to respond in real time without touching production. A parallel test brings recovery systems up and runs processing alongside production, comparing results, switching nothing over. A full interruption test actually stops production and makes the recovery capability carry the load.
Two points recur in questions. The ladder is a sequence — nobody runs a full interruption test on an untested plan. And a parallel test proves recovery works without risking production, which makes it the usual best answer when a question weighs assurance against operational risk. Full interruption is for items that explicitly demand the most conclusive validation available.
Six traps worth knowing
Life safety overrides everything. If a scenario involves any physical threat and an option covers evacuation or personnel safety, that is the answer. No volume of data or revenue outranks it. This is an absolute, not a tie-breaker — and it is the one place where an option containing an absolute is likely correct rather than a distractor.
Senior management is accountable, always. They approve the policy, fund the programme and own the residual risk. Execution can be delegated; accountability cannot. Options placing ultimate responsibility with IT, a coordinator, or a third-party provider are usually distractors.
Backups are not a disaster recovery plan. A backup is a control. The plan is what says what to restore, in what order, by whom, and to where. The related trap is assurance: an untested backup provides none. When a question asks for the greatest concern about a recovery capability, "never tested" and "not updated since the last major change" beat almost anything else on the list.
Governance precedes technology. Where a question offers a policy or ownership fix alongside a technical one, and the underlying gap is governance, the technical option loses however sound it looks. This is the single most reliable pattern in the domain.
Reciprocal agreements are cheap and unreliable. Rarely enforceable, capacity rarely verified, and the partner's own needs may collide with yours precisely when you need them. Attractive on cost, weak on assurance.
During return-to-normal, move the least critical systems back first. Counterintuitive until you see why: the point is to validate the rebuilt environment using something you can afford to have fail. Critical services move last.
Why examiners keep asking
The definitions are not difficult. The distinction persists on syllabi because it reveals a professional instinct: separating what the business needs from how technology delivers it, and keeping those in the right order.
Candidates who lead with the technical answer get marked down — not because the technology is wrong, but because it answers a question nobody has asked yet.
Preparing for CISA? Business resilience is Domain 4, the joint-largest on the exam. See the CISA exam prep guide, or read how hard the CISA exam really is.