RTO vs RPO Decision Matrix
A simple guide to choose when to make RTO tighter, when to make RPO tighter, and what ranges fit common system types. Short sentences. Clear terms.
1. RTO and RPO in plain English
RTO = Recovery Time Objective. It answers: “How long can we be down?”
RPO = Recovery Point Objective. It answers: “How much data can we lose?” We measure that loss as time since the last good backup or replica.
Example: RTO = 2 hours means you aim to bring the system back within 2 hours. RPO = 15 minutes means you accept losing up to 15 minutes of new data.
For deeper background, read Disaster Recovery Basics: RTO vs RPO. To turn targets into money, use the RTO/RPO Impact Calculator.
2. When to tighten RTO vs RPO
“Tighter” means a smaller number (faster recovery, or less data loss). Tighter targets cost more to build and run. Spend where impact is highest.
- Tighten RTO when every minute offline hurts sales, safety, or trust— even if you can rebuild the data later.
- Tighten RPO when lost records are hard or expensive to recreate— payments, orders, ledgers, regulated data.
- Tighten both for systems that take money and must stay online (for example, checkout with payment capture).
- Keep both loose for low-impact tools (internal wiki, draft analytics) so budget goes to critical systems first.
3. Decision matrix by system type
These ranges are planning starters—not legal requirements. Adjust for your industry, peak traffic, and compliance rules.
| System type | Suggested RTO | Suggested RPO | Why |
|---|---|---|---|
| Customer-facing checkout | 5–30 minutes | 0–5 minutes | Shoppers leave quickly. Lost carts and orders hurt revenue and trust. Prioritize fast restore and near-zero data loss. |
| Payments / ledger | 15–60 minutes | 0–1 minute (near-zero) | A short outage may be OK if no money movement is lost. Data integrity usually matters more than speed alone. |
| Analytics / reporting | 4–24 hours | 1–24 hours | Dashboards can wait. Pipelines can often replay from source systems. Do not overspend here first. |
| Internal wiki / docs | 8–72 hours | 4–24 hours | Work slows, but revenue rarely stops. Daily (or better) backups are often enough. |
| Email / notifications | 1–4 hours | 15–60 minutes | Users notice quickly. Lost password resets and OTP messages hurt trust. Queues with retries can soften RPO if sends are idempotent. |
| CI / build pipeline | 2–8 hours | 1–4 hours | Deploys pause; customers may still be fine. Artifact loss hurts more than a short outage — protect build history. |
| Identity / SSO | 5–30 minutes | 0–15 minutes | When login dies, many apps die with it. Treat identity as a shared critical path. |
How to use the matrix
1. List your systems.
2. Match each one to a row (or a close mix).
3. Set draft RTO/RPO.
4. Check cost with the formulas below.
5. Spend first on the systems with the highest yearly impact.
4. More scenarios (how teams actually decide)
The matrix is a starting map. Real choices mix money, compliance, and how hard data is to rebuild. Use these short scenarios to pressure-test your draft targets.
Scenario 1 — Flash sale storefront
A retail site runs a two-hour campaign. Every minute offline burns ads and carts. Tighten RTO hard (often under 15–30 minutes). Also keep RPO near zero for orders and payments. Analytics can wait until after the sale.
Scenario 2 — B2B SaaS with overnight batch jobs
Customers work 09:00–18:00. A two-hour outage at 02:00 may be annoying but not fatal. A two-hour outage at 10:00 is different. Set a daytime RTO tighter than a nighttime RTO if your contract allows tiered targets. Keep RPO tight for tenant configuration and billing records even at night.
Scenario 3 — Healthcare appointment system
Staff can sometimes book by phone during a short outage (RTO can be moderate). Losing yesterday’s schedule updates is dangerous (RPO must stay tight). Compliance rules may force both targets lower than a pure money model suggests.
Scenario 4 — Data warehouse for executives
Dashboards can be hours late. Prefer replaying from source systems over expensive synchronous replication. Put budget into source OLTP reliability first. Loose RTO/RPO here is often the correct call.
Scenario 5 — Marketplace with two-sided chat
Search can degrade. Chat and order state cannot silently drop messages. Split the product: looser targets for browse cache, tighter RPO for messaging and orders. One company-wide RTO of “1 hour” hides this split and wastes money.
Scenario 6 — Regulated ledger, low traffic
Few users, but every posting must be correct. You might accept a longer RTO if a read-only mode covers audits. You almost never accept a loose RPO. Spend on synchronous or near-sync replication before you spend on fancy multi-region active-active.
Scenario 7 — Startup with one engineer on call
Ambitious targets without people fail. Pick RTO/RPO you can rehearse. A documented 4-hour restore you have practiced beats a paper 15-minute target nobody can hit. Revisit targets after you hire or buy managed failover.
5. Annualized cost formulas
These are the same formulas used on TechImpact.online. They turn targets into a yearly estimate.
RTO annual loss = rto_hours × incidents_per_year × revenue_per_hour
RPO annual cost = rpo_hours × incidents_per_year × rebuild_cost_per_hour
Total annualized impact = RTO annual loss + RPO annual cost
revenue_per_hour = money you lose (or fail to earn) while the system is down. rebuild_cost_per_hour = cost to recreate or recover one hour of lost data (people time, vendor fees, manual re-entry).
6. Worked examples
Example A — Tighten RTO (checkout)
System: Customer-facing checkout
Current RTO: 4 hours · Incidents/year: 2 · Revenue/hour: $8,000
RTO yearly impact = 4 × 2 × 8,000 = $64,000
If you cut RTO to 1 hour: 1 × 2 × 8,000 = $16,000
Saving: $48,000 per year (before counting RPO).
Checkout also needs a tight RPO. Lost orders are hard to rebuild from memory.
Example B — Tighten RPO (payments)
System: Payments ledger
Current RPO: 2 hours · Incidents/year: 3 · Rebuild cost/hour: $5,000
RPO yearly impact = 2 × 3 × 5,000 = $30,000
If you cut RPO to 0.25 hours (15 minutes): 0.25 × 3 × 5,000 = $3,750
Saving: $26,250 per year on data rebuild exposure.
Here RPO is the priority. A brief outage with zero lost payments may be better than long uptime with missing transactions.
Example C — Combined total
RTO = 2 hours · RPO = 1 hour · Incidents = 3/year
Revenue/hour = $5,000 · Rebuild/hour = $2,000
RTO yearly = 2 × 3 × 5,000 = $30,000
RPO yearly = 1 × 3 × 2,000 = $6,000
Combined estimate: $36,000 per year
Example D — Identity outage (shared dependency)
System: Company SSO used by 12 apps
RTO: 0.5 hours · Incidents/year: 2 · Revenue/hour at risk: $20,000
RTO yearly = 0.5 × 2 × 20,000 = $20,000
The “system” looks small, but blast radius is large. Model revenue for all dependent apps, not only the login service bill.
Example E — Analytics can wait
System: Executive dashboard warehouse
RTO: 24 hours · RPO: 12 hours · Incidents/year: 4
Revenue/hour ≈ $200 (mostly delayed decisions) · Rebuild/hour ≈ $150
RTO yearly = 24 × 4 × 200 = $19,200
RPO yearly = 12 × 4 × 150 = $7,200
Combined: $26,400 — often less than the cost of making this tier “four nines.”
Run your own numbers in the RTO/RPO Impact Calculator.
7. Common pitfalls
- One target for everything. Checkout and the wiki should not share the same spend.
- Confusing SLA with RTO/RPO. A cloud uptime SLA is about provider credits. Your recovery objectives come from your design and drills.
- Paper targets without restores. If you never restore a backup, your RPO is a wish.
- Ignoring dependencies. DNS, identity, payments, and a single region can dominate risk.
- Optimizing only RTO. Fast empty databases still fail the business if RPO was ignored.
8. FAQ
Should one RTO/RPO cover the whole company?
No. Set targets per system or workflow. Checkout and the wiki do not need the same spend.
Does a cloud SLA guarantee my RTO or RPO?
No. An SLA is about provider credits for their uptime promise. Your RTO/RPO come from your architecture, backups, and runbooks.
Does RPO = 5 minutes mean backups every 5 minutes?
Yes—or continuous replication that can recover to a point no older than 5 minutes.
Where can I learn more?
Start with Disaster Recovery Basics: RTO vs RPO, then model cost in the RTO/RPO Impact Calculator.