Skip to content
SaaS Development4 July 2026 · 7 min readUpdated 30 September 2026

SaaS Disaster Recovery in 2026: RPO, RTO, and What Actually Matters

A SaaS disaster recovery plan starts with a restore you have timed. RPO and RTO in plain terms, Postgres PITR compared with daily backups and branch restore, and how to restore one tenant from a shared database.

SaaS Disaster Recovery in 2026: RPO, RTO, and What Actually Matters

If you run a multi-tenant SaaS and can't say when you last restored a backup, start there. This week, restore your most recent backup into a throwaway database and time how long it takes. That one exercise tells you more about your SaaS disaster recovery plan than any policy document, because it replaces "we have backups" with a measured recovery time and a measured amount of lost data.

What do RPO and RTO actually mean for a SaaS team?

A glowing isometric hourglass valve with cyan data particles falling and a bright amber marker ring frozen partway down the glass, representing one fixed, recoverable point in time

RPO is the most data you can afford to lose, measured backward in time from the incident. RTO is the longest you can afford to be down before the business impact becomes real. Both are decisions you make, then test.

RPO follows from how often you capture data. With a daily backup, a failure just before the next run can lose almost a day of writes. That is arithmetic, not a vendor claim. For an early-stage B2B SaaS without real-time financial transactions, I'd start the conversation at 24 hours and move it down only when a customer contract or a revenue number says so. That is my recommendation, not a standard.

Point-in-time recovery (PITR) shortens RPO to seconds. The PostgreSQL documentation describes it as a file-system-level base backup plus archived write-ahead log (WAL) files: you restore the base backup and replay the WAL up to a chosen stopping point, which can be a time, a named restore point or a transaction ID. It also says that to recover you need a continuous sequence of archived WAL files going back at least to the start of the base backup, and that you should set up and test WAL archiving before you take the first base backup.

Managed platforms package the same idea. Supabase's backup docs say PITR combines physical backups with WAL archiving, and that by default the WAL files are backed up at two-minute intervals. This suggests a realistic data-loss window of minutes rather than seconds on that setup, because the last unarchived WAL segment is what you can lose. The docs also say that once you enable PITR, Supabase stops taking daily backups, since PITR gives finer granularity. Switching on PITR therefore replaces one mechanism with another, and your retention window becomes the PITR window you pay for.

How do Postgres PITR, daily backups and branch restore compare?

The products differ in what you can restore to, how far back, and what they cost. The table uses each provider's own documentation; retention windows and prices change, so check the linked pages before you budget.

ApproachHow it recoversRetentionCost and limits
Self-managed WAL archivingBase backup plus WAL replay to a targetWhatever you archiveStorage plus your engineering time
Supabase daily backupsRestore from a daily backup7 days on Pro, 14 on Team, up to 30 on EnterpriseIncluded; stops when PITR is enabled
Supabase PITRBase backups plus WAL, archived every two minutes by default7, 14 or 28 daysAbout $100, $200 or $400 per month; Pro plan or higher; not covered by the spend cap
Neon instant restoreRewinds a branch timeline to a timestamp in the history windowFree 6 hours; Launch default 1 day, max 7; Scale default 1 day, max 30Retained history is billed as instant restore storage on paid plans; root branches only

The sources are Supabase's backup docs, its PITR pricing page, Neon's branch restore overview and Neon's restore window settings. Neon describes a restore as a complete overwrite of the database timeline, not a merge, and keeps the previous state as an automatic backup branch. Supabase's pricing page points out that the PITR add-on is not covered by the Spend Cap, so it can appear as a line item even when everything else is capped.

Why is a nightly backup not a disaster recovery plan?

A single amber-lit storage crate spotlighted among rows of identical dim, dusty, sealed cyan crates stacked in a warehouse corridor

A backup is a file, and a disaster recovery plan is a tested procedure that turns that file into a running application inside a time window someone agreed to. Treating the first as the second is the most common gap.

Consider a two-person team whose restore script was written for last year's schema. The backup job has run every night, so the dashboard is green. Nobody has run the restore since two migrations ago, and the first real attempt happens during an outage. That is a hypothetical, but it shows the failure mode: the file exists and the procedure doesn't work.

Real data points the same way. In Veeam's 2026 Data Trust and Resilience Report, based on a survey of more than 900 security leaders, only 28% of organizations hit by ransomware said they fully recovered all affected data, and 44% said they recovered less than 75% of it. This is a ransomware survey, not a SaaS database study, so read it as evidence that confidence and recovery are different things, not as a benchmark for your stack.

How do you restore a single tenant from a shared database backup?

A mechanical arm lifting one glowing amber data pod out of a transparent tank full of identical cyan data pods submerged in fluid, water droplets falling

In a pooled schema you restore the whole backup into a temporary database first, then extract and reload only that tenant's rows. Postgres has no command to restore one customer out of a shared cluster.

The PostgreSQL documentation says continuous archiving "can only support restoration of an entire database cluster, not a subset." AWS's guidance on multi-tenant backup and recovery, published in December 2022, says it prefers "segregation during recovery in the pooled model to defer the cost of segregation until it is required." Its approach is to create a temporary database from a snapshot, let the tenant's records be selected there, and keep the cost down by deleting everything except the tenant's data or by caching the recovered records. It also notes that exporting multiple tables needs consistency, which a restored snapshot already provides.

The alternative is database-per-tenant, where one tenant restores cleanly. My view is that this multiplies backup jobs, monitoring and patching by tenant count, and that most early products should accept the slower single-tenant restore. That is a recommendation and depends on what your contracts promise.

On Callidus, the React, Firebase and Stripe Connect clinic platform, tenant data lives under tenants/{tenantId}/… in Firestore, and audit logs carry a three-tier TTL: 90 days for routine events, one year for financial events and two years for GDPR-sensitive ones, with no expiry on the erasure record. That is a retention policy and not a recovery design, but it answers the same question as your PITR window: how far back do you need to reach, and who pays for the storage?

What does a database backup not contain?

Check the gaps before you need them. Supabase's backup docs state that database backups do not include objects stored through the Storage API, because the database only holds metadata about those objects. If your product stores uploads, your database restore brings back references to files that may need their own recovery path. List everything your application needs to run, then confirm each item has a backup and a restore step.

What RPO and RTO should an early-stage SaaS target?

Start with an RPO of 24 hours and an RTO of a few hours, then tighten only when a contract or revenue figure justifies the cost. That is my recommendation for early-stage products, not an industry standard.

Chasing a sub-minute RPO before any customer requires it spends engineering time on a guarantee nobody is paying for. The SaaS MVP stack choice should include the recovery story: which backup product, which retention window, and who runs the drill. If you use PostgreSQL on a managed platform, the table above shows what each step up in protection costs.

What belongs in a recovery runbook?

A runbook is useful only if someone with less context than its author can follow it under pressure. Write down these items and keep them next to the code.

  1. The exact command or console path for each restore type: full database, point in time, single tenant.
  2. Where the credentials live and who can reach them, including a second person.
  3. The steps that turn a restored database back into a working product: connection strings, migrations status, background jobs and webhooks that may have fired twice or not at all.
  4. How you decide to declare an incident and who tells customers.
  5. The date of the last successful drill and the time it took.

The third item catches the most surprises. A restored database is not a running product until the app points at it and the side effects that happened after the restore point are accounted for. Emails sent, payments captured and webhooks delivered after your restore point don't disappear when the database rewinds, so decide in advance who reconciles them and how.

How often should you run a recovery drill?

Pick a cadence you will keep, and record the measured recovery time each time.

  1. Each quarter, restore the latest backup into a throwaway environment and confirm the application boots against it. A file that decompresses is not a working restore.
  2. After every schema migration, check that your restore procedure still matches the current tables.
  3. Once a year, rehearse a bad case: a corrupted primary or a single-tenant restore under a time limit.
  4. Write down the RTO you measured, not the one you hoped for, and show that number to whoever owns the risk.
  5. Rotate who runs the drill, so the runbook is tested by someone other than its author.

Which of these five could you complete by Friday? Pick the quarterly restore, schedule it in your calendar now, and write the measured time at the top of your runbook.

Free resource

Free SaaS MVP Scope Template

A Notion document with the full feature checklist, MVP vs. nice-to-have table, pre-build questions, and cost signals — so you walk into any developer call knowing exactly what to ask for.

Get the template →
DL

Dusko Licanin

Full-Stack Developer · Banja Luka, Bosnia

Full-stack developer shipping SaaS MVPs, web apps, and mobile apps using AI-augmented workflows — without agency coordination overhead. Live portfolio: BookBed, Callidus, Pizzeria Bestek.

Frequently Asked Questions

What is point-in-time recovery in Postgres?

Point-in-time recovery restores a Postgres cluster to a chosen moment instead of only to the time of the last backup. The PostgreSQL documentation describes it as a base backup plus continuously archived write-ahead log files, replayed up to a target such as a timestamp, a named restore point or a transaction ID. The continuous archiving chapter also says it restores an entire database cluster, not a subset, and that you should set up and test WAL archiving before taking the first base backup.

What should a SaaS backup strategy include?

A SaaS backup strategy needs automated backups, a written restore procedure and a restore you have actually timed, not only a scheduled job. It should also list what the database backup leaves out: Supabase's backup docs state that objects stored through the Storage API are not included, only their metadata. Add a retention policy you chose on purpose, a named owner for the drill and a quarterly restore into a throwaway environment. Skipping the drill is how a team finds out mid-incident that the restore script no longer matches the schema.

How do you restore a single tenant from a shared database?

In a shared schema you restore the whole backup into a temporary database and then extract the tenant's rows. PostgreSQL's documentation says continuous archiving restores an entire cluster, not a subset. AWS's multi-tenant backup guidance says it prefers segregation during recovery in the pooled model, to defer the cost until it is needed. Database-per-tenant makes single-tenant restores simpler but multiplies backup jobs and monitoring by tenant count, so weigh that against what your contracts promise.

What RPO and RTO should an early-stage SaaS set?

A reasonable starting point for an early-stage SaaS is an RPO of 24 hours and an RTO of a few hours, then tighten it when a contract or revenue figure justifies the cost. That is a recommendation, not a standard. The number matters less than whether you have tested it: in Veeam's 2026 Data Trust and Resilience Report, only 28% of ransomware-hit organizations said they fully recovered all affected data. That is a ransomware survey, so read it as a caution about untested recovery, not as a benchmark.

How often should you test a disaster recovery plan?

Run a full restore every quarter and re-check the procedure after each schema migration that changes the tables it depends on. Once a year, rehearse a worse case such as a corrupted primary or a single-tenant restore with a time limit. Record the recovery time you measured each time and rotate who runs the drill, so the runbook is tested by someone other than its author. Also check whether the restore brings back everything the app needs, such as stored files, which database backups may not contain.