If you run a multi-tenant SaaS and can't say when you last restored a backup, start there. This week, restore your most recent backup into a throwaway database and time how long it takes. That one exercise tells you more about your SaaS disaster recovery plan than any policy document, because it replaces "we have backups" with a measured recovery time and a measured amount of lost data.
What do RPO and RTO actually mean for a SaaS team?

RPO is the most data you can afford to lose, measured backward in time from the incident. RTO is the longest you can afford to be down before the business impact becomes real. Both are decisions you make, then test.
RPO follows from how often you capture data. With a daily backup, a failure just before the next run can lose almost a day of writes. That is arithmetic, not a vendor claim. For an early-stage B2B SaaS without real-time financial transactions, I'd start the conversation at 24 hours and move it down only when a customer contract or a revenue number says so. That is my recommendation, not a standard.
Point-in-time recovery (PITR) shortens RPO to seconds. The PostgreSQL documentation describes it as a file-system-level base backup plus archived write-ahead log (WAL) files: you restore the base backup and replay the WAL up to a chosen stopping point, which can be a time, a named restore point or a transaction ID. It also says that to recover you need a continuous sequence of archived WAL files going back at least to the start of the base backup, and that you should set up and test WAL archiving before you take the first base backup.
Managed platforms package the same idea. Supabase's backup docs say PITR combines physical backups with WAL archiving, and that by default the WAL files are backed up at two-minute intervals. This suggests a realistic data-loss window of minutes rather than seconds on that setup, because the last unarchived WAL segment is what you can lose. The docs also say that once you enable PITR, Supabase stops taking daily backups, since PITR gives finer granularity. Switching on PITR therefore replaces one mechanism with another, and your retention window becomes the PITR window you pay for.
How do Postgres PITR, daily backups and branch restore compare?
The products differ in what you can restore to, how far back, and what they cost. The table uses each provider's own documentation; retention windows and prices change, so check the linked pages before you budget.
| Approach | How it recovers | Retention | Cost and limits |
|---|---|---|---|
| Self-managed WAL archiving | Base backup plus WAL replay to a target | Whatever you archive | Storage plus your engineering time |
| Supabase daily backups | Restore from a daily backup | 7 days on Pro, 14 on Team, up to 30 on Enterprise | Included; stops when PITR is enabled |
| Supabase PITR | Base backups plus WAL, archived every two minutes by default | 7, 14 or 28 days | About $100, $200 or $400 per month; Pro plan or higher; not covered by the spend cap |
| Neon instant restore | Rewinds a branch timeline to a timestamp in the history window | Free 6 hours; Launch default 1 day, max 7; Scale default 1 day, max 30 | Retained history is billed as instant restore storage on paid plans; root branches only |
The sources are Supabase's backup docs, its PITR pricing page, Neon's branch restore overview and Neon's restore window settings. Neon describes a restore as a complete overwrite of the database timeline, not a merge, and keeps the previous state as an automatic backup branch. Supabase's pricing page points out that the PITR add-on is not covered by the Spend Cap, so it can appear as a line item even when everything else is capped.
Why is a nightly backup not a disaster recovery plan?

A backup is a file, and a disaster recovery plan is a tested procedure that turns that file into a running application inside a time window someone agreed to. Treating the first as the second is the most common gap.
Consider a two-person team whose restore script was written for last year's schema. The backup job has run every night, so the dashboard is green. Nobody has run the restore since two migrations ago, and the first real attempt happens during an outage. That is a hypothetical, but it shows the failure mode: the file exists and the procedure doesn't work.
Real data points the same way. In Veeam's 2026 Data Trust and Resilience Report, based on a survey of more than 900 security leaders, only 28% of organizations hit by ransomware said they fully recovered all affected data, and 44% said they recovered less than 75% of it. This is a ransomware survey, not a SaaS database study, so read it as evidence that confidence and recovery are different things, not as a benchmark for your stack.
How do you restore a single tenant from a shared database backup?

In a pooled schema you restore the whole backup into a temporary database first, then extract and reload only that tenant's rows. Postgres has no command to restore one customer out of a shared cluster.
The PostgreSQL documentation says continuous archiving "can only support restoration of an entire database cluster, not a subset." AWS's guidance on multi-tenant backup and recovery, published in December 2022, says it prefers "segregation during recovery in the pooled model to defer the cost of segregation until it is required." Its approach is to create a temporary database from a snapshot, let the tenant's records be selected there, and keep the cost down by deleting everything except the tenant's data or by caching the recovered records. It also notes that exporting multiple tables needs consistency, which a restored snapshot already provides.
The alternative is database-per-tenant, where one tenant restores cleanly. My view is that this multiplies backup jobs, monitoring and patching by tenant count, and that most early products should accept the slower single-tenant restore. That is a recommendation and depends on what your contracts promise.
On Callidus, the React, Firebase and Stripe Connect clinic platform, tenant data lives under tenants/{tenantId}/… in Firestore, and audit logs carry a three-tier TTL: 90 days for routine events, one year for financial events and two years for GDPR-sensitive ones, with no expiry on the erasure record. That is a retention policy and not a recovery design, but it answers the same question as your PITR window: how far back do you need to reach, and who pays for the storage?
What does a database backup not contain?
Check the gaps before you need them. Supabase's backup docs state that database backups do not include objects stored through the Storage API, because the database only holds metadata about those objects. If your product stores uploads, your database restore brings back references to files that may need their own recovery path. List everything your application needs to run, then confirm each item has a backup and a restore step.
What RPO and RTO should an early-stage SaaS target?
Start with an RPO of 24 hours and an RTO of a few hours, then tighten only when a contract or revenue figure justifies the cost. That is my recommendation for early-stage products, not an industry standard.
Chasing a sub-minute RPO before any customer requires it spends engineering time on a guarantee nobody is paying for. The SaaS MVP stack choice should include the recovery story: which backup product, which retention window, and who runs the drill. If you use PostgreSQL on a managed platform, the table above shows what each step up in protection costs.
What belongs in a recovery runbook?
A runbook is useful only if someone with less context than its author can follow it under pressure. Write down these items and keep them next to the code.
- The exact command or console path for each restore type: full database, point in time, single tenant.
- Where the credentials live and who can reach them, including a second person.
- The steps that turn a restored database back into a working product: connection strings, migrations status, background jobs and webhooks that may have fired twice or not at all.
- How you decide to declare an incident and who tells customers.
- The date of the last successful drill and the time it took.
The third item catches the most surprises. A restored database is not a running product until the app points at it and the side effects that happened after the restore point are accounted for. Emails sent, payments captured and webhooks delivered after your restore point don't disappear when the database rewinds, so decide in advance who reconciles them and how.
How often should you run a recovery drill?
Pick a cadence you will keep, and record the measured recovery time each time.
- Each quarter, restore the latest backup into a throwaway environment and confirm the application boots against it. A file that decompresses is not a working restore.
- After every schema migration, check that your restore procedure still matches the current tables.
- Once a year, rehearse a bad case: a corrupted primary or a single-tenant restore under a time limit.
- Write down the RTO you measured, not the one you hoped for, and show that number to whoever owns the risk.
- Rotate who runs the drill, so the runbook is tested by someone other than its author.
Which of these five could you complete by Friday? Pick the quarterly restore, schedule it in your calendar now, and write the measured time at the top of your runbook.
