You have a pile of documents, a customer who wants answers from them, and a vague sense that you need a vector database. Before you provision anything, check one thing: how many tokens your documents add up to, and whether they fit in a single prompt. That number decides more than any architecture diagram.
When Is RAG Worth Building for a B2B SaaS Product?
RAG is worth building for a SaaS product when your corpus is too large for a prompt, changes faster than you can retrain anything, and a wrong answer costs you a renewal or a support escalation. Drop one of those three conditions and something cheaper usually wins.
What RAG buys you

RAG gives you grounded answers over data the model never saw in training, with a citation the user can open, and no retraining when the data changes. OpenAI's optimizing accuracy guide draws the same line: optimize the context when the model lacks proprietary or up-to-date knowledge, and optimize the model itself when the problem is inconsistent format, tone or reasoning.
It does not make the model smarter, and it does not teach it your product's voice. It narrows what the model may look at and hands the user a source. In a B2B SaaS product that source link is the feature, because a wrong answer about a customer's contract terms lands with their legal team, not in a thumbs-down counter.
RAG vs a long context window: where the line moved

The vector database began as a workaround for small context windows, and that constraint has loosened. Anthropic's models overview lists a 1M-token context window for Claude Sonnet 5.5, and notes that 200k tokens is roughly 150k words. A product manual, a changelog and a policy handbook can each fit with room to spare.
The arithmetic is short enough to do yourself. On Anthropic's pricing page, Sonnet 5.5 input costs $2 per million tokens, a 5-minute cache write costs 1.25x that, and a cache read costs 0.1x. So a 200,000-token knowledge base costs $0.40 to send once uncached, $0.50 to write into the cache, and $0.04 per follow-up question that hits the cache. Those are my calculations from the published rates, and they assume the cached prefix stays identical between requests, because the caching docs require an exact prefix match and a default lifetime of five minutes.
Consider a knowledge base of a few hundred pages with modest daily traffic. Four cents a question means the pipeline you would build around it (chunking, embedding, an index, a refresh job) has to save more than it costs to own. For a small corpus it rarely does. At a scale where the corpus no longer fits in a window, or where thousands of queries a day make even cached long prompts expensive, retrieval starts to win. Latency is a separate axis. A very long prompt takes longer to process than a few retrieved passages, so speed can justify RAG even when the money says otherwise. Measure that on your own data rather than trusting a rule of thumb.
When to skip RAG entirely

Skip it when the corpus fits in a prompt, when the answer lives in one database row, or when nobody has asked the question yet. The first case is now much wider than it was: documentation, a handbook and an onboarding guide are tens of thousands of tokens, not millions.
The second case is the one that costs the most for the least return. "How many seats are left on my plan" is a SELECT statement. Sending it through an embedding model, a similarity search and a language model is an expensive way to lose precision on something your database answers exactly. The broader argument is in which AI features in a SaaS product pay for themselves, and retrieval belongs in the same bucket: worth it when the alternative is a person reading documents, not worth it when the alternative is a WHERE clause.
The third case looks like diligence. A team builds retrieval before it knows what users ask. Ship a plain search box first and log the queries. That log is your requirements document, and it costs nothing.
RAG vs Fine-Tuning: When Should You Use Each?
Prompting is the right first version, RAG is the right answer when knowledge is large and changing, and fine-tuning is for shaping behaviour, not for storing facts. This matches how OpenAI's guide separates in-context memory from learned memory.
| Approach | What it changes | Handles data that changes | Best fit |
|---|---|---|---|
| Just prompting | Nothing, you paste context in | Yes, as long as it fits the window | Small corpora and every first version |
| RAG | What the model sees per question | Yes, at the price of a re-embed job | Large corpora, citations, per-tenant data |
| Fine-tuning | How the model behaves | No, knowledge is frozen at training time | Fixed tone, output format, consistent reasoning |
This is a recommendation, not a benchmark. I'm not giving build-cost ranges here, because they vary too much by team and scope to state honestly. If you want a number, estimate the engineering hours for your own chunking, indexing and refresh work, then compare it with the per-question cost above.
These options combine. A common pattern is a prompt that pins the format and tone you want, retrieval that supplies the facts, and fine-tuning only if prompt instructions stop holding the format steady. Start at the top of the table and move down only when a specific failure forces you to, and write that failure down before you do.
Why Does Pure Vector Search Miss Exact Strings?
Dense embeddings match meaning, not characters, so they can rank an article about general troubleshooting above the one page containing your customer's exact error code. Support search is full of exact strings.
The evidence for this is old and consistent. The 2021 BEIR benchmark paper tested retrieval models across 18 datasets and reported that BM25 is a strong baseline, while dense retrieval models often underperform other approaches when they meet data they weren't trained on. Your product's vocabulary, with its SKUs, error codes and internal feature names, is exactly that kind of data.
The fix is hybrid retrieval: keyword search and vector search, with the two ranked lists merged. Elastic's reciprocal rank fusion documentation describes RRF as combining result sets with different relevance indicators, with a default rank constant of 60. Qdrant offers the same fusion over sparse and dense vectors, available since v1.10.0.
If you already run Postgres, you may not need a second system. The pgvector README says to use it together with Postgres full-text search for hybrid search, and the PostgreSQL text search docs describe the tsvector and tsquery types and ranking. Keeping vectors in the database that already holds your tenants' rows also keeps one place where access rules live. I describe the React and Supabase setup on its own stack page.
How Often Should You Refresh Embeddings?
Refresh embeddings when the text they describe changes, when you change your chunking, or when you switch embedding model, and set a separate cadence per content type instead of one global schedule.
Consider a two-person team that edits its pricing docs weekly and its architecture overview twice a year. One schedule for both either wastes money re-embedding the slow content or serves outdated prices from the fast content. Separate cadences remove both problems.
Staleness is quiet. A retrieval system whose documents drifted out of date keeps returning confident, well-scored, wrong answers, and no latency graph shows it. I'd track two things separately: making a new document searchable, which is cheap, and regenerating vectors after an edit or a model change, which is the operation people schedule carelessly.
The embedding bill itself is small. OpenAI's pricing page lists text-embedding-3-small at $0.02 per million tokens and text-embedding-3-large at $0.13, with batch pricing about 50% lower. Take a hypothetical corpus of 10,000 documents at 1,000 tokens each: that is 10 million tokens, or about $0.20 on the small model for a full re-embed. The real cost of a refresh is the job, the monitoring and the person who owns it. Use the embeddings guide for the practical limits, such as the 8,192-token input maximum on both text-embedding-3 models.
A build order to follow
Ship the smallest thing that answers a real question, then measure before building the next layer. This is my recommendation, and each step earns the next.
- Ship a search box with no model behind it and log every query for two weeks.
- Answer the top queries with a cached prompt that holds the relevant documents whole.
- Add keyword search when exact strings appear in the logs.
- Add embeddings only when paraphrased questions show up that keywords can't match.
- Fuse the two lists and tune against your own logs, not against a public benchmark.
- Measure retrieval separately from generation, so a drop in quality is visible before customers report it.
For each step, decide in advance what result would make you stop. If a cached prompt answers the top queries acceptably, the later steps stay unbuilt, and that is a good outcome.
Many teams start at step 4. Steps 1 and 6 are what tell you whether the rest worked. The setup I default to for the first version is in my SaaS MVP stack.
Here is your next action. Export the last thousand searches your users ran and read them yourself. How many are exact strings, how many does page one of your docs already answer, and how many came back empty?
