Your OpenAI invoice is not a token-price problem, it is a call-count problem. Before you compare models, open your usage dashboard and try to name the three features that spend the most. If you can't, that gap is the first thing to fix, and the rest of this post shows the numbers and the order to build the controls in.
How Do You Reduce OpenAI API Costs in a SaaS?

Send each request to the cheapest model that can do the job, keep the prompt prefix stable so caching applies, cap output length, and tag every call by feature so you can see where money goes.
Start with the current rates. These come from OpenAI's pricing page, Standard mode, per million tokens, checked on 30 September 2026:
| Model | Input | Cached input | Output | Suggested role |
|---|---|---|---|---|
| gpt-5.6-luna | $0.20 | $0.02 | $1.20 | Default. Classification, extraction, tagging, short replies, anything you can pin to a JSON schema. |
| gpt-5.6-terra | $2.00 | $0.20 | $12.00 | Escalation. Multi-step reasoning, long documents, output a customer reads line by line. |
| gpt-5.6-sol | $4.00 | $0.40 | $20.00 | Rare. Its price is promotional (see below). |
Two things matter more than the absolute numbers.
Why does output length matter more than prompt length?
Output tokens cost five to six times more than input tokens in every row: $1.20 against $0.20 on Luna, $12.00 against $2.00 on Terra, $20.00 against $4.00 on Sol. Most cost advice obsesses over shrinking the prompt, the cheaper half of each call. A hard cap on output tokens plus a structured JSON response usually saves more than an afternoon of prompt editing. That's my recommendation, not a measured benchmark, so test it against your own logs.
Here is a labelled hypothetical. Consider a feature that sends 2,000 input tokens and gets 300 back. On Luna that is about $0.00076 per call. On Terra, ten times the per-token rate, it is about $0.0076. If a user triggers it 100 times a month, that is roughly 8 cents on Luna and 76 cents on Terra. Model choice matters, but only by a fixed multiple. Call count is the number that can grow without limit.
Is the Sol price permanent?
No. The pricing page states that "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." If your unit economics only work at the promotional rate, you have a countdown rather than a margin.
Route first, escalate on evidence

A common failure is having no escalation signal, so a team promotes the whole feature to a larger model after one bad demo and never demotes it. Escalation should be triggered by the response, not by a feeling. Three signals that are cheap to check: the model returned a schema-invalid object, a required field came back null or empty, or your own validator rejected the answer. When one fires, retry once on Terra, log the escalation, and read that log weekly. If a prompt escalates 40% of the time, it needs a better prompt or a bigger default, not a router.
Why Is My OpenAI Prompt Cache Not Hitting?

Prompt caching reuses the longest matching prefix of your request, so any content that changes early in the prompt, such as a timestamp or a username, prevents everything after it from being cached.
OpenAI's prompt caching guide says cached reads cost 0.1x the uncached input price for GPT-5.6 models, and that caching needs at least 1,024 tokens on GPT-5.6 and later. On Luna that is $0.02 instead of $0.20 per million tokens. The guide reports hits in `usage.input_tokens_details.cached_tokens`, and a persistent zero there is your signal that the prefix is changing or too short.
Patterns that break the prefix (illustrative, not observed in any specific codebase):
- A system prompt that opens with `Today is {{date}}.` so every request starts differently.
- Injecting the user's name and plan tier above the instructions instead of below them.
- Building a tool definition in a different order on each request.
The 1,024-token minimum has a practical consequence for small features. A 300-token classification prompt can't be cached at all. If the feature runs at high volume, adding genuinely useful few-shot examples to push the stable part past the threshold can be cheaper than leaving it short. Check the arithmetic for your own volume before doing it.
Order the prompt like this: static instructions, examples, tool definitions, conversation, then this request's variable content last. Once that's stable, a semantic cache can remove calls entirely. In a 2024 paper on GPT Semantic Cache, Regmi and Pun report cutting API calls by up to 68.8% across their test query categories. Treat that as a ceiling from one experiment on repetitive queries. Prompt caching makes each call cheaper, semantic caching prevents some calls, and the second is a system you maintain, so do them in that order.
The feature nobody uses is still billing you
Per-feature cost attribution catches failures the other controls miss, because a runaway call site raises no errors. A public example is Arpit Gupta's write-up of a $1,800 bug: a batch report generator had also been wired into an autosave hook that fires every 30 seconds by default. The author reports the OpenAI bill rising from $620 to $2,480 in 23 days with zero errors, and dropping 61% the following month after the call moved back to the manual export button.
The lesson is that every request looked correct. Only cost grouped by feature over time showed the problem. So tag each call with a feature name and tenant id before shipping. On a multi-tenant product (if the term is new, here is what B2B SaaS means), the tenant tag also shows which customer's usage the others are subsidising.
Tier limits arrive before your budget does
OpenAI's rate limits guide currently lists usage tiers by total credit purchases: Free, then Build at $5, Launch at $100 and Grow at $500, with monthly usage limits of $100, $500, $5,000 and $200,000. Tiers upgrade automatically as purchases cross each threshold. Check which tier your account is in before a launch, because your own throughput ceiling is set there and not by your budget.
What Should You Charge for an AI Feature?
Price it as a variable-cost feature from day one, because inference turns a fixed cost into a per-use one and changes the margin your business model assumed.
Monetizely's summary of Bessemer Venture Partners' analysis puts fast-scaling AI SaaS startups at roughly 25% gross margin early on and steadier ones near 60%, against 80-90% for traditional SaaS. Those figures describe companies whose whole product is inference. If ChatGPT is one feature inside a product that already does something useful, your blended margin moves much less. The risk sits in between: an AI feature sold as the headline reason to upgrade, priced flat, and used hardest by customers on the cheapest tier.
Three patterns worth considering (my recommendations):
- A monthly allowance per seat, with overage priced at a multiple of your marginal cost. Most users never reach the ceiling and it stops the extreme ones.
- Credits in something the customer already counts, such as documents processed or minutes of transcript, rather than tokens.
- The AI feature on a higher tier only for the first release, which limits exposure while you learn real usage.
On Callidus OS, a multi-tenant clinic SaaS built on React, Firebase and Stripe Connect, I worked solo over roughly ten weeks. The transferable point from any per-tenant product is that a variable cost you don't meter per tenant on day one is hard to attribute later. Meter first and decide the price afterwards.
Which Guardrails Should You Build First?
Build them in this order, because each step makes the next one measurable.
- Set a hard spend limit in the OpenAI dashboard, below the amount that would hurt.
- Pick Luna as the default and write the escalation rule as a function in code.
- Put static content first and per-request content last, then confirm hits through `cached_tokens`.
- Cap output tokens at each call site.
- Tag every call with feature and tenant.
- Add a per-tenant rate limit in your own application layer. Yours can return a useful message and a retry window, where the provider's limit returns a raw 429.
- Only then add semantic caching, batching or a fallback provider.
Steps 1 to 6 are an afternoon of work. Step 7 is a project most products never need. For the surrounding habits, see the AI-augmented development workflow I use, which AI solutions earn their place in a business, and the SaaS MVP stack I recommend. If you're unsure what counts as SaaS for your category, the glossary entry is a quick check.
Open your usage dashboard now and write down your top three features by spend. Which one of them would you cap first?
