On this page
- Where does an AI agent's API bill actually go?
- How do you measure agent spend before optimizing it?
- Which optimizations cut the most, ranked?
- What is prompt caching and how much does it save?
- How does model routing cut costs without hurting quality?
- Can a smaller model do the job for less?
- How do you stop context from eating your budget?
- When should you batch instead of paying full price?
- How do you stop an AI agent from running up a huge bill?
- Is it cheaper to self-host an LLM than pay API costs?
- What is the 30-minute audit that finds your biggest savings?
Last Updated: October 8, 2026
Key takeaways:
- Providers bill AI agents per token, and agent loops multiply the meter: system prompts, tool schemas and history ride along on every turn, so a 25-turn agent can burn 500k tokens on one task.
- Prompt caching is the biggest cheap lever: Anthropic charges 0.1x for cached reads (90 percent off) and OpenAI discounts cached input 50-90 percent depending on the model.
- Batch and flex tiers cut every token bill by 50 percent at OpenAI, Anthropic and Google in exchange for slower processing, up to 24 hours for batch jobs.
- Model routing cuts costs 30-85 percent by sending routine tasks to small models, and TensorZero's 2025 study shows fine-tuned small models beat frontier models on narrow tasks at 5-30x lower cost.
- Guardrails prevent disasters: iteration caps, per-run spend ceilings, provider timeouts and fallback chains stop the runaway loops that hit 40 dollars in minutes.
API providers charge for AI agents the same way utilities charge for power: per unit consumed, metered on every call. The unit is the token, and an agent consumes them in loops. According to J.SERVO's 2026 token-cost analysis, the combined techniques in this playbook cut typical AI agent bills by 50-70 percent without dropping output quality. We run agent stacks on our own infrastructure at Flowtivity, and the fixes below are the ones that actually moved our numbers, including one incident where a single bloated session was re-billing a 254k-token context on every single turn.
Where does an AI agent's API bill actually go?
An agent's bill is driven by what rides along with each request, not by what the user typed. On every turn the agent re-sends the system prompt, the schemas for every tool it can call, the conversation history, and the results of previous tool calls. According to CData's 2026 analysis of Claude API costs, the actual spend concentrates in context carried across turns and tool-call overhead, while most cost-cutting advice targets the one part that is smallest: the user's prompt. Output tokens and reasoning tokens bill at 3-5x the input rate across frontier models, so an agent that thinks out loud pays premium prices to talk to itself.

How do you measure agent spend before optimizing it?
Measure per request, per agent, per model, before touching anything. Gateways like Helicone, LiteLM and OpenRouter log tokens, cost and latency on every call, and according to Helicone's cost-optimization guide, teams that attribute spend per feature consistently find that a handful of workflows drive most of the bill. The number that matters most is the turn multiplier: average model calls per task times average prompt tokens per call. A 25-turn agent carrying a 20k-token prompt is a 500k-token task before the answer is written. Halve the multiplier and you halve the bill at any per-token price.

Which optimizations cut the most, ranked?
Cache first, route second, right-size third. The table below ranks the ten levers by typical saving against the effort to implement them, using verified provider rates and published benchmarks. According to LMSYS's RouteLLM research, routing between a strong and weak model cuts costs by over 2x without compromising response quality, and their best configurations reached 85 percent savings on MT Bench. No single lever halves a bill on its own, but stacking three of the low-effort ones reliably does.
| Optimization | Typical saving | Effort | Trade-off |
|---|---|---|---|
| Prompt caching (reads) | 90% on cached prefix | Low | 1.25x writes, 5-min TTL churn |
| Batch API / flex tier | 50% on all tokens | Low | Up to 24h latency |
| Model routing per task | 30-85% | Medium | Router to build and maintain |
| Distilled small model | 5-30x cheaper | High | Dataset curation + training |
| Semantic response cache | 30-50% on chat traffic | Medium | False positives ship wrong answers |
| Prompt compression | Up to 20x smaller prompts | Medium | About 1.5% quality loss |
| Context trim + compaction | 84% fewer tokens in case studies | Low-Med | Info loss if over-trimmed |
| Reasoning effort control | 2-5x fewer thinking tokens | Low | Quality risk on hard steps |
| Tool schema pruning | Scales with tool count | Low | Manual audit per agent |
| Self-host open model | $0 marginal per token | High | Hardware + operations |
What is prompt caching and how much does it save?
Prompt caching lets the provider store the unchanging front of your prompt and bill re-reads of it at a steep discount. According to Anthropic's documentation, cache reads cost 0.1x the normal input rate, a 90 percent discount, while cache writes cost a 1.25x premium and the cache expires after a 5-minute TTL by default. OpenAI's caching is automatic and discounts cached input by 50-90 percent depending on the model. The engineering rule that unlocks it: put stable content first. System prompt, tool schemas and reference documents go at the front where the shared prefix forms; anything volatile goes last. Get the order wrong and you pay write premiums for a cache that never hits.

How does model routing cut costs without hurting quality?
Routing sends each request to the cheapest model that can handle it and reserves premium models for the steps that earn them. The RouteLLM study from LMSYS found that trained routers selecting between a strong and weak model reduce costs by over 2 times without compromising response quality, with up to 85 percent savings in their best configurations. In production, a static map captures most of the win without any machine learning: classification, extraction and summarization go to small models, while planning, code and architecture decisions go to frontier models. Every agent should also carry an explicit fallback chain, which is routing on failure: when a lane dies, the run continues on the next tier instead of hanging or dying.

Can a smaller model do the job for less?
Yes, and by a wider margin than most teams expect. According to TensorZero's 2025 distillation study, small models fine-tuned on programmatically curated outputs from a large teacher beat the large model's performance on the target task at 5-30x lower inference cost, with some configurations reaching 24.1x lower cost per success than GPT-4o. The recipe: collect traces where the frontier model got the task right, filter them programmatically, fine-tune a small model, and deploy it on that one task. This is a high-effort lever, so it belongs on your highest-volume, narrowest task, usually the one classification or extraction step that runs thousands of times a day.
How do you stop context from eating your budget?
Treat context as a budget and manage it like one. Trim tool results before they enter the history: extract the two or three fields the next step needs instead of carrying the whole payload forward. Keep the history window bounded, compact old turns into summaries, and offload durable facts to an external notes file the agent can reload on demand. Published context-engineering case studies report 84 percent token reduction from removing stale material and loading tools on demand. Tool schemas deserve the same discipline, because according to the MCP token overhead calculators, every tool definition is billed on every single request whether or not it gets called.

When should you batch instead of paying full price?
Whenever the result is not needed in the next few seconds. OpenAI, Anthropic and Google all sell a 50 percent discount on their batch APIs in exchange for a 24-hour completion window, and OpenAI plus Google also run flex tiers that cut synchronous calls roughly in half for slower processing. According to Datablist's OpenAI flex guide, you flip one service tier parameter and pay flex rates on the same models. Cron jobs, daily digests, evaluation runs, backfills and bulk classification have no business paying synchronous prices. In our own stack, the nightly content digests and scheduled agent runs are the first candidates for the batch lane: they already tolerate minutes of delay, so 24 hours costs nothing.

How do you stop an AI agent from running up a huge bill?
Wrap every run in guardrails, because agents fail expensively. According to Particula's runaway-spend playbook, a max-iteration guard plus a per-run spend ceiling would prevent most of the incidents where an agent burns 40 dollars in minutes on a retry loop. The full ring: an iteration cap per run, a token or dollar ceiling enforced at the gateway, provider timeouts so a dead lane fails over in seconds rather than hanging, capped retries with a circuit breaker, and real-time alerts at 2x the daily median. We learned this one the hard way in October 2026: a model server on our DGX Spark went down, and our OpenClaw gateway hung every model call for about 135 seconds before failing over. Setting provider timeoutSeconds to 180 and an explicit fallback chain turned an open-ended hang into a bounded, priced failure.

Is it cheaper to self-host an LLM than pay API costs?
For steady, high-volume internal work, yes, because marginal token cost goes to zero. Our dual DGX Spark setup runs DeepSeek V4 Flash at roughly 60 tokens per second with a 1M-token context and no per-token meter at all, which is why our agent loops for summaries, drafts and triage run on hardware we already own. The API still wins for elastic spikes and frontier-reasoning tasks, and DeepSeek's own rates show how cheap metered tokens have become: about 0.30 to 1.20 dollars per million tokens at peak with half-price off-peak windows. The honest rule of thumb: self-host the boring, high-volume middle of your workload, keep the API for the spikes and the hard reasoning, and let a fallback chain route between them.
What is the 30-minute audit that finds your biggest savings?
Run the six steps in the HowTo list above this section and you will know exactly which levers apply to your stack. In practice the first audit almost always surfaces the same three findings: tools that are defined but never called, a session or history setting that lets context grow unbounded, and at least one deferrable workload paying synchronous prices. Fix those three and most stacks are 40-60 percent cheaper before anyone touches a model choice. Then automate the rest: point agents at a gateway with per-request cost logging, set the guardrails once, and review the spend dashboard weekly while the stack is young.
The providers will keep charging per token because that is their business model. Your job is to make sure every token they meter is one your agent needed to spend. Measure first, cache what repeats, route what is routine, right-size what is narrow, and cap everything that loops. That is the whole playbook, and according to the benchmarks above, it is worth 50-90 percent of whatever you are spending today.
One email a month, no noise
Practical AI notes for Australian businesses. Unsubscribe anytime.