Skip to main content
TACAVAR
•Build in Public

Agent Token-Cost Forecasting: Price the Run First

Same task, same model, up to 30x token spread. How agent token-cost forecasting works, where it breaks, and how to wire a budget stop so runs halt at budget.

Every LLM bill is read after the run. The interesting number is the one before it: what will this task cost if the agent behaves the way agents behave. Most teams cannot answer, because they treat an agent run like an API call with a knowable price. It is not. The same task, given to the same model, can vary by more than an order of magnitude in token consumption across runs.

Recent research puts a number on the variance — up to 30x between executions of an identical task. That spread is not noise. It is structural, and it is why cost planning for agents keeps failing at the invoice stage. The fix is upstream: forecast consumption before and during execution, then treat the forecast as a control input, not a report.

Why agent token consumption is unpredictable

Three forces compound, and none of them exist in a single-request workload.

First, the agent is conditional. It chooses its next step based on tool feedback and intermediate results. A task that resolves in one edit can, on the next run, take five rounds of retries before it converges. The plan is not fixed; the model writes it as it goes.

Second, context grows, and growth is multiplicative, not additive. Every tool output, every intermediate result, every failed attempt lands in the context. Each later call re-reads everything that came before it. A segment that consumed 2,000 tokens early in the run does not cost 2,000 tokens once — it costs 2,000 tokens times the number of subsequent calls that carried it forward. Early bloat is the most expensive bloat, because it is re-paid by every call after it.

Third, generation is nondeterministic. Same input, different output lengths, different numbers of verification cycles. Multiply the three forces and the 30x spread stops being surprising.

What forecasting actually looks like

The research frontier here is worth naming: a 2026 paper called TokenCast (arXiv:2609.35760) treats forecasting as a live estimation problem rather than a pre-flight guess. The mechanics matter more than the specifics, because they describe what any serious cost-planning layer needs.

The approach decomposes an agent run into execution segments, and learns a cost representation for each one: the tokens it consumes directly, plus the context growth it introduces. Composing adjacent segments yields a cumulative estimate — including the re-read cost described above. As execution unfolds and real evidence arrives (a tool returned, a step completed), the forecast refreshes. Two properties make this practical rather than academic:

  • No extra LLM calls. The forecaster rides alongside execution. Prediction overhead averaged 32.8 milliseconds per run on SWE-bench Verified.
  • It pays for itself in budget control. In offline budget-control replay, forecasting-driven budget policy used 21.3% fewer tokens than a fixed-budget policy at matched trace completion.

That last number is the whole argument. A fixed budget — "stop at 200k tokens" — is a blunt instrument that either truncates good runs or lets bad ones burn to the cap. A live forecast converts the budget from a wall into a steering signal.

Where forecasts break

A forecast is a model of the run, and it fails in the same places the run is hardest to predict.

  • Branching tool loops. When a step's output decides how many more steps exist, one unexpected tool result can multiply the remaining path. Forecasts tuned on typical trajectories under-count the tail. The tail is where the money is.
  • Retries. A retry is a full re-execution of a segment with a longer context than the first attempt. Cost models that treat attempts as independent underestimate retry-heavy runs badly.
  • Context accumulation itself. The re-read multiplier means small errors in early-run estimates compound. A 10% underestimate at the start of the run can be a 30% underestimate by the end.
  • Silent growth. Long tool outputs and verbose intermediate reasoning inflate every subsequent call without any single step looking expensive. Per-step dashboards do not show this. Only a cumulative forecast does.

The operational conclusion: do not buy a number once. Refresh it as the run produces evidence, and expect the revision to matter more than the initial estimate.

From forecast to control: the budget stop

A forecast that only produces a dashboard is a report. The useful version closes a loop.

Tacavar runs a tiered model gateway in production — free-tier primary, paid fallbacks, routing by task type, falling back by cost. The routing layer (which we've written about as cost-per-success) decides where a call goes. The forecasting layer decides whether the run should continue at all. The wiring is direct:

  1. Forecast before dispatch. Estimate the run's expected token envelope from task shape and historical traces. Route cheap tasks to the free tier confidently; route wide-forecast tasks to the capable model first, because a cheap-model retry chain is the most expensive way to fail.
  2. Refresh mid-run. Every completed segment updates the remaining-cost estimate. A run that was cheap at step 3 can be expensive at step 9; the route and the go/no-go decision should reflect that.
  3. Stop at budget, not at failure. When projected consumption crosses the envelope, halt the run — cleanly, with state preserved — instead of letting it die at the provider's rate limit mid-retry. This is the same discipline as a trading kill switch: the stop is not a feature bolted onto the strategy; it is part of the architecture. We cover the trading side in trading-bot-kill-switch-architecture, and the principle transfers exactly. A budget stop is a kill switch whose trigger is money.
  4. Record forecast vs. actual. Every run produces one datapoint for the next forecast. The cost model compounds the same way the systems do.

This is also why forecasting belongs upstream of the cost-arithmetic work on this site: zero-cost inference via free tiers only works if you know a task fits the free tier before you send it, the coffee-budget infrastructure only stays under budget if consumption is bounded during the run, and the multi-model founder stack only routes correctly when expected cost is an input to the route.

The operator's minimum viable forecast

You do not need a research-grade forecaster to stop being surprised by invoices. A working baseline:

  • Log per-run totals and per-segment consumption for every agent run. You cannot forecast a distribution you have never observed.
  • Classify tasks by historical envelope (median, p95). Most tasks cluster; the tail is task-specific and predictable by task type.
  • Set budget stops at the p95 with a hard ceiling above it. Kill at the ceiling, degrade (switch models, trim context, summarize) at the p95.
  • Recompute remaining budget after every tool call against actual context length, not the original plan.

The gap between a fixed cap and a live forecast was 21.3% of tokens in the research replay. In production terms, that is the difference between a cost plan and a cost hope.

Agents are systems, and systems that spend without forecasting do not endure. Forecast the run, wire the forecast into the stop, and let the budget be a constraint the architecture respects — not a line item you discover later.

You built it. We optimize it.