Agentic Workflows Need a Cost Budget for Every Step
Agents / Workflow design / Inference costs

Agentic Workflows Need a Cost Budget for Every Step

Show how autonomous workflows multiply spend through repeated model calls and unnecessary steps, then explore controls such as smart harnesses, routing, caching, and explicit workflow discipline. The post would focus on operational design for agents rather than treating the final API bill as an unavoidable surprise.

An agentic workflow does not have a single API cost. It has an execution-graph cost. A conventional request is relatively predictable: one input, one model call, one billable result. An agent may plan, retrieve, call tools, verify its work, retry failures, grow its context, and hand off to sub-agents. Each edge in that graph can add latency and spend, while parallel branches multiply both.

That changes the operating question from “What does this request cost?” to “What may this task cost before it finishes?” The answer must be designed before deployment, not inferred from the first production invoice. Assign every step an expected cost and a ceiling; account for repetitions, context growth, retries, and fan-out; then define what happens when the allowance runs out. A practical budget makes autonomy bounded rather than open-ended.

The final API bill is a lagging indicator. It tells you what was charged after the workflow has already made its planning, retry, routing, and context decisions. By then, an expensive loop is a production fact, not a design concern.

The workflow—not the individual request—is the useful cost unit. Its graph includes model calls, tool calls, retries, parallel branches, growing context, and idle or unnecessary steps. A cheap request repeated across six hops can cost more than one deliberate call to a larger model. A timeout can multiply both inference and tool spend without improving the result.

A practical estimate starts at the step level:

E[Cworkflow]=∑ipi×E[ni]×mcontext,i×mretry,iE[C_{workflow}] = \sum_i p_i \times E[n_i] \times m_{context,i} \times m_{retry,i}

Here, pip_i is the price of step ii, and E[ni]E[n_i] is its expected number of executions. The context multiplier captures repeated or expanding payloads; the retry multiplier captures failures and recovery attempts. Fan-out belongs in E[ni]E[n_i]: if a step launches four parallel agents, its expected repetitions are four, not one.

This model is intentionally simple. Its value is visibility. Teams can assign a cost to every hop, measure actual repetition, and find which multiplier is driving variance. Token price alone cannot show whether the workflow is spending money on useful work or merely replaying its own state.

A left-to-right workflow diagram showing one user task branching into planner, retrieval, tool, verifier, and retry loops. Use progressively thicker c

Consider a support agent handling a refund request. A frontier model classifies the ticket, builds a plan, chooses a routine lookup tool, interprets its result, and writes the response. Sending the same system instructions and account context at each hop turns one task into several expensive calls.

Then the lookup times out. The harness retries with the full prompt. Two parallel agents independently investigate the same order, and both pass their findings back to the planner. The final answer may be correct, but its cost was determined by duplication and recovery behavior, not by the initial token estimate.

Each action needs an expected cost, latency, and value:

  • Classification: low latency and modest value; usually not worth a frontier model.
  • Planning: higher cost, justified only when the request is ambiguous.
  • Tool selection and execution: predictable cost, but exposed to timeout and retry multipliers.
  • Verification: valuable for high-impact refunds, wasteful when it repeats unchanged checks.
  • Final phrasing: often low risk and suitable for a cheaper route.

A useful estimate is:

E[Cstep]=Ccall×E[repetitions]×Mcontext×MretryE[C_{step}] = C_{call} \times E[repetitions] \times M_{context} \times M_{retry}

The same accounting applies to latency and outcome value. A parallel branch should exist only if its expected quality gain exceeds its added cost and delay. Token counts are an input to that decision, not the decision itself.

Put a Budget Object on Every Run

A workflow needs a budget object at creation time, then a child budget for each step. At minimum, record:

  • workflow, run, tenant, and user identity
  • model route and pricing version
  • input and output tokens, tool calls, retries, and cache hits
  • elapsed time, dollar spend, token spend, and remaining budget
  • status, stop reason, and quality outcome

The budget should track money and tokens separately. A token ceiling will not protect you from an unexpectedly expensive route; a dollar ceiling will not explain context growth.

Reserve before execution. A planner step might reserve the expected model call plus its allowed retries; a fan-out step reserves capacity for its maximum branches. After execution, reconcile the reservation against actual usage and refund the difference. Record the variance—large gaps indicate bad estimates or an uncontrolled loop.

Use soft warnings for approaching limits and hard ceilings for enforcement. A warning can trigger a cheaper model, compressed context, or human review. A ceiling must stop new work, including retries and child-agent launches. Do not let an in-flight tool call silently create more budget.

When the budget is exhausted, return a typed partial result: completed steps, unresolved objective, spend, and the reason execution stopped. Preserve resumable state when safe, but never pretend the task succeeded. A clear budget-exhausted outcome is operationally useful; an unexplained timeout is not.

A cutaway view of a budget ledger attached to an agent run. Show a top-level dollar and token allowance divided into step-level envelopes for planning

Make the Harness the Cost-Control Layer

The harness—not the model—should own the workflow’s cost behavior. Encode the run as an explicit state machine: plan → retrieve → act → verify → complete, with defined transitions and failure states. This makes loops visible and testable instead of leaving termination to a model’s judgment.

Every loop needs a bound: maximum iterations, tool calls, elapsed time, and spend. Termination should depend on observable conditions—a validated result, a satisfied schema, or an exhausted deadline—not on another free-form “try again.” A verifier that cannot establish progress should stop or escalate.

Validate tools before execution. Check arguments against schemas, reject unavailable or unsafe operations, and make side effects idempotent where possible. Deduplicate equivalent retrievals, plans, and tool calls; an agent should not pay twice because two branches discovered the same work. Keep compact structured state—decisions, identifiers, results, and unresolved questions—instead of replaying the full transcript on every hop. Replayed context increases tokens without necessarily improving the next decision.

Escalate ambiguity and high-impact actions to a person. A human approval step can be cheaper than an unbounded investigation or an incorrect refund, deletion, or production change. The harness should preserve the task objective while narrowing waste: fewer calls, shorter context, and earlier stopping when additional work has no measurable value. Better orchestration is not reduced autonomy; it is autonomy with operating limits.

Route by difficulty and risk

Most steps do not require a frontier model. Use code for deterministic preprocessing and validation; use retrieval plus a small model for classification, extraction, and straightforward tool selection. Reserve the expensive model for ambiguous reasoning, conflicting evidence, or actions where an incorrect decision has material impact.

Routing needs explicit escalation signals: low confidence, disagreement between checks, missing required fields, novel inputs, or a risk threshold breach. A cheap model can propose an answer, while a validator checks schema, citations, policy, or domain constraints. If validation fails, escalate with the compact state rather than replaying the entire conversation. If the frontier route is unavailable, return a bounded partial result or request human review; silently substituting a weaker model is usually worse.

Evaluate routing by successful-task cost, not token price. A cheap route that fails, retries, or causes an expensive correction is not cheap. Track quality, latency, escalation rate, and total spend per completed task, then compare those outcomes against an all-frontier baseline. The right router minimizes expected cost subject to a defined success and risk target.

A decision-tree illustration routing the same workflow through code, a small model, a mid-tier model, or a frontier model. Show confidence and risk th

Caching is a control on repeated work, not a dashboard vanity metric. Separate the mechanisms:

  • Prompt or prefix caching avoids reprocessing stable system instructions and schemas.
  • Retrieval caching reuses results for the same corpus, query, and authorization scope.
  • Tool-response caching reuses deterministic, safe-to-repeat results such as metadata lookups.
  • Semantic reuse carries forward a validated plan or extraction instead of asking another model to recreate it.

Each cache needs an explicit invalidation policy. Data changes, permission changes, tool side effects, model or prompt revisions, and time-sensitive facts can all make a hit unsafe. Never share entries across tenants or users without enforcing the same privacy and authorization boundary. Stale context is especially dangerous when it drives a purchase, access decision, or other irreversible action.

Track cache hits inside the workflow ledger. If a step has hit rate hh, its expected cost is roughly (1−h)C(1-h)C only when misses and hits have comparable outcomes; refreshes, validation, and stale-result failures can change that. Budget for those exceptions, and measure cache impact as cost per successful workflow—not as an isolated percentage.

Retries and fan-out need explicit limits, not good intentions. Set a maximum retry count per step, use exponential backoff, and open a circuit breaker when failures indicate an unhealthy dependency. Retried tool calls must be idempotent or protected by request keys; otherwise a timeout can duplicate a payment, ticket, or deployment. Give every run a deadline and cap parallel branches. Require approval before expensive, irreversible, or high-volume actions.

Track the controls that determine whether the workflow is economical: cost per successful task, step-level cost variance, retry rate, cache-hit rate, fan-out width, and budget-exhaustion rate. Alert on changes in distributions, not only averages. A rising retry rate or widening step variance usually appears before the monthly bill does.

Implementation checklist

Map the execution graph before deployment. Assign a budget to every hop, then instrument a ledger that records actual usage. Define stop conditions, route work by difficulty, cache stable inputs, and cap retries and fan-out. Review cost beside task quality and latency—not after them. Autonomy is an operational design choice, not a default. Predictable spend comes from making every step earn its place.

ShareLinkedIn

Keep reading

The Cheapest AI May Be the One That Stops Talking
inference costs / specialized models / machine control

The Cheapest AI May Be the One That Stops Talking

TypeSafe AI’s Jev model replaces text generation with typed probabilistic decisions for machine interaction, with the company claiming major speed and cost advantages over language models. The post would examine whether narrower outputs—not simply smaller models—are the more important path to cheaper production AI, while clearly separating company claims from independent evidence.

The “AI” Label Is Becoming a Product Risk
AI products / product strategy / risk

The “AI” Label Is Becoming a Product Risk

AI is simultaneously used as a research field, a product category, a marketing term, and a source of social claims about work, education, and inequality. This post argues that overloaded language creates expectations no implementation can reliably meet.

Kolibri’s 1M-Token Context Window Is Only Half the Sovereignty Story
Kolibri / open-weight models / mixture of experts

Kolibri’s 1M-Token Context Window Is Only Half the Sovereignty Story

Aleph Alpha’s Kolibri combines a 78B-parameter English-German mixture-of-experts architecture with 3B active parameters, a 1M-token context window, downloadable weights, and Apache 2.0 licensing. The post would focus on how model architecture, licensing, and local weight access fit together in the push for sovereign open-weight AI.

← All posts