Functional AI is coming soon. Join the waitlist for early access.

// blog

7 min read

How Much Does an AI Agent Cost to Run in Production?

Concrete production cost breakdown for AI agents — real model prices, token math, tool call fees, and three levers that cut your inference bill in production.

Running an AI agent in production costs $0.0003 to over $0.50 per completed task — a range that spans three orders of magnitude depending on model choice, loop depth, and how well you manage context size. That wide spread makes agent cost one of the few infrastructure concerns that can break a SaaS business model if left unexamined.

What You're Actually Paying For

Every production AI agent has three billing layers:

LLM inference (tokens) is the dominant spend for most teams. Model providers charge per million tokens of input and output. As of late 2026, published rates range from $0.10/M tokens for compact models to $30/M for frontier reasoning workloads — OpenAI pricing, Anthropic pricing, Google AI pricing.

Tool call fees stack on top of token costs. OpenAI's web search tool costs $10 per 1,000 calls ($0.01 per call); Anthropic charges the same. Google Search grounding on Gemini 2.5 models costs $35 per 1,000 prompts. An agent that runs five web searches per task adds $0.05 before any token math.

Compute and orchestration is typically a rounding error for API-based agents. Serverless function invocations cost fractions of a cent per call. GPU costs only dominate if you self-host models at scale.

For the SaaS teams building on provider APIs, inference tokens are where to focus attention — and where surprises tend to land.

Why Agents Cost More Than You Expect

Here is the insight that changes every cost estimate: agents do not make one API call. They loop.

A 2025 peer-reviewed study of agentic coding workflows — "How Do AI Agents Spend Your Money?" — analyzed eight frontier LLMs on SWE-bench Verified and found that agentic tasks consume orders of magnitude more tokens than equivalent code reasoning or code chat interactions, with input tokens — not output — driving the bulk of the cost. The same study found token usage varies by up to 30× across runs of the same task on identical inputs.

The reason input dominates: context accumulation. Each loop iteration re-sends the full conversation history plus tool outputs from previous steps. By step five, the model may be processing 20,000 tokens of accumulated context just to decide its next action. OpenRouter's analysis of 100 trillion real-world tokens documents that average prompt length grew nearly fourfold between early 2024 and late 2025 — from ~1,500 tokens per request to over 6,000 — driven largely by agentic and programming workloads.

That 30× variance does not smooth out at scale. It surfaces as billing spikes.

Real Cost Per Task: Current Model Prices

Using published API rates (October 2026):

Agent typeModelTotal inputTotal outputTool callsCost/task
Lightweight classifierGemini 2.5 Flash-Lite~3K tokens~200 tokensnone~$0.0004
Lightweight classifierClaude Haiku 4.5~3K tokens~200 tokensnone~$0.004
Support triage (5 steps)Claude Sonnet 5.5~20K tokens~1K tokens2 web searches~$0.07
Research/coding agentClaude Opus 5.5~80K tokens~4K tokens5 tool calls~$0.45
Research/coding agentGPT-5.6 Sol~80K tokens~4K tokens5 tool calls~$0.57

Token estimates reflect cumulative context across all steps. Pricing sources: OpenAI, Anthropic, Google. Web search billed at $10/1K calls.

A well-scoped production agent — five reasoning steps, a stable system prompt, two tool calls — lands around $0.05–$0.10 per task on a mid-tier model. At 100,000 tasks/month, that is $5,000–$10,000 in model API spend before any optimization. At one million tasks, it becomes a board-level line item.

Three Optimization Levers

The majority of production agent spend is avoidable. Here are the three highest-leverage interventions, each achievable in a single sprint.

1. Prompt caching

Anthropic charges $0.20/M tokens for cached input reads on Claude Sonnet 5.5 — a 90% discount versus the $2/M uncached rate. OpenAI applies automatic 50% prefix caching to eligible prompts with no code changes required. The cache warms up once the prompt exceeds 1,024 tokens and stays warm across requests with a stable prefix.

The math on system prompt caching alone — a 5,000-token system prompt sent with 100,000 daily agent calls:

  • Uncached (Sonnet 5.5): 5,000 × 100,000 × $2/M = $1,000/day
  • 70% cache hit rate (Anthropic): (3,500 × 100,000 × $0.20/M) + (1,500 × 100,000 × $2/M) = $70 + $300 = $370/day
  • Daily saving: $630 (~$19,000/month from one prompt)

To capture this: keep your system prompt static and long; put variable per-request content further down in the context so the stable prefix caches reliably across sessions.

2. Model routing

Not every step in an agent loop justifies a frontier model. Planning steps — deciding which tool to call, forming a retrieval query, structuring a response outline — work well on compact, fast, inexpensive models. Execution steps requiring precision or synthesis justify the premium.

Routing half of your agent steps from Claude Sonnet 5.5 ($2/M input) to Claude Haiku 4.5 ($1/M input) cuts input cost on those steps by 50%. Combined with caching on the system prompt, blended effective input rates fall well below the listed price of any single model.

Functional AI's Agentic Function Runtime is designed to encode routing logic in fn.yml alongside quality eval thresholds — version-controlled, testable, and applied consistently across every Agentic Function invocation rather than wired into application code. Join the waitlist to see how this works end to end. Trust your agents in production.

3. Context window management

Re-sending accumulated context is the largest avoidable cost in multi-step agent loops. CockroachLabs' production analysis of agentic AI costs finds that agentic workflows can consume 5 to 30 times more tokens per task than a standard chatbot query — and a significant portion of that difference is context the model already processed in prior steps.

The practical fix: anchored iterative summarization. Instead of re-transmitting raw tool outputs and conversation history on every iteration, the agent summarizes completed steps into a compact working-memory block and carries that forward. The summarization call costs a fraction of re-sending the full history. The savings compound across every subsequent step.

The Budget Enforcement Problem

Optimization improves average cost. It does not protect you from outliers.

A multi-agent research pipeline entered an infinite loop between two agents in November 2025. Neither had a session budget ceiling. Monitoring alerts fired — but no one responded in time. The pipeline ran for 264 hours. Final cost: $47,000 — from a system budgeted at under $200/month.

The lesson is architectural: cost monitoring is not cost enforcement. An agent with a hard per-session limit of $0.50 that terminates immediately when that ceiling is reached cannot generate a $47,000 bill regardless of what happens in the loop. An agent with a monitoring dashboard and no hard ceiling can.

Effective production budget control needs three layers:

  1. Session ceiling — maximum tokens or dollars per individual agent run, enforced at termination
  2. Daily aggregate cap — catch drift before it runs for days
  3. Loop circuit breaker — detect repeated identical tool calls, a reliable signal the agent is stuck

When per-session spend limits are a first-class concept in the runtime rather than application code you wrote and maintain, runaway loops are caught early. That structural guarantee is part of the design intent behind Functional AI's Agentic Function Runtime.

What This Looks Like at Scale

Using the mid-complexity estimate (~$0.07/task before caching) on Claude Sonnet 5.5 with two web searches per task:

Monthly task volumeRaw costWith system prompt caching (est. 63% input savings)
10,000 tasks$700~$580
100,000 tasks$7,000~$5,800
1,000,000 tasks$70,000~$58,000

Model routing to cheaper steps reduces cost further. The exact savings depend on your system prompt size, cache hit rate, and step composition — but the savings from caching alone are calculable before you write a line of optimization code.

The expensive mistake is assuming the pilot cost is the production cost. Agent workloads compound in ways that single-turn chatbot economics do not.

FAQ

Q: What is the average cost per AI agent task in production?

For a mid-complexity agent — five reasoning steps, two tool calls, run on Claude Sonnet 5.5 — expect approximately $0.05–$0.10 per completed task before caching. Lightweight classification agents on compact models (Gemini 2.5 Flash-Lite, Claude Haiku 4.5) run $0.0004–$0.004. Complex research or coding agents on frontier models (Claude Opus 5.5, GPT-5.6 Sol) can reach $0.45–$0.57 per task.

Q: Why do AI agents cost more to run than standard chatbots?

Agents run in loops. Each loop iteration re-sends the full conversation history and tool outputs accumulated so far, causing token counts to compound across steps. A 2025 arXiv study (2604.22750) found that agentic coding tasks consume orders of magnitude more tokens than equivalent code reasoning or chat interactions, with input tokens — not outputs — driving the majority of cost.

Q: How much does prompt caching reduce AI agent costs?

Significantly, on the input side. Anthropic's cache read rate for Claude Sonnet 5.5 is $0.20/M tokens — a 90% discount versus the $2/M uncached input rate. OpenAI's automatic prefix caching gives a 50% discount on cached prefixes. A stable 5,000-token system prompt cached at a 70% hit rate drops from $1,000/day to $370/day at 100,000 daily calls on Anthropic.

Q: What causes unexpected cost spikes in production AI agents?

Three main causes: (1) context accumulation without pruning — token counts grow with every loop step; (2) token usage variance — the same task can consume up to 30× more tokens on one run versus another (documented in arXiv 2604.22750); (3) infinite loops in multi-agent systems with no per-session budget ceiling.

Q: How do you set budget limits for production AI agents?

Three enforcement layers: a per-session token or dollar ceiling that terminates the run immediately when reached; a daily or hourly aggregate cap; and a circuit breaker that detects repeated identical tool calls and halts the loop. These need to be enforced mechanically — alert thresholds alone require human response time that runaway loops cannot wait for.

// private beta

Ship agents you can trust.

Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.

No spam. One email when Functional AI launches.

Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.