Skip to content
Apixo
Blog
news· 3 min read· via Towards AI

Why Agent Costs Depend on Context Loops Rather Than Model Prices

Switching from Opus to Haiku rarely yields the expected savings in tool-heavy agents. Token re-injection, retries, and context loops drive the bulk of API expenses.

Why Agent Costs Depend on Context Loops Rather Than Model Prices

Switching an AI agent from Anthropic’s Opus at $4 per million input tokens to Haiku at $1 per million input tokens theoretically suggests a 75 percent cost reduction. In production, however, teams running tool-heavy agent loops frequently observe far smaller savings. The reason is that the published per-token rate is rarely the primary factor in the total bill. Instead, the repetitive loop executing around the model—re-ingesting history, executing tool calls, and handling retries—drives the majority of token consumption.

Where tokens accumulate in agentic loops

An agentic workflow functions as a multi-step cycle: the model plans an action, executes a tool, reads the output, and determines whether to proceed or terminate. Every stage represents an independent API transaction. While generating a tool call requires relatively few output tokens, the data returned by external tools must be passed back into the context window for subsequent turns.

Because Anthropic's Messages API resends the entire conversation history with each turn, every tool result is billed repeatedly. If a task demands six tool calls before completion, the result of the first tool is not billed once; it is re-submitted and re-billed as input tokens across all five subsequent calls. A single failure or retry compounds this effect, as the system must re-transmit the complete system prompt, tool schemas, and prior history just to repeat a step.

Extended thinking introduces an additional billing distinction. Thinking tokens are classified as output tokens and tracked within usage.output_tokens_details.thinking_tokens. On Opus 5.5 and newer, previous thinking blocks remain in the conversation context and are billed again as input tokens during subsequent turns. Conversely, on Sonnet and Haiku, these thinking blocks are stripped out between turns and do not carry forward, creating differing cost curves across tiers.

Architectural levers for lowering agent expenses

To manage these escalating expenses, developers must rely on architectural controls rather than solely changing model tiers. Several mechanisms provide measurable reductions:

  • Context editing: Anthropic’s clear_tool_uses_20250919 edit type automatically replaces older tool results with placeholders once input tokens exceed a configured threshold. In an Anthropic documentation benchmark, clearing stale tool outputs dropped input tokens on a single turn from 70,000 down to 25,000—a 64 percent reduction achieved without altering the prompt, task, or model tier.
  • Prompt caching: Writing to the cache costs 1.25 times the standard input rate for a 5-minute time-to-live (TTL) or 2 times for a 1-hour TTL. However, cache hits cost only 10 percent of the base input rate on most models, and 5 percent on Opus. Reusing system prompts and tool schemas across long runs yields significant compounding savings.
  • Batch processing: Asynchronous, non-interactive requests receive a flat 50 percent discount on both input and output tokens across Haiku, Sonnet, and Opus, making it ideal for background summarization or re-indexing tasks.
  • Token counting: Before executing workflows, teams can use Anthropic's token-counting endpoint, which is free to call, to forecast token volume across prompts and tool schemas.

What it means for developers

For engineering teams, managing agent costs resembles application profiling rather than vendor price negotiations. Merely downgrading model tiers without addressing context inflation leaves the most expensive parts of the loop untouched.

A disciplined approach requires tracking explicit metrics returned in the Messages API usage object: input_tokens, output_tokens, cache_creation_input_tokens, cache_read_input_tokens, and thinking_tokens. Correlating these metrics with individual loop stages reveals whether retries, tool result accumulation, or context drift are driving the bill.

When exploring different operational setups, developers can try top AI models cheaply through one API at https://apixoai.online, allowing them to test architectural optimizations across various models without maintaining separate integrations. Ultimately, cost efficiency in agent engineering stems from pruning context, caching static schemas, and instrumenting each loop iteration before attempting to resolve high bills by swapping models.


Source: Why Your Agent Loop Costs More Than Its Model Price Tag — Towards AI. Written by the Apixo team from that report.

#ai-news#ai-agents#llm-finops#anthropic#claude#prompt-caching
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading