Skip to content
Apixo
Blog
news· 3 min read· via Towards AI

Claude Code Tool Search Slashes Agent Token Overhead by Nearly Half

Benchmarks reveal Claude Code's lazy tool loading cuts upfront token consumption by 64% and entire task context by 46%, though custom proxy setups require manual configuration.

Claude Code Tool Search Slashes Agent Token Overhead by Nearly Half

As developer tools incorporate increasingly complex Model Context Protocol (MCP) toolkits, the context window required just to initialize an AI agent has ballooned. In benchmark tests conducted on Claude Code version 2.1.285, Anthropic's tool search mechanism—a lazy-loading approach for tool definitions—reduced upfront prompt payload by 64% in a multi-server setup and cut total token consumption across a full task by 46%.

Traditionally, Claude Code loads every connected tool definition eagerly, inserting thousands of tokens into the context window before a user even enters a prompt. Tool search alters this lifecycle. Instead of transmitting complete schemas on every turn, the runtime supplies full definitions for only 11 core tools—including the ToolSearch mechanism itself—and substitutes deferred tools with a compact list of tool names. When the underlying model determines it needs a capability, it searches for it, prompting the system to load the full schema on demand.

Benchmarking eager versus deferred tool loading

Measurements conducted using OpenAI's o200k_base tokenizer demonstrated how steep the tool definition overhead can be. In a standard setup with five MCP servers connected—including filesystem, memory, everything, sequential-thinking, and Microsoft Playwright servers totaling 65 MCP tools—eager mode submitted 27,184 tokens on the first request. Enabling tool search slashed that initial payload to 9,722 tokens, marking a 64% reduction.

Even in configurations without third-party MCP servers, tool search halved initial overhead from 17,546 tokens down to 8,756 tokens, because Claude Code 2.1.285 defers 13 of its 23 built-in utilities by default. For the MCP tools alone, schema weight dropped from 9,638 tokens eagerly loaded to just 966 tokens for the deferred list—a 90% reduction on external tool context.

Across a complete task—instructing the agent to record a reminder using the memory server's create_entities tool—the savings settled at 46%. Eager mode finished in two requests totaling 54,422 tokens. Lazy mode required an extra round trip for the model to search and load the tool definition, concluding in three requests at 29,543 tokens. Once loaded, the requested tool schema added 140 tokens to subsequent turns, still leaving per-request context 64% lower than in eager mode.

While tool search functions automatically against Anthropic's native endpoints, tests revealed that setting a custom endpoint via ANTHROPIC_BASE_URL disables the feature by default. In proxied environments, Claude Code 2.1.285 falls back to eager loading, sending all tool schemas over the wire and negating the token savings.

This behavior occurs because Claude Code disables tool search for non-first-party hosts unless explicitly instructed otherwise. Developers routing traffic through local servers, custom gateways, or reverse proxies must set the environment variable ENABLE_TOOL_SEARCH=true (or update their ~/.claude/settings.json) to retain lazy loading, provided their gateway correctly forwards tool_reference blocks.

What it means for developers

For engineering teams orchestrating multi-tool agents, tool search addresses one of the most persistent hidden costs in production: paying recurring token context fees for tools that are rarely invoked during a session.

The trade-off is latency versus context expenditure. Lazy loading introduces an extra round trip when a tool is first discovered. However, because subsequent requests carry only the newly loaded tool's schema rather than an entire registry, cumulative token savings compound heavily during extended debugging sessions or complex workflows.

Developers managing agent toolchains across different providers also face the challenge of varying token economics. For teams testing and running models across providers, developers can try top AI models cheaply through one API at https://apixoai.online, simplifying access to Claude, GPT, and other architectures through a single endpoint.

Ultimately, as agent ecosystems scale past dozens of MCP tools, lazy loading mechanisms like tool search will be essential to prevent context exhaustion and keep per-task token consumption under control.


Source: Claude Code Tool Search Nearly Halves Your Context Bill — Towards AI. Written by the Apixo team from that report.

#ai-news#claude-code#anthropic#mcp#ai-agents#token-optimization
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading