Track State, Not History: Why LLM Agent Memory is a Costly Liability
A new study shows that tracking structured state instead of full conversational history can slash LLM agent token costs by 94% while dramatically improving accuracy.

For developers building autonomous LLM agents, a common and costly obstacle is token bloat. As an agent proceeds through a multi-step task, appending every tool execution, observation, and reasoning step to its context window causes token usage to snowball. This history-heavy approach creates an $O(T^2)$ cost curve, quickly hitting rate limits and running up massive bills.
A recent research preprint offers a compelling alternative: tracking structured state instead of full execution history. The paper, titled SKILL.state (arXiv 2608.26263) and published in late August 2026 by researchers at Google and Purdue University, demonstrates that discarding reasoning traces and only passing an immutable skill specification, a structured JSON state, and the latest observation can reduce cumulative token consumption by up to 94 percent.
The Efficiency of State Over History
The core mathematical advantage of the SKILL.state architecture is that it flattens the cumulative cost curve to $O(T)$. By validating and updating a structured state patch at each step and then throwing away the conversational history, the prompt size remains relatively constant.
In the paper's synthetic Warehouse environment tests using Gemini-3-Flash over a horizon of $T = 100$ steps, the stateful agent consumed 65,408 cumulative tokens compared to 1,062,387 tokens for the traditional baseline. This represents a 16.2-fold reduction in token usage. The benefits are not limited to cost; at an equal token budget of roughly 1,800 prompt tokens per step, the structured state approach achieved an accuracy of 0.94, compared to 0.52 for a capped summary, 0.22 for ReAct with LLMLingua, and just 0.18 for a truncated window.
However, implementing this in production reveals practical challenges that benchmarks often overlook. Pierre-Laurent Medori, who runs engineering at the app platform GoodBarber, recently put these concepts to the test on his company's Model Context Protocol (MCP) server. Using the claude-haiku-4-5 model, Medori replayed automated translation tasks to compare traditional ReAct loops against structured state tracking.
The Reality of State Tracking in Production
Medori’s experiments highlighted just how heavy tool schemas can be. A test application with 62 tools required 17,763 tokens just for the tool definitions, while a larger app with 77 tools took 23,174 tokens. When an agent enters a ReAct loop, it opens every turn with these massive schemas, and any list it calls adds more data that the history never forgets. Medori observed that a standard agent run peaked at 694,434 prompt tokens in a 60-second window.
While caching can make this transcript cheaper—serving up to 93 percent of the prompt from cache—it does not make the context window any smaller or more reliable.
To address this, Medori attempted to build a custom runtime based on the SKILL.state concept, where the agent's state was tracked as a 56-token JSON object. However, the implementation proved highly complex. The model frequently struggled with state validation, inventing tool arguments, or prematurely overwriting state variables. According to the Google and Purdue paper, smaller open-weight models suffer from premature state overwrites in 68 percent of their failures. On claude-haiku-4-5, the failures stemmed from the runtime's own lack of validation.
What it means for developers
For software engineers designing agentic workflows, the transition from history tracking to state tracking requires a fundamental shift in how agent memory is managed.
First, developers must establish a strict verification rule: an agent's state is not what the model reports, but what the external environment has actively confirmed. Instead of relying on a simple "200 OK" status code or the model's internal belief, the runtime must perform deterministic, offline checks—such as verifying that a created database record exists, contains the correct parameters, and matches expected length limits—before committing any state changes.
Second, developers should carefully audit their tool definitions. Because tool schemas are sent with every turn, trimming inventories and utilizing token-counting endpoints is crucial to keeping the initial prompt small.
Third, there are limits to discarding history. When the deliverable of a task is the sequence of events itself—such as a forensic log analysis—the history cannot be thrown away. In these scenarios, developers must maintain a ledger rather than a state schema.
Testing these complex state-validation loops across different LLMs can quickly become expensive. Developers can try top AI models cheaply through one API at https://apixoai.online, which provides a cost-effective way to benchmark how different models like Claude, Gemini, and GPT handle structured state validation and schema compliance without managing multiple platform accounts.
Ultimately, moving state tracking out of the LLM's conversational memory and into a validated, external system of record is the key to building reliable, cost-efficient agents.
Source: Your Agent’s Memory Is a Liability: Track State, Not History — Towards AI. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

