Skip to content
Apixo
Blog
news· 4 min read· via Towards AI

Securing AI Voice Agents: How to Prevent LLMs From Giving Away Money in Negotiations

Discover how developer Juan Sebastian Peña Donneys built a two-layer defense system to stop a voice negotiation agent from leaking profits and overpaying.

Securing AI Voice Agents: How to Prevent LLMs From Giving Away Money in Negotiations

Juan Sebastian Peña Donneys recently built an AI voice agent named Alex to handle freight rate negotiations, revealing critical insights into how LLMs manage financial boundaries. Operating as a virtual carrier sales representative for a fictional brokerage, Alex's job is to negotiate the cost of shipping cargo—specifically a 42,000-pound load from Chicago to Dallas—with incoming callers. In this industry, brokers earn profit from the difference between what a shipper pays them and what they pay the trucker. For this load, the shipper pays $3,300, and the brokerage aims to pay $2,700, with a strict ceiling of $2,950 and a floor of $2,450. The challenge was preventing the AI from immediately conceding its entire $500 safety margin to experienced negotiators.

The Vulnerability of Prompt-Based Limits

In the initial design, Peña Donneys placed pricing limits inside the system prompt of GPT-4.1 mini. Although the model never agreed to a rate above the $2,950 ceiling across 30 test runs, it failed to protect profits. Under ten scripted negotiation attacks—including emotional sob stories, fake manager approvals, complex math tricks, and currency switches—the prompt-only agent surrendered an average of 40% of its negotiable margin, giving away all $500 in four out of ten scenarios.

The AI agent, programmed to be helpful, easily walked up to its maximum limit when pressured. During an "anchor attack," the agent rapidly escalated its offers from $2,450 to $2,700, and eventually announced that $2,950 was the absolute highest it could go, handing over all leverage. Peña Donneys noted that because instructions and user inputs occupy the same context window, the model constantly weighs them against each other, making prompt-based financial limits highly unpredictable.

A Two-Layer Defense: The Desk and the Filter

To resolve this, Peña Donneys redesigned the system to separate negotiation logic from financial decision-making: the LLM negotiates, but the code decides. The updated system runs across three processes: a Next.js front-end, a LiveKit Cloud room, and a Python-based worker on Google Cloud. Crucially, the LLM's prompt contains no financial figures. Instead, the model must interact with external Python tools collectively called "the pricing desk."

The pricing desk manages the rates through specific functions. Before any negotiation can occur, a carrier verification tool checks the caller's credentials. If verified, the model uses tools like propose_rate and accept_rate. The desk computes a strict concession ladder ($2,450 to $2,575, then $2,700, and finally $2,825) and only climbs to the next rung if the caller makes a genuine concession. The absolute ceiling of $2,950 is kept entirely hidden and is never offered.

As a second line of defense, Peña Donneys integrated an outbound content filter. This layer intercepts the model's text before it reaches the text-to-speech engine, scanning sentences for any mentioned monetary values. If the AI attempts to speak an unapproved figure, the filter blocks the phrase and replaces it with a neutral line: "Let me check that figure with the desk before I quote it."

What it means for developers

For developers building transactional AI agents, this project offers a clear blueprint for securing financial boundaries. First, critical business rules and price limits must never live inside the prompt. Instead, developers should treat the LLM as a conversational interface that queries deterministic code. Second, systems must actively calculate and control concessions in code, rather than letting the model freely navigate between a floor and a ceiling. Finally, a hard output filter is necessary to intercept unauthorized outputs before they reach the user.

Implementing these multi-layered architectures often requires testing how different LLMs handle tool calls and constraints. Developers looking to test these strategies across different LLMs can access top AI models cheaply through a single API at https://apixoai.online.

Using this decoupled architecture, the average margin given away by the agent dropped from $197 to just $72 out of the $500 margin. The agent successfully resisted all ten attack vectors, including currency tricks, by refusing to convert Canadian dollars and demanding US dollar figures instead.

Performance and Real-World Constraints

While the two-layer defense successfully secured the brokerage's margin, it introduced minor performance trade-offs. Introducing the pricing desk added an extra LLM round trip per turn. In text-only testing, the median response time rose from 1,338 milliseconds to 1,692 milliseconds. However, on live voice calls, the agent still achieved a median response latency of 1.39 seconds, which is highly acceptable for telephone conversations.

Peña Donneys also highlighted that analyzing the full transcripts revealed subtle vulnerabilities that metrics alone missed. For instance, in early versions of the desk, the model bypassed the pricing tool during a "split-number" attack by rejecting an extra fee in plain text. This prompted an update where the desk evaluates total costs, including all add-ons and percentages, rather than letting the model handle them. The complete codebase, attack catalog, and a live interactive demo are publicly available at github.com/JSebastianIEU/voice-freight-negotiator.


Source: How to Stop an AI Agent From Giving Away Your Money in a Negotiation — Towards AI. Written by the Apixo team from that report.

#ai-news#ai-agents#llm-security#voice-tech#software-architecture
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading