Skip to content
Apixo
Blog
news· 4 min read· via Towards AI

Why Local LLM Tool Calling Fails: The Hidden Runtime Parser Bottleneck

A deep dive into how local runtimes like Ollama, llama.cpp, vLLM, and SGLang parse tool calls, revealing why popular models like Llama and Phi often output plain text instead of structured data.

Why Local LLM Tool Calling Fails: The Hidden Runtime Parser Bottleneck

On October 6, 2026, an analysis of the main branches of four major local LLM runtimes—Ollama, llama.cpp, vLLM, and SGLang—revealed a hidden friction point in local AI agent development: tool-calling success heavily depends on the runtime's specific output parser. While developers often assume local models inherently support tool calling, the runtime must translate raw text outputs back into structured JSON or function calls. The study, conducted by Local Agent Lab, looked at 136 named tool-call parsers across these runtimes and found significant discrepancies in model family coverage.

The Parser Gap in Local Runtimes

Out of 41 model families that have a dedicated tool-call parser in at least one runtime, only eight are supported across all four. These universally supported families are Qwen, DeepSeek, Gemma, Mistral, GPT-OSS, Cohere Command, LFM2, and Muse Glimmer. Surprisingly, prominent model families like Meta's Llama and Microsoft's Phi are not universally covered. Llama has dedicated parsers in only vLLM and SGLang, while Phi has a dedicated parser solely in vLLM (specifically for phi4_mini_json).

In total, the four runtimes ship 136 named tool-call parsers: vLLM leads with 53, followed by SGLang with 42, Ollama with 25, and llama.cpp with 16. For the 41 model families identified, 8 are supported everywhere, 7 in three runtimes, 10 in two, and 16 in just a single runtime (such as Granite, Nemotron, and Jamba).

How Runtimes Choose and Execute Parsers

Because LLMs only generate plain text, they rely on specific wire formats (like XML tags, [TOOL_CALLS] markers, or custom JSON structures) to signal a tool call. If a runtime's parser fails to recognize these markers, the tool call is output as standard text, causing the agent to miss the call entirely.

The four analyzed runtimes take different approaches to selecting these parsers:

  • vLLM requires manual selection using the --tool-call-parser flag. It refuses to enable automatic tool choices (--enable-auto-tool-choice) unless a parser is explicitly named.
  • SGLang allows developers to manually name a parser or pass auto. The automated option reads the chat template and uses detection rules for 28 of its 42 parsers, though Llama 3 is notably absent from these auto-detection rules.
  • Ollama extracts the parser name directly from the model's configuration. If none is specified, it defaults to a generic parser that pulls the tool-call tag from the model's template.
  • llama.cpp matches the chat template against its 16 dedicated handlers. If no match is found, it dynamically constructs a parser from the template itself. For instance, Llama 3.x relies on this template-derived parser rather than a dedicated handler.

Discrepancies Between Code and Documentation

The analysis also highlighted that official documentation is frequently out of sync with the actual codebase. In vLLM, for instance, only 29 of its 53 parsers are documented on its tool-calling page. Eighteen parsers—including phi4_mini_json, kimi_k3, minimax_m3, and llama4_json—do not appear in the documentation at all.

SGLang's dedicated tool-parser page lists only 14 of its 42 parsers, though 41 of them can be found in its server-arguments reference. Meanwhile, llama.cpp's documentation names 15 "native" formats, but the source code reveals that only three (Functionary v3.2, GPT-OSS, and LLM-jp) actually utilize a dedicated handler. The remaining 12 rely on the template-derived parser.

What it means for developers

For developers building local AI agents, these findings mean that tool-calling failures are often a runtime configuration issue rather than a model capability issue. If a local model outputs a tool call as plain text instead of executing it, developers should verify if the runtime is using the correct parser or falling back to an incompatible generic path.

To avoid these local runtime headaches, developers can try top AI models cheaply through one API at https://apixoai.online, which simplifies access to models like Claude, GPT, Gemini, and DeepSeek without needing to manage local parsers, runtimes, or complex configurations.

When working locally, developers must inspect the runtime's source tree directly to find the most accurate list of supported parsers. For vLLM, this means checking vllm/tool_parsers/__init__.py. For projects requiring highly reliable tool calling, choosing one of the eight universally supported model families, such as Qwen or DeepSeek, remains the safest bet across different local runtimes.


Source: Local LLM Tool Calling Depends on Your Runtime’s Parser — Towards AI. Written by the Apixo team from that report.

#ai-news#llm#tool-calling#vllm#llama-cpp#ollama#sglang
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading