Prime Intellect Debuts Prime Inference for Frontier Open Models
Prime Intellect launches Prime Inference, bringing serverless and reserved serving capabilities to open-source frontier models like GLM-5.3.

Prime Intellect has officially launched Prime Inference, a dedicated model serving platform designed specifically for frontier open-source artificial intelligence models. Operating on Prime's own GPU infrastructure across multiple datacenters, the service provides both on-demand serverless endpoints and reserved capacity. Prior to its public release, Prime Intellect stress-tested the platform internally, processing nearly one trillion tokens per day across workloads such as reinforcement learning rollouts, synthetic data generation, evaluation benchmarks, and long-running autonomous coding agents.
Prime Inference serves as the production deployment layer of Prime Intellect’s open training ecosystem, joining existing post-training tools like prime-rl, sandboxes, and verifiers. By hosting deployed models directly, the platform captures live production traces that can subsequently be fed back into training pipelines. According to company reports, the platform’s GLM-5.3 endpoint ranks among the fastest available on OpenRouter, maintaining a near-zero tool-call error rate and achieving 100 percent uptime since launch through automated cross-datacenter failover routing.
Under the hood: Architecture and optimizations
The serving stack powering Prime Inference combines several open-source technologies, including NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, developed in partnership with Inferact and NVIDIA. The platform is engineered specifically to target heavy agentic workloads, where a typical agent turn appends approximately 6,000 tokens to an existing 140,000-token prompt context.
To manage these high-context demands, the stack decouples prefill and decode execution onto separate GPU groups. NVIDIA Dynamo handles session routing while vLLM executes model passes on each group, utilizing NIXL to transfer computed key-value (KV) states. Internal testing shows this prefill/decode disaggregation reduces p90 inter-token latency by nearly 40 percent. Additionally, Dynamo's cache-aware router evaluates cached prefix overlap against active queues to maintain session affinity on specific decoders, while Mooncake introduces a secondary KV cache tier located in host DRAM.
When benchmarking GLM-5.3 on NVIDIA GB200 NVL72 hardware, Prime Intellect targeted an interactivity baseline of 100 end-to-end tokens per second per user. A 1:4 prefill-to-decode ratio served 66 concurrent sessions per prefill group at 101 tokens per second per user and 100 output tokens per second per GPU. Using a DEP8 prefill topology yielded roughly five times the usable prefix-cache capacity compared to TEP8 configurations. Furthermore, cutting the prefill budget from 8,000 to 4,000 tokens per step per GPU lowered median queue wait times from 550 ms to 110 ms, while median time to first token (TTFT) improved by approximately 20 percent.
To optimize memory efficiency, the platform applies NVFP4 KV cache compression, shrinking each Multi-head Latent Attention (MLA) cache row from 576 bytes to 352 bytes and raising token capacity per decoder from 1.09 million to 1.63 million tokens. A custom native sparse-MLA kernel executed in roughly 12.0 microseconds for 15 query tokens, compared to 17.7 microseconds for staged execution. Meanwhile, updating the KV layout to BLHNC reduced transfer descriptors from 19,559 to around 1,940, reducing mean transfer time from 146 ms to 78 ms.
What it means for developers
For software engineers building agentic workflows, Prime Inference provides full compatibility with the OpenAI SDK via a unified endpoint at https://api.pinference.ai/api/v1, featuring team-level usage tracking and consolidated billing. Addressing a common failure point where agent tools fail due to schema mismatches or broken argument formats, Prime Intellect built a structural-tag builder into Dynamo for GLM's tool format and applied token masking in vLLM via xgrammar to enforce schema rules. Parsing bugs, including < incorrectly decoding into < within code outputs, were also resolved.
As open-source AI infrastructure advances, developers looking to test and build across a broader selection of state-of-the-art systems—including Claude, GPT, Gemini, Grok, and DeepSeek—can access top AI models cheaply through one API at https://apixoai.online.
Market context and future roadmap
In the competitive landscape for hosting models like GLM-5.3, alternative inference providers such as Together AI, Fireworks AI, and Baseten offer serverless endpoints priced around $1.40 per million input tokens and $4.40 per million output tokens. While Prime Intellect has not yet fully published its individual per-model rates in its documentation, it distinguishes itself by providing reserved capacity options alongside standard serverless access.
Currently operating on NVIDIA Blackwell hardware, Prime Intellect has listed NVIDIA Vera Rubin support as coming soon. The company's future technical roadmap also includes adding batch inference capabilities and single-click dedicated GPU deployments.
Source: Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models — MarkTechPost. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

