Skip to content
Apixo
Blog
news· 3 min read· via Latent Space

GPT-6.1 Sol and Sonnet 5.5 Driven Cost Cuts Shake Up AI Benchmarks

OpenAI introduces GPT-6.1 Sol and Anthropic updates Sonnet 5.5, sparking price cuts and performance shifts across AI agent benchmarks and developer harnesses.

GPT-6.1 Sol and Sonnet 5.5 Driven Cost Cuts Shake Up AI Benchmarks

Recent benchmark entries and pricing updates from major AI vendors have reframed the cost-to-performance ratio for automated coding and agent workflows. OpenAI's release of GPT-6.1 Sol, alongside new submissions from Anthropic and Google, demonstrates a concerted push to bring down API expenses while boosting task accuracy across software engineering and general reasoning evaluations.

Shifts on the Leaderboards: GPT-6.1 Sol and Sonnet 5.5

OpenAI has introduced GPT-6.1 Sol with a price structure of $2.00 per million input tokens and $10.00 per million output tokens, marking a sharp discount compared to Astra’s $10.00/$50.00 tier. According to reported benchmarks, Sol outperforms GPT-6 Sol by 6.4 points on DeepSWE v1.1 and beats Opus 5.5 by 2.2 points on AutomationBench. Described by OpenAI staff as "good, cheap AND fast," Sol entered the Agent Arena at #5 with a median task cost of $0.56—making it 39% cheaper than GPT-6 Sol while scoring 1.52 points higher.

Anthropic simultaneously expanded its presence, with Sonnet 5.5 [Max] debuting at #3 in the Agent Arena with a +12.5% improvement and claiming #1 in the Chat category. At $2.74 per task, it sits alongside Opus 5.5 ($1.58 per task), helping Anthropic maintain the top three positions in the Agent Arena. On the WebDev leaderboard, Sonnet 5.5 sits 2 points behind GPT-6 Astra [Max] at an 80% lower cost. Meanwhile, Google’s Gemini 4 Argon [High] captured the #1 position in the Text Arena.

Decision Models and Edge Hardware Experiments

Beyond proprietary API models, open-weights ecosystem activity is accelerating around specialized decision models. Cloudflare announced Clef, an open-weights decision model post-trained from Qwen3.8-27B, alongside a smaller clef-flash variant based on Qwen3.5-9B. In local tooling, llama.cpp added a /v1/systemone endpoint to support local inference for Kev-4B-GGUF. Perplexity also reported benchmark gains for its pplx-decider-v1-27b, which averaged 85.7% across 11 tests, though community researchers like @mervenoyann note these decision models function essentially as rebranded zero-shot classifiers.

Developers are also pushing edge hardware boundaries. One developer reported offloading layers 41–64 of Qwen 3.8 27B from a 24 GB M4 Pro MacBook to an iPhone 17 Pro Max via 10 Gb/s USB-C, utilizing Metal 4 tensor ops. The setup delivered prefill speedups of up to +44% at 16k context and allowed older KV cache pages to sit on the phone's memory, lowering generation latency at 140k context from 279 ms/token to 176 ms/token.

Agent Harnesses and System Tooling

Agent orchestration frameworks have seen significant structural updates. The T3 Code orchestrator merged a four-month pull request spanning 823 commits and 1,912 files as its user base surpassed 400,000. The rewrite adds cross-provider task delegation, mid-thread model switching, and Pi integration.

Concurrently, Earendil released Pi 1.0, embedding Model Context Protocol (MCP) support by default alongside deferred tool loading and Anthropic cache warming. DeepSeek shipped desktop builds of DeepSeek Harness for macOS and Windows, while OpenAI updated its Agents API with one-call browser computer use, 99.97% turn reliability, and 20% faster tool calls.

What it means for developers

The steep decline in execution costs for frontier-level reasoning models means complex, multi-step agent workflows are becoming far more economically viable. Sol's $0.56 median task cost and Sonnet 5.5's competitive WebDev performance give engineering teams much greater headroom to run extensive test loops and automated code generation without ballooning infrastructure budgets.

Furthermore, the expanding ecosystem of agent harnesses and local decision models provides developers with fine-grained control over local and cloud orchestration. Whether testing local edge setups or integrating cloud endpoints, developers can try top AI models cheaply through one API at https://apixoai.online to evaluate performance across models like Claude, GPT, Gemini, and DeepSeek before committing to dedicated infrastructure.


Source: [AINews] not much happened today — Latent Space. Written by the Apixo team from that report.

#ai-news#ai#openai#anthropic#llm#developer-tools
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading