Akka Evaluates Spec-Driven AI Workflows Across 65 Open Source Ports
Akka tested a spec-driven workflow using 65 open-source projects, analyzing token use, model performance, and architectural outcomes.

Akka recently completed an extensive evaluation of a specification-driven workflow designed for AI-assisted software porting. The experiment spanned 65 open source projects, measuring key metrics such as specification structure, context management, model selection, effort configuration, automated validation, token consumption, and runtime performance. Across the initiative, the initial tranche consumed 9.41 billion tokens over 99.3 hours. Ultimately, Akka reported improvements in either lines of code or performance in 57 out of the 65 total ports.
The Delivery Harness and Workflow
The evaluation utilized a structured delivery harness operating in cycles of discovery, specification, porting, benchmarking, and improvement. During the discovery phase, the system examined code, models, schemas, and runtime behavior. For implementation, testing, and review, the process leveraged Claude combined with Akka Specify. A unified benchmark runner then evaluated tests, code size, and latency.
Initially, Akka analyzed all 65 projects by generating specifications and carrying out up to 10% of each project's surface area. Based on system characteristics and measurable outcomes, 10 projects were selected for complete implementation. According to the findings, structured specifications containing claims, evidence, and typed behavior enhanced first-pass implementations, although context gaps persisted regarding cross-component decisions.
Validation relied on original unit and integration tests alongside auditors checking serialization, security, error handling, PII, idempotency, and architectural boundaries. Guardrails were introduced as failures revealed new issues, though additional exit conditions raised overall porting costs. Performance outcomes varied by category: applications, frameworks, and libraries generally saw performance gains, whereas infrastructure and tooling experienced median performance degradation. Specific results included extreme variances, such as a reported 143,333 times improvement for Dify—though workloads differed—and Netflix Metaflow running approximately 100 times slower.
Model Efficiency and Engineering Discussion
The experiment compared different models and effort levels. Sonnet averaged 61 minutes per port compared to 120 minutes for Opus, while Opus consumed roughly 40% fewer tokens. Interestingly, higher effort settings increased token consumption without consistently driving better efficiency.
These results sparked technical discussions among engineers on LinkedIn. Aaditya noted that smaller models followed specifications closely, whereas larger models tended to improvise, mirroring modernization patterns he had observed. He also questioned whether reductions in code volume stemmed from removing dead code or language differences. Rick Bryce, Head of Marketing at Avahi, suggested that strict constraints might heavily influence final outcomes, leaving open questions about whether model capability or specification constraints drive porting efficiency.
What it means for developers
As organizations experiment with large language models for complex software engineering tasks like code porting, managing token consumption and maintaining strict specification guidelines are crucial. Developers looking to experiment with these capabilities can try top AI models cheaply through one API at https://apixoai.online, allowing flexible model selection and cost-effective testing for automated workflows. Future focus areas highlighted by the research include interface enumeration, test ingestion, provenance tracking, differential testing, and adversarial testing.
Source: Akka Tests Spec-Driven AI Delivery Across 65 Open Source Projects — InfoQ AI & ML. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

