NVIDIA Launches DGX Spark 64GB Desktop AI System for Local Workloads
NVIDIA has expanded its desktop AI lineup with a 64GB configuration of the DGX Spark, offering a local alternative to metered cloud APIs for developers running open models and agents.

NVIDIA has introduced a new 64GB configuration of its DGX Spark desktop artificial intelligence system, built in partnership with OEMs including Acer, ASUS, Dell, Gigabyte, HP, and MSI. The hardware is designed to give developers a dedicated desk-side environment for running open-source models and autonomous agents locally, circumventing the ongoing costs associated with metered cloud APIs. Continuous token consumption for agentic workflows—driven by tool calls, retries, and multi-step reasoning—has escalated significantly since early 2026, making local hardware an increasingly appealing alternative to pay-per-token cloud services. For developers seeking flexible cloud access without maintaining dedicated hardware, platforms like https://apixoai.online provide a convenient way to try top AI models cheaply through one API.
Hardware Specifications and Unified Memory Architecture
The 64GB variant retains the core architecture of the original system, utilizing the NVIDIA GB10 Grace Blackwell superchip. This combines a 20-core Arm CPU—comprising ten Cortex-X925 and ten Cortex-A725 cores—with a Blackwell GPU featuring fifth-generation Tensor Cores. The system delivers up to one petaFLOP of FP4 AI compute with sparsity. Instead of the 128GB found in the initial model, this configuration integrates 64GB of unified LPDDR5x memory operating at 273 GB/s bandwidth.
The defining design element is the coherent unified memory architecture enabled by NVLink-C2C, which connects the CPU and GPU at five times the bandwidth of PCIe Gen 5 without requiring weight duplication between system RAM and VRAM. Running on a standard wall outlet, the device measures 150 by 150 by 50.5 millimeters and weighs 1.2 kilograms. It includes self-encrypting NVMe M.2 storage options ranging from 1TB to 4TB, a ConnectX-7 network interface card supporting 200GbE, Wi-Fi 7, and Bluetooth 5.3. Out of the box, the Ubuntu-based NVIDIA DGX OS includes pre-configured software stacks such as PyTorch, Jupyter, Ollama, and single-command installations for NVIDIA NemoClaw and NVIDIA OpenShell guardrails.
Optimized Models and Practical Use Cases
The 64GB memory footprint accommodates 30B to 35B parameter open-weight models that suit local execution. Key compatible models include Meta's Muse Glimmer, a 29.6B dense text-and-image model running a quantized footprint of roughly 17GB; NVIDIA's Nemotron 3.5 Lightning, a 30B mixture-of-experts model; and Alibaba's Qwen3.8-27B.
Developers can leverage this setup for five primary applications:
- Always-on personal agents: Running background tasks like triaging repository issues or summarizing research papers securely.
- Local fine-tuning: Utilizing QLoRA on a 70B model or executing distributed fine-tuning for custom coding assistants.
- Model evaluations: Testing newly released Hugging Face open-weight models instantly via Ollama or vLLM without incurring API fees.
- Multi-model agent teams: Combining models like Bonsai 2 as a router with Glimmer for reasoning within the unified memory pool.
- Edge prototyping: Validating vision transformers locally before deploying them to NVIDIA Jetson devices.
What it means for developers
For developers building complex agentic systems that burn through tokens via continuous tool calls and long context windows, owned hardware removes per-token financial friction. The inclusion of ConnectX-7 networking enables two DGX Spark 64GB units to cluster directly using a QSFP cable without requiring an external switch. Clustering two 64GB systems yields 128GB of pooled memory and up to 1.7 times the performance of a single 128GB DGX Spark, doubling combined bandwidth to 546 GB/s. The NVIDIA Sync management software and Cluster Assistant simplify multi-node setups across local networks or Tailscale meshes. While the system is optimized for long-input and short-output tasks rather than high-concurrency chat serving, it provides a robust platform for local development, privacy-focused agent operations, and scalable multi-node computation.
Source: NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference — MarkTechPost. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

