Skip to content
Apixo
Blog
news· 3 min read· via NVIDIA Technical Blog

NVIDIA Introduces DIN Deploy for Local C++ AI Integration

Discover DIN Deploy, an open-source C++ sample collection leveraging ONNX Runtime and NVIDIA TensorRT RTX for efficient local AI application deployment.

NVIDIA Introduces DIN Deploy for Local C++ AI Integration

Integrating artificial intelligence capabilities into local software environments demands a portable model format, dependable runtime engines, and performance acceleration capable of scaling across diverse hardware architectures. To bridge this divide, NVIDIA has introduced Do Inference Now (DIN) Deploy, an open-source library consisting of practical C++ samples designed to streamline local application development.

By uniting ONNX Runtime with the NVIDIA TensorRT RTX execution provider, this release assists software engineers in transitioning seamlessly from raw model checkpoints to native, hardware-accelerated applications operating on both Windows and Linux systems. Developers can similarly access this underlying ONNX Runtime API functionality through WinML 2.0.

Architecture from Export to C++ Implementation

Every individual sample within the DIN Deploy repository originates with a Python exporter script responsible for fetching a model checkpoint directly from Hugging Face and transforming it into a standardized ONNX artifact. The application layer itself consists of a native C++ command-line interface constructed on top of ONNX Runtime.

Separating the model conversion phase from the actual deployment logic allows engineering teams to integrate an exported model into local software without needing model-specific runtimes. The vast majority of the sample code relies directly on standard ONNX Runtime session and tensor APIs in C++. Specialized vendor code, such as CUDA APIs and execution kernels, remains isolated entirely to optional accelerated pathways. Execution providers compatible with the necessary ONNX Runtime tensor APIs can execute the shared code base seamlessly, while ONNX Runtime's copy tensor API helps maintain data locality.

For specialized image workflows, the FLUX.2 sample leverages graphics interop capabilities introduced in ONNX Runtime version 1.25, utilizing Vulkan and DirectX interfaces for efficient sampling during preprocessing and postprocessing stages. The repository supplies preconfigured CMake presets targeting Windows and Linux platforms, including specific variants for Arm64 architectures, though DirectX remains exclusive to Windows environments.

Supported AI Workloads and Performance

DIN Deploy covers a broad spectrum of AI tasks, including automatic speech recognition with both offline and streaming capabilities. OpenAI Whisper manages offline transcription tasks across multiple model sizes, whereas NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming handle real-time streaming pipelines. Furthermore, Meta SAM 2.1 samples enable interactive image and video masking, transforming model outputs into precise segmentation masks suited for computer vision applications.

For generative tasks, the FLUX.2-klein-4B sample delivers prompt-driven image generation incorporating Vulkan and DirectX graphics interop. This setup permits applications to combine GPU-resident resources with cross-vendor shader interfaces while employing post-training quantization via the NVIDIA Model Optimizer to generate optimized ONNX assets.

Measurements on DGX Spark highlight the performance benefits of GPU acceleration over CPU execution across these workloads. For instance, openai/whisper-large-v3-turbo achieved 58.5x real-time speed on GPU compared to 3.8x on CPU. Similarly, nvidia/parakeet-tdt-0.6b-v3 reached 206.41x real-time performance on GPU versus 14.44x on CPU, while facebook/sam2.1-hiera-base-plus attained 38.3 frames per second on GPU compared to 0.5 frames per second on CPU.

What it means for developers

For developers building high-performance native applications, DIN Deploy simplifies the path from cloud-hosted model weights to local execution. By relying on a unified ONNX Runtime API and standardizing model formats, teams can minimize vendor lock-in and deploy accelerated pipelines on Windows and Linux. Whether you are building speech recognition, computer vision, or image generation tools, these open-source C++ samples provide a robust blueprint for production-ready local AI. For tasks requiring cloud-based inference, developers can try top AI models cheaply through one API at https://apixoai.online.

Getting started involves utilizing the repository's CMake presets, which automatically fetch ONNX Runtime and TensorRT RTX by default. Engineers can build the project, export their chosen models, and execute the command-line interface via TensorRT RTX or adapt the pipeline implementations directly into their own codebases.


Source: Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples — NVIDIA Technical Blog. Written by the Apixo team from that report.

#ai-news#nvidia#c#onnx#ai#tensorrt
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading