AIToday
Top Companies' AI MovesAI Business & IndustryTop Companies AIPublished: Oct 2, 2026, 06:30 JST

NVIDIA's DIN Deploy ships C++ local AI samples

NVIDIA's DIN Deploy ships C++ local AI samples

3 Key Points

  1. What happened

    NVIDIA published DIN Deploy, open-source C++ samples that combine ONNX Runtime with the TensorRT RTX execution provider, with support for Windows and Linux.

  2. Why it matters

    Developers get a shared code path with vendor-specific CUDA code only in optional accelerated paths, so the same samples can run across systems that support the required tensor APIs.

  3. What to watch

    The samples are only as portable as the execution providers that support ONNX Runtime's tensor APIs. Watch the DGX Spark results, where whisper-large-v3-turbo ran 58.5x real time on GPU versus 3.8x on CPU.

WHO IT HITSDevelopers building native, hardware-accelerated applications in C++ — particularly those targeting Windows and Linux — can use the samples as a starting point for local AI features without a model-specific runtime.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

DIN Deploy is aimed at a gap NVIDIA describes in its own framing: adding AI models to local applications needs a portable model format, a reliable runtime, and acceleration that works across target systems. The collection addresses that by splitting each sample into a Python exporter that converts a model checkpoint into an ONNX artifact and a native C++ CLI built on ONNX Runtime.

The repository covers several task types at once — automatic speech recognition, interactive masking for images and video, and prompt-driven image generation. Performance numbers measured on DGX Spark show GPU acceleration far ahead of CPU for these workloads, with openai/whisper-large-v3-turbo at 58.5x real time on GPU versus 3.8x on CPU, and facebook/sam2.1-hiera-base-plus at 38.3 FPS on GPU versus 0.5 FPS on CPU.

How useful the samples prove in practice is likely to hinge on whether execution providers support the required ONNX Runtime tensor APIs, since that is the condition NVIDIA sets for the shared code to run. Developers on Windows and Linux, including Arm64 variants, can start from the repository's CMake presets, which download ONNX Runtime and TensorRT RTX by default.

FAQ
What AI tasks do the DIN Deploy samples cover?
They cover automatic speech recognition with offline and streaming pipelines, image and video masking with Meta SAM 2.1, and prompt-driven image generation with the FLUX.2-klein-4B sample.
How much faster is GPU than CPU on these samples?
On DGX Spark, openai/whisper-large-v3-turbo ran 58.5x real time on GPU versus 3.8x on CPU, and facebook/sam2.1-hiera-base-plus ran 38.3 FPS on GPU versus 0.5 FPS on CPU.
Do I need to change my application code to use a quantized model?
The FLUX.2 sample shows post-training quantization with NVIDIA Model Optimizer produces a quantized ONNX model. Because ONNX interfaces stay unchanged, the quantized model is a drop-in replacement requiring no application-code changes.
Top Companies AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAmphenol orders hit $10.7 billion as AI demand runs hot