Skip to main content

NVIDIA AI Dynamo

NVIDIA AI Dynamo is a distributed runtime for large language models: prefill (reading the prompt) and decode (writing the answer) can be split across separate GPUs and workers, with a router deciding which worker handles which request.

In Cordatus it uses the same launch wizard as any other application, with a rail built for its extra decisions.

An aggregated Dynamo run: runtime, router, GPU pipeline, worker arguments, and the ten containers it creates

 


Launching​

Steps 1 and 2 are the same as any application — see the Launch Guide:

  1. Select Device
  2. Select Version
  3. Advanced Configuration, where the rail is Dynamo-specific.

The rail​

SectionWhat it holds
ModelThe model to serve, from the model hub, User models or a Custom name/URL.
RuntimeAggregated or disaggregated, and the router.
GPU pipelineWhich GPUs each worker gets.
AccessThe chat interface.
OverridesContainer options, environment variables, notebook.
Worker argumentsEngine flags, per worker class.
Review & startThe frontend's command and every container that will be created.

The Ready to start card at the foot of the rail names anything still missing — most often "a worker still needs a GPU".


Model​

Three sources, as elsewhere: Model hub (Hugging Face, Ollama or NGC), User models (already on this device) and Custom (a model name or a URL you paste, for example nemotron-mini:4b).


Runtime​

Processing mode​

ModeBehaviour
AggregatedOne worker handles prefill and decode. Lower communication overhead, simplest to run.
DisaggregatedPrefill and decode run on separate GPUs — more flexible, at the cost of a KV cache transfer between them.

Router​

Which worker gets a request:

RouterBehaviour
kvRoutes on KV score — reuses cached prefixes where possible. The default.
round-robinSequential and balanced.
randomA random worker.

KV connector​

How the KV cache is moved or offloaded. Available options depend on the mode:

  • Aggregated — none, lmcache, kvbm
  • Prefill workers (disaggregated) — lmcache, kvbm, nixl, none
  • Decode workers (disaggregated) — nixl

lmcache and kvbm add optional offload to CPU RAM or disk; where they do, the pane exposes the cache sizes in GB RAM, GB CPU cache and GB disk cache.


GPU pipeline​

A drag-and-drop board rather than a list of dropdowns.

  • Available GPUs on one side, each showing its utilisation and free memory.
  • Prefill and Decode worker columns on the other (in aggregated mode, a single Workers column).
  • Drag a GPU onto a worker to assign it; Unassign this GPU puts it back.
  • Add prefill worker / Add decode worker create more workers; Remove this worker deletes one.

The pane refuses to be complete while a worker has no GPU ("No GPU yet — assign one from Available GPUs"), and tells you when everything is allocated ("Every GPU is assigned to a worker."). At least one worker is required.

Each worker's subtitle shows its tensor-parallel size, TP=n.


Access​

A single switch: Chat interface — Open WebUI, started beside the Dynamo frontend with an admin account created for you.

With it off, the model is reachable on the Dynamo frontend's own port.


Overrides​

The same four tabs as any application — Container, Environment, Engine flags and Notebook — with the same lock markers on values the application fixes, and the same Suggest with AI action that fills all three at once and reports what it adjusted to fit your hardware.


Worker arguments​

Engine flags, split by which pass they apply to:

  • flags for every worker (aggregated), or
  • flags for the pass that reads the prompt (prefill) and flags for the pass that writes the answer (decode).

If the image publishes none, the pane says so rather than showing an empty editor.


Review & start​

Dynamo creates several containers, so the review pane says which command it is showing: the frontend's command, with the other containers listed beneath it under Containers to be created. Blockers are listed first and each one jumps to the pane that fixes it.


After launching​

Every Dynamo container appears as one container group on the Containers page. Open the group to reach each worker's logs, parameters and ports individually.

tip

The frontend is the container to call. Its Ports tab gives you the local address and lets you generate a public URL.