NVIDIA AI Dynamo
NVIDIA AI Dynamo is a distributed runtime for large language models: prefill (reading the prompt) and decode (writing the answer) can be split across separate GPUs and workers, with a router deciding which worker handles which request.
In Cordatus it uses the same launch wizard as any other application, with a rail built for its extra decisions.
Launching
Steps 1 and 2 are the same as any application — see the Launch Guide:
- Select Device
- Select Version
- Advanced Configuration, where the rail is Dynamo-specific.
The rail
| Section | What it holds |
|---|---|
| Model | The model to serve, from the model hub, User models or a Custom name/URL. |
| Runtime | Aggregated or disaggregated, and the router. |
| GPU pipeline | Which GPUs each worker gets. |
| Access | The chat interface. |
| Overrides | Container options, environment variables, notebook. |
| Worker arguments | Engine flags, per worker class. |
| Review & start | The frontend's command and every container that will be created. |
The Ready to start card at the foot of the rail names anything still missing — most often "a worker still needs a GPU".
Model
Three sources, as elsewhere: Model hub (Hugging Face, Ollama or NGC), User models (already on
this device) and Custom (a model name or a URL you paste, for example
nemotron-mini:4b).
Runtime
Processing mode
| Mode | Behaviour |
|---|---|
| Aggregated | One worker handles prefill and decode. Lower communication overhead, simplest to run. |
| Disaggregated | Prefill and decode run on separate GPUs — more flexible, at the cost of a KV cache transfer between them. |
Router
Which worker gets a request:
| Router | Behaviour |
|---|---|
| kv | Routes on KV score — reuses cached prefixes where possible. The default. |
| round-robin | Sequential and balanced. |
| random | A random worker. |
KV connector
How the KV cache is moved or offloaded. Available options depend on the mode:
- Aggregated —
none,lmcache,kvbm - Prefill workers (disaggregated) —
lmcache,kvbm,nixl,none - Decode workers (disaggregated) —
nixl
lmcache and kvbm add optional offload to CPU RAM or disk; where they do, the pane exposes the
cache sizes in GB RAM, GB CPU cache and GB disk cache.
GPU pipeline
A drag-and-drop board rather than a list of dropdowns.
- Available GPUs on one side, each showing its utilisation and free memory.
- Prefill and Decode worker columns on the other (in aggregated mode, a single Workers column).
- Drag a GPU onto a worker to assign it; Unassign this GPU puts it back.
- Add prefill worker / Add decode worker create more workers; Remove this worker deletes one.
The pane refuses to be complete while a worker has no GPU ("No GPU yet — assign one from Available GPUs"), and tells you when everything is allocated ("Every GPU is assigned to a worker."). At least one worker is required.
Each worker's subtitle shows its tensor-parallel size, TP=n.
Access
A single switch: Chat interface — Open WebUI, started beside the Dynamo frontend with an admin account created for you.
With it off, the model is reachable on the Dynamo frontend's own port.
Overrides
The same four tabs as any application — Container, Environment, Engine flags and Notebook — with the same lock markers on values the application fixes, and the same Suggest with AI action that fills all three at once and reports what it adjusted to fit your hardware.
Worker arguments
Engine flags, split by which pass they apply to:
- flags for every worker (aggregated), or
- flags for the pass that reads the prompt (prefill) and flags for the pass that writes the answer (decode).
If the image publishes none, the pane says so rather than showing an empty editor.
Review & start
Dynamo creates several containers, so the review pane says which command it is showing: the frontend's command, with the other containers listed beneath it under Containers to be created. Blockers are listed first and each one jumps to the pane that fixes it.
After launching
Every Dynamo container appears as one container group on the Containers page. Open the group to reach each worker's logs, parameters and ports individually.
The frontend is the container to call. Its Ports tab gives you the local address and lets you generate a public URL.