Quickstart
Get a model serving on your own hardware in about 10 minutes. This page is the short version; every section links to the full guide.
Prerequisites
- A Cordatus account with the container permissions
- At least one device connected to Cordatus and showing as Online
- A GPU on that device (recommended for LLM applications)
- The Cordatus Client installed and running on it
No device yet? Do the Device Hub quickstart first — nothing in LLM Launcher works without one.
The fastest route: AI Launch (about 4 minutes)
If you only want a model running, skip the wizard entirely.
- Go to LLM Launcher → LLM Models.
- Search for a small model —
Qwen/Qwen3-8Bis a good first one. - Press AI Launch on its row.
- Watch the six planning phases. When it says Plan ready, read the summary line: Ready to launch on {device}.
- Open Details if you want to see the engine, image, GPUs, quantization and context it chose, and Why these settings for the evidence behind each.
- Turn on Chat interface under the plan, and copy the credentials it shows.
- Press Launch.
That is the whole thing. Skip to Verify your deployment.
The full route: the launch wizard (about 10 minutes)
1. Browse the applications (1 minute)
LLM Launcher → Applications. Filter by device type — Server - Workstation, Jetson or DGX Spark — and pick an engine:
- vLLM — high-throughput LLM inference
- TensorRT-LLM — NVIDIA's optimised engine
- Ollama — simple local models
- llama.cpp — GGUF weights, including on Jetson
- NVIDIA Dynamo — distributed multi-GPU
- NVIDIA VSS — video search and summarisation
Open one with View Details and look at its Versions tab.
2. Start the wizard (1 minute)
Press Start Single Device Application.
Step 1 — Select Device. Pick a connected device. The footer confirms Deploying to {device}.
Step 2 — Select Version. Only compatible images are listed. Each is marked:
- Downloaded — already on the device
- Will Download — pulled on start
- Incompatible — with the reason, for example a driver version mismatch
Press Next.
3. Configure (4 minutes)
Step 3 is a rail. Work down it.
Profile (top of the screen) — press Quick start. That switches on the chat interface and touches nothing else.
Model
- Stay on the Model hub card.
- Pick a small model — 7B or 8B for a first deployment.
- Read the fit line under the catalogue. You want "{model} fits your selection comfortably."
Compute
- Select a GPU, or switch on Use all.
- Leave Host limits on Auto.
Access — already set by the Quick start profile. Copy the credentials it shows; they are displayed once.
Overrides — skip it. The application's defaults are already correct.
Review & start — read the Resulting command and the Containers to be created. If a blocker is listed, click it: it takes you to the pane that fixes it.
4. Launch (1 minute, plus the image pull)
Press Start Environment.
Behind the scenes Cordatus connects to the device, pulls the image if it is not there, creates the containers, configures networking and volumes, and starts the serving layer alongside.
A first pull can take several minutes. The container shows as Downloading while it happens.
Verify your deployment
- Go to LLM Launcher → Containers.
- Find the new container group and check its state.
- Open the group, then Container Informations on the engine.
- Logs — you should see the model loading.
- Ports — copy the local address.
Test it:
curl http://<device-ip>:<port>/v1/models
Then open the Open WebUI container's address in a browser and sign in with the admin credentials you copied.
Expected result
- The engine container is running, with model-loading messages in its logs
- The Open WebUI container is running, and you can chat with the model
- Both appear as one group under Containers
- If you enabled the gateway, a LiteLLM container (plus its PostgreSQL, Redis and Prometheus) is running too, and the model is registered into it
What's next
Check a model fits before deploying it (5 minutes)
LLM Launcher → VRAM Calculator:
- Search for a bigger model, for example Llama-2-70B.
- Choose a GPU — or Select Device GPUs to use your real hardware.
- Set the sequence length and batch size you actually intend to serve.
- Switch Calculation Type to With Activation.
- Compare quantizations until it says VRAM is Sufficient.
Use models already on your device (10 minutes)
- Devices → your device → Models → Storage Paths — register the HuggingFace, Ollama and NIM paths.
- LLM Models → User Models → Explore Models on Your Devices — pick the device, Start scanning, choose target engines, Add {n} model(s).
- Press AI Launch on one of them — the plan is made for the device that stores it, and nothing is downloaded.
Download a model onto a device (5 minutes)
LLM Models → Download on any row: pick the device, the registered directory and the quantization, add a token if the repository is gated, and let the device fetch it.
Serve it properly (10 minutes)
LLM Launcher → Serving Layer: deploy a LiteLLM gateway, register your running model into it, and issue scoped virtual keys for your clients instead of handing out the master key.
Go further (15–30 minutes)
- NVIDIA AI Dynamo — aggregated vs disaggregated, routers, KV connectors, and a drag-and-drop GPU pipeline. → NVIDIA AI Dynamo
- NVIDIA VSS — a five-service video pipeline where each service can be new, existing or remote. → NVIDIA VSS
- Sparkrun — multi-node inference across a DGX Spark cluster. → Sparkrun Clusters
- Jetson AI Lab — NVIDIA's catalogue, launched on a Jetson in one step. → Jetson AI Lab
Manage what you have (5 minutes)
Start and stop containers individually or by group, duplicate one to try a variant, generate public URLs, read the API usage examples, and delete what you no longer need.
Troubleshooting
The launch button is disabled Open Review & start. Every blocker is listed there, and each one is a button that jumps to the pane that fixes it. The most common are "no GPU selected" and "no model selected".
Container won't start
- Is the device still Online?
- Is there disk space for the image?
- Read the container's Logs tab — the error is almost always there.
Out of VRAM
- Check it in the VRAM Calculator with With Activation.
- Lower the quantization (BF16 → FP8 → INT4), shorten the sequence length, reduce the batch size, or add GPUs.
- The wizard's fit advisor warns about this before you start: "fits, but only just" means little room is left for the KV cache.
Model not found
- Custom model: check the repository id exactly as it appears on huggingface.co.
- User Models: check the model paths are registered on the destination device too.
- Check the volume mapping in the Container overrides tab.
Gated model / 401 The row in LLM Models says whether a token is required. Add one from your token vault, or paste it in the Environment overrides tab.
Open WebUI shows no models
- Is the engine container actually running?
- If it is pointed at a gateway, is the model registered into the gateway? Check Serving Layer → Serving.
AI Launch says "Could not read the model's files on the device" The registered model path on that device is unreadable. Fix the path; the plan fell back to Hugging Face and would otherwise download the weights again.
Quick reference
Application types
| Type | Use case | Complexity | Setup time |
|---|---|---|---|
| AI Launch | Any Hugging Face model, planned for you | Lowest | 2 min |
| Standard applications | Basic containers | Low | 5 min |
| LLM engines | Model inference with full control | Medium | 5 min |
| NVIDIA Dynamo | Multi-GPU distributed | High | 10 min |
| NVIDIA VSS | Video analysis pipeline | High | 15 min |
| Sparkrun | Multi-node DGX Spark | High | 20 min (plus cluster setup) |
Rough GPU memory guide
| Model size | Quantization | Minimum VRAM | Typical GPU |
|---|---|---|---|
| 7–8B | INT4 | 4–6 GB | RTX 3090, RTX 4090 |
| 7–8B | INT8 | 8–10 GB | RTX 4090, A10 |
| 13B | INT4 | 8–10 GB | RTX 4090, A10 |
| 13B | INT8 | 14–16 GB | A10, A100 40GB |
| 70B | INT4 | 40–50 GB | A100 80GB, 2× A100 40GB |
| 70B | INT8 | 80–90 GB | A100 80GB, 2× A100 80GB |
These are weights only. Add the KV cache for the context length and batch size you intend to serve — that is what the VRAM Calculator is for.
Where things are
| I want to… | Go to |
|---|---|
| Browse and launch applications | LLM Launcher → Applications |
| See what is running | LLM Launcher → Containers |
| Find or download a model | LLM Launcher → LLM Models |
| Use a model already on a device | LLM Models → User Models |
| Run a gateway or chat interface on its own | LLM Launcher → Serving Layer |
| Set up multi-node inference | LLM Launcher → Sparkrun Clusters |
| Run a Jetson AI Lab model | LLM Launcher → Jetson AI Lab |
| Check whether a model fits | LLM Launcher → VRAM Calculator |
| Push your own image | LLM Launcher → Private Registry |
| Register model paths | Devices → device → Models |