Skip to main content

Quickstart

Get a model serving on your own hardware in about 10 minutes. This page is the short version; every section links to the full guide.

Pick a model, let AI Launch plan it, start it, then chat with it in the interface it created

 


Prerequisites​

  • A Cordatus account with the container permissions
  • At least one device connected to Cordatus and showing as Online
  • A GPU on that device (recommended for LLM applications)
  • The Cordatus Client installed and running on it
tip

No device yet? Do the Device Hub quickstart first — nothing in LLM Launcher works without one.


The fastest route: AI Launch (about 4 minutes)​

If you only want a model running, skip the wizard entirely.

  1. Go to LLM Launcher → LLM Models.
  2. Search for a small model — Qwen/Qwen3-8B is a good first one.
  3. Press AI Launch on its row.
  4. Watch the six planning phases. When it says Plan ready, read the summary line: Ready to launch on {device}.
  5. Open Details if you want to see the engine, image, GPUs, quantization and context it chose, and Why these settings for the evidence behind each.
  6. Turn on Chat interface under the plan, and copy the credentials it shows.
  7. Press Launch.

That is the whole thing. Skip to Verify your deployment.

→ AI Launch


The full route: the launch wizard (about 10 minutes)​

1. Browse the applications (1 minute)​

LLM Launcher → Applications. Filter by device type — Server - Workstation, Jetson or DGX Spark — and pick an engine:

  • vLLM — high-throughput LLM inference
  • TensorRT-LLM — NVIDIA's optimised engine
  • Ollama — simple local models
  • llama.cpp — GGUF weights, including on Jetson
  • NVIDIA Dynamo — distributed multi-GPU
  • NVIDIA VSS — video search and summarisation

Open one with View Details and look at its Versions tab.

→ Launch Guide

2. Start the wizard (1 minute)​

Press Start Single Device Application.

Step 1 — Select Device. Pick a connected device. The footer confirms Deploying to {device}.

Step 2 — Select Version. Only compatible images are listed. Each is marked:

  • Downloaded — already on the device
  • Will Download — pulled on start
  • Incompatible — with the reason, for example a driver version mismatch

Press Next.

3. Configure (4 minutes)​

Step 3 is a rail. Work down it.

Profile (top of the screen) — press Quick start. That switches on the chat interface and touches nothing else.

Model

  1. Stay on the Model hub card.
  2. Pick a small model — 7B or 8B for a first deployment.
  3. Read the fit line under the catalogue. You want "{model} fits your selection comfortably."

Compute

  1. Select a GPU, or switch on Use all.
  2. Leave Host limits on Auto.

Access — already set by the Quick start profile. Copy the credentials it shows; they are displayed once.

Overrides — skip it. The application's defaults are already correct.

Review & start — read the Resulting command and the Containers to be created. If a blocker is listed, click it: it takes you to the pane that fixes it.

→ Launch Guide

4. Launch (1 minute, plus the image pull)​

Press Start Environment.

Behind the scenes Cordatus connects to the device, pulls the image if it is not there, creates the containers, configures networking and volumes, and starts the serving layer alongside.

A first pull can take several minutes. The container shows as Downloading while it happens.


Verify your deployment​

  1. Go to LLM Launcher → Containers.
  2. Find the new container group and check its state.
  3. Open the group, then Container Informations on the engine.
  4. Logs — you should see the model loading.
  5. Ports — copy the local address.

Test it:

curl http://<device-ip>:<port>/v1/models

Then open the Open WebUI container's address in a browser and sign in with the admin credentials you copied.

→ Containers


Expected result​

  • The engine container is running, with model-loading messages in its logs
  • The Open WebUI container is running, and you can chat with the model
  • Both appear as one group under Containers
  • If you enabled the gateway, a LiteLLM container (plus its PostgreSQL, Redis and Prometheus) is running too, and the model is registered into it

What's next​

Check a model fits before deploying it (5 minutes)​

LLM Launcher → VRAM Calculator:

  1. Search for a bigger model, for example Llama-2-70B.
  2. Choose a GPU — or Select Device GPUs to use your real hardware.
  3. Set the sequence length and batch size you actually intend to serve.
  4. Switch Calculation Type to With Activation.
  5. Compare quantizations until it says VRAM is Sufficient.

→ VRAM Calculator

Use models already on your device (10 minutes)​

  1. Devices → your device → Models → Storage Paths — register the HuggingFace, Ollama and NIM paths.
  2. LLM Models → User Models → Explore Models on Your Devices — pick the device, Start scanning, choose target engines, Add {n} model(s).
  3. Press AI Launch on one of them — the plan is made for the device that stores it, and nothing is downloaded.

→ User Models

Download a model onto a device (5 minutes)​

LLM Models → Download on any row: pick the device, the registered directory and the quantization, add a token if the repository is gated, and let the device fetch it.

→ LLM Models

Serve it properly (10 minutes)​

LLM Launcher → Serving Layer: deploy a LiteLLM gateway, register your running model into it, and issue scoped virtual keys for your clients instead of handing out the master key.

→ Serving Layer

Go further (15–30 minutes)​

  • NVIDIA AI Dynamo — aggregated vs disaggregated, routers, KV connectors, and a drag-and-drop GPU pipeline. → NVIDIA AI Dynamo
  • NVIDIA VSS — a five-service video pipeline where each service can be new, existing or remote. → NVIDIA VSS
  • Sparkrun — multi-node inference across a DGX Spark cluster. → Sparkrun Clusters
  • Jetson AI Lab — NVIDIA's catalogue, launched on a Jetson in one step. → Jetson AI Lab

Manage what you have (5 minutes)​

Start and stop containers individually or by group, duplicate one to try a variant, generate public URLs, read the API usage examples, and delete what you no longer need.

→ Containers


Troubleshooting​

The launch button is disabled Open Review & start. Every blocker is listed there, and each one is a button that jumps to the pane that fixes it. The most common are "no GPU selected" and "no model selected".

Container won't start

  • Is the device still Online?
  • Is there disk space for the image?
  • Read the container's Logs tab — the error is almost always there.

Out of VRAM

  • Check it in the VRAM Calculator with With Activation.
  • Lower the quantization (BF16 → FP8 → INT4), shorten the sequence length, reduce the batch size, or add GPUs.
  • The wizard's fit advisor warns about this before you start: "fits, but only just" means little room is left for the KV cache.

Model not found

  • Custom model: check the repository id exactly as it appears on huggingface.co.
  • User Models: check the model paths are registered on the destination device too.
  • Check the volume mapping in the Container overrides tab.

Gated model / 401 The row in LLM Models says whether a token is required. Add one from your token vault, or paste it in the Environment overrides tab.

Open WebUI shows no models

  • Is the engine container actually running?
  • If it is pointed at a gateway, is the model registered into the gateway? Check Serving Layer → Serving.

AI Launch says "Could not read the model's files on the device" The registered model path on that device is unreadable. Fix the path; the plan fell back to Hugging Face and would otherwise download the weights again.


Quick reference​

Application types​

TypeUse caseComplexitySetup time
AI LaunchAny Hugging Face model, planned for youLowest2 min
Standard applicationsBasic containersLow5 min
LLM enginesModel inference with full controlMedium5 min
NVIDIA DynamoMulti-GPU distributedHigh10 min
NVIDIA VSSVideo analysis pipelineHigh15 min
SparkrunMulti-node DGX SparkHigh20 min (plus cluster setup)

Rough GPU memory guide​

Model sizeQuantizationMinimum VRAMTypical GPU
7–8BINT44–6 GBRTX 3090, RTX 4090
7–8BINT88–10 GBRTX 4090, A10
13BINT48–10 GBRTX 4090, A10
13BINT814–16 GBA10, A100 40GB
70BINT440–50 GBA100 80GB, 2× A100 40GB
70BINT880–90 GBA100 80GB, 2× A100 80GB

These are weights only. Add the KV cache for the context length and batch size you intend to serve — that is what the VRAM Calculator is for.

Where things are​

I want to…Go to
Browse and launch applicationsLLM Launcher → Applications
See what is runningLLM Launcher → Containers
Find or download a modelLLM Launcher → LLM Models
Use a model already on a deviceLLM Models → User Models
Run a gateway or chat interface on its ownLLM Launcher → Serving Layer
Set up multi-node inferenceLLM Launcher → Sparkrun Clusters
Run a Jetson AI Lab modelLLM Launcher → Jetson AI Lab
Check whether a model fitsLLM Launcher → VRAM Calculator
Push your own imageLLM Launcher → Private Registry
Register model pathsDevices → device → Models

Get help​