Skip to main content

Launch Guide

This guide covers the full launch wizard: the path you take when you want to choose the image, the GPUs, the flags and the serving layer yourself. If you would rather have Cordatus decide, use AI Launch instead — and note that you can switch from an AI Launch plan into this wizard at any point with Edit in Advanced.

From the Applications catalogue to a running engine, gateway and chat interface

 


1. Find the application​

LLM Launcher → Applications lists every ready-to-run application. The page gives you:

  • Search in the subheader.
  • Filter by device type — Server - Workstation, Jetson, DGX Spark.
  • For Jetson, an extra Select Jetson group… filter, so you only see what your own Jetson groups can run.
  • Active filter chips under the header, each removable, plus Clear All.
  • Cards showing the application name, its publisher logo, a short description and up to three tags (the rest behind +{n} more). Clicking a tag filters by it.
  • Load More Applications, or infinite scroll as you reach the bottom.

Click a card, or its View Details footer, to open the application.

Search, then filter by tag — the chips show what is active

 


2. The application detail page​

The header shows the name, the type pill, the description, its tags, and two compatibility badges: Server-Workstation (x86_64) and Jetson (arm64), green when supported and red when not.

Two actions sit on the right:

  • Start Single Device Application (or Start Single Device Environment) — the wizard described below.
  • Start Sparkrun — only on engines that support it, for multi-node DGX Spark runs. See Sparkrun Clusters.

Four tabs:

README / Description​

The application's README, with a sidebar containing:

  • At a glance — category, latest release, number of versions.
  • Device compatibility — the same two badges, with the architecture beside each.
  • Eligible devices — every device of yours that can run it, with an online/offline dot, and a Start button underneath.

Versions​

A sortable, paginated table of every image version:

ColumnMeaning
TagThe image tag.
SizeUncompressed size, when the version has a single variant.
CreatedPublication date.
Required Version(s)The minimum NVIDIA driver version, or a Jetpacks ({n} Version(s)) badge you can hover to list them.
DevicesServer - Workstation, or Jetson ({n} Group(s)) — hover to see which groups.

Versions with several hardware variants are collapsed into one row with an expand arrow; opening it shows each variant with its own size, date, driver requirement and device list.

Customize Columns (the icon beside the search box) lets you show, hide and drag-reorder these columns. Your choice is remembered.

Containers​

Everything you have already launched from this application, in the same two-tab layout as the main Containers page: Containers and Sparkrun Workloads.

Models​

Only for engine applications: the model catalogue filtered to models this engine can serve, with a Deploy button that drops you straight into the wizard with the model pre-selected.


3. Step 1 — Select Device​

The wizard opens as a three-step accordion. Each step's header keeps showing what you chose, so the decisions above the open one stay readable.

Pick the device to deploy to. Only connected devices can be selected; the footer confirms Deploying to device.

Next stays disabled until a connected device is chosen.


4. Step 2 — Select Version​

The version list is filtered to what your selected device can actually run. Compatibility is checked strictly: an image built for a newer JetPack is never offered to an older device.

Each version shows:

  • Downloaded — already on the device, so starting is immediate.
  • Will Download — pulled to the device on start.
  • Incompatible — with the reason, for example "Driver version mismatch — your device: X, required: Y".

You also get a search box, a Filter by image properties control and a reload button.

If nothing matches, the list says "No compatible versions found for your device" rather than showing you images that would fail.

note

When launching from the Private Registry there is no version step — the tag you started from is the version. The wizard goes straight from device to advanced configuration.


5. Step 3 — Advanced Configuration​

This is where the redesign is most visible. Instead of a long scroll of panels, the step is a vertical rail: one entry per decision, each with its own subtitle summarising the current answer and a mark showing whether it is settled.

Setup
● Model Qwen/Qwen3-8B ✓
● Compute GPUs · memory · CPU ✓
● Access chat UI · gateway ✓
● Overrides 3 options · notebook on ✎
● Review & start command · containers ›

Ready to start
✓ GPU selected
✓ Model selected
✓ Configuration valid

The Ready to start card is pinned to the foot of the rail, next to the button it gates. If something is missing it names it.

Profiles​

Above the rail, four presets:

ProfileEffect
Quick startChat interface only.
Serve an APIGateway + chat interface.
Gateway onlyGateway and keys, no chat interface.
Engine onlyNo extra service.

A profile only fills in the Access section. It never touches your model, GPU or override choices.


5.1 Model​

The first pane, for engine applications. Two source cards:

  • Model hub — the catalogue (Hugging Face / Ollama / NVIDIA NIM depending on the engine).
  • User models — models already on this device. See User Models.
  • Custom — a name or URL you paste yourself.

The GPU fit advisor​

One line above the catalogue that reads the selected model against the device's GPUs:

  • "Pick a model first — this panel then says how many GPUs it needs."
  • "{model} fits your selection comfortably."
  • "{model} fits, but only just." — little room is left for the KV cache, so a long context or many concurrent requests may fail.
  • "Your selection is not enough — {model} needs at least {n} GPUs."
  • "{model} does not fit this device." — the weights alone would fill every GPU together.

It also warns when the device's GPUs report different amounts of VRAM, because tensor parallelism splits evenly and the smallest card bounds what fits.

note

The estimate is made from weight size only. The context length you choose still has to fit alongside it — use the VRAM Calculator when the margin matters.

The catalogue​

Each row shows the tag, Quant, Size, and a badge:

  • On device — the weights are already there.
  • Will download — they will be fetched on start.
  • Fit on selection — it fits the GPUs you have selected.

You can search, sort, and switch between the fitting models and All models. When you pick a model with several quantizations, the row expands so you can pick the quantization to launch.

Model transfer​

If you choose a model from User models that lives on a different device, Cordatus offers to transfer it:

  • Model Transfer Required — normal case; shows model size and source path.
  • Resume Model Transfer — an earlier attempt was incomplete; shows how much is already there and how many files are missing or partial.
  • Corrupted Files Detected — checksum verification failed and the affected files will be re-downloaded.

Progress is per-file, with total bytes and file counts. Use API Only skips the transfer when the model is reachable over an API instead.


5.2 Compute​

GPUs​

Every GPU on the device as a card, showing VRAM FREE, UTIL and TEMP. Select individual GPUs, or switch on Use all. At least one GPU must be selected before the launch is allowed.

Host limits​

CPU and memory caps for the container:

  • Auto — the default. Leaves the host free to schedule other work: the container is capped at the usable amount, with the Host Reserved CPU and RAM kept back for the system.
  • Custom — set the CPU limit and Memory limit by hand, between the shown minimum and the usable maximum.
  • Turning a limit off entirely is possible and is spelled out on the card: "No CPU cap — the container may use all {n} cores."
tip

Keep the Host Reserved values unless you have a reason not to. They are what stops a heavy engine from starving the device's own agent.


5.3 Access​

Only present when the application supports a serving layer. Two cards, each a switch:

  • Chat interface — Open WebUI, with an admin account created for you.
  • API gateway — LiteLLM: one OpenAI-compatible endpoint with keys, quotas and request logs.

With neither on, only the engine's own port is published.

Which gateway?​

If the device is already running a LiteLLM gateway, Cordatus finds it and offers to reuse it — the model is registered into it and no new credentials are needed. You can still choose Deploy a new gateway instead, which brings up its own PostgreSQL, Redis and Prometheus containers and runs database migrations on first boot (a few minutes).

Special cases the picker reports honestly:

  • more than one gateway running → you choose which one the model is registered into;
  • a gateway whose master key is not in its environment → Cordatus cannot register into it, so a new one is deployed;
  • the device could not be queried → a new gateway is deployed if none is running.

Credentials​

One card with every secret this deploy hands out — Open WebUI admin address and password, the LiteLLM master key, the model name clients send — each with Reveal, Copy and a Regenerate all action.

warning

These are shown once, right after the environment starts. Cordatus does not store them. The addresses appear afterwards in each container's Ports tab.

Accounts and keys​

  • Open WebUI accounts — extra accounts created once Open WebUI has booted, so people sign in with their own e-mail instead of sharing the admin account.
  • LiteLLM virtual keys — created on the gateway once it is up, each with a name, an optional budget and an optional this model only scope. Hand these out instead of the master key.

Advanced access options​

Collapsed by default:

  • Model name clients send — the public alias. Keeping it separate from the real weights lets you swap them later without breaking clients.
  • Image versions for LiteLLM and Open WebUI — pinned releases first, then pre-releases, then moving tags such as rc, dev and main-stable, whose contents change under you.
  • Connect Open WebUI to — the LiteLLM gateway (so it sees every model the gateway serves) or the Engine directly.
  • Create a public URL for Open WebUI and/or for the gateway.

5.4 Overrides​

One card with four tabs, each showing a live count:

TabWhat it holds
Containerdocker run options — ports, volumes, network, restart policy, devices.
EnvironmentEnvironment variables, including tokens from your vault.
Engine flagsEngine-specific arguments for this image.
NotebookJupyter settings, when the image supports it.

A legend says which rows you may change: ✎ you can edit vs 🔒 fixed by the application. Anything the application fixes cannot be removed.

Ports and volumes​

Ports show what is free on the host and warn when one is already in use. Volumes are picked with a host-folder browser that reads the device's real filesystem.

caution

Host folders are exposed to the container as-is. Make sure you know what you are mounting.

Tokens​

Environment variables can take a value from your token vault — Choose a token, or Add a new token inline, which stores it and makes it reusable in other environments. Gated Llama repositories are the usual reason to need one.

Suggest with AI​

Suggest with AI fills the Container, Environment and Engine-flag tabs at once from a resolver that reads the model against the selected hardware. Its report appears above the tabs, not inside one of them, because it rewrites all three:

  • "Model won't fit on the selected hardware."
  • "Adjustments applied to fit your hardware:" followed by each adjustment.
  • "Inferred from {base model}" when the configuration was derived from a related model.
  • Insights — hardware suggestions and qualitative observations, in one list.

Pressing it again turns it off and restores the static defaults.

Notebook​

When available, Jupyter runs inside the same container, so notebooks share the model, the mounted folders and the GPUs. Protection is either an Access token (generated for you, asked once per browser) or a Password.

caution

An unprotected notebook gives anyone who can reach the host on that port a shell-capable notebook. Use it only on an isolated network.


5.5 Review & start​

The last pane before anything runs:

  • Summary — every decision as a label/value row, each with a pencil that jumps back to the pane that owns it. Anything unset reads not set.
  • Resulting command — the actual command, assembled from the same options, variables and arguments the device is sent. Copy puts it on your clipboard.
  • Containers to be created — the full list, including the ones the serving layer brings along.
  • Blockers — if anything is missing, it is listed first, and each entry is a button that takes you to the pane that fixes it.

6. Starting​

The footer carries the verdict next to the button it governs: the device you are deploying to, a readiness pill, and:

ButtonWhen
Back / NextMoving between the three steps.
Review & startOn the rail path, takes you to the Review pane.
Start EnvironmentStarts the deployment. Disabled while a blocker stands.
CancelCloses the wizard.

After starting:

  1. Cordatus authorises the operation on the device.
  2. If the image is not present, it is pulled — visible as the container's Downloading state.
  3. The containers appear on the Containers page, grouped.
  4. If a serving layer was configured, the Save your serving layer credentials dialog appears. Copy them; they are shown once.

Multi-component applications​

Two applications use the same rail with a different set of panes:

  • NVIDIA AI Dynamo — Model, Runtime, GPU pipeline, Access, Overrides, Worker arguments, Review. → NVIDIA AI Dynamo Guide
  • NVIDIA VSS — one numbered rail entry per service in the pipeline, then Review. → NVIDIA VSS Guide

Docker-Compose based applications get a Compose services / Compose profiles / Configuration variables layout instead; services in the selected profiles start with the environment, and the primary service always runs.