What is LLM Launcher?
LLM Launcher is the part of Cordatus you use to put a model into service. It covers everything between "I want to run this model" and "here is the endpoint my application can call": finding the model, downloading it onto a device, choosing the hardware, working out the engine and its flags, starting the containers, and putting a chat interface and an API gateway in front of them.
It was previously called Application Hub. The menu group, the pages inside it and the launch flow have all been reworked; this section documents the current interface.
Who is it for?
- AI engineers and data scientists who need a served model quickly, without hand-writing
docker runlines and engine flags. - ML platform / DevOps teams running several engines across a fleet of devices and needing one place to see, start, stop and delete them.
- Researchers comparing quantizations, context lengths and engines on the hardware they actually have.
- Teams running on-premise, where the model, the gateway and the chat interface all have to live on their own machines.
The pages in this section
| Page | What it is for |
|---|---|
| Containers | Everything running (or stopped) on your devices, grouped, with logs, parameters, ports and Sparkrun workloads. → Containers |
| Applications | The catalogue of ready-to-run AI applications — vLLM, TensorRT-LLM, Ollama, NVIDIA Dynamo, NVIDIA VSS and more. → Launch Guide |
| Environments | Development environments (PyTorch, TensorFlow, CUDA, TensorRT…) rather than inference engines. → Environments |
| Private Registry | Your own private container repositories and tags. → Private Registry |
| LLM Models | Model catalogue: Hugging Face, Ollama, NVIDIA NIM, and the models already stored on your own devices. → LLM Models |
| Sparkrun Clusters | DGX Spark clusters for multi-node inference. → Sparkrun Clusters |
| Serving Layer | An API gateway (LiteLLM) and a chat interface (Open WebUI), managed on their own. → Serving Layer |
| Jetson AI Lab | NVIDIA's Jetson AI Lab catalogue, launchable on your Jetson devices in one step. → Jetson AI Lab |
| VRAM Calculator | Work out whether a model fits before you deploy it. → VRAM Calculator |
What can I do with LLM Launcher?
-
Find and download models
- Browse Hugging Face, Ollama and NVIDIA NIM from one catalogue
- Filter by the GPUs or the Jetson module you actually own
- Download weights straight onto a device, into a registered model directory
- Register models the device already stores, and transfer them between devices
-
Deploy AI applications
- vLLM, TensorRT-LLM, Ollama, llama.cpp and other inference engines
- NVIDIA AI Dynamo (distributed LLM runtime)
- NVIDIA VSS (Video Search & Summarization)
- Docker-Compose applications and your own private images
-
Let Cordatus plan the launch for you
- AI Launch ranks your devices, chooses the engine, image, GPUs, quantization and flags
- Every choice comes with the evidence behind it, and blockers are stated before you commit
-
Or configure everything yourself
- GPU selection, host CPU and RAM limits
- Model source: model hub, your own models, or a custom name/URL
- Container options, environment variables, engine flags, notebook
- The exact command shown before it runs
-
Serve the model to clients
- An OpenAI-compatible LiteLLM gateway with keys, quotas and request logs
- An Open WebUI chat interface, with extra accounts
- Public URLs for reaching either from outside the network
-
Calculate VRAM requirements before deploying anything
-
Monitor and control containers
- Live logs, parameters and ports
- Start, stop, delete and duplicate, individually or by group
- Generate, regenerate and close public URLs
Two ways to launch
Cordatus gives you two routes to the same result. Both end with running containers you manage from the Containers page.
1. AI Launch — the system decides
You pick a model. Cordatus ranks your devices, reads the model's configuration, chooses the engine, the container image, the GPUs, the quantization and the flags, and shows you the plan with the evidence behind every choice before anything starts.
Use it when you want a working endpoint and do not want to make ten decisions to get one.
2. The launch wizard — you decide
Three steps (device → version → advanced configuration), and inside the third step a vertical rail with one pane per decision: Model, Compute, Access, Overrides, and finally Review & start, which shows the exact command that will run and every container that will be created.
Use it when you need a specific image, specific flags, a specific GPU split, or a multi-component deployment such as Dynamo or VSS.
The two are not separate worlds. From an AI Launch plan you can press Edit in Advanced and the wizard opens with everything the planner worked out already filled in.
How does LLM Launcher work?
1. Find a model (or skip straight to an application)
From LLM Models, browse Hugging Face, Ollama, NVIDIA NIM or User Models. Filter by your Jetson module, by a GPU model with a count, or by reading the real GPUs off a connected device (From Device). Each row says whether the model is open or needs an access token.
2. Download it onto a device (optional)
Download on a model row fetches the weights on the device, not through your browser:
- Device — where to put it.
- Directory — one of the model directories registered on that device. If none is registered yet, the dialog registers one for you; when the directory needs elevated rights it says exactly what will be created as root and what will not.
- Quantization / Profile — when a repository ships the same model exported several ways, only the one you pick is downloaded plus the shared files, or choose Every quantization.
- Token — a saved token or one you paste; optional for open models.
- sudo password, only when the target directory needs it. A password already stored on the device is used instead, and is never sent to your browser.
Cordatus reports honestly instead of failing generically: a download already in flight on that device, a directory owned by another user, or a file list that could not be read (in which case the whole repository is fetched).
3. Select an application and a version
From Applications, filter by device type (Server-Workstation, Jetson, DGX Spark), open one, and look at its Versions tab. Each version shows its size, date, minimum driver or JetPack version, and which device groups it supports — and, once a device is chosen, whether it is Downloaded or Will Download.
4. Choose the device and the version
The wizard's first two steps. Compatibility is checked strictly, so an image built for a newer JetPack is never offered to an older device.
5. Configure
The third step is a rail, not a wall of panels:
Model — model hub, User models or Custom, with a GPU fit advisor that says whether the selected GPUs can hold the model and how many it would take. Model transfer is offered automatically when the weights live on another device.
Compute — which GPUs (or Use all), and Host limits for CPU and RAM. Auto keeps the Host Reserved values back for the system; Custom lets you set the caps.
Access — the serving layer: an Open WebUI chat interface and/or a LiteLLM gateway, with credentials, virtual keys, extra accounts and public URLs. Four Profiles (Quick start / Serve an API / Gateway only / Engine only) set this pane in one click and touch nothing else.
Overrides — one card with four tabs:
- Container — port mappings (with free/in-use warnings) and volume bindings through a file browser over the device's real disks, plus network, restart policy and devices
- Environment — variables, including values from your token vault
- Engine flags — engine-specific arguments
- Notebook — Jupyter, protected by a token or a password
Suggest with AI fills the first three at once and reports what it adjusted to fit your hardware, or says outright that the model will not fit.
Review & start — the summary with a pencil beside every row, the resulting command, and the full list of containers that will be created, including the ones the serving layer brings along.
6. Launch and monitor
Start the deployment, watch the image pull if it is not already on the device, and pick the containers up on the Containers page — or on the application's own Containers tab. If a serving layer was configured, copy its credentials from the dialog that appears; they are shown once.
7. Manage what is running
Logs, parameters and ports in real time; start, stop, delete and duplicate; public URLs; API usage examples for a running engine.
Application types
Standard applications
Simple containers with GPU, CPU and RAM configuration:
- single-container deployment
- direct GPU assignment
- the standard override tabs
LLM engine applications
Inference engines with model management:
- model selection from the hub, User Models or a custom name/URL
- automatic volume configuration per engine
- model transfer between devices, with resume and checksum repair
- the GPU fit advisor
- an optional serving layer (chat interface and/or gateway)
- engine-specific arguments, optionally filled by the resolver
NVIDIA AI Dynamo
Distributed LLM runtime for multi-GPU deployments:
- processing modes — Aggregated / Disaggregated
- routers —
kv(KV-score aware),round-robin,random - KV connectors —
lmcache,kvbm,nixl,none, with optional RAM/disk offload sizes - a drag-and-drop GPU pipeline for assigning GPUs to prefill and decode workers
- multi-container orchestration behind one frontend
NVIDIA VSS (Video Search & Summarization)
A five-service pipeline: VSS (with an optional Event Reviewer), VLM, LLM, Embed and Rerank. Each service can be a new container, a container already running on the device, or a remote OpenAI-compatible endpoint.
Sparkrun (multi-node)
Recipes launched across a DGX Spark cluster rather than as Docker containers on one device.
Development environments
PyTorch, TensorFlow, CUDA, TensorRT and the rest — same flow, no model step, with Jupyter inside the same container.
Your own images
Anything you push to the Private Registry, optionally marked as an engine so Cordatus can launch it as one.
Models on your devices
Model path configuration
Model paths are per device, under Devices → the device → Models → Storage Paths:
- HuggingFace cache path
- Ollama models path
- NVIDIA NIM cache path
- Custom paths
Downloaded Models then lists what was found under each, with a Deploy button per model. Removing a path removes it from settings only — the files are not deleted.
Adding models to your library
- Explore Models on Your Devices — Cordatus asks the device for its HuggingFace, Ollama and NIM directories and reads each one. Models already in your library are marked Added, the rest New; pick target engines and import them. Rescan cleans up anything the device no longer stores.
- Add a model manually — for a model in a custom path: device, path, name, type, tag and engines.
Model transfer
- Transfers run device-to-device over your own network
- Interrupted transfers resume, and corrupted files are detected by checksum and re-downloaded
- Progress is shown per file, with totals
- Use API Only skips the copy when the model is reachable over an API
- Volume mapping is configured automatically once it completes
Deploying a user model
AI Launch plans directly for the device that holds the model; Manual → Run manually with… opens the wizard with the model pre-selected and the correct paths already mapped.
VRAM Calculator
Calculate before you deploy
- Select a model — search Hugging Face; parameters are fetched automatically. GGUF repositories get a per-file quantization selector with approximate bits-per-weight.
- Choose the GPU — from the catalogue, or read the real GPUs off one of your devices with Select Device GPUs.
- Configure — GPU count, quantization, sequence length, batch size, GPU memory utilization, and the calculation type (20% Overhead or With Activation).
- Read the result — a doughnut chart of Model Weights / KV Cache / Overhead or Activations / Free VRAM, the detailed metrics, a usage bar, and a VRAM is Sufficient / Insufficient verdict.
Use cases
- Test hardware requirements before purchasing GPUs
- Optimize quantization, sequence length and batch size for the hardware you have
- Plan multi-GPU deployments
- Compare different model configurations side by side
Container management
Operations
- Start / Stop — individually, per child, or for a whole group
- Delete — type
DELETEto confirm; volumes are listed first and can be skipped - Duplicate — a new launch pre-filled with this container's configuration
- Open — the group panel, with a search box for large groups
Disabled actions always say why: Container(s) already stopped, Device connection unavailable, You don't have permission to duplicate.
Batch operations
Select rows with the checkboxes and use Delete Selected. Group actions apply to every container in the group.
Container information
- Logs — live output, with Copy
- Parameters — every value the container was created with
- Ports — local and global addresses, with copy and open
Public URLs
Generate a public URL for any exposed port, then regenerate, close or reassign it. Newly exposed ports can be added from the same tab.
API usage examples
For a running engine: a quick call, an OpenAI-compatible client snippet, and Open with WebUI.
Sparkrun workloads
Multi-node runs live on their own tab, with per-node logs, access URLs, engine arguments, environment variables (masked) and the recipe YAML.
Key concepts
Application — a container image registered in Cordatus, with its versions, supported device classes and published options.
Environment — a development image rather than an inference engine. Same launch flow, no model step.
Version / chipset tag — a specific image tag for a specific device class, shown as Downloaded or Will Download.
Container group — several containers deployed together as one unit.
Serving layer — the optional chat interface (Open WebUI) and API gateway (LiteLLM) that come up beside a model.
Sparkrun workload — a multi-node inference run on a DGX Spark cluster; not a Docker container, so it is listed on its own tab.
User model (User Models) — a model already stored on one of your devices and registered with Cordatus.
Overrides — the container options, environment variables, engine flags and notebook settings layered on top of what the application fixes. Fixed rows are locked and cannot be removed.
Quantization — BF16/FP16 (full precision), INT8/FP8 (half the memory), INT4/FP4 (a quarter), and the GGUF variants with their own bits-per-weight.
VRAM components — model weights, KV cache, activation memory or overhead, and free VRAM.
Best practices
Resource management
- Keep the Host Reserved CPU and RAM values; they are what keeps the device responsive
- Leave 20–30% free VRAM for unexpected load
- Use the VRAM Calculator before production deployments
- Read the GPU fit advisor line before choosing a model — "fits, but only just" means a long context or many concurrent requests may fail
Model management
- Register the HuggingFace, Ollama and NIM paths on every device you deploy to; transfer, scanning and "already on device" detection all depend on it
- Keep parameter count and quantization filled in — the fit advisor and VRAM estimates use them
- Rescan after deleting models on a device
- Use meaningful names for custom models
Deployment
- Test with a small model first
- Use With Activation in the VRAM Calculator for an accurate estimate
- Read Review & start before committing — it shows the exact command and every container
- Watch the logs on the first deploy of a new image
- Copy serving layer credentials immediately; they are shown once
Multi-GPU and multi-node
- Use NVIDIA Dynamo for distributed inference on one device, and Sparkrun for multi-node
- Assign every worker a GPU before starting — the pipeline pane will not report ready otherwise
- Choose the router to match the workload:
kvfor prefix reuse,round-robinfor even load - Watch GPU utilisation across all workers after launch
What is new in this release
If you have used the previous Application Hub, these are the changes that matter:
- AI Launch. A whole launch path where the plan is produced for you, with per-decision evidence, blockers and warnings.
- The advanced step is a rail, not a wall of panels. Model, Compute, Access, Overrides and Review, each with its own pane and completion mark. The readiness checklist now sits beside the Start button instead of a screen away.
- Review & start. The resulting command and the full list of containers shown before you commit.
- GPU fit advisor under the model catalogue.
- Serving layer as a first-class thing — with the model, reused if already running, or managed on its own from the Serving Layer page.
- Profiles — Quick start / Serve an API / Gateway only / Engine only.
- Model download onto a device, with quantization selection and honest failure reporting.
- Sparkrun Clusters and Sparkrun Workloads for multi-node DGX Spark inference.
- Jetson AI Lab, with an express run panel that resolves image, model and engine arguments from the published recipe.
- LLM Models rebuilt with four source tabs, GPU- and Jetson-aware filtering, and per-model AI Launch.
- Environments, Containers, Applications and VRAM Calculator redesigned around the same card, table and subheader components.
Where to go next
- New to the platform → Quickstart
- Just want a model running → AI Launch
- Need full control → Launch Guide
- Already running things → Containers
- Serving it to clients → Serving Layer