Skip to main content

What is LLM Launcher?

LLM Launcher is the part of Cordatus you use to put a model into service. It covers everything between "I want to run this model" and "here is the endpoint my application can call": finding the model, downloading it onto a device, choosing the hardware, working out the engine and its flags, starting the containers, and putting a chat interface and an API gateway in front of them.

It was previously called Application Hub. The menu group, the pages inside it and the launch flow have all been reworked; this section documents the current interface.

Every page in the LLM Launcher menu group, in order

 


Who is it for?​

  • AI engineers and data scientists who need a served model quickly, without hand-writing docker run lines and engine flags.
  • ML platform / DevOps teams running several engines across a fleet of devices and needing one place to see, start, stop and delete them.
  • Researchers comparing quantizations, context lengths and engines on the hardware they actually have.
  • Teams running on-premise, where the model, the gateway and the chat interface all have to live on their own machines.

The pages in this section​

PageWhat it is for
ContainersEverything running (or stopped) on your devices, grouped, with logs, parameters, ports and Sparkrun workloads. → Containers
ApplicationsThe catalogue of ready-to-run AI applications — vLLM, TensorRT-LLM, Ollama, NVIDIA Dynamo, NVIDIA VSS and more. → Launch Guide
EnvironmentsDevelopment environments (PyTorch, TensorFlow, CUDA, TensorRT…) rather than inference engines. → Environments
Private RegistryYour own private container repositories and tags. → Private Registry
LLM ModelsModel catalogue: Hugging Face, Ollama, NVIDIA NIM, and the models already stored on your own devices. → LLM Models
Sparkrun ClustersDGX Spark clusters for multi-node inference. → Sparkrun Clusters
Serving LayerAn API gateway (LiteLLM) and a chat interface (Open WebUI), managed on their own. → Serving Layer
Jetson AI LabNVIDIA's Jetson AI Lab catalogue, launchable on your Jetson devices in one step. → Jetson AI Lab
VRAM CalculatorWork out whether a model fits before you deploy it. → VRAM Calculator

What can I do with LLM Launcher?​

  • Find and download models

    • Browse Hugging Face, Ollama and NVIDIA NIM from one catalogue
    • Filter by the GPUs or the Jetson module you actually own
    • Download weights straight onto a device, into a registered model directory
    • Register models the device already stores, and transfer them between devices
  • Deploy AI applications

    • vLLM, TensorRT-LLM, Ollama, llama.cpp and other inference engines
    • NVIDIA AI Dynamo (distributed LLM runtime)
    • NVIDIA VSS (Video Search & Summarization)
    • Docker-Compose applications and your own private images
  • Let Cordatus plan the launch for you

    • AI Launch ranks your devices, chooses the engine, image, GPUs, quantization and flags
    • Every choice comes with the evidence behind it, and blockers are stated before you commit
  • Or configure everything yourself

    • GPU selection, host CPU and RAM limits
    • Model source: model hub, your own models, or a custom name/URL
    • Container options, environment variables, engine flags, notebook
    • The exact command shown before it runs
  • Serve the model to clients

    • An OpenAI-compatible LiteLLM gateway with keys, quotas and request logs
    • An Open WebUI chat interface, with extra accounts
    • Public URLs for reaching either from outside the network
  • Calculate VRAM requirements before deploying anything

  • Monitor and control containers

    • Live logs, parameters and ports
    • Start, stop, delete and duplicate, individually or by group
    • Generate, regenerate and close public URLs

Two ways to launch​

Cordatus gives you two routes to the same result. Both end with running containers you manage from the Containers page.

1. AI Launch — the system decides​

You pick a model. Cordatus ranks your devices, reads the model's configuration, chooses the engine, the container image, the GPUs, the quantization and the flags, and shows you the plan with the evidence behind every choice before anything starts.

Use it when you want a working endpoint and do not want to make ten decisions to get one.

→ AI Launch

2. The launch wizard — you decide​

Three steps (device → version → advanced configuration), and inside the third step a vertical rail with one pane per decision: Model, Compute, Access, Overrides, and finally Review & start, which shows the exact command that will run and every container that will be created.

Use it when you need a specific image, specific flags, a specific GPU split, or a multi-component deployment such as Dynamo or VSS.

→ Launch Guide

tip

The two are not separate worlds. From an AI Launch plan you can press Edit in Advanced and the wizard opens with everything the planner worked out already filled in.


How does LLM Launcher work?​

1. Find a model (or skip straight to an application)​

From LLM Models, browse Hugging Face, Ollama, NVIDIA NIM or User Models. Filter by your Jetson module, by a GPU model with a count, or by reading the real GPUs off a connected device (From Device). Each row says whether the model is open or needs an access token.

→ LLM Models

2. Download it onto a device (optional)​

Download on a model row fetches the weights on the device, not through your browser:

  • Device — where to put it.
  • Directory — one of the model directories registered on that device. If none is registered yet, the dialog registers one for you; when the directory needs elevated rights it says exactly what will be created as root and what will not.
  • Quantization / Profile — when a repository ships the same model exported several ways, only the one you pick is downloaded plus the shared files, or choose Every quantization.
  • Token — a saved token or one you paste; optional for open models.
  • sudo password, only when the target directory needs it. A password already stored on the device is used instead, and is never sent to your browser.

Cordatus reports honestly instead of failing generically: a download already in flight on that device, a directory owned by another user, or a file list that could not be read (in which case the whole repository is fetched).

→ LLM Models — downloading

3. Select an application and a version​

From Applications, filter by device type (Server-Workstation, Jetson, DGX Spark), open one, and look at its Versions tab. Each version shows its size, date, minimum driver or JetPack version, and which device groups it supports — and, once a device is chosen, whether it is Downloaded or Will Download.

4. Choose the device and the version​

The wizard's first two steps. Compatibility is checked strictly, so an image built for a newer JetPack is never offered to an older device.

5. Configure​

The third step is a rail, not a wall of panels:

Model — model hub, User models or Custom, with a GPU fit advisor that says whether the selected GPUs can hold the model and how many it would take. Model transfer is offered automatically when the weights live on another device.

Compute — which GPUs (or Use all), and Host limits for CPU and RAM. Auto keeps the Host Reserved values back for the system; Custom lets you set the caps.

Access — the serving layer: an Open WebUI chat interface and/or a LiteLLM gateway, with credentials, virtual keys, extra accounts and public URLs. Four Profiles (Quick start / Serve an API / Gateway only / Engine only) set this pane in one click and touch nothing else.

Overrides — one card with four tabs:

  • Container — port mappings (with free/in-use warnings) and volume bindings through a file browser over the device's real disks, plus network, restart policy and devices
  • Environment — variables, including values from your token vault
  • Engine flags — engine-specific arguments
  • Notebook — Jupyter, protected by a token or a password

Suggest with AI fills the first three at once and reports what it adjusted to fit your hardware, or says outright that the model will not fit.

Review & start — the summary with a pencil beside every row, the resulting command, and the full list of containers that will be created, including the ones the serving layer brings along.

→ Launch Guide

6. Launch and monitor​

Start the deployment, watch the image pull if it is not already on the device, and pick the containers up on the Containers page — or on the application's own Containers tab. If a serving layer was configured, copy its credentials from the dialog that appears; they are shown once.

→ Containers

7. Manage what is running​

Logs, parameters and ports in real time; start, stop, delete and duplicate; public URLs; API usage examples for a running engine.


Application types​

Standard applications​

Simple containers with GPU, CPU and RAM configuration:

  • single-container deployment
  • direct GPU assignment
  • the standard override tabs

→ Launch Guide

LLM engine applications​

Inference engines with model management:

  • model selection from the hub, User Models or a custom name/URL
  • automatic volume configuration per engine
  • model transfer between devices, with resume and checksum repair
  • the GPU fit advisor
  • an optional serving layer (chat interface and/or gateway)
  • engine-specific arguments, optionally filled by the resolver

→ Launch Guide · User Models

NVIDIA AI Dynamo​

Distributed LLM runtime for multi-GPU deployments:

  • processing modes — Aggregated / Disaggregated
  • routers — kv (KV-score aware), round-robin, random
  • KV connectors — lmcache, kvbm, nixl, none, with optional RAM/disk offload sizes
  • a drag-and-drop GPU pipeline for assigning GPUs to prefill and decode workers
  • multi-container orchestration behind one frontend

→ NVIDIA AI Dynamo

NVIDIA VSS (Video Search & Summarization)​

A five-service pipeline: VSS (with an optional Event Reviewer), VLM, LLM, Embed and Rerank. Each service can be a new container, a container already running on the device, or a remote OpenAI-compatible endpoint.

→ NVIDIA VSS

Sparkrun (multi-node)​

Recipes launched across a DGX Spark cluster rather than as Docker containers on one device.

→ Sparkrun Clusters

Development environments​

PyTorch, TensorFlow, CUDA, TensorRT and the rest — same flow, no model step, with Jupyter inside the same container.

→ Environments

Your own images​

Anything you push to the Private Registry, optionally marked as an engine so Cordatus can launch it as one.

→ Private Registry


Models on your devices​

Model path configuration​

Model paths are per device, under Devices → the device → Models → Storage Paths:

  • HuggingFace cache path
  • Ollama models path
  • NVIDIA NIM cache path
  • Custom paths

Downloaded Models then lists what was found under each, with a Deploy button per model. Removing a path removes it from settings only — the files are not deleted.

Adding models to your library​

  1. Explore Models on Your Devices — Cordatus asks the device for its HuggingFace, Ollama and NIM directories and reads each one. Models already in your library are marked Added, the rest New; pick target engines and import them. Rescan cleans up anything the device no longer stores.
  2. Add a model manually — for a model in a custom path: device, path, name, type, tag and engines.

Model transfer​

  • Transfers run device-to-device over your own network
  • Interrupted transfers resume, and corrupted files are detected by checksum and re-downloaded
  • Progress is shown per file, with totals
  • Use API Only skips the copy when the model is reachable over an API
  • Volume mapping is configured automatically once it completes

Deploying a user model​

AI Launch plans directly for the device that holds the model; Manual → Run manually with… opens the wizard with the model pre-selected and the correct paths already mapped.

→ User Models


VRAM Calculator​

Calculate before you deploy​

  1. Select a model — search Hugging Face; parameters are fetched automatically. GGUF repositories get a per-file quantization selector with approximate bits-per-weight.
  2. Choose the GPU — from the catalogue, or read the real GPUs off one of your devices with Select Device GPUs.
  3. Configure — GPU count, quantization, sequence length, batch size, GPU memory utilization, and the calculation type (20% Overhead or With Activation).
  4. Read the result — a doughnut chart of Model Weights / KV Cache / Overhead or Activations / Free VRAM, the detailed metrics, a usage bar, and a VRAM is Sufficient / Insufficient verdict.

Use cases​

  • Test hardware requirements before purchasing GPUs
  • Optimize quantization, sequence length and batch size for the hardware you have
  • Plan multi-GPU deployments
  • Compare different model configurations side by side

→ VRAM Calculator


Container management​

Operations​

  • Start / Stop — individually, per child, or for a whole group
  • Delete — type DELETE to confirm; volumes are listed first and can be skipped
  • Duplicate — a new launch pre-filled with this container's configuration
  • Open — the group panel, with a search box for large groups

Disabled actions always say why: Container(s) already stopped, Device connection unavailable, You don't have permission to duplicate.

Batch operations​

Select rows with the checkboxes and use Delete Selected. Group actions apply to every container in the group.

Container information​

  • Logs — live output, with Copy
  • Parameters — every value the container was created with
  • Ports — local and global addresses, with copy and open

Public URLs​

Generate a public URL for any exposed port, then regenerate, close or reassign it. Newly exposed ports can be added from the same tab.

API usage examples​

For a running engine: a quick call, an OpenAI-compatible client snippet, and Open with WebUI.

Sparkrun workloads​

Multi-node runs live on their own tab, with per-node logs, access URLs, engine arguments, environment variables (masked) and the recipe YAML.

→ Containers


Key concepts​

Application — a container image registered in Cordatus, with its versions, supported device classes and published options.

Environment — a development image rather than an inference engine. Same launch flow, no model step.

Version / chipset tag — a specific image tag for a specific device class, shown as Downloaded or Will Download.

Container group — several containers deployed together as one unit.

Serving layer — the optional chat interface (Open WebUI) and API gateway (LiteLLM) that come up beside a model.

Sparkrun workload — a multi-node inference run on a DGX Spark cluster; not a Docker container, so it is listed on its own tab.

User model (User Models) — a model already stored on one of your devices and registered with Cordatus.

Overrides — the container options, environment variables, engine flags and notebook settings layered on top of what the application fixes. Fixed rows are locked and cannot be removed.

Quantization — BF16/FP16 (full precision), INT8/FP8 (half the memory), INT4/FP4 (a quarter), and the GGUF variants with their own bits-per-weight.

VRAM components — model weights, KV cache, activation memory or overhead, and free VRAM.


Best practices​

Resource management​

  • Keep the Host Reserved CPU and RAM values; they are what keeps the device responsive
  • Leave 20–30% free VRAM for unexpected load
  • Use the VRAM Calculator before production deployments
  • Read the GPU fit advisor line before choosing a model — "fits, but only just" means a long context or many concurrent requests may fail

Model management​

  • Register the HuggingFace, Ollama and NIM paths on every device you deploy to; transfer, scanning and "already on device" detection all depend on it
  • Keep parameter count and quantization filled in — the fit advisor and VRAM estimates use them
  • Rescan after deleting models on a device
  • Use meaningful names for custom models

Deployment​

  • Test with a small model first
  • Use With Activation in the VRAM Calculator for an accurate estimate
  • Read Review & start before committing — it shows the exact command and every container
  • Watch the logs on the first deploy of a new image
  • Copy serving layer credentials immediately; they are shown once

Multi-GPU and multi-node​

  • Use NVIDIA Dynamo for distributed inference on one device, and Sparkrun for multi-node
  • Assign every worker a GPU before starting — the pipeline pane will not report ready otherwise
  • Choose the router to match the workload: kv for prefix reuse, round-robin for even load
  • Watch GPU utilisation across all workers after launch

What is new in this release​

If you have used the previous Application Hub, these are the changes that matter:

  • AI Launch. A whole launch path where the plan is produced for you, with per-decision evidence, blockers and warnings.
  • The advanced step is a rail, not a wall of panels. Model, Compute, Access, Overrides and Review, each with its own pane and completion mark. The readiness checklist now sits beside the Start button instead of a screen away.
  • Review & start. The resulting command and the full list of containers shown before you commit.
  • GPU fit advisor under the model catalogue.
  • Serving layer as a first-class thing — with the model, reused if already running, or managed on its own from the Serving Layer page.
  • Profiles — Quick start / Serve an API / Gateway only / Engine only.
  • Model download onto a device, with quantization selection and honest failure reporting.
  • Sparkrun Clusters and Sparkrun Workloads for multi-node DGX Spark inference.
  • Jetson AI Lab, with an express run panel that resolves image, model and engine arguments from the published recipe.
  • LLM Models rebuilt with four source tabs, GPU- and Jetson-aware filtering, and per-model AI Launch.
  • Environments, Containers, Applications and VRAM Calculator redesigned around the same card, table and subheader components.

Where to go next​