The vLLM engine (GPU concurrent serving)
vLLM is one of Skulk's served engines: instead of loading the model
in-process, the worker launches an external vllm serve subprocess and proxies
its OpenAI HTTP API, the same managed-server-plus-proxy shape as the
llama_server engine. Its continuous batching and paged attention support concurrent GPU workloads.
Latency and throughput depend on the model, GPU, context, request mix and memory
budget; benchmark the intended workload before choosing an engine.
It coexists with the other engines rather than replacing them: MLX owns
Apple Silicon, and the llama.cpp engines remain the GGUF paths. vLLM is
GPU-only in Skulk's scope (vllm-cuda on NVIDIA, vllm-rocm on AMD CDNA).
When a model runs on vLLM
Model support and live node support must agree:
- The model card declares and ranks the engines that can serve it in
compatible_backends. A card that lists a vLLM backend is a vLLM candidate. - The node advertises a vLLM backend, which it does only when
SKULK_VLLM_BINpoints at a usablevllmCLI and a GPU backend resolves (declared viaSKULK_VLLM_BACKENDS, or inferred from the observed GPU vendor). A node without the binary is never a placement candidate for vLLM cards.
Signed engine-support claims can also establish compatibility for an exact artifact, engine build, capability and hardware class. Skulk applies its runner limits after that match. See Model capabilities.
When several nodes qualify, placement prefers the card's higher-ranked backend, so the card is where the "this model is better on vLLM than on llama.cpp here" judgment lives.
Current scope
The engine serves single-node text generation with tool calling. Its boundaries are enforced loudly rather than degraded silently:
- Tool calling works on parser-pinned cards. A card that pins
vllm_tool_call_parserin its[runtime]section (the registry's Qwen2.5 vLLM cards pinhermes) launches the server with vLLM's native tool-call parsing, and a tool-enabled request runs non-streamed so the caller receives the assembled call, the same shape as the llama.cpp engines. A tool attempt cut short bymax_tokensreportslengthinstead of surfacing an incomplete call. Cards without a pinned parser reject tool requests with a clear error rather than silently dropping them; there is no family-default fallback, because one model family can span generations with different tool wire formats. - Reasoning is split on parser-pinned cards. vLLM only separates
reasoning_contentfromcontentwhen the server is launched with a reasoning parser, so a card pinsvllm_reasoning_parserin its[runtime]section (a Muse Glimmer card pinsmuse_glimmerfor both parsers) and the runner passes it as--reasoning-parser. Explicit only, no family fallback, for the same reason as the tool-call parser. An unpinned reasoning model streams its thinking inline. Muse Glimmer's always-on reasoning is steered through the template'sreasoning_strengthkwarg, which the runner derives from the request'sreasoning_effort. - Per-token logprobs are rejected with a clear error: the OpenAI SSE proxy does not surface them, and Skulk refuses to silently omit what you asked for.
- Multi-node placement is refused. vLLM's own tensor and pipeline parallelism are not wired into Skulk placement.
- Reasoning is best-effort. Thinking controls (
enable_thinking,reasoning_effort) are forwarded so the model behaves as requested, and separated reasoning deltas are parsed into thinking chunks when the server emits them; on models where vLLM needs a family-specific reasoning parser to split thinking from content, the thinking text can arrive inline in the content stream instead.
The served context window is sized to the memory the cluster admitted for the
instance (passed as --max-model-len), never blindly to the model's full
trained context.
Setup
The easiest path is the one-command installer's flag on an NVIDIA Linux node:
curl -fsSL https://raw.githubusercontent.com/Foxlight-Foundation/Skulk/main/install.sh | bash -s -- --with-vllm
This creates a dedicated virtual environment at ~/.skulk/vllm-env with
Skulk's validated dependency matrix (a pinned vLLM release, a compatible
transformers, and the matching CUDA torch backend; several GB of wheels) and
records SKULK_VLLM_BIN=~/.skulk/vllm-env/bin/vllm in ~/.skulk/skulk.env,
which the service wrappers source. The separate venv is not an accident: Skulk's
own environment and vLLM currently require conflicting dependency versions, so
vLLM must never be installed into Skulk's venv. Skulk drives its CLI purely as
an external process.
Already have vLLM installed some other way? Point SKULK_VLLM_BIN at its CLI
before launching Skulk. Confirm the supported GPU backend, engine build, model
parser pins and compiler/Python development prerequisites with skulk doctor;
a CLI path alone does not prove that the server can load a given model.
Concurrency behavior and knobs
The vLLM runner dispatches concurrently: it keeps multiple requests in flight
against the one vllm serve process at once, which is what lets the server's
continuous batching actually engage and decode them together.
SKULK_VLLM_MAX_CONCURRENT_REQUESTS(default 32) bounds how many generations the runner keeps in flight; requests beyond it queue in the runner's bounded pool. This is a client-side admission bound, not the server's batch width (vLLM batches up to its own--max-num-seqs).- vLLM's share of GPU memory (
--gpu-memory-utilization) is sized to the instance's placement. The runner passes the fraction of the device that holds the memory Skulk reserved for the model (weights with overhead, the KV cache for the served window, and a small floor), rounded up to the next ten-thousandth of the device total CUDA reports. vLLM spends whatever part of that share weights and runtime memory leave on KV cache. A vLLM model can therefore share a GPU with other models and stays within its reservation. The share never drops below the weights, the served window's KV cache at the model's own geometry (read from itsconfig.json, counting every layer at its widest attention head) and about a gigabyte of vLLM's own runtime memory, so wide-head models and very small models still fit their window. It never exceeds the previous fixed share of 0.90. SKULK_VLLM_GPU_MEMORY_UTILIZATIONpins a fixed share instead. Set it on a GPU dedicated to vLLM to give the KV cache the rest of the device, for more concurrent long requests; other models then find that GPU full. A model whoseconfig.jsondoes not name its layer count, head width and KV-head count, or a GPU whose memory total cannot be read, gets the fixed 0.90 share with a warning in the node log.
Operationally: server startup on a large model can take a couple of minutes
(weight load, compilation, CUDA-graph capture) and is allowed a generous health
deadline; the server's own log is written to a deterministic per-runner file
under the system temp directory for postmortems. Cancelling a request aborts
its proxied HTTP connection, which stops the server-side generation; if the
runner process itself dies, the kernel reaps the vllm serve child so it never
orphans GPU memory.
Honest performance framing
Continuous batching can improve aggregate throughput under concurrency, but it does not promise constant latency as load rises. Single-request performance also depends on quantization, hardware and speculative decoding. Use the card's backend ranking with measurements from the workload you expect to serve.
vLLM runs card-driven speculative decoding for checkpoints that ship
native multi-token-prediction heads (Qwen3.6 among them): the card's
vllm_spec_method = "mtp" and vllm_spec_num_tokens map to vLLM's
--speculative-config, engaging the model's own prediction heads with no
separate draft model. Measured on an A100-80GB, this roughly doubles
single-stream decode on Qwen3.6-27B-FP8 (2.01x at depth 2, 77%
acceptance). Vendor schemes with a separately published speculator use the
same fields: vllm_spec_method = "dflash" plus vllm_spec_draft_repo
pairs Poolside's Laguna models with their block-parallel DFlash drafter
(vLLM 0.25.1 or later), with the drafter repo resolved through vLLM's own
Hugging Face cache at engine start (measured 1.35x single-stream on an
A100-80GB; gains on other GPUs require measurement). Deep speculative depths need more scheduler budget
than vLLM's defaults provide, so for carded depths of 8 or more the runner
pins both --max-num-batched-tokens and --max-num-seqs explicitly, since
vLLM's own defaults for them vary by version and hardware; shallow MTP depths
run with vLLM's defaults untouched. DFlash speculators also JIT their kernels
through NVRTC at engine start and need a CUDA 12.8+ toolchain on the node.
See Speculative Decoding for the other engines'
mechanisms.