Release Notes 2.0.0
Status: release candidate. Publication remains pending the automated
fresh-install matrix, human acceptance, and promotion to main.
Skulk 2.0.0 turns a cluster into a platform. Signed capabilities install from a publisher's catalog and appear in the topology beside the machines that run them, an operator app reaches the cluster from anywhere through a relay that cannot read its traffic, and the fabric now generates video and music as well as text, images, and speech. Models arrive with signed cards from the model registry instead of shipping inside Skulk, and Skulk itself can explain the cluster's state and propose fixes for an operator to approve. Tool calling behaves the same way on every engine, served models take the context window a placement asks for, and the dashboard is redesigned.
This release changes how models are authorized and requires upgrading every node together, so read the upgrade notes before updating.
Highlights
-
Capabilities. Plugins now add capabilities to the fabric: services, tools, and complete applications built on the cluster's models and compute. A managed plugin runs under a per-host plugin manager, separately from Skulk, so a failing plugin withdraws only its own capabilities. Prepare a host once with
skulk-plugin-service setup, connect a publisher's signed catalog under Plugins → Browse, review a release's signed permissions, screens, and whether any action can spend money, then install it with one consent. Updates keep an installation's settings and saved work, and an uninstalled plugin can be removed completely. In the Cluster view, each capability node appears as a satellite of the machine that runs it, with a flyout that opens its screens and actions. Plugins can stream as well, for example live audio, through the same discovery and stream calls as in-process providers. Plugins are managed from a browser on the host itself, over Tailscale, or from a paired device granted plugin access. Skulk Video Studio, a video workspace for MiniMax H3, shows what a plugin can deliver. A plugin is not sandboxed, so install only from publishers you trust. -
Operate a cluster from your phone. The Skulk operator app for iOS and Android, distributed separately, pairs with a cluster by QR code and reaches it from anywhere through a relay, with no VPN on the phone. One API node acts as the operator gateway:
skulk operator configure-relay --provisioning-file <file>installs the route the relay service supplies, the gateway connects outward to the relay, and the app's TLS session ends at the gateway, so the relay forwards encrypted bytes it cannot read. Invitations are created under Settings → Devices & pairing on the gateway's dashboard, opened on that machine or over Tailscale, or withskulk operator pair; one invitation can pair up to 20 devices for as long as 90 days, and any paired device can be revoked at once. Every request a device makes is checked against the permissions it was granted, and if the gateway or relay is down only remote access stops. -
Video generation.
POST /v1/videosis an OpenAI-shaped job API for audio and video: create a job from a prompt, from first and last frames, or from image, video, and audio references, follow its progress, then download the finished MP4 and its thumbnail. MiniMax H3 renders through a managed ComfyUI engine on NVIDIA CUDA and AMD ROCm nodes, with timed keyframes, turbo adapters, style embeddings, ControlNet guides derived from ordinary footage (pose, depth, or edges), and masked regeneration.GET /v1/modelspublishes every setting a video card accepts, and a finished job records exactly what it ran with. Video models are opt-in per node withSKULK_ENABLE_VIDEO_MODELS=true. -
Music generation.
POST /v1/musicturns a text prompt, with optional lyrics, into a song as an asynchronous job: follow its status, then download the WAV. The signed catalog offers MiniMax Music 3 (lyrics required) and ACE-Step 1.5 Turbo (lyrics optional) for 10 to 60 second targets. Their engine, audio.cpp, is not part of the Skulk install: the first time a music model is mounted, one eligible worker downloads, verifies, and caches a pinned engine package. A model mounts only where a signed support claim covers that exact model, engine build, and hardware. At release, signed claims cover both models on AMD Strix Halo (Vulkan), NVIDIA GB10, and single NVIDIA GPUs of compute capability 8.9 on Linux; MiniMax Music alone on Apple Silicon (Metal); and ACE-Step alone on Linux CPUs. Music is available through the API only in this release. -
Models bring their own cards. Skulk no longer ships model cards. A node's catalog is the signed model registry, the card each installed model keeps beside its files, and any custom cards an operator adds. An installed model's card and a hashed manifest of its files live with the bytes, so the model stays servable with no network at all, and a node's existing cache can move into the model store without another Hugging Face download. Nodes verify the registry's signed catalog against a root embedded in Skulk and refresh it at most once a minute, so new and corrected cards arrive without a Skulk release, and during an outage a node keeps using the last verified catalog for up to 30 days. Publishing a card, or explicitly adding one, is what authorizes its repository code; there is no separate approval step. Before any repository code runs, Skulk still verifies the card's signed identity, its pinned revisions, the installed card record, and the file manifest. A signed card can also name one exact artifact inside a repository that holds many quantizations.
-
Skulk answers for itself. With Intelligent Fabric turned on in Settings, the cluster keeps a resident model placed as a hidden system instance and answers questions about its own health, models, downloads, and diagnostics through Ask Skulk, the Skulk chat target, or any OpenAI-compatible client using the model
skulk/steward. Every turn starts from a fresh observation of the cluster, and Skulk investigates through read-only tools before answering. It can propose placing a model, stopping or restarting an instance, or cancelling a download, but nothing happens until an operator approves that exact proposal, which expires after ten minutes and is checked again against the live cluster first. Skulk speaks with its own signature voice when a streaming speech model offers it. The resident model moves up to a better one when capacity allows, is placed again after node loss or master failover, and is off by default. -
Tool calling behaves the same on every engine. A model that calls several tools at once returns all of them in one response. A call that follows reasoning is recognized, and a call the model only considered while reasoning is never carried out.
tool_choicemeans the same thing whichever engine serves the model:"none"guarantees no call, and naming one function guarantees the model cannot call another. A request that offers no tools never receives call markup as answer text. Calls are recognized when their opening marker arrives split across streamed pieces or after a sentence of prose, and Llama, Mistral, GLM, and Gemma 4 call formats now work on the engines that previously returned them as raw text. The Muse Glimmer family is supported on the MLX, llama-server, and vLLM engines. -
Context windows you choose. Served engines (llama-server, in-process llama.cpp, and vLLM) reserve a window's memory when a model loads. They now default to
inference.served_context_tokens(32,768, editable in Settings) unless a placement asks for more, and the dashboard's placement dialog shows the largest window the chosen nodes can hold and roughly how much memory it reserves. Unified-memory nodes (Macs, Strix Halo, and GB10) size served GGUF windows from live available memory instead of a fixed 8,192 tokens, and vLLM now starts beside other resident models with its GPU share sized to its placement. -
Engines. The managed llama.cpp engine advances from b10092 to b10753, bringing upstream's newer architectures and fixes to the served MTP path, and its CUDA wheel now ships for Linux ARM64 as well, so Grace Blackwell and GB10 nodes serve through CUDA instead of Vulkan. vLLM's validated version moves from 0.25.1 to 0.28.0, and Apple nodes move to MLX 0.32.2 with matching mlx-lm, mlx-vlm, and mlx-audio. GGUF vision cards can pin one exact multimodal projector and run through llama-server on CUDA, ROCm, Vulkan, or CPU, including multi-node RPC placements. A separately signed engine-support claim can make an artifact placeable on the exact engine build and hardware class it was tested on without replacing its card, and placement falls through a card's ordered engine preferences when the preferred engine or node is unavailable.
-
A redesigned dashboard. The dashboard comes in Night (dark) and the new Noon Ridge (light), chosen under Settings → Appearance. Settings opens as a resizable drawer that keeps unsaved edits, Chat's model chooser sits in the composer, and Find Models filters supported models or all of Hugging Face by family, task, fit, store presence, and readiness. The Cluster view draws each node as a gauge of memory, compute, and health with its vendor mark, and Node information shows its chip, operating system, version, observed links, and any health problem with its fix. A new Integrations page writes ready-to-paste setup for tools such as Claude Code, Codex, OpenCode, Open WebUI, and n8n from the models that are ready now. A Hugging Face token entered once in Settings reaches every node, and a gated download that fails says why and which node needs a token.
Reliability
- A runner's last events are always delivered when it exits, so stopping a model no longer stalls and records a timeout.
- A single-node placement whose runner dies is relaunched, and fails with a recorded reason if it keeps dying, instead of staying dead behind a live instance.
- A node no longer crashes when a peer restart changes the master.
- Staged models that were serving survive a restart, an update, or a won election instead of being copied from the store again.
- The dashboard no longer opens to a blank page after a Skulk update.
- An NVIDIA GB10 now reports its GPU memory, measured through CUDA and counting reclaimable page cache as free, and a Strix Halo counts a loaded Vulkan model once, so both admit models that fit.
- Linux runners no longer load the unused MLX library while stopping, which could crash a runner as it exited.
- A better election proposal that arrives after the timeout now corrects the result, so two nodes can no longer both stay master.
- A download reports complete only once the model can load.
- Text requests go to ready instances first and fail with
instance_unavailablewhen none is viable, instead of waiting on a failed instance. - Placement skips nodes whose data plane has no peers, which previously loaded models that never returned output.
- vLLM retries a lost port race, and the CUDA engine loads on fresh hosts.
kill -USR1 <pid>writes every thread's Python stack from a node or runner process to its log without stopping it.
Security
- mlx-lm runs a model configuration's
model_filecode only when the card allows repository code (CVE-2026-5843). Before, a card that did not allow it could still run that code. - Adding models and starting downloads now require direct operator access, as described in the upgrade notes.
First-install acceptance
Automated qualification covers the exact physical-fleet and clean NVIDIA matrix. The final usability pass follows the Human Release Qualification guide against the same candidate commit. Any product, installer, default, or dashboard fix creates a new candidate and repeats the automated gate.
Upgrade notes
- All nodes in a cluster must run the same Skulk version before serving workloads, and 2.0.0 changes the messages nodes exchange, so upgrade the whole fleet together: 1.5.x and 2.0.0 nodes do not interoperate.
- Skulk ships no model cards, and nodes need HTTPS access to the model
registry (
https://registry.foxlight.ai/, orSKULK_MODEL_REGISTRY_URL). Start each upgraded node online once: a model downloaded under 1.5.x has no card record until the registry recognizes it, and until then it stays unlisted (skulk doctorreports it underinstalled-card-records). A node that has never reached the registry and holds no installed or custom models starts with an empty catalog and logs why.GET /v1/modelsnow reports each model'scatalog_source:registry,installed, orcustom. - Reading or launching an unknown Hugging Face model no longer fetches its
card implicitly. Add the model first; an addition without a revision
resolves
mainonce to an immutable commit. - A custom card that runs repository code without an immutable
source_revisionfails closed until it is added again, and separately hosted companion repositories (processors, vision weights, assistant and draft models) need their own revisions. - Repository code runs only for a card that signed publication or an explicit addition authorized, and only at its pinned revision. In 1.5.x, a card fetched from Hugging Face allowed repository code by default.
- Plugins are admitted only when their
skulk_requiresincludes the running version, so plugin releases built for 1.5.x are refused. A managed plugin release also fits one exact Skulk build, so after any Skulk update install the publisher's release for the new build; until then the plugin reads Needs attention. - Intelligent Fabric is off by default. Turning it on downloads and places a resident model that uses real memory, and ordinary deletion of that instance is refused while the mode is on.
- Without
context_tokens, served placements (llama-server, in-process llama.cpp, and vLLM) now get a 32,768-token window. Discrete-GPU placements used to receive their full memory fit, and unified-memory or CPU placements a fixed 8,192 tokens. Raiseinference.served_context_tokensin Settings, or request a window per placement withcontext_tokensonPOST /place_instance. MLX is unchanged. - Rerun the installer to update engines.
install.sh --with-vllminstalls vLLM 0.28.0 (vllm==0.28.0+cu129), which Skulk never upgrades by itself, and a vLLM node needs a C++ compiler and Python headers (skulk doctorchecks both). Online NVIDIA nodes install the new CUDA llama-server wheel at startup; rerun the installer for the Vulkan wheel, or Skulk falls back to the checksum-verified upstream archive. Offline nodes download nothing, so provision them before upgrading.llama-serverandggml-rpc-servermust match, so upgrade multi-node RPC pools together. - vLLM's GPU share now follows its placement. On a GPU dedicated to vLLM, set
SKULK_VLLM_GPU_MEMORY_UTILIZATIONto keep giving the KV cache the rest of the device. - A configuration that enables the model store with a blank
store_hostorstore_pathis refused at startup and in Settings. - Streaming chat completion chunks carry
"object": "chat.completion.chunk", as the OpenAI format documents. Clients that validated the old value need updating. - Adding a model (
POST /models/add, the dashboard's Add & Download) and starting a node download (POST /download/start) now require direct operator access: the node itself, a private LAN or Tailscale peer that sends no proxy-forwarding headers, or a paired device through the operator gateway. Before, any caller that reached the API could use them; requests through a reverse proxy now receive 403. The card a download request carries must also match the current catalog exactly. - A Hugging Face token saved in Settings is now written to every node and
used ahead of an
hf auth logintoken file. Only anHF_TOKENexported when a node starts keeps precedence on that node, and a blank field never clears a saved token. - Music needs no setting. The first mount on a node downloads the engine
package from Foxlight's package index as well as the weights; an offline
node, or one started with
SKULK_NO_ENGINE_AUTOPROVISION=1, can use only a package it already cached. Music cards place only from the signed catalog. - Remote access stays off until a gateway is configured, and it adds no
skulk.yamlkeys. Reusable pairing invitations need a current operator app build; the default five-minute QR also works with earlier builds.
The complete change inventory remains in the repository changelog.