Skip to main content

Instances

Placement previews, launch flows, instance lookup, and lifecycle management for running models.

📄️ Create an instance from a fully specified placement

Create an instance from an already computed placement object when you want exact control instead of Skulk picking the placement. Every embedded shard card must identify the same model alias as the assignment and exactly match the effective card already present in the authorized local catalog. A positive contextTokenLimit caps the requested window: admission may lower it for current resources but never raises it. Nonpositive explicit limits are rejected during placement. Master admission requires observed VRAM for concrete GPU shards and accounts for existing placements. Omitted non-RPC backends resolve from node engine telemetry before admission; missing evidence is refused. Text-to-music placements require one specified node; the API prepares audio.cpp there and verifies a ready signed build claim before accepting the command. A refused acknowledged command retains placement_failed evidence in the instance failure history.

📄️ Quick-launch a model placement

Place and launch a model with Skulk choosing a valid concrete placement from the requested sharding, instance metadata, and minimum-node constraints. The placement is validated against the current cluster state before the command is forwarded: an impossible placement returns 400 with the specific reason (no connected cycle, exclusions removed every candidate, a node cannot fit its shard with runtime headroom, ...). If node memory info is still being gathered (cluster just formed), the request waits up to 15 seconds for it before returning 503; retry shortly in that case. Model cards may declare several open backend tags in preference order; the planner filters current engine/build, topology, and capacity blockers before ranking and automatically falls through to the next launchable candidate. Placement accepts only models already present through signed publication, an installed card record, or an authenticated add; it never discovers an unknown Hugging Face repository as a side effect. Placement failures retain a readable error message and expose a stable category in the X-Skulk-Placement-Failure response header. If capacity changes after acknowledgement, master admission retains placement_failed in instanceFailures under the acknowledged instance ID.

📄️ Preview valid placements for a model

Return candidate placements for a model before launch. This is the best first step when you want to see what Skulk can place on the current node or cluster. Pass `excluded_node_ids` (repeatable) to mirror the `excluded_nodes` field on POST /place_instance and preview against the post-exclusion topology. Besides the planner's ranked pick per placement shape, the response includes per-host single-node previews marked `alternative: true` for every other host that passes admission, so heterogeneous fleets expose the full set of valid hosts rather than only the ranking winner. Previews explain current planner choices; ordinary launch requests remain adaptive and do not reserve or replay one preview. Unavailable previews include a stable error_code.