Skip to main content

Quick-launch a model placement

POST 

/place_instance

Place and launch a model with Skulk choosing a valid concrete placement from the requested sharding, instance metadata, and minimum-node constraints. The placement is validated against the current cluster state before the command is forwarded: an impossible placement returns 400 with the specific reason (no connected cycle, exclusions removed every candidate, a node cannot fit its shard with runtime headroom, ...). If node memory info is still being gathered (cluster just formed), the request waits up to 15 seconds for it before returning 503; retry shortly in that case. Model cards may declare several open backend tags in preference order; the planner filters current engine/build, topology, and capacity blockers before ranking and automatically falls through to the next launchable candidate. Placement accepts only models already present through signed publication, an installed card record, or an authenticated add; it never discovers an unknown Hugging Face repository as a side effect. Placement failures retain a readable error message and expose a stable category in the X-Skulk-Placement-Failure response header. If capacity changes after acknowledgement, master admission retains placement_failed in instanceFailures under the acknowledged instance ID.

Request​

Responses​

Successful Response