PlacementPreview
card_digest object
Canonical authorized-card content digest used to construct a launchable preview; compare with approved model requirements before submitting the exact instance. Null when no instance is present.
- string
- null
Possible values: Value must match regular expression ^[a-f0-9]{64}$
Possible values: [Tensor, Pipeline]
Possible values: [MlxRing, MlxJaccl, LlamaRpc]
instance object
- MlxRingInstance
- MlxJacclInstance
- LlamaRpcInstance
- null
shardAssignments objectrequired
runnerToShard objectrequired
property name* object
- PipelineShardMetadata
- CfgShardMetadata
- TensorShardMetadata
- RpcDonorShardMetadata
modelCard objectrequired
The persisted, declarative metadata Skulk holds for one model.
This is the model-card interface: the single source of truth for how a
model is sized, sharded, placed, and run. It is created once (from a
HuggingFace repo or hand-authored), broadcast cluster-wide, and read by the
planner (placement), the downloader (sizing + which files to fetch), and the
worker runner (engine + behavior). As a CamelCaseModel it is camelCase on
the wire and strict (extra="forbid"), so every node in a cluster must run
the same Skulk version (a stale node rejects newer fields).
Two layers live here: the card (this declarative metadata) and the
normalized resolved capability profile derived from it plus family defaults
(see capabilities.py and website/docs/model-capabilities.md). The
optional reasoning / modalities / tooling / runtime /
vision / placement sub-configs refine that resolution; when absent,
conservative family defaults apply.
The selectable artifact alias. Historically this was always the upstream Hugging Face repository id; registry cards may use a distinct alias so two exact files or quants from one repository remain separate artifacts.
sourceRepository object
Upstream Hugging Face repository that owns the artifact bytes. None
means it is identical to model_id for legacy and locally generated cards.
- string
- null
storageSize objectrequired
On-disk size of the weights this card loads (for a GGUF card, just the selected quant's shard group, not every quant the repo hosts). The planner uses this for memory-fit and placement-width decisions.
0Number of transformer layers. Drives pipeline sharding (how layers split across nodes) and KV-cache sizing.
Possible values: > 0
Model hidden dimension, used in memory/KV-cache estimates.
Possible values: > 0
Whether the model may be served with tensor parallelism (Sharding.Tensor).
GGUF/llama.cpp cards set this False (single-node engine).
numKeyValueHeads object
KV-head count for grouped-query attention, used in KV-cache sizing. None
when unknown/not applicable.
- integer
- null
Possible values: > 0
ggufCacheGeometry object
Artifact-derived attention and recurrent cache dimensions for GGUF admission.
This is intrinsic model metadata. Slot count, speculation and runtime buffer
overhead are separate engine inputs. An absent value means unknown, never
zero recurrent cost. Registry geometry is an explicit projection of the
separately signed header target retained in registry_gguf_metadata.
- GgufCacheGeometry
- null
Target full-attention layers.
Possible values: >= 0
Target recurrent layers.
Possible values: >= 0
Embedded MTP attention layers.
Possible values: >= 0
Key elements per token per layer.
Possible values: > 0
Value elements per token per layer.
Possible values: > 0
Convolution-state elements per recurrent layer and row.
Possible values: >= 0
Recurrent-state elements per recurrent layer and row.
Possible values: >= 0
The task types this model serves (TextGeneration, TextEmbedding,
TextToImage, ImageToImage, TextToSpeech, SpeechToText,
SpeechTranslation, TextToMusic, TextToVideo, ImageToVideo,
ReferenceToVideo); selects which runner handles it.
Possible values: [TextGeneration, TextToImage, ImageToImage, TextEmbedding, TextToSpeech, SpeechToText, SpeechTranslation, TextToMusic, TextToVideo, ImageToVideo, ReferenceToVideo]
components object
For multi-component models (e.g. a diffusion stack), the per-component
weight layout. None for a single-weights model.
- object[]
- null
Logical name of this component (e.g. text_encoder, transformer).
Repo-relative subdirectory holding this component's weights.
storageSize objectrequired
On-disk size of this component's weights.
0nLayers object
Layer count for this component when it is shardable; None otherwise.
- integer
- null
Possible values: > 0
Whether this component may be split across nodes (vs. loaded whole).
safetensorsIndexFilename object
The component's *.safetensors.index.json filename when sharded across
files; None for a single-file component.
- string
- null
Model family token (e.g. qwen3, gemma4) used to pick family-specific
defaults during capability resolution. Empty when not classified.
Human quantization label (e.g. 4bit, Q4_K_M); informational.
The upstream base model id when this is a quant/finetune of another; empty if not applicable.
ggufFile object
For GGUF (llama.cpp) models: the repo-relative path of the weights file the
runner loads (the selected quant's first shard). Resolved once at card creation
(preferring a quant over BF16) so the download fetches only that quant and the
runner loads deterministically, instead of each layer re-globbing/guessing.
None for non-GGUF (safetensors/MLX) cards.
- string
- null
artifactBundle object
Exact signed file selection and engine working directory for a v2 card.
Legacy cards omit this field and preserve repository-wide tensor downloads.
- ArtifactBundleConfig
- null
Content-derived immutable identity of the normalized file bundle.
root object
Repository-relative directory used as the engine's model root.
- string
- null
files object[]required
Every repository file required to install and run the artifact.
Canonical repository-relative POSIX path.
Exact upstream byte size at the card's immutable source revision.
objectId object
Optional algorithm-qualified Hub object identity used for verification.
- string
- null
Total bytes transferred for the complete bundle.
sourceRevision object
Immutable Hugging Face commit for this card's model artifacts.
None preserves the historical behavior of resolving the repository's
mutable main branch. Curated or operator-authored cards should set this
to a full commit hash when the exact artifact has been qualified, so an
upstream file replacement cannot silently change what the store and workers
execute.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
Free-form capability tags carried for compatibility/auxiliary use; the
structured reasoning/modalities/tooling configs are authoritative
for capability resolution.
[]The model's advertised maximum context length in tokens (0 if unknown).
The admission ceiling is the smaller of this and what fits in memory.
0Whether the model uses classifier-free guidance (relevant to some image / diffusion models).
falsePassed to the model loader: whether to execute the repo's custom Python.
Defaults True to match upstream loaders; set False to refuse it.
trueMarks an operator-added custom card (not from the curated catalog). Excluded from the persisted card file so it is recomputed per environment.
falseMarks an unsigned custom card owned by the temporary qualification lifecycle.
The registry service credential may clean up only cards carrying this marker; ordinary operator-owned custom cards remain outside its deletion authority.
falsegeneratorRevision object
Revision of the machine card generator that produced this card
(:data:CARD_GENERATOR_REVISION at generation time), persisted in the card
file. None means hand-authored (or generated before revisions existed):
such a card keeps full override precedence over a signed or installed card
for the same id. A stamped card older than the current generator is
superseded by that card at load time, since a generated card is a cache of Hugging
Face metadata plus generator logic, not operator intent.
- integer
- null
vision object
Optional vision (image-input) configuration; None for text-only models.
- VisionCardConfig
- null
imageTokenId object
Token id the model uses as the image placeholder in the prompt. Required by
the MLX vision path (which splices image embeddings at this token); None
is allowed for a llama.cpp-only vision GGUF, whose chat handler inserts image
features itself and never reads this. MLX cards always set it (from
config.json).
- integer
- null
Vision model-type tag (from config.json's vision_config), selecting
the image processor (MLX) or chat handler (llama.cpp). Empty when a bare GGUF
repo only signals vision via its mmproj projector; the llama.cpp runner
then falls back to its general multimodal handler.
Repo holding the vision-tower weights when separate from the LM; empty if bundled with the main weights.
weightsRevision object
Immutable commit for a separate weights_repo.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
imageToken object
The literal image placeholder string, when distinct from image_token_id.
- string
- null
processorRepo object
Repo providing the image processor/preprocessor config, if not the main repo.
- string
- null
processorRevision object
Immutable commit for processor_repo. Signed registry cards require
this whenever a separate processor repository can supply executable code.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
boiTokenId object
Begin-of-image token id, for families that bracket image spans.
- integer
- null
eoiTokenId object
End-of-image token id, for families that bracket image spans.
- integer
- null
projectorFile object
Exact repository-relative GGUF projector selected for served vision.
The projector is pinned by the owning card's immutable source_revision.
Legacy GGUF vision cards may omit this field and continue to use the
in-process llama.cpp compatibility path.
- string
- null
projectorSize object
Exact byte size of projector_file at source_revision.
- integer
- null
Possible values: > 0
reasoning object
Optional reasoning/thinking configuration (toggle, budget, format, default
effort); None falls back to family defaults.
- ReasoningCardConfig
- null
supportsToggle object
Whether the model can have reasoning turned on/off per request.
- boolean
- null
supportsBudget object
Whether the model accepts a reasoning-effort/budget control.
- boolean
- null
format object
How reasoning is marked in the output stream: none, token_delimited
(special tokens), or channel_delimited (a separate reasoning channel).
- ReasoningFormat
- null
Reasoning marker formats used by model families.
Possible values: [none, token_delimited, channel_delimited]
defaultEffort object
Reasoning effort applied when the request does not specify one.
- string
- null
Possible values: [none, minimal, low, medium, high, xhigh]
disabledEffort object
The effort value that means "reasoning off" for this model.
- string
- null
Possible values: [none, minimal, low, medium, high, xhigh]
modalities object
Optional extra-modality flags (audio input, native multimodal); None
falls back to family defaults.
- ModalitiesCardConfig
- null
supportsAudioInput object
Whether the model accepts audio input.
- boolean
- null
supportsNativeMultimodal object
Whether the model natively interleaves modalities (vs. a bolt-on adapter).
- boolean
- null
audio object
Optional speech-serving configuration (TTS/STT kind, audio formats,
streaming/realtime support, voices, reference audio, translation, sample
rates); None for non-speech models.
- AudioCardConfig
- null
kind object
Speech serving kind: tts for text-to-speech or stt for speech-to-text.
- AudioCardKind
- null
Speech model kind declared by a model card's [audio] section.
Possible values: [tts, stt]
defaultResponseFormat object
Default encoded audio response format for TTS requests.
- AudioResponseFormat
- null
Audio response formats supported by the speech serving API.
Possible values: [mp3, wav, flac, ogg, opus, pcm]
Encoded audio formats this model can produce for TTS requests.
supportsStreaming object
Whether a validated Skulk runtime path can stream partial speech or transcripts.
- boolean
- null
supportsRealtime object
Whether the model exposes a realtime session interface.
- boolean
- null
supportsVoiceListing object
Whether the model can enumerate voices through a voice-listing API.
- boolean
- null
Stable built-in voice identifiers exposed by the model.
voiceCatalog object[]
Optional display and language metadata for every declared built-in voice.
Model-specific voice identifier accepted by speech synthesis.
Human-readable voice name shown by clients.
Ordered BCP 47 language tags for which this voice is a preferred match.
[]referenceProfile object
Bundled reference profile used to condition models without built-in voices.
- string
- null
defaultVoice object
Built-in voice used when a TTS request omits an explicit voice.
- string
- null
supportsReferenceAudio object
Whether the model accepts managed reference audio for voice conditioning.
- boolean
- null
supportsTranslation object
Whether the model can translate speech instead of only transcribing it.
- boolean
- null
Supported output or input sample rates in hertz.
music object
Text-to-music family, lyric requirement, and qualified target-duration bounds.
- MusicCardConfig
- null
Architecture family used by the runner's fixed option translator.
Possible values: [minimax_music3, ace_step_1_5]
Whether lyrics are required, permitted, or unsupported.
Possible values: [required, optional, unsupported]
Shortest generation target accepted for this artifact.
Possible values: > 0
Longest generation target accepted for this artifact, at most 120 seconds.
Possible values: > 0
languageModelGguf object
MiniMax language-model component selected by this exact artifact.
- string
- null
rvqDepthDecoderGguf object
MiniMax RVQ decoder component selected by this exact artifact.
- string
- null
flowTransformerGguf object
MiniMax flow-transformer component selected by this exact artifact.
- string
- null
video object
Optional audio-video generation contract (modes, duration and frame
grid, canvas rules, audio output, reference bounds, sampling defaults,
pinned companions); None for models that do not generate video.
- VideoCardConfig
- null
Generation modes this artifact serves; each implies a ModelTask.
Shortest output duration the model supports.
Possible values: > 0
4Longest output duration the model supports.
Possible values: > 0
15Output frame rate.
Possible values: > 0
24Frame counts must satisfy count % frame_grid_multiple == frame_grid_offset.
Possible values: > 0
1Residue a valid frame count leaves modulo frame_grid_multiple.
0Width and height must be multiples of this many pixels.
Possible values: > 0
1defaultShortEdge object
Trained short-edge resolution used when a request gives no size.
- integer
- null
Possible values: > 0
maxPixels object
Largest width times height the model serves at native quality.
- integer
- null
Possible values: > 0
Advertised aspect ratios as W:H strings; empty means unconstrained.
Whether generated video carries a synchronized audio track.
falseaudioSampleRate object
Sample rate of generated audio in hertz.
- integer
- null
Possible values: > 0
audioChannels object
Channel count of generated audio.
- integer
- null
Possible values: > 0
Sampling steps used when a request and its companions do not decide.
Possible values: > 0
20videoShift object
Trained video sigma shift, when the sampler exposes one.
- number
- null
audioShift object
Trained audio sigma shift, when the sampler exposes one.
- number
- null
referenceLimits object
Reference bounds; required when ref2va is among the modes.
- VideoReferenceLimits
- null
Maximum reference images per request.
0Maximum reference video clips per request.
0Maximum reference audio clips per request.
0maxFiles object
Maximum files across every reference type; None means the sum.
- integer
- null
clipMinSeconds object
Minimum duration of one reference video or audio clip.
- integer
- null
Possible values: > 0
clipMaxSeconds object
Maximum duration of one reference video or audio clip.
- integer
- null
Possible values: > 0
totalClipSeconds object
Maximum combined duration of reference clips of one type.
- integer
- null
Possible values: > 0
companions object[]
Pinned adapters, patches, embeddings, and graph templates.
What the companion is; selects how an engine applies it.
Possible values: [lora, model_patch, embedding, graph_template, preprocessor]
Stable companion identifier callers and engines refer to.
Canonical repository-relative POSIX path of the companion file.
repo object
Repository hosting the companion; None means the card's artifact
repository, whose source_revision then also pins this file.
- string
- null
revision object
Immutable commit of repo; required whenever repo is external.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
sizeBytes object
Exact upstream byte size at the pinned revision when known.
- integer
- null
Generation modes the companion applies to; empty means every mode.
steps object
For distillation adapters, the sampling step count they were trained for.
- integer
- null
Possible values: > 0
strength object
Default application strength when the engine supports one.
- number
- null
videoShift object
Video sigma shift the companion expects, when it differs from the card.
- number
- null
audioShift object
Audio sigma shift the companion expects, when it differs from the card.
- number
- null
role object
For preprocessor companions, what the weights do; required for them and refused on every other kind.
- VideoPreprocessorRole
- null
What a preprocessor companion's weights do in deriving a guide video.
Possible values: [pose_estimator, person_detector, depth_estimator]
license object
The license its hosting repository declares, as a lowercase SPDX-style
identifier (mit, apache-2.0), shown to the operator beside the
card's own license; companions in the card's repository fall under that.
- string
- null
license object
Optional operator-facing license facts surfaced by the catalog and user interfaces; never enforced by download or placement.
- LicenseCardConfig
- null
Human-readable license name.
url object
Where the license text lives.
- string
- null
spdxId object
SPDX identifier when one exists; custom community licenses have none.
- string
- null
notice object
Short operator-facing note, for example a territorial scope or an application requirement.
- string
- null
displayName object
Product attribution the license requires user interfaces to show prominently, for example the model's brand name.
- string
- null
tooling object
Optional tool-calling configuration (support, call format, builtin tools);
None falls back to family defaults.
- ToolingCardConfig
- null
supportsToolCalling object
Whether the model supports function/tool calling.
- boolean
- null
toolCallFormat object
The wire format the model emits tool calls in (generic, gemma4,
gpt_oss, dsml), selecting the output parser.
- ToolCallFormat
- null
Tool-call output formats emitted by model families.
Possible values: [generic, gemma4, gpt_oss, dsml, atem]
builtinTools object
Builtin tools Skulk advertises to this model (e.g. web_search,
open_url, extract_page).
- BuiltinToolType (string)[]
- null
Builtin tool contracts that Skulk can advertise to model families.
Possible values: [web_search, open_url, extract_page]
runtime object
Optional runtime-behavior configuration (prompt renderer, output parser,
MTP/speculative-decoding sidecar, MLX knobs); None falls back to defaults.
- RuntimeCapabilityCardConfig
- null
promptRenderer object
How prompts are rendered for this model (tokenizer chat template,
gemma4, dsml); None uses the family default.
- PromptRendererType
- null
Prompt renderer strategies supported by the runtime.
Possible values: [tokenizer, gemma4, dsml]
outputParser object
How model output is parsed (generic, gemma4, gpt_oss,
deepseek_v32), e.g. for reasoning/tool-call extraction; None uses the
family default.
- OutputParserType
- null
Output parser strategies supported by the runtime.
Possible values: [generic, gemma4, gpt_oss, deepseek_v32, muse_glimmer]
metalFastSynch object
Per-model override for the MLX MLX_METAL_FAST_SYNCH flag.
None means "no opinion" — fall through to the cluster default
selected by the runner. Set explicitly to False for models that
deadlock under FAST_SYNCH on the ring backend (e.g. gemma-4 with
multimodal load: the Metal command queue wedges in
pipeline_last_eval_output, transitively starves WindowServer,
and trips the macOS kernel watchdog into a panic). Set explicitly
to True for models that have been measured to benefit and are
known to be safe under the deployment's collective backend.
- boolean
- null
mtpHeads object
True when native MTP prediction heads are available via sidecar.
Set alongside mtp_sidecar_repo. When false or absent, the runner
skips sidecar loading and uses standard autoregressive generation.
- boolean
- null
mtpMaxDepth object
Maximum draft depth the MTP heads support.
Start at 1 for Apple Silicon. Deeper values can be evaluated via profiling but are unlikely to amortize on Metal due to near-linear verify-pass scaling.
- integer
- null
mtpSidecarRepo object
Hugging Face repo ID containing the published mtp.safetensors sidecar.
Example: "FoxlightAI/qwen3-5-7b-instruct-mtp-q4k"
The sidecar is downloaded alongside the base model weights and loaded
into the runner for speculative decoding. Produced by SWP.
- string
- null
mtpSidecarRevision object
Immutable commit for a separate mtp_sidecar_repo.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
mtpNormConvention object
How the sidecar stores its RMSNorm weights.
"zero_centered" means deviation-from-1 (the raw Qwen3.5 checkpoint
convention — the runner applies a +1.0 shift at load, mirroring what
mlx-lm's sanitize() does for trunk weights). "actual_scale" means
the stored value is the scale itself (DeepSeek convention). None
falls through to the family default keyed off the detected sidecar
layout. Override per card when a publisher changes conventions — getting
this wrong measured 0% draft acceptance on Qwen3.5-2B (issue #192).
- string
- null
Possible values: [zero_centered, actual_scale]
mtpConcatOrder object
Concatenation order of the MTP fc projection input.
"embed_first" = fc(concat([enorm(embed(t_next)), hnorm(h)])) —
verified for Qwen3.5 (72.4% offline agreement, issue #192).
"hidden_first" is the inherited DeepSeek assumption (unverified).
None falls through to the family default keyed off the detected
sidecar layout.
- string
- null
Possible values: [embed_first, hidden_first]
speculativeMultiNode object
Whether speculation may run on multi-node placements of this model.
None (default) places no restriction. Set False for models
where multi-node speculation is measured SLOWER than plain distributed
decode: the 2026-06-06 benchmark matrix found gemma-4-26B-A4B (MoE)
at 30.2 tok/s plain vs 28.2 with MTP on a 2-node pipeline (-7%), while
single-node MTP on the same model measures 2.2x — fast sharded MoE
decode plus modest acceptance makes the per-round draft+verify
overhead net negative. Single-node speculation is unaffected by this
knob. The decision is card-driven so every rank makes the same
speculate-or-not choice (the distributed agreement collective requires
rank symmetry).
- boolean
- null
assistantModelRepo object
Hugging Face repo ID of a companion assistant (drafter) model.
Gemma 4 does speculative decoding differently from the Qwen3/DeepSeek
mtp.* heads: instead of embedded prediction heads, it pairs the target
with a separate small gemma4_assistant model (e.g.
"mlx-community/gemma-4-26B-A4B-it-assistant-bf16") that cross-attends
over the target's KV cache. When set, the assistant repo is downloaded
alongside the base model. Mutually exclusive with the mtp_* fields.
NOTE: consuming the assistant for speculative generation requires the
gemma4_assistant drafter from mlx-vlm >= 0.5.0 and is not yet wired into
the runner — declaring it here only pre-downloads it. See the Gemma 4 MTP
initiative in the foxlight-docs hub (Phase C).
- string
- null
assistantModelRevision object
Immutable commit for a separate assistant_model_repo.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
servedSpecType object
Speculative-decoding mode for the llama_server (served-backend) engine.
Maps to a llama-server --spec-type token in the runner
(_SPEC_TYPE_FLAG): draft_mtp -> draft-mtp (usually the model's own built-in
MTP heads; a separate draft is optional, e.g. Gemma 4's assistant;
Qwen3.6/DeepSeek/GLM/Kimi/Nemotron bake theirs in),
draft_eagle3 -> draft-eagle3 (an EAGLE-3 head), draft_simple ->
draft-simple (a separate draft model), draft_dflash ->
draft-dflash (a separate block-parallel DFlash speculator GGUF via
served_spec_draft_repo/served_spec_draft_file; llama-server >=
b10092), ngram -> ngram-cache (prompt-lookup), none/None
plain decoding. Only the served engine reads
this; the in-process mlx and llama_cpp engines ignore it (MLX
speculation is the mtp_* / assistant_model_repo fields above).
- string
- null
Possible values: [none, draft_mtp, draft_eagle3, draft_simple, draft_dflash, ngram]
servedSpecNMax object
Max draft tokens per step for the served engine (--spec-draft-n-max).
Must be a positive integer (validated at card load so a bad value fails fast
rather than producing an undefined --spec-draft-n-max at the server).
None uses the llama-server default (3). Acceptance falls off with depth
(per-position acceptance drops), so 2-3 is the usual sweet spot; tune per
card from measured acceptance.
- integer
- null
Possible values: > 0
servedSpecDraftRepo object
Hugging Face repo of a separate draft GGUF for the served engine.
Some served speculative modes need a second model passed to llama-server via
--model-draft, NOT built-in heads: draft_simple (a vocab-matched small
draft model) and draft_eagle3 (an EAGLE-3 head) always require one, and
Gemma 4 draft_mtp uses its assistant as a separate draft GGUF (llama.cpp
PR #23398) rather than baking heads into the base. Qwen3.6/DeepSeek/GLM
draft_mtp leave this unset (heads are in the base GGUF). When set, the
draft GGUF is downloaded as a companion alongside the base and passed as
--model-draft. Pairs with served_spec_draft_file.
- string
- null
servedSpecDraftRevision object
Immutable commit for a separate served_spec_draft_repo.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
servedSpecDraftFile object
Repo-relative GGUF filename of the served draft model (in
served_spec_draft_repo), e.g. "mtp-gemma-4-31B-it.gguf". Required when
served_spec_draft_repo is set; selects the exact draft quant the runner
passes to --model-draft.
- string
- null
vllmSpecMethod object
Speculative-decoding method for the vllm served engine.
Maps to vLLM's --speculative-config method key. "mtp" engages
the checkpoint's own native multi-token-prediction heads (vLLM resolves
the matching drafter architecture, e.g. Qwen3_5MTP, with no separate
draft model); requires a checkpoint that ships MTP heads
(mtp_num_hidden_layers in its config). "dflash" engages a
separate block-parallel DFlash speculator (Poolside's scheme, vLLM
= 0.25.0) and requires
vllm_spec_draft_reponaming the drafter. Only the vllm engine reads this;served_spec_typeremains the llama_server equivalent. The vocabulary starts deliberately narrow and grows as methods are validated live.
- string
- null
Possible values: [mtp, dflash]
vllmSpecNumTokens object
Draft tokens per step for vLLM speculative decoding
(--speculative-config num_speculative_tokens).
Requires vllm_spec_method. Positive; acceptance falls per position
(measured on Qwen3.6-27B-FP8: 86% at position 0, 69% at position 1), so
2 is the usual sweet spot for single-layer MTP heads, which re-run their
one layer for deeper positions. Block-parallel drafters (dflash)
predict a whole block at once, so vendor-recommended depths run much
deeper (Poolside ships 15 for a block size of 16). None uses vLLM's
method default.
- integer
- null
Possible values: > 0
vllmSpecDraftRepo object
Hugging Face repo of a separate draft/speculator model for vLLM
speculative decoding (--speculative-config model).
Required by draft-model methods (dflash); must be unset for
mtp, whose drafter lives inside the target checkpoint. The vllm
engine passes the repo id through to vllm serve, which resolves it
from its own Hugging Face cache at engine start (the target model still
stages through the Skulk model store; staging the draft through the
store as a pinned companion is a follow-up).
- string
- null
vllmSpecDraftRevision object
Immutable commit supplied to vLLM for vllm_spec_draft_repo.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
vllmToolCallParser object
vLLM server-side tool-call parser name for this model
(--tool-call-parser, paired with --enable-auto-tool-choice).
Engine-specific platform knob, so it lives in runtime beside the
vllm_spec_* fields rather than in tooling (which stays model
truth). This explicit field is the ONLY source: one family string can
span tool-call generations with different wire formats, so there is no
family fallback, and an unset field launches the server without tool
support (tool requests are rejected loudly). Names follow vLLM's
parser registry (hermes, llama3_json, mistral, pythonic, deepseek_v3,
openai, ...); pin only pod-validated names.
- string
- null
vllmReasoningParser object
vLLM server-side reasoning parser name for this model
(--reasoning-parser).
Same doctrine as vllm_tool_call_parser: an engine-specific platform
knob, explicit only, no family fallback. Without it vLLM streams a
reasoning model's thinking inline with its answer (the server only splits
reasoning_content when a parser is configured), so a card whose model
always reasons (Muse Glimmer's to=self channel) must pin the parser
that vLLM registers for the family (muse_glimmer, qwen3,
deepseek_r1, openai_gptoss, ...). Pin only pod-validated names.
- string
- null
placement object
Where the model is allowed to run and which backend is preferred: the
compatible_backends hard filter and backend_preference soft score the
planner uses to route the model to suitable nodes.
Hard constraint: only route to nodes whose advertised backends intersect
this set. Making the implicit {"mlx"} explicit is what enables future
heterogeneous (llama_cpp / rocm / cuda) routing.
minVramGib object
Hard constraint: planner gates on node available memory when set.
- number
- null
maxContextTokens object
Soft: caps the placement-time KV budget check (see #145) when set.
- integer
- null
maxPipelineSplitLayer object
Largest layer boundary at which a pipeline rank may begin.
Some architectures end with layers that reuse KV produced by earlier
concrete layers. Keeping every split at or before this boundary ensures the
final rank owns those producers as well as their dependent tail. None
allows the planner to split at any ordinary layer boundary.
- integer
- null
Possible values: >= 1
Soft, ordered preference among the node's backend tags (e.g.
("llama_cpp-vulkan", "llama_cpp-rocm")).
Unlike compatible_backends (a hard filter on which nodes are eligible),
this only ranks eligible nodes/devices: the planner prefers a node that
advertises an earlier-listed tag, and the runner picks the earliest-listed
backend the chosen node actually has. The same model runs on any compatible
backend, but their performance differs per model, so this captures "fastest
on Vulkan, ROCm is an acceptable fallback" while still degrading gracefully
to a node that only offers the fallback. Order is significant and preserved;
an empty tuple means no preference (use the node's default).
registryCardId object
Immutable content-derived registry card id, or None for local cards.
- string
- null
Possible values: Value must match regular expression ^card_[a-z2-7]{52}$
registrySnapshotId object
Signed registry snapshot that supplied this runtime card.
- string
- null
registryProvenance object
Audited registry origin, kept separate from immutable artifact identity.
- string
- null
Possible values: [foxlight, agent, community]
registryArchitecture object
Trusted upstream architecture identity used for support-matrix joins.
- string
- null
registryArtifactFormat object
Exact signed artifact format used for support-matrix joins.
- string
- null
registryCapabilityClaims object[]
Open signed model/artifact capability claims, independent of engines.
Open namespaced intrinsic capability identifier.
Possible values: Value must match regular expression ^[a-z0-9][a-z0-9._:-]{0,199}$
Whether the claim describes the model or selected artifact.
Possible values: [model, artifact]
Evidence state without an engine-support implication.
Possible values: [claimed, observed, complete, incomplete, unknown]
Evidence channel that produced the claim.
Possible values: [upstream_structured, artifact_manifest, agent_analysis]
Source confidence from zero through one.
Possible values: >= 0 and <= 1
Bounded evidence references.
Possible values: <= 20
[]reviewer_model object
Agent reviewer identity, if any.
- string
- null
Possible values: <= 300 characters
Declared input modalities.
Possible values: <= 20
[]Declared output modalities.
Possible values: <= 20
[]details object
Open capability-specific evidence details.
Open capability-specific evidence details.
registryGgufMetadata object
Exact signed header evidence used for this runtime geometry projection.
It participates in the full-card authorization digest. A changed projection cannot silently reuse an approval for a different memory contract.
- RegistryGgufArtifactMetadata
- null
Possible values: >= 3 characters and <= 512 characters
Possible values: Value must match regular expression ^[0-9a-f]{40}$
Possible values: non-empty and <= 4096 characters
header objectrequired
Bounded facts read from one exact artifact's complete GGUF metadata area.
The digest covers the file prefix through the last metadata value, including the GGUF preamble. It is evidence identity, not a full-artifact checksum. Consumers own architecture-specific interpretation and engine allocation.
Possible values: non-empty and <= 128 characters
scalars objectrequired
Architecture-relative integer GGUF fields, without defaults.
Whether an explicit attention.recurrent_layers field exists.
Possible values: Value must match regular expression ^[0-9a-f]{64}$
Possible values: >= 24 and <= 16777216
resolvedBackend object
- string
- null
resolvedEngineBuild object
Exact music engine build observed by placement and required at sidecar launch.
- string
- null
llamaServerSettings object
Node serving settings captured by placement for memory admission.
- LlamaServerSettings
- null
Operator-selected concurrent slots before model-specific limits.
Possible values: > 0
16Whether the node permits the model card's speculative mode.
truefalseshouldTimeout object
- number
- null
Possible values: >= 0
Possible values: >= 0
Possible values: >= 0
nodeToRunner objectrequired
contextTokenLimit object
- integer
- null
The caller's per-placement node exclusions, stamped at placement time: repair re-placements (memory refusal, download failure) reconstruct their intent from the instance and keep honoring these instead of widening eligibility back to the full topology. Sorted for deterministic replicated events; empty for instances replayed from older event logs.
systemRole object
Marks a fabric-maintained system placement. "steward" is the intelligent-fabric resident: the master re-establishes exactly one such placement while intelligent-fabric mode is enabled, the dashboard hides it from ordinary instance surfaces, and the API refuses ordinary deletion while the mode is on. None (the default, and the value on every instance replayed from older event logs) means a normal user placement. Repair re-placements re-stamp the role so it survives node loss.
- string
- null
stewardrequestedContextTokens object
The caller's requested context window at placement, if any. Repair re-placements carry it forward so a replacement keeps the window the operator chose rather than falling back to the fleet default.
- integer
- null
Possible values: >= 256 and <= 1048576
hostsByNode objectrequired
property name* object[]
memory_delta_by_node object
- object
- null
max_context_tokens object
Largest context window this placement can hold (memory fit, card maximum, engine caps); a context_tokens request above it is refused. Null when no instance is present or no ceiling applies.
- integer
- null
default_context_tokens object
Window this placement gets when context_tokens is omitted: the fleet default for engines that reserve at load, otherwise the maximum.
- integer
- null
Whether the engine reserves the whole window's KV memory when the model loads (llama-server, in-process llama.cpp, vLLM), so the chosen window costs memory whether or not requests use it.
falsekv_bytes_per_token object
Estimated KV-cache bytes one token of window costs across the placement, for showing what a window reserves. Null when the card's attention geometry is unknown.
- integer
- null
error object
- string
- null
error_code object
Stable placement failure category, or null for a launchable preview. model_code_approval_required is retained only for older nodes; current authorization policy does not emit it. Backend and hardware identifiers remain open strings elsewhere.
- string
- null
Possible values: [no_valid_placement, placement_info_pending, model_code_approval_required, model_card_identity_mismatch]
trust_requirement objectdeprecated
Deprecated compatibility detail from the retired secondary model-approval ceremony; current previews return null.
- string
- null
compatibility_source object
Truth source that admitted the selected backend, or null on error.
- string
- null
Possible values: [card, signed_engine_support]
Active signed support claims applicable to this placement.
compatibility_detail object
Operator-readable model, artifact, engine/build, or platform gap.
- string
- null
True for a per-host alternative to the planner's ranked pick: a single-node placement on a host that passes admission but lost the ranking. On heterogeneous fleets the ranked winner (often the largest GPU) would otherwise hide every other valid host.
false{
"card_digest": "string",
"model_id": "string",
"sharding": "Tensor",
"instance_meta": "MlxRing",
"instance": {
"instanceId": "string",
"shardAssignments": {
"modelId": "string",
"runnerToShard": {},
"nodeToRunner": {}
},
"contextTokenLimit": 0,
"excludedNodes": [
"string"
],
"systemRole": "steward",
"requestedContextTokens": 0,
"hostsByNode": {},
"ephemeralPort": 0
},
"memory_delta_by_node": {},
"max_context_tokens": 0,
"default_context_tokens": 0,
"reserves_context_at_load": false,
"kv_bytes_per_token": 0,
"error": "string",
"error_code": "no_valid_placement",
"compatibility_source": "card",
"support_claim_ids": [
"string"
],
"compatibility_detail": "string",
"alternative": false
}