PlaceInstanceParams
Possible values: [Tensor, Pipeline]
PipelinePossible values: [MlxRing, MlxJaccl, LlamaRpc]
MlxRing1Optional. Node IDs the master should treat as if absent when scoring candidate cycles for this placement. Empty list = consider all nodes. Already-running instances on the listed nodes are not affected — exclusion is per-placement, not cluster-wide.
context_tokens object
Optional context window for this placement, in tokens. The placer honors it up to the largest window the chosen nodes hold (max_context_tokens in the placement preview) and refuses a larger request with 400. When omitted, llama-server, in-process llama.cpp and vLLM placements take the fleet default (inference.served_context_tokens, 32768 unless changed) because they reserve the whole window's memory at load; MLX keeps the full memory fit.
- integer
- null
Possible values: >= 256 and <= 1048576
{
"excluded_nodes": [],
"instance_meta": "MlxRing",
"min_nodes": 1,
"model_id": "mlx-community/Llama-3.2-1B-Instruct-4bit",
"sharding": "Pipeline"
}