Skip to main content

Speculative Decoding (MTP)

This guide explains Skulk's speculative decoding feature from an operator's point of view: what it does, which models use it, what speedups to expect, when it turns itself off, and how to confirm it is working.

The short version:

  • speculative decoding can accelerate supported model and hardware combinations
  • it is automatic (there is nothing to configure)
  • it pays off the most for dense models sharded across multiple nodes
  • it deliberately turns itself off in a few honest cases (see below)
  • you can verify it is active from the runner logs

What It Is​

Speculative decoding (we also call it MTP, for multi-token prediction) is a way to generate several tokens per model step instead of one. A small, cheap "drafter" proposes a short run of likely next tokens, and the full model verifies all of them in a single forward pass, keeping the longest correct prefix. When the drafter guesses well, you get multiple tokens for roughly the cost of one. Because the verify step accepts or rejects against the real model, the output quality is the model's own: greedy requests produce a valid greedy continuation, and sampled requests preserve the model's output distribution exactly. It is a pure speedup, not a quality trade-off. (On some Qwen models the greedy text can differ token-for-token from a non-speculative run while remaining equally greedy; see Known Limitations.)

Skulk runs speculative decoding on single-node, tensor-parallel, and pipeline (sharded) placements through one shared decode loop. You do not enable it, size it, or tune it for normal use: if a model ships with a drafter and the placement supports it, it activates on its own.

Engines: MLX, llama.cpp, and vLLM​

Skulk selects speculative execution from the model card and the available engine. These paths have different placement and companion requirements:

  • MLX (Apple Silicon, in-process): the drafters and models in the table below. Skulk owns the generation loop, the multi-node ring, and the speculative decode across sharded placements. This is the engine the speedup numbers and multi-node discussion on this page describe.
  • Served / llama_server (GPU nodes, including AMD): llama.cpp's native MTP, reached by launching llama-server --spec-type draft-mtp and proxying it (native MTP lives in the server app, not the in-process binding). This is how speculative decoding works on an AMD node. It is single-node and applies to GGUF models whose card declares served_spec_type = draft_mtp, in two shapes: baked-in MTP heads (Qwen3.5 / Qwen3.6 MTP GGUFs) and a base plus a separate --model-draft GGUF (Gemma 4 31B). See AMD Strix Halo nodes for enabling it (SKULK_LLAMA_SERVER_BIN) and the served MTP cards.

Everything below (the MLX drafter table, the multi-node speedups, and the turn-itself-off cases) is about the MLX engine. The served engine's speculation is configured on the model card (served_spec_type, served_spec_n_max) and, like MLX MTP, is carded off per model when a pairing does not pay. That is the per-model opt-out pattern across engines: a model whose measured drafter-and-target pairing nets negative carries the disable on its own card, rather than a global switch.

The vLLM engine runs its own speculative decoding for checkpoints that ship native multi-token-prediction heads (Qwen3.6 among them): the card's vllm_spec_method = "mtp" and vllm_spec_num_tokens fields map to vLLM's --speculative-config, and vLLM resolves the matching drafter architecture from the checkpoint itself, with no separate draft model. Measured on an A100-80GB, this roughly doubles single-stream decode on Qwen3.6-27B-FP8 (about 2x at depth 2, 70-83% draft acceptance). The same per-model carding rule applies: a model whose measured pairing does not pay simply omits the fields. Models without native heads placed on vLLM decode plain; that engine's headline win remains aggregate throughput under concurrent load (continuous batching), so the two mechanisms address different bottlenecks and now both exist on the served path.

The same card fields absorb vendor speculation schemes that use a separate published speculator instead of in-checkpoint heads. Poolside's Laguna models ship a block-parallel DFlash drafter as its own Hugging Face repo; the card declares vllm_spec_method = "dflash" with vllm_spec_draft_repo naming that drafter (mapped to the speculative-config model key), and vLLM (0.25.1 or later, which serves both the Laguna architecture and its DFlash drafter natively) fetches and runs it. No vendor fork or engine-specific code is involved: a new scheme is a new card declaration. Block-parallel drafters propose a whole block per step, so their carded depths run much deeper than MTP's (the Laguna XS 2.1 FP8 card uses the vendor-recommended 15 for a 16-token block).

One served-engine degradation behavior worth knowing: a card-declared separate draft GGUF is a best-effort companion at download time, so if the draft file is missing on disk when the model loads, the served runner logs a warning and serves the model without speculation rather than failing the placement. If a served MTP model decodes at plain speed, check the runner log for that warning before suspecting the model.

Served-engine speedups (AMD / llama.cpp)​

Measured on an AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151) serving GGUF models through the Vulkan llama-server, native MTP on (--spec-type draft-mtp) versus off. Both arms run through Skulk's production API with the same protocol as the MLX table: greedy decoding, 200-token completions, median of 3 runs, with throughput in decode tokens per second; the off arm is the identical GGUF served in plain decode (the node's SKULK_LLAMA_SERVER_FORCE_NO_SPEC benchmarking knob), so the gain is attributable to speculation alone.

ModelClassPlainWith MTPGain
Qwen3.5-9B-MTPdense, small55.676.2+37%
Qwen3.6-27B-MTPdense, mid20.035.6+78%
Qwen3.6-35B-A3B-MTPMoE (A3B)90.795.8+6%
gemma-4-31B (+ draft)dense, draft-model17.425.2+45%

The shape mirrors the MLX results: the dense mid-size model gains the most (+78%), because its slower base decode gives speculation the most to amortize; the MoE model gains the least (+6%), because its small active-parameter count already makes decode memory-bound-fast, so the per-round draft and verify overhead nets little. The Gemma row uses the other MTP shape (a separate --model-draft GGUF rather than baked-in heads) and still pays (+45%), confirming both served MTP shapes work on the Radeon backend.

Which Models Ship With It​

These models carry a drafter in their model card and use speculative decoding automatically. "Sidecar" drafters are MTP heads trained alongside the model; "assistant" drafters are a small companion model that cross-attends the target's cache.

ModelDrafterTypeDepth
mlx-community/gemma-4-e2b-it-8bitmlx-community/gemma-4-E2B-it-assistant-bf16assistant2
mlx-community/gemma-4-e4b-it-8bitmlx-community/gemma-4-E4B-it-assistant-bf16assistant2
mlx-community/gemma-4-12B-it-4bitmlx-community/gemma-4-12B-it-assistant-bf16assistant2
mlx-community/gemma-4-31b-it-4bitmlx-community/gemma-4-31B-it-assistant-bf16assistant2
mlx-community/gemma-4-26b-a4b-it-4bitmlx-community/gemma-4-26B-A4B-it-assistant-bf16assistant1 (single-node only)
mlx-community/Qwen3.5-9B-MLX-4bitFoxlightAI/qwen3-5-9b-base-mtpsidecar1
mlx-community/Qwen3.5-27B-4bitFoxlightAI/qwen3-5-27b-mtpsidecar1
mlx-community/Qwen3.6-27B-4bitFoxlightAI/qwen3-6-27b-mtpsidecar1
mlx-community/Qwen3.5-2B-4bitFoxlightAI/qwen3-5-2b-base-mtpsidecar1

The drafter weights are companion repos. Skulk fetches and stages them alongside the target model, so you do not download or reference them directly. Skulk Weights Publisher extracts some native heads into sidecars; model authors publish other assistant models directly. Signed cards pin external companions to immutable revisions. A sidecar or assistant is a dependency of its base model, not a separate model to place or chat with. The signed registry can update these bindings without a Skulk software release; consult the selected card for its exact contract.

What Speedups To Expect​

The numbers below are measurements on base Apple M4 hardware. They describe these model artifacts, placements, and request settings, rather than guaranteed throughput on another machine. Memory bandwidth, context, workload, engine version, and interconnect can change both absolute throughput and the benefit of speculation.

Protocol: production API, greedy decoding, 200-token completions, median of 3 runs per arm on the same live instance.

ConfigurationHardwarePlainWith MTPGain
gemma-4-E2B-8bit, single nodeM4 24GB37.754.0+43%
gemma-4-E4B-8bit, single nodeM4 24GB19.525.4+30%
Qwen3.5-9B-MLX-4bit, single nodeM4 24GB21.328.8+35%
gemma-4-12B-4bit, 2-node pipeline2× M4 16GB8.415.1+81%
gemma-4-31B-4bit dense, 2-node pipeline2× M4 16GB5.37.35+38%
Qwen3.5-27B-4bit dense, 2-node pipeline2× M4 16GB6.310.5+67%
Qwen3.5-9B-MLX-4bit, 2-node tensor-parallel2× M4 16GB16.721.8+31%

These ratios hold up under longer generations and sampling. At 1000 tokens the 12B 2-node pipeline still measures +60% (8.3 → 13.3) and Qwen 9B single still +28% (21.4 → 27.4); at temperature 0.7 the 12B pipeline is +54% and Qwen 9B is +21%. The 200-token greedy table is not flattering the feature by much.

Where It Shines: Dense Models Sharded Across Nodes​

The biggest wins are dense models split across a pipeline (the +67% to +81% rows above). When a model is sharded, every decoded token has to cross the inter-node links, and that hop latency is what makes pipelined decode slow. A speculative round crosses those hops once regardless of how many tokens it verifies, so every accepted draft amortizes exactly the latency that sharding adds. This is also the favourable case in general: the bigger the target model relative to the nodes it runs on, the more speculation pays, which is precisely Skulk's cluster pitch of running a big model across several smaller machines.

A practical corollary: shard to the smallest node count that fits the model. Over-sharding costs MTP headroom because each verify round then pays an extra network traversal: the 31B drops from +38% on 2 nodes to +17% on 3 nodes. Skulk's placement already prefers the smallest cycle that fits, so the default does the right thing.

Where It Turns Itself Off (And Why)​

Speculative decoding is honest about when it does not help. In these cases Skulk falls back to plain decode rather than slowing you down:

  • Multi-node MoE placements. Sparse (mixture-of-experts) models like gemma-4-26b-a4b-it-4bit already decode fast when sharded, because sharding halves the active-parameter bandwidth bottleneck. At that point the per-round draft+verify overhead nets slightly negative: measured -7% (30.2 to 28.2 tok/s) on a 2-node pipeline. The card gates this with speculative_multi_node = false, so these models run plain decode when sharded but keep speculation on a single node, where the same model measures ~2.2x (16 to 35.1 tok/s).
  • Sampled requests at higher depth. Any request with temperature > 0 forces draft depth to 1. Acceptance under sampling still preserves the output distribution exactly, but deeper chains stop paying, so the loop caps depth automatically.
  • Requests with repetition penalties. A request that sets a repetition penalty disables speculation for that request. This only affects the individual request that asked for the penalty.

How To Verify It Is Active​

The simplest signal is the runner log. While a supported model generates, the drafting rank periodically emits an acceptance line:

MTP acceptance so far: 137/180 (76%)

A non-zero acceptance rate confirms that speculation is running. Acceptance alone does not prove a speedup: draft and verification work also cost time. The other signal is throughput: compare the runner's generated N tokens @ X tok/s figure for a supported model against the plain-decode numbers in the table above.

Tuning​

Draft depth (how many tokens the drafter proposes per round) is a per-model field on the model card (mtp_max_depth), set from direct measurement on each carded artifact. The shipped defaults are measured optima:

  • Gemma assistant cards use depth 2
  • Qwen sidecar cards use depth 1

You can override depth with a custom model card, but the defaults are not guesses: they are the measured best for each model. Deeper is not better. On this hardware, verifying up to 2 candidates per step is effectively free, but each additional candidate beyond that costs a meaningful fraction of a full forward pass (the "verify-width cliff"). Past depth 2 the extra width costs more than the declining odds of the deeper guesses being accepted pay back, so a larger depth can be measurably slower, not faster. Trust the shipped values unless you are running your own depth sweep on your own hardware.

Known Limitations​

  • Speculative decoding currently runs one generation at a time. Models with an active drafter use a sequential generator, so concurrent requests to that model queue and run strictly first-in-first-out. (Gemma 4 models use the sequential path regardless of speculation, so this applies to every model in the table above; models outside these constraints batch concurrent requests.) The queueing is correct and stable (a 4-way concurrent test completed cleanly in FIFO order with no failures and no interleaving) but it means throughput on these models does not currently scale with concurrent callers.
  • Non-streaming errors return a truncated body. A non-streaming request that fails part-way through generation terminates promptly but returns an empty or truncated body under a 200 status (the status line is already on the wire when the failure lands) rather than a clean error document. Treat an unparseable non-streaming body as a failure and retry, or use streaming requests, which surface first-class error events.
  • Greedy MTP output on some Qwen models is semantically greedy but not byte-identical to the same model decoding without speculation. This affects Qwen models built on a recurrent/state-space attention design, whose running state makes single-step and speculative decode take slightly different but equally valid greedy paths. The text is a valid greedy generation; it may differ token-for-token from the non-MTP path.