Skip to main content

KV Cache Backends

Skulk includes several opt-in KV cache backends for MLX text generation. These backends are intended for long-context and memory-pressure experiments, while preserving existing behavior unless explicitly enabled.

Choose a backend in Settings​

Open Settings → Inference → KV Cache Backend, choose a backend and select Save changes. The change takes effect on the next model launch; it does not replace the cache of a running model. This is the normal path for packaged-app users. Start with Default and validate another backend against your workload.

An explicit SKULK_KV_CACHE_BACKEND launch override disables the selector; the panel explains that override. Advanced bit widths and retained edge-layer counts are not exposed in Settings. MLX Quantized requires the SKULK_KV_CACHE_BITS launch setting. The environment examples below are for headless/source deployments and advanced launch configuration, not prerequisites for using the dashboard selector.

Current Status​

  • default: existing behavior (no cache quantization)
  • mlx_quantized: MLX LM built-in QuantizedKVCache
  • turboquant: correctness-first TurboQuant-inspired KV cache for standard KVCache layers
  • turboquant_adaptive: keeps outer KV layers in FP16 and applies TurboQuant to middle KV layers
  • optiq: rotation-based KV cache via mlx-optiq; uses randomized orthogonal rotations with Lloyd-Max quantization and rotated-space attention for compatible attention layouts

The effective backend follows the dashboard/configuration choice unless a SKULK_KV_CACHE_BACKEND environment override is set. The default configuration uses default.

Advanced launch configuration examples​

mlx-optiq​

SKULK_KV_CACHE_BACKEND=optiq \
SKULK_OPTIQ_BITS=4 \
SKULK_OPTIQ_FP16_LAYERS=4 \
uv run skulk

The optiq backend uses mlx-optiq's rotation-based vector quantization, which eliminates per-key rotation overhead at inference time via rotated-space attention. It keeps the first and last N KV layers in FP16 for adaptive quality.

TurboQuant Adaptive​

SKULK_KV_CACHE_BACKEND=turboquant_adaptive \
SKULK_TQ_K_BITS=3 \
SKULK_TQ_V_BITS=4 \
SKULK_TQ_FP16_LAYERS=4 \
uv run skulk

This mode keeps the first and last 4 KV layers in normal FP16-style cache and applies TurboQuant only to the middle KV layers. Validate output quality and memory use with the exact model and context length you intend to serve.

Available Environment Variables​

VariableBackendsDefaultDescription
SKULK_KV_CACHE_BACKENDalldefaultBackend selection
SKULK_KV_CACHE_BITSmlx_quantized(required)Bit width for MLX quantized cache
SKULK_OPTIQ_BITSoptiq4Bit width for mlx-optiq cache
SKULK_OPTIQ_FP16_LAYERSoptiq4Edge layers kept in FP16
SKULK_TQ_K_BITSturboquant, turboquant_adaptive3Key quantization bits
SKULK_TQ_V_BITSturboquant, turboquant_adaptive4Value quantization bits
SKULK_TQ_FP16_LAYERSturboquant_adaptive4Edge layers kept in FP16

Invocation Examples​

Default behavior:

SKULK_KV_CACHE_BACKEND=default uv run skulk

mlx-optiq (rotation-based):

SKULK_KV_CACHE_BACKEND=optiq SKULK_OPTIQ_BITS=4 SKULK_OPTIQ_FP16_LAYERS=4 uv run skulk

MLX quantized KV cache:

SKULK_KV_CACHE_BACKEND=mlx_quantized SKULK_KV_CACHE_BITS=4 uv run skulk

TurboQuant adaptive:

SKULK_KV_CACHE_BACKEND=turboquant_adaptive SKULK_TQ_K_BITS=3 SKULK_TQ_V_BITS=4 SKULK_TQ_FP16_LAYERS=4 uv run skulk

Practical Expectations​

Quantization can reduce the memory used by standard KV layers, with additional compute and possible output-quality changes. Bit width, retained edge layers, attention layout and context length determine the trade-off; there is no universal quality or speed ranking. Compare a representative workload against default before choosing a setting. Model weights and unchanged recurrent or rotating caches are not compressed by these switches.

Supported Cache Layouts​

All quantized backends (optiq, turboquant, mlx_quantized) compress only standard KVCache entries and preserve these cache types unchanged:

  • ArraysCache
  • RotatingKVCache

Mixed cache layouts are supported:

  • KVCache + ArraysCache
  • KVCache + RotatingKVCache
  • KVCache + ArraysCache + RotatingKVCache

Current Limitations​

  • All quantized KV cache backends force sequential generation (no batch/history mode)
  • Optiq checks the observed attention geometry before patching: non-power-of-two head dimensions or detected grouped-query attention (different query/KV head counts) log a warning and fall back to the default cache. Check logs to confirm whether quantization actually engaged.
  • The optiq backend requires mlx-optiq to be installed (pip install mlx-optiq)
  • The optiq backend's patch_attention() monkey-patches MLX's SDPA, so avoid switching between optiq and other backends within the same process lifetime without a restart

Implementation reference​

The accepted backend names and defaults live in src/skulk/worker/engines/mlx/constants.py. Cache conversion, compatibility checks and Optiq attention patching live in src/skulk/worker/engines/mlx/cache.py. These settings affect MLX runners; they do not configure llama.cpp, vLLM, speech or video engines.