KV Cache Backends
Skulk includes several opt-in KV cache backends for MLX text generation. These backends are intended for long-context and memory-pressure experiments, while preserving existing behavior unless explicitly enabled.
Choose a backend in Settings
Open Settings → Inference → KV Cache Backend, choose a backend and select Save changes. The change takes effect on the next model launch; it does not replace the cache of a running model. This is the normal path for packaged-app users. Start with Default and validate another backend against your workload.
An explicit SKULK_KV_CACHE_BACKEND launch override disables the selector; the
panel explains that override. Advanced bit widths and retained edge-layer counts
are not exposed in Settings. MLX Quantized requires the
SKULK_KV_CACHE_BITS launch setting. The environment examples below are for
headless/source deployments and advanced launch configuration, not prerequisites
for using the dashboard selector.
Current Status
default: existing behavior (no cache quantization)mlx_quantized: MLX LM built-inQuantizedKVCacheturboquant: correctness-first TurboQuant-inspired KV cache for standardKVCachelayersturboquant_adaptive: keeps outer KV layers in FP16 and applies TurboQuant to middle KV layersoptiq: rotation-based KV cache via mlx-optiq; uses randomized orthogonal rotations with Lloyd-Max quantization and rotated-space attention for compatible attention layouts
The effective backend follows the dashboard/configuration choice unless a
SKULK_KV_CACHE_BACKEND environment override is set. The default configuration
uses default.
Advanced launch configuration examples
mlx-optiq
SKULK_KV_CACHE_BACKEND=optiq \
SKULK_OPTIQ_BITS=4 \
SKULK_OPTIQ_FP16_LAYERS=4 \
uv run skulk
The optiq backend uses mlx-optiq's rotation-based vector quantization, which eliminates per-key rotation overhead at inference time via rotated-space attention. It keeps the first and last N KV layers in FP16 for adaptive quality.
TurboQuant Adaptive
SKULK_KV_CACHE_BACKEND=turboquant_adaptive \
SKULK_TQ_K_BITS=3 \
SKULK_TQ_V_BITS=4 \
SKULK_TQ_FP16_LAYERS=4 \
uv run skulk
This mode keeps the first and last 4 KV layers in normal FP16-style cache and applies TurboQuant only to the middle KV layers. Validate output quality and memory use with the exact model and context length you intend to serve.
Available Environment Variables
| Variable | Backends | Default | Description |
|---|---|---|---|
SKULK_KV_CACHE_BACKEND | all | default | Backend selection |
SKULK_KV_CACHE_BITS | mlx_quantized | (required) | Bit width for MLX quantized cache |
SKULK_OPTIQ_BITS | optiq | 4 | Bit width for mlx-optiq cache |
SKULK_OPTIQ_FP16_LAYERS | optiq | 4 | Edge layers kept in FP16 |
SKULK_TQ_K_BITS | turboquant, turboquant_adaptive | 3 | Key quantization bits |
SKULK_TQ_V_BITS | turboquant, turboquant_adaptive | 4 | Value quantization bits |
SKULK_TQ_FP16_LAYERS | turboquant_adaptive | 4 | Edge layers kept in FP16 |
Invocation Examples
Default behavior:
SKULK_KV_CACHE_BACKEND=default uv run skulk
mlx-optiq (rotation-based):
SKULK_KV_CACHE_BACKEND=optiq SKULK_OPTIQ_BITS=4 SKULK_OPTIQ_FP16_LAYERS=4 uv run skulk
MLX quantized KV cache:
SKULK_KV_CACHE_BACKEND=mlx_quantized SKULK_KV_CACHE_BITS=4 uv run skulk
TurboQuant adaptive:
SKULK_KV_CACHE_BACKEND=turboquant_adaptive SKULK_TQ_K_BITS=3 SKULK_TQ_V_BITS=4 SKULK_TQ_FP16_LAYERS=4 uv run skulk
Practical Expectations
Quantization can reduce the memory used by standard KV layers, with additional
compute and possible output-quality changes. Bit width, retained edge layers,
attention layout and context length determine the trade-off; there is no universal
quality or speed ranking. Compare a representative workload against default
before choosing a setting. Model weights and unchanged recurrent or rotating
caches are not compressed by these switches.
Supported Cache Layouts
All quantized backends (optiq, turboquant, mlx_quantized) compress only standard KVCache entries and preserve these cache types unchanged:
ArraysCacheRotatingKVCache
Mixed cache layouts are supported:
KVCache+ArraysCacheKVCache+RotatingKVCacheKVCache+ArraysCache+RotatingKVCache
Current Limitations
- All quantized KV cache backends force sequential generation (no batch/history mode)
- Optiq checks the observed attention geometry before patching: non-power-of-two head dimensions or detected grouped-query attention (different query/KV head counts) log a warning and fall back to the default cache. Check logs to confirm whether quantization actually engaged.
- The optiq backend requires
mlx-optiqto be installed (pip install mlx-optiq) - The optiq backend's
patch_attention()monkey-patches MLX's SDPA, so avoid switching between optiq and other backends within the same process lifetime without a restart
Implementation reference
The accepted backend names and defaults live in
src/skulk/worker/engines/mlx/constants.py. Cache conversion, compatibility
checks and Optiq attention patching live in
src/skulk/worker/engines/mlx/cache.py. These settings affect MLX runners;
they do not configure llama.cpp, vLLM, speech or video engines.