Skip to main content

Node doctor

skulk doctor audits a node's environment against the same facts snapshot that Skulk's capability pipeline uses: which GPUs the node can see, which inference engines are usable, whether declared configuration matches observed hardware, and whether storage has headroom. Every non-OK verdict states its consequence for serving and the exact remediation.

# Full audit
uv run skulk doctor

# Apply safe idempotent remediations first, then re-audit
uv run skulk doctor --fix

# Machine-readable output
uv run skulk doctor --json

Exit codes: 0 when everything is OK, 2 when only DEGRADED verdicts remain, 1 when any FAIL remains.

Verdicts:

  • OK: the contract holds.
  • DEGRADED: serving works, but below the hardware's capability or with reduced observability.
  • FAIL: serving is broken or misconfigured in a way that will visibly hurt.

The startup fast path runs the same detection automatically: every node logs its facts summary and capability conflicts at launch, and conflicts surface as nodeHealth reasons on GET /state and in the dashboard topology view, so a degraded node is loud even if nobody runs the doctor.

Checks​

Inference engine availability (engine-available)​

Verifies at least one inference engine is usable: in-process MLX on macOS, an importable llama-cpp-python build, a llama-server binary (SKULK_LLAMA_SERVER_BIN), or a vllm CLI (SKULK_VLLM_BIN). A node with none advertises no backends and can only participate as management. Supports --fix.

ComfyUI video engine (comfy-engine)​

When video models are enabled (SKULK_ENABLE_VIDEO_MODELS), verifies the served ComfyUI video engine is configured (SKULK_COMFY_BIN plus SKULK_COMFY_ROOT) or provisioned as the managed install under the engines directory. A Linux NVIDIA or AMD node without one is degraded: video cards never place there. Management nodes and nodes with video models disabled pass. Supports --fix.

Capability conflicts (capability-conflicts)​

Runs backend derivation over the node facts snapshot and surfaces every observation-vs-declaration conflict: a GPU that no engine would use (silent CPU serving), degraded NVIDIA detection (missing nvidia-ml-py or a driver mismatch), an engine binary override pointing at an unusable path, or a declared backend the observed hardware cannot support. Supports --fix.

Model storage (models-storage)​

Verifies the models directory exists, is writable, and has download headroom (warns under 10 GB free, fails at 2 GB or less). Supports --fix.

Installed model cards (installed-card-records)​

Verifies every complete model in the model directories and in the model store's canonical and staging directories carries its card record (.skulk/installed-card.json, or the detached record kept for a read-only model directory), the record that keeps a downloaded model servable without the network. Records are checked by file size, never hashed, so the check stays fast on large stores. A model downloaded before these records existed gets one when Skulk starts with network access and recognizes it. Incomplete downloads are counted, not flagged.

Dashboard assets (dashboard-assets)​

Reports whether the built web dashboard is present. The API serves without it; headless workers are expected to run this way.

Hugging Face token (hf-token)​

Reports whether this node can authenticate to Hugging Face, and whether it is the node that needs to. A token entered in any node's dashboard Settings propagates over the encrypted cluster fabric to every node, and joining nodes adopt it at bootstrap, so one entry covers the fleet; this check verifies it actually arrived on the node that performs downloads (the model store host when a store is configured, otherwise this node itself). Without one, public models still download and only gated or private repositories fail.

vLLM build prerequisites (vllm-prerequisites)​

When a vLLM engine is configured, verifies the node can actually compile its kernels. vLLM JITs Triton and torch.compile kernels at runtime, shelling out to a C++ compiler (Inductor drives g++, so gcc alone is not enough) against the Python development headers; neither is a dependency of the vLLM wheel. Without them the node advertises vLLM capacity and accepts placements, then fails every engine start with an InductorError.