LlamaServerSettings
Node settings that change llama-server's persistent cache allocation.
parallelSlotsParallelslots (integer)
Operator-selected concurrent slots before model-specific limits.
Possible values: > 0
Default value:
16speculationEnabledSpeculationenabled (boolean)
Whether the node permits the model card's speculative mode.
Default value:
trueLlamaServerSettings
{
"parallelSlots": 16,
"speculationEnabled": true
}