Skip to main content

LlamaServerSettings

Node settings that change llama-server's persistent cache allocation.

parallelSlotsParallelslots (integer)

Operator-selected concurrent slots before model-specific limits.

Possible values: > 0

Default value: 16
speculationEnabledSpeculationenabled (boolean)

Whether the node permits the model card's speculative mode.

Default value: true
LlamaServerSettings
{
"parallelSlots": 16,
"speculationEnabled": true
}