Skip to main content

VideoCardConfig

Declarative audio-video generation contract for a model card.

The section states model truth: which modes exist, the duration and frame grid the model was trained on, canvas rules, audio output, reference bounds, sampling defaults, and pinned companion artifacts. Engines map these facts onto their own mechanisms; nothing here names an engine.

modesVideoMode (string)[]required

Generation modes this artifact serves; each implies a ModelTask.

Possible values: [t2va, fl2va, ref2va]

minSecondsMinseconds (integer)

Shortest output duration the model supports.

Possible values: > 0

Default value: 4
maxSecondsMaxseconds (integer)

Longest output duration the model supports.

Possible values: > 0

Default value: 15
fpsFps (integer)

Output frame rate.

Possible values: > 0

Default value: 24
frameGridMultipleFramegridmultiple (integer)

Frame counts must satisfy count % frame_grid_multiple == frame_grid_offset.

Possible values: > 0

Default value: 1
frameGridOffsetFramegridoffset (integer)

Residue a valid frame count leaves modulo frame_grid_multiple.

Default value: 0
canvasMultipleCanvasmultiple (integer)

Width and height must be multiples of this many pixels.

Possible values: > 0

Default value: 1
defaultShortEdge object

Trained short-edge resolution used when a request gives no size.

anyOf
integer

Possible values: > 0

maxPixels object

Largest width times height the model serves at native quality.

anyOf
integer

Possible values: > 0

aspectRatiosstring[]

Advertised aspect ratios as W:H strings; empty means unconstrained.

Default value: []
audioOutputAudiooutput (boolean)

Whether generated video carries a synchronized audio track.

Default value: false
audioSampleRate object

Sample rate of generated audio in hertz.

anyOf
integer

Possible values: > 0

audioChannels object

Channel count of generated audio.

anyOf
integer

Possible values: > 0

defaultStepsDefaultsteps (integer)

Sampling steps used when a request and its companions do not decide.

Possible values: > 0

Default value: 20
videoShift object

Trained video sigma shift, when the sampler exposes one.

anyOf
number
audioShift object

Trained audio sigma shift, when the sampler exposes one.

anyOf
number
referenceLimits object

Reference bounds; required when ref2va is among the modes.

anyOf
maxImagesMaximages (integer)

Maximum reference images per request.

Default value: 0
maxVideosMaxvideos (integer)

Maximum reference video clips per request.

Default value: 0
maxAudioClipsMaxaudioclips (integer)

Maximum reference audio clips per request.

Default value: 0
maxFiles object

Maximum files across every reference type; None means the sum.

anyOf
integer
clipMinSeconds object

Minimum duration of one reference video or audio clip.

anyOf
integer

Possible values: > 0

clipMaxSeconds object

Maximum duration of one reference video or audio clip.

anyOf
integer

Possible values: > 0

totalClipSeconds object

Maximum combined duration of reference clips of one type.

anyOf
integer

Possible values: > 0

companions object[]

Pinned adapters, patches, embeddings, and graph templates.

  • Array [
  • kindVideoCompanionKind (string)required

    What the companion is; selects how an engine applies it.

    Possible values: [lora, model_patch, embedding, graph_template, preprocessor]

    nameName (string)required

    Stable companion identifier callers and engines refer to.

    pathPath (string)required

    Canonical repository-relative POSIX path of the companion file.

    repo object

    Repository hosting the companion; None means the card's artifact repository, whose source_revision then also pins this file.

    anyOf
    string
    revision object

    Immutable commit of repo; required whenever repo is external.

    anyOf
    string

    Possible values: Value must match regular expression ^[0-9a-f]{40}$

    sizeBytes object

    Exact upstream byte size at the pinned revision when known.

    anyOf
    integer
    modesVideoMode (string)[]

    Generation modes the companion applies to; empty means every mode.

    Possible values: [t2va, fl2va, ref2va]

    Default value: []
    steps object

    For distillation adapters, the sampling step count they were trained for.

    anyOf
    integer

    Possible values: > 0

    strength object

    Default application strength when the engine supports one.

    anyOf
    number
    videoShift object

    Video sigma shift the companion expects, when it differs from the card.

    anyOf
    number
    audioShift object

    Audio sigma shift the companion expects, when it differs from the card.

    anyOf
    number
    role object

    For preprocessor companions, what the weights do; required for them and refused on every other kind.

    anyOf
    VideoPreprocessorRole (string)

    What a preprocessor companion's weights do in deriving a guide video.

    Possible values: [pose_estimator, person_detector, depth_estimator]

    license object

    The license its hosting repository declares, as a lowercase SPDX-style identifier (mit, apache-2.0), shown to the operator beside the card's own license; companions in the card's repository fall under that.

    anyOf
    string
  • ]
  • VideoCardConfig
    {
    "modes": [
    "t2va"
    ],
    "minSeconds": 4,
    "maxSeconds": 15,
    "fps": 24,
    "frameGridMultiple": 1,
    "frameGridOffset": 0,
    "canvasMultiple": 1,
    "defaultShortEdge": 0,
    "maxPixels": 0,
    "aspectRatios": [
    "string"
    ],
    "audioOutput": false,
    "audioSampleRate": 0,
    "audioChannels": 0,
    "defaultSteps": 20,
    "videoShift": 0,
    "audioShift": 0,
    "referenceLimits": {
    "maxImages": 0,
    "maxVideos": 0,
    "maxAudioClips": 0,
    "maxFiles": 0,
    "clipMinSeconds": 0,
    "clipMaxSeconds": 0,
    "totalClipSeconds": 0
    },
    "companions": [
    {
    "kind": "lora",
    "name": "string",
    "path": "string",
    "repo": "string",
    "revision": "string",
    "sizeBytes": 0,
    "modes": [
    "t2va"
    ],
    "steps": 0,
    "strength": 0,
    "videoShift": 0,
    "audioShift": 0,
    "role": "pose_estimator",
    "license": "string"
    }
    ]
    }