VideoCardConfig
Declarative audio-video generation contract for a model card.
The section states model truth: which modes exist, the duration and frame grid the model was trained on, canvas rules, audio output, reference bounds, sampling defaults, and pinned companion artifacts. Engines map these facts onto their own mechanisms; nothing here names an engine.
Generation modes this artifact serves; each implies a ModelTask.
Possible values: [t2va, fl2va, ref2va]
Shortest output duration the model supports.
Possible values: > 0
4Longest output duration the model supports.
Possible values: > 0
15Output frame rate.
Possible values: > 0
24Frame counts must satisfy count % frame_grid_multiple == frame_grid_offset.
Possible values: > 0
1Residue a valid frame count leaves modulo frame_grid_multiple.
0Width and height must be multiples of this many pixels.
Possible values: > 0
1defaultShortEdge object
Trained short-edge resolution used when a request gives no size.
- integer
- null
Possible values: > 0
maxPixels object
Largest width times height the model serves at native quality.
- integer
- null
Possible values: > 0
Advertised aspect ratios as W:H strings; empty means unconstrained.
[]Whether generated video carries a synchronized audio track.
falseaudioSampleRate object
Sample rate of generated audio in hertz.
- integer
- null
Possible values: > 0
audioChannels object
Channel count of generated audio.
- integer
- null
Possible values: > 0
Sampling steps used when a request and its companions do not decide.
Possible values: > 0
20videoShift object
Trained video sigma shift, when the sampler exposes one.
- number
- null
audioShift object
Trained audio sigma shift, when the sampler exposes one.
- number
- null
referenceLimits object
Reference bounds; required when ref2va is among the modes.
- VideoReferenceLimits
- null
Maximum reference images per request.
0Maximum reference video clips per request.
0Maximum reference audio clips per request.
0maxFiles object
Maximum files across every reference type; None means the sum.
- integer
- null
clipMinSeconds object
Minimum duration of one reference video or audio clip.
- integer
- null
Possible values: > 0
clipMaxSeconds object
Maximum duration of one reference video or audio clip.
- integer
- null
Possible values: > 0
totalClipSeconds object
Maximum combined duration of reference clips of one type.
- integer
- null
Possible values: > 0
companions object[]
Pinned adapters, patches, embeddings, and graph templates.
What the companion is; selects how an engine applies it.
Possible values: [lora, model_patch, embedding, graph_template, preprocessor]
Stable companion identifier callers and engines refer to.
Canonical repository-relative POSIX path of the companion file.
repo object
Repository hosting the companion; None means the card's artifact
repository, whose source_revision then also pins this file.
- string
- null
revision object
Immutable commit of repo; required whenever repo is external.
- string
- null
Possible values: Value must match regular expression ^[0-9a-f]{40}$
sizeBytes object
Exact upstream byte size at the pinned revision when known.
- integer
- null
Generation modes the companion applies to; empty means every mode.
Possible values: [t2va, fl2va, ref2va]
[]steps object
For distillation adapters, the sampling step count they were trained for.
- integer
- null
Possible values: > 0
strength object
Default application strength when the engine supports one.
- number
- null
videoShift object
Video sigma shift the companion expects, when it differs from the card.
- number
- null
audioShift object
Audio sigma shift the companion expects, when it differs from the card.
- number
- null
role object
For preprocessor companions, what the weights do; required for them and refused on every other kind.
- VideoPreprocessorRole
- null
What a preprocessor companion's weights do in deriving a guide video.
Possible values: [pose_estimator, person_detector, depth_estimator]
license object
The license its hosting repository declares, as a lowercase SPDX-style
identifier (mit, apache-2.0), shown to the operator beside the
card's own license; companions in the card's repository fall under that.
- string
- null
{
"modes": [
"t2va"
],
"minSeconds": 4,
"maxSeconds": 15,
"fps": 24,
"frameGridMultiple": 1,
"frameGridOffset": 0,
"canvasMultiple": 1,
"defaultShortEdge": 0,
"maxPixels": 0,
"aspectRatios": [
"string"
],
"audioOutput": false,
"audioSampleRate": 0,
"audioChannels": 0,
"defaultSteps": 20,
"videoShift": 0,
"audioShift": 0,
"referenceLimits": {
"maxImages": 0,
"maxVideos": 0,
"maxAudioClips": 0,
"maxFiles": 0,
"clipMinSeconds": 0,
"clipMaxSeconds": 0,
"totalClipSeconds": 0
},
"companions": [
{
"kind": "lora",
"name": "string",
"path": "string",
"repo": "string",
"revision": "string",
"sizeBytes": 0,
"modes": [
"t2va"
],
"steps": 0,
"strength": 0,
"videoShift": 0,
"audioShift": 0,
"role": "pose_estimator",
"license": "string"
}
]
}