Skip to main content

Try LTX 2.3 Quality: Audio to Video in the Workbench

Run this model interactively, tune parameters, and compare outputs.
Model ID: ltx-2-3-quality-audio-to-video LTX 2.3 Quality (Audio to Video) is the high-quality preset of Lightricks LTX-2.3 on fal, generating video driven by an input audio track, a text prompt, and an optional starting image. It runs a distilled DiT workflow with a quality preset control, synchronizing motion such as lip movement and gesture to the supplied audio. When match audio length is enabled, the number of frames is derived from the audio duration and frame rate; otherwise a fixed frame count is used. An optional first-frame image can be conditioned with an adjustable strength, and the workflow can run from text and audio alone when no image is provided. It supports up to 481 frames at 1 to 60 FPS and is well suited for singing, talking-head, and performance clips. *Quantization is specific to the inference provider and the model may be offered with different quantization levels by other providers.

Example request

Use the Workbench as a request builder: configure parameters for this model in the UI, then open the API tab to copy the exact cURL or Python call.
This blocks until the video is ready (typically 5-15 minutes). Prefer Async or Async with SSE for anything beyond quick experimentation.See the video generation reference for more details.

Fetch model details

The models endpoint returns the full model object, including its json_request_schema.

Request parameters

Required parameters

Optional parameters