Lightricks audio-first lip sync model

LTX 2.3 LipSync starts from the voice

Use it when the audio performance is the anchor and the visual should follow with crisp mouth shapes and controlled framing.

digital human From 10 credits for 5 seconds at 480p.

Avatar Generation Input

Generate an audio-driven talking head with natural lip sync, optional portrait guidance, and 480p to 1080p output.

From 10 credits for 5 seconds at 480p.

Enable
Enable

Your generated avatar will appear here.

Upload a portrait image, add voice audio, and click 'Generate Avatar' to start.

Examples

LTX 2.3 lip-sync example

Hover a card to preview, then open the player to watch with sound and review the full prompt.

Business news briefing

A precise newsread with numbers and fast consonants to stress lip-sync accuracy.

View prompt

Professional news studio, stable head-and-shoulders framing, crisp enunciation, minimal head movement, natural blinking.

Model context

Why is LTX 2.3 different from other avatar models?

LTX 2.3 LipSync is audio-first: the speech track is required, while the portrait can be optional. That makes it useful when timing, numbers, consonants, and mouth shapes are the core test.

It is a good fit for news-style reads, briefings, explainer clips, and any talking-head output where lip sync matters more than broad body animation.

Model strengths

Lip-sync work that fits LTX 2.3

Use it when the audio performance is the anchor and the visual should follow with crisp mouth shapes and controlled framing.

  • News, business, and explainer reads where pronunciation and timing need close review.
  • Audio-first talking heads where a default or optional portrait can support the speech.
  • Short clips that benefit from resolution choices and restrained head movement.

Prompt guide

Preparing an LTX 2.3 lip-sync clip

Finalize the audio first

Clean timing and pronunciation before generating. The mouth performance follows the track.

Use a stable portrait if supplied

A centered head-and-shoulders image with direct eye contact gives the model a clean face to animate.

Prompt restraint

For accurate lip sync, ask for minimal head movement, natural blinking, and stable framing.

Check hard syllables

Review numbers, plosives, and fast consonants because they expose lip-sync errors quickly.

Frequently asked questions

Why is LTX 2.3 called audio-first?

Speech audio is the required input. A portrait can guide the face, but the voice timing drives the output.

What clips work best?

Short talking-head clips with clear speech, stable framing, and modest facial motion work best.

Can I use it without a portrait?

The model can run with audio as the required input, but supplying a clear portrait gives more control over the visible speaker.

When should I use Wan 2.2 Avatar instead?

Use Wan 2.2 Avatar when the clip needs longer portrait-and-audio avatar motion with more body presence.