Talking avatars and lip sync
Make a portrait talk from an image and an audio file, or give an existing video a new voice with lip-sync models.
What this video covers
In LUVI, you can make a portrait talk, or give an existing video a new voice. Both jobs are handled by lip-sync models.
Search the model list for "lipsync". Models that take an image and audio make a talking avatar; models that take a video match its lips to a new voice.
Let's start with a talking avatar. Here, we're using OmniHuman one point five.
Two slots open in the Reference section: one for the image and one for the audio file. Both are required.
For the image, use a portrait where the face is clearly visible. Just right-click it in your library and choose Use as Reference.
No audio file yet? You can make one in LUVI with a text-to-speech model, which is how we made this line. The audio goes into its own slot the same way.
The video runs as long as the audio, and OmniHuman takes up to sixty seconds of it. The price is worked out per second of audio and appears once both slots are filled.
In the prompt, describe the emotion, the pace and the delivery. We asked for a warm tone, a smile, and eyes on the camera.
Press Generate. The result lands in your library as a video that speaks with the audio you gave it.
Now let's give the same video a new voice, say a Turkish one. For that, we pick Sync Lipsync; VEED Lipsync is a cheaper option.
Here, the slots are a video and an audio file, and there's no prompt. Put the video and the new audio in; each can be up to sixty seconds long.
Here too, the price goes by the seconds of audio. Press Generate, and the lips are re-synced to the new audio.
In short: an image and audio for a talking avatar, a video and audio for a new voice, and the price follows the length of the audio.
More from Going further
All tutorials
1:43Tutorial 14Editing your images
Edit an image with an editing model: put it in the reference slot, say what should change and what should stay, keep the aspect ratio and combine several images.
1:31Tutorial 15aYour first workflow
Build your first chain in Workflow: add nodes, pick models, wire an image into a video, check the price and run it.
1:50Tutorial 15bBatch generation in Workflow
Run one line with many inputs using Batch, Dice and Stack, and see how each one affects the price.
1:49Tutorial 15cWorkflow operators
Shape the text before it reaches the Generator with the Merge, Caption, Style and Pose operators.
1:49Tutorial 15dManaging workflow runs
The Preflight Check, running part of a graph, groups, cancelling, saving workflows and the panels on the right of the canvas.
1:24Tutorial 16aAI voice-over
Turn text into speech in LUVI: pick a voice-over model and a voice, set the stability and language, and generate.
1:26Tutorial 16bAI music
Make background music, or a full song with your own lyrics, in LUVI with MiniMax Music 2.6.
1:06Tutorial 16cTranscribing audio
Turn a recording into text with xAI STT v1, with speaker labels and filler words when you need them.
1:23Tutorial 18aGenerating 3D models
Turn an image or a description into a 3D model with Hunyuan 3D or HI3D, and see how extra views and parameters shape the result.
1:17Tutorial 18bThe 3D viewer and downloads
Inspect your 3D model in the viewer with lenses, lighting, wireframe and texture views, then download it as a GLB.