Tutorial 17Going further

Talking avatars and lip sync

Make a portrait talk from an image and an audio file, or give an existing video a new voice with lip-sync models.

00:00 / 02:06

What this video covers

In LUVI, you can make a portrait talk, or give an existing video a new voice. Both jobs are handled by lip-sync models.

Search the model list for "lipsync". Models that take an image and audio make a talking avatar; models that take a video match its lips to a new voice.

Let's start with a talking avatar. Here, we're using OmniHuman one point five.

Two slots open in the Reference section: one for the image and one for the audio file. Both are required.

For the image, use a portrait where the face is clearly visible. Just right-click it in your library and choose Use as Reference.

No audio file yet? You can make one in LUVI with a text-to-speech model, which is how we made this line. The audio goes into its own slot the same way.

The video runs as long as the audio, and OmniHuman takes up to sixty seconds of it. The price is worked out per second of audio and appears once both slots are filled.

In the prompt, describe the emotion, the pace and the delivery. We asked for a warm tone, a smile, and eyes on the camera.

Press Generate. The result lands in your library as a video that speaks with the audio you gave it.

Now let's give the same video a new voice, say a Turkish one. For that, we pick Sync Lipsync; VEED Lipsync is a cheaper option.

Here, the slots are a video and an audio file, and there's no prompt. Put the video and the new audio in; each can be up to sixty seconds long.

Here too, the price goes by the seconds of audio. Press Generate, and the lips are re-synced to the new audio.

In short: an image and audio for a talking avatar, a video and audio for a new voice, and the price follows the length of the audio.

Try it yourself

Create a free LUVI account and follow along with the video.

Get started

More from Going further

All tutorials