Transcribing audio
Turn a recording into text with xAI STT v1, with speaker labels and filler words when you need them.
What this video covers
In this part, we're transcribing a recording. Pick xAI STT v1 in the model picker.
This model doesn't take a prompt: the prompt box is disabled, and an Audio File slot opens under Reference.
You can upload a file from your computer, or right-click one in your library and choose Use as Reference. That's how we add the voice-over from the first part.
Once the file is in, the price appears; it's based on how long the file is.
The language is detected automatically, or you can pick it from the list yourself.
If more than one person is speaking, switch on Speaker labels. To keep fillers like "um" and "uh" in the text, switch on Keep filler words.
Press Generate. The transcript arrives in seconds, shown with its word count.
Press Copy, and the text goes to your clipboard.
The transcript is saved to your library, too.
To sum up: for voice-over, write the text; for music, describe the style; for a transcript, give it an audio file.
More from Going further
All tutorials
1:43Tutorial 14Editing your images
Edit an image with an editing model: put it in the reference slot, say what should change and what should stay, keep the aspect ratio and combine several images.
1:31Tutorial 15aYour first workflow
Build your first chain in Workflow: add nodes, pick models, wire an image into a video, check the price and run it.
1:50Tutorial 15bBatch generation in Workflow
Run one line with many inputs using Batch, Dice and Stack, and see how each one affects the price.
1:49Tutorial 15cWorkflow operators
Shape the text before it reaches the Generator with the Merge, Caption, Style and Pose operators.
1:49Tutorial 15dManaging workflow runs
The Preflight Check, running part of a graph, groups, cancelling, saving workflows and the panels on the right of the canvas.
1:24Tutorial 16aAI voice-over
Turn text into speech in LUVI: pick a voice-over model and a voice, set the stability and language, and generate.
1:26Tutorial 16bAI music
Make background music, or a full song with your own lyrics, in LUVI with MiniMax Music 2.6.
2:06Tutorial 17Talking avatars and lip sync
Make a portrait talk from an image and an audio file, or give an existing video a new voice with lip-sync models.
1:23Tutorial 18aGenerating 3D models
Turn an image or a description into a 3D model with Hunyuan 3D or HI3D, and see how extra views and parameters shape the result.
1:17Tutorial 18bThe 3D viewer and downloads
Inspect your 3D model in the viewer with lenses, lighting, wireframe and texture views, then download it as a GLB.