Grok Speech·par xAISpeech to text

xAI STT v1

xAI STT v1 — transcribes an audio file to text. 24+ languages with automatic language detection, speaker diarization and filler-word control. The result is a downloadable text file.

Paramètres

ParamètreTypePar défautPlage ou options
Raw audio format
audio_format

Format hint for raw/headerless audio. Container formats (mp3, wav, …) are auto-detected — leave on 'auto'.

selectautoauto, pcm, mulaw, alaw
Channels
channels

Number of audio channels (2–8). Only required for multichannel raw audio; auto-detected for container formats.

slider22 – 8
Speaker labels
diarize

Speaker diarization — each word carries the detected speaker id.

selectfalsetrue, false
Keep filler words
filler_words

When true, filler words ("um", "uh") are kept; when false they are removed.

selectfalsetrue, false
Language
language

ISO 639-1 language code (Filipino uses 'fil'). The model transcribes any supported language regardless; leave empty to auto-detect.

select, ar, cs, da, nl, en, fil, fr, de, hi, id, it, ja, ko, mk, ms, fa, pl, pt, ro, ru, es, sv, th, tr, vi
Per-channel transcript
multichannel

Transcribe each audio channel independently.

selectfalsetrue, false
Sample rate
sample_rate

Sample rate in Hz. Only required for raw audio (pcm, mulaw, alaw).

select80008000, 16000, 22050, 24000, 44100, 48000
Number normalization
text_normalization

Converts spoken-form numbers and currency into written form (e.g. "one hundred dollars" to "$100").

selectfalsetrue, false

Entrées

Fichiers que ce modèle accepte en plus du prompt.

EntréeAccepteFichiers max.
Audio File
audio
audio1Obligatoire

Autres modèles de cette famille