Grok Voice STT 1.0
by ZenMux · listed Aug 2026
Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.
Current rates USD per 1M tokens
Input—
Output—
Cache read—
Cache write—
Against the market
On the record
- Model id
x-ai/grok-voice-stt-1.0- Context window
- 15K tokens
- Max output
- 15K tokens
- Input modalities
- audio
- Output modalities
- text
- Reasoning
- no
- Tool calls
- no
- Open weights
- no
- Knowledge cutoff
- —
- Released
- 2026-08-04
- Last updated
- 2026-08-04
- Provider docs
- docs.zenmux.ai