xAI released Grok Voice Transcribe 2.0, its latest speech-to-text model.
Today we’re releasing Grok Voice Transcribe 2.0, our latest speech-to-text model. Across our real-world evaluations, Grok Voice Transcribe 2.0 is one of the most accurate transcription models available today and twice as accurate as Grok Voice Transcribe 1.0, at the same price.
What was announced
- Grok Voice Transcribe 2.0 is available now via the Speech-to-Text API.
- xAI states it ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
- Internal evaluations show lower word error rates than version 1.0 on telephony, conversational, credentials, and short multilingual phrases.
- Features include batch and streaming, word-level timestamps, speaker diarization, multichannel support, key-term biasing, text formatting, filler-word removal, and smart turn detection.
- Pricing is unchanged: $0.10 per hour batch and $0.20 per hour streaming.
- Version 2.0 will become the default; 1.0 will be deprecated in the coming weeks (pin
grok-voice-transcribe-1.0to stay on it). - Atlassian Loom is cited as using the model for video transcription.
Context
Grok Voice Transcribe is the standalone speech-to-text API built on the same audio foundation that powers Grok Voice, Tesla vehicle assistants, and other production voice systems. The 2.0 release focuses on real-world noisy, multilingual, and telephony audio rather than clean studio speech.
Limits of this report
This article reports only the claims and figures published on the x.ai news page. It does not independently verify leaderboard rankings, word-error-rate numbers, or production deployment details beyond what the official post states.