Google's Gemini 3.5 Transcribe enters the voice API market at 2.6% error rate
As seen on the 24/7 Wall St. homepage on August 26, 2026.
Alphabet is now selling speech-to-text at 0.40s latency after speech ends, pushing Gemini straight into the voice API market that Whisper and Deepgram have owned.
Google has released Gemini 3.5 Transcribe, ranking #5 on AA-WER at 2.6%, alongside Gemini 3.5 Transcribe Live, achieving 4.0% AA-WER Streaming at 0.40s after speech end Google’s Gemini 3.5 Transcribe release comprises two API offerings: Gemini 3.5 Transcribe Live for continuous https://t.co/y81oauk7Sj
- Replies1
- Reposts0
- Likes0
Continue ReadingShow less
Google launched two speech-to-text products under the Gemini 3.5 Transcribe name: a standard transcription model and Gemini 3.5 Transcribe Live, built for continuous, real-time use. Both are available as API offerings, meaning developers can start building against them immediately.
The standard model posted a 2.6% word error rate on Artificial Analysis's AA-WER benchmark, placing it fifth overall and putting Google in the conversation with the best available transcription options.
The Live variant responds just 0.40 seconds after speech ends, the figure that matters most for voice assistants, live captioning, and any application where a user is waiting on a response.
Alphabet is moving directly into territory that dedicated speech API providers have held. Bundling competitive transcription into the Gemini platform gives existing Google Cloud customers a reason to consolidate their voice workloads rather than route them to third-party services.
Mentioned: GOOGL