StepFun's StepAudio 3 ASR takes the top non-streaming speech-to-text spot with 1.7% WER

As seen on the 24/7 Wall St. homepage on September 22, 2026.

A Chinese startup just took the top spot in non-streaming speech-to-text, cutting its own word-error rate from 4.7% to 1.7% in one release. Transcription pricing power is getting harder to defend for US incumbents.

StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming Speech to Text with 1.7% WER, a notable improvement on StepAudio 2.5 ASR (4.7%) StepAudio 3 ASR is StepFun's new Speech to Text model, available through the StepFun API for non-streaming https://t.co/hF17aMAFo7
  • Replies6
  • Reposts1
  • Likes50
Continue ReadingShow less

StepFun, a Chinese AI startup, released StepAudio 3 ASR on September 22 and it immediately claimed the number one position on the AA-WER Index for non-streaming speech-to-text. The benchmark that matters here is word error rate, where lower is better, and StepAudio 3 ASR posted a 1.7% WER.

That is a sharp reduction from the 4.7% WER of predecessor StepAudio 2.5 ASR, meaning StepFun nearly tripled its accuracy in a single model generation.

Sponsored

Twelve Tabs, One Thesis

Your Research Resets Every Morning

The quote page in one tab. Filings in another. A chart you rebuilt from scratch, a transcript you never went back and found, a screener whose settings you will redo next week. Nothing you built yesterday is still there.

AlphaSpace replaces all of it with one screen you arrange yourself. Earnings calendar, estimate versus actual, the call transcript, live news, your own charts, every panel wired to whatever ticker you click. Close the browser and it is all still sitting there tomorrow.

See What a Built View Looks Like →

(Sponsor)

StepAudio 3 ASR is available now through the StepFun API for non-streaming use cases. Non-streaming transcription, where an audio file is processed in full rather than in real time, is the dominant workload for enterprise transcription pipelines, legal and medical documentation, and media captioning.

The result puts pressure on established speech-to-text providers. When a newer entrant claims the top accuracy ranking on a widely watched index, it becomes harder for incumbents to justify premium pricing purely on quality grounds. Accuracy parity or superiority from a lower-cost competitor tends to shift procurement conversations quickly in enterprise software.