FEATUREDHacker News Frontpage· rssEN18:32 · 09·17
→Wispr introduces Canto: a real-time speech model built for real-world dictation
Wispr released Canto, a real-time speech model that achieved the lowest word error rate on 10 hours of real-world dictations from over 2,300 speakers, beating models from Google, OpenAI, AssemblyAI, and Deepgram. On a 3-hour challenge set with noise, low volume, and short utterances, Canto led among real-time models but trailed Gemini 3.1 Pro, a large multimodal model unfit for low-latency use. Canto was pretrained on millions of hours of speech and text, then fine-tuned with supervised learning and GRPO reinforcement learning to optimize full-transcript quality. On public benchmarks, Canto tied for first on LibriSpeech and was competitive but not leading on FLEURS and Common Voice; the post notes those datasets consist mostly of read speech, which differs from spontaneous dictation.
#Wispr#Wispr Advanced Interfaces Lab#Google
why featured
Featured · importance 72 · hook + knowledge
editor take
Wispr's Canto beats Google and OpenAI on real-world dictation, but hold off on calling it a clean sweep.
sharp
The reason to click: the test set isn't the usual read speech. It's 10 hours of real dictation from over 2,300 people using various mics in their actual environments. Canto posted the lowest word error rate, ahead of Google, OpenAI, AssemblyAI, and Deepgram. On a 3-hour challenge set built to stress-test models, it led among real-time models but trailed Gemini 3.1 Pro overall—a large multimodal model that isn't built for low-latency use.
I'd discount this a bit. The eval is Wispr's own, built from its user base. They enforced speaker separation between train and test, but the app distribution and labeling standards aren't reproducible from the outside. On public benchmarks, Canto tied for first only on LibriSpeech; it didn't lead on FLEURS or Common Voice, which the team acknowledges are read-speech datasets that differ from spontaneous dictation.
The training recipe is worth noting: pretraining on millions of hours of speech and text, then supervised fine-tuning plus GRPO reinforcement learning to optimize full-transcript quality. But the post doesn't give model size, inference latency, or pricing—the details that determine whether this actually works in production. For now, this reads more like a technical credential for Wispr Flow than a general-purpose speech model ranking.
HKR breakdown
hook ✓knowledge ✓resonance —