How do you choose speech-to-text software?
Speech-to-text software — dictation software — turns your speech into text and types it into whatever application you are using. What separates products: where the audio is processed (cloud or your device), what happens to your words afterwards (a rule-based clean-up or a generative rewrite), whether use is limited, which platforms are covered, and whether you eventually own it.
The five questions that decide the choice
One: where does the audio go? Cloud dictation uploads your voice; on-device dictation runs the model on your own hardware, so audio never leaves. Two: what happens to your words after recognition — a rule-based clean-up, or a generative rewrite? Three: are there limits — daily caps, per-minute metering?
Four: does it cover the platforms you actually use, or just one? Five: what is the pricing shape — free built-in, one-time purchase, subscription — and do you own anything when you stop paying?
Accuracy, the spec everyone asks about first, has largely converged for everyday dictation on modern engines — the questions above now separate products more than raw accuracy does.
Rule-based clean-up vs generative rewrite
Every dictation product does something to the raw transcript before typing it. There are two families.
A rule-based (deterministic) clean-up applies fixed rules — remove stutters, fillers and repeated words; restore contractions and proper nouns; case sentences; punctuate — with no language model involved. The same audio yields the same text every time, and every word that survives is a word you said.
A generative rewrite runs the transcript through a language model that infers what you meant and how it should read: it can restructure a list, drop hedging, and match tone to the app you are in. The output reads smoother; it is also, by design, allowed to differ from what you said.
They optimize for different things: a rewrite for text that reads smoothly, a rule-based pass for text that is exactly what you said. Dictating for a living — messages, notes, code comments, where the exact words carry the meaning? Every sentence a rewrite touches is one more to re-read; a rule-based pass is one less.
In July 2026 one cloud dictation vendor traced accuracy complaints to its rewrite default having changed words users had not asked it to touch — the failure mode of the generative family in one sentence.
ClassEve's Lven Instant uses a rule-based pass and no rewrite; Lven Cloud does the same unless you turn on its AI polish. The rule list and the test corpus that every platform's implementation is held to are published under Research, so what the pass does and does not touch is checkable rather than asserted.
The landscape
Built-in options (Windows voice typing, phone keyboards) are free and generally cloud-processed. Professional suites like Dragon add deep voice-command control at professional prices. A newer generation — including ClassEve's Lven Instant — runs transcription entirely on-device: private by architecture, unlimited by design, and working offline. The comparisons page puts them side by side, row by row.