It is fast enough that the text is there when you stop talking. On consumer-grade hardware the shipped engine runs about six times faster than you speak, with a median of 0.93 seconds per clip and a 3.76% word error rate on clean speech — measured, dated and published under Research.
On the same hardware whisper-large-v3 ran 10 to 30 times slower per clip and never reached realtime.
It works in every application with a text field. The floating bubble sits at the edge of the screen; the global hotkey works from any window. Where an application refuses automated input, the transcript waits on the clipboard so nothing is lost.
It stays light. Nothing is metered because nothing goes to a server, about 600 MB of memory is all the engine needs, and no GPU is involved.
Dictate in 25 European languages, detected automatically.