CLASSEVE
RouteResearch
Research note · measured Jul 23, 2026 · published Aug 9, 2026

The engine Lven Instant ships vs. Whisper-class models: a measurement on consumer-grade hardware

The engine Lven Instant ships runs at RTF 0.166 on a consumer-grade CPU, with a 3.76% word error rate over 73 clean-speech clips. whisper-large-v3 never reached realtime: RTF 1.14 on the GPU and 2.46 on the CPU, 10 to 30 times slower per clip. It was one point more accurate on the 24 clips both engines ran.

Setup

Hardware: consumer-grade, CPU and GPU. Test set: 73 clips of clean read English speech (the librispeech_asr_dummy set). Metrics: word error rate (WER), real-time factor (RTF — processing time divided by audio duration; below 1 means faster than realtime), and median per-clip latency.

Candidates: the shipping Lven Instant engine on CPU, against whisper-large-v3 (fp16 on the consumer-grade GPU, and fp32 on CPU) and distil-large-v3.

Results

Engine / configurationWERRTFMedian per clip
Lven Instant engine — CPU3.76% (full 73)0.1660.93 s
whisper-large-v3 — fp16, consumer-grade GPU1.79% (paired first 24; Lven Instant: 2.79%)1.149.2 s (3.36 GB VRAM)
whisper-large-v3 — fp32, CPU—2.4627.6 s
distil-large-v3 — consumer-grade GPU4.11% (full 73)1.04—

No Whisper configuration reached realtime on this hardware: whisper-large-v3 ran at RTF 1.14 on the GPU and 2.46 on the CPU, and distil-large-v3 at 1.04 with a higher error rate than our engine on the full set (4.11% vs 3.76%). whisper-large-v3 was one point more accurate on the 24 clips both engines ran (1.79% vs 2.79%).

Conclusion

For live dictation on consumer hardware, the constraint is not accuracy in isolation — it is accuracy at an RTF far enough below 1 that text lands as you speak. On this hardware, our engine transcribes at RTF 0.166 with competitive accuracy; the Whisper-class models do not reach realtime at all.

That is why Lven Instant ships this engine on its hot path.

Scope.

  • Different CPUs and GPUs will shift every number; measure on your own hardware before deciding.
  • Clean read English speech — noisy, accented, or conversational audio can reorder the accuracy results.
  • Each engine was run in its practical deployment mode rather than a matched-precision ablation.
  • The paired-WER comparison for whisper-large-v3 covers the first 24 clips only.