Most software hides its trade-offs. You get one experience, tuned for an imaginary average user, and if your computer is older or newer than that average you simply get a worse deal without ever being told why.
Offline speech recognition cannot hide it, because the model runs on your processor rather than somebody else's. So PeekoType puts the choice in front of you: five model sizes, one dropdown, and a single honest trade between speed and accuracy.
The short version: Tiny is instant and rough. Base is the sensible default. Small is where most people are happiest. Medium is noticeably sharper if your machine can take it. Large is the most accurate of the lot and really wants a graphics card. Changing is one dropdown, and each size downloads once.
What actually changes
Bigger models have seen more and can hold more context, which shows up in three specific places rather than uniformly.
Accents. The clearest difference. A strong regional accent that a small model finds tricky is usually handled comfortably a size or two up.
Specialist vocabulary. Medical, legal and technical terms improve markedly with size, because rarer words benefit most from a model that has seen more language. You can close much of this gap far more cheaply with custom vocabulary, which is worth doing whichever size you run.
Imperfect audio. Background noise, a distant microphone, a recording made on a phone in a cafe. Larger models cope better with all of it. Though a decent headset does more for you than any size increase, as we cover in how accurate is voice typing.
On clean audio, a clear speaker and everyday vocabulary, the gap between sizes is much smaller than people expect. Which is why Base is a perfectly respectable place to live.
Picking yours
Tiny. For genuinely old hardware, or when speed matters more than polish. Quick notes, reminders, rough capture. It is startlingly fast and you will notice the errors.
Base. The default, and a good one. On a typical laptop it keeps up with natural dictation without fuss. Start here.
Small. The sweet spot for most people. Meaningfully better with accents and unusual words, still comfortably quick on any machine made in the last few years. If Base is almost right but not quite, this is your next stop.
Medium. High accuracy, and it wants a reasonably fast processor. Worth it if you dictate technical material daily and your machine can take it.
Large. The most accurate of the five. Best on a machine with a dedicated graphics card. On a laptop without one you will be waiting, and that waiting breaks the rhythm of dictation more than a few extra errors would.
The mistake people make
Choosing Large immediately because bigger sounds better, on a laptop that cannot really run it, and concluding that offline dictation is slow.
It is not slow. It is being asked to do heavy work on light hardware. Drop to Small, and the same laptop feels immediate. Dictation is a rhythm, and a model that keeps up with you at 93 per cent beats one that makes you wait at 96 per cent, because the waiting costs you the sentence you were about to say.
If your computer has a few years on it, our guide to voice typing on older PCs goes into this in more depth.
Live dictation and file transcription use the same setting
Worth knowing, because it changes how you might use it. The model you pick applies to both live dictation and file transcription.
That opens a useful pattern. Keep a smaller model for everyday dictation, where responsiveness matters most, then switch up to Medium or Large before running a batch of recordings overnight, where you are not waiting at the keyboard and accuracy is worth the extra time. Switch back the next morning.
Downloads and disk space
Each size downloads once, the first time you select it, and then stays on your machine. That first fetch is the only moment an internet connection is needed. After that everything runs locally, as described in how offline speech recognition works.
Try two sizes on your own voice before settling. Ten minutes of real dictation tells you more than any specification table, and the answer genuinely differs from machine to machine.