There is a moment that catches almost everyone the first time they use proper offline dictation. You unplug the internet, out of sheer disbelief, and carry on talking. The words keep appearing. Nothing breaks, nothing stalls, no little spinner tells you it is thinking about it somewhere in Virginia.
That moment is worth understanding, because it is not a gimmick. Offline speech recognition has quietly become good enough to be the sensible default rather than the compromise, and the reasons why are genuinely interesting.
The short version: offline speech recognition downloads an AI model once, then runs it on your own processor. Your voice is turned into text on your desk, not on a server. It works with no internet connection, it keeps confidential material confidential, and on clear speech it lands around 90 to 95 per cent accurate. PeekoType does exactly this, in 100 languages, for £39 once.
What actually happens when you speak
Cloud dictation follows a path most people never think about. Your microphone captures audio, your computer packages it up, and that package travels across the internet to a data centre. A very large model there converts it to text, and the text travels back. It usually takes under a second, which is why the journey stays invisible.
Offline recognition deletes the middle of that journey. The model sits on your hard drive. Your audio goes from the microphone into your own memory, gets processed by your own processor, and comes back as text without ever touching a network card. The round trip is measured in millimetres rather than thousands of miles.
The clever part is the model itself. It was trained on an enormous volume of multilingual speech, which is expensive and slow and happened once, on somebody else's hardware. What you download is the finished result. Running a trained model is far lighter work than training one, which is precisely why a laptop can do it and why it does not need to phone home.
Why accuracy stopped being the trade-off
For years the honest answer to "should I use offline dictation" was "only if privacy matters more than accuracy". That answer is out of date.
Two things changed. Models got dramatically more efficient, so the quality that once needed a rack of servers now fits in a few hundred megabytes. And ordinary computers got much better at the specific kind of arithmetic these models need. The result is that on clear speech, in a reasonably quiet room, a good offline model now performs about as well as the cloud service you were paying monthly for.
Where offline still asks something of you is choice. PeekoType ships five model sizes, and picking one is a straight trade between speed and precision. Tiny is startlingly quick and fine for notes. Large is the most accurate of the lot and happiest on a machine with a graphics card. Most people settle in the middle and never think about it again.
The privacy argument, without the hand-waving
Plenty of software claims to respect your privacy. Offline processing is one of the few claims you can check yourself: pull out the network cable and see whether the feature still works. If it does, the audio was never going anywhere.
This matters more than it sounds. Dictation is not like other software. People dictate things they would never type into a search box. A solicitor talks through a client's case. A therapist writes up a session. A GP records notes about a patient. A founder mutters a half-formed strategy into a document at eleven at night. With cloud dictation every one of those sentences becomes somebody else's data, governed by a privacy policy that can be rewritten with an email.
With on-device processing the question does not arise. There is no upload, so there is no retention policy to read, no human review programme to opt out of, and no international transfer to justify. If you are working under UK GDPR this simplifies your life enormously, which we go into properly in our piece on GDPR-compliant voice typing.
Speed, and the thing nobody mentions
Offline recognition removes the network round trip, and on short bursts of dictation you can feel it. There is no waiting for a handshake, no retry when the wifi wobbles, no degradation because you happen to be on hotel broadband.
The thing nobody mentions is reliability of a different sort. Cloud services change. Prices go up, free tiers shrink, features move behind a higher plan, and occasionally a product is simply discontinued. A model sitting on your own drive does none of that. It works the same on a Tuesday in three years as it does today, whether or not anybody is still selling it.
What offline recognition is genuinely good at
- Working anywhere. Trains, planes, rural cottages, secure buildings with no guest wifi. If your laptop turns on, dictation works.
- Confidential material. Client files, patient notes, HR matters, anything under NDA. See our guides for solicitors and therapists and clinicians.
- Long sessions. No per-minute billing means a two-hour dictation costs exactly the same as a two-line one, which is to say nothing.
- Multilingual work. The same model handles 100 languages, so switching between them costs nothing extra. More on that in multilingual voice typing.
- Older hardware. Choose a smaller model and a machine you were about to replace suddenly has a new job.
Where it asks a little patience
Honesty is more useful than a sales pitch, so here are the genuine trade-offs.
The first run needs an internet connection, once, to fetch the model. After that you are free, but you cannot skip that step. Larger models take longer to transcribe on modest hardware, so if you pick Large on an old laptop you will be waiting. And a heavily accented speaker in a noisy room will still challenge any system, cloud or local, though a decent headset does more for accuracy than any setting.
None of these are dealbreakers. They are just the shape of the technology, and knowing the shape helps you get the best from it.
Getting the most out of it
A few things make a disproportionate difference. Use a headset microphone rather than the one built into your laptop lid, because input quality feeds everything downstream. Set the language correctly before you start, since a model expecting English will struggle with French. Speak at a normal pace, as over-enunciating actually makes recognition worse, not better. And teach it the words it keeps missing using custom vocabulary, which is the single highest-return five minutes you can spend.
If you are completely new to dictation, our five-minute setup guide for beginners walks through the whole thing from scratch.
So is offline the right choice?
If you dictate anything confidential, work somewhere with unreliable internet, dislike subscriptions, or simply prefer your software to keep working regardless of what happens to a company you have never met, then yes. The old reason to choose cloud dictation was accuracy, and that reason has largely evaporated.
PeekoType is built around this idea from the ground up. Everything happens on your machine: dictation, file transcription with speaker labels, translation, read-aloud voices and live captions for video calls. There is no account, no telemetry and no cloud service behind it, because there is nothing for one to do.
It is £39 once, with a free 14-day trial and no card required. Try it with the internet switched off and see what you think.