We build the on-device one. Here is the case for both, written so you can disagree with us.
Every voice dictation tool makes the same architectural choice: transcribe on your machine, or send the audio somewhere that transcribes it for you. Most of the difference in how these products feel traces back to that one decision.
The honest summary is that cloud dictation is usually more accurate and on-device dictation is more private and more dependable. If accuracy on difficult audio is the thing you care about most, a hosted service is a reasonable choice and we would rather say so than pretend otherwise.
Model size drives transcription accuracy more than anything else, and a server is not constrained by your laptop's memory or battery. Spok3n ships Whisper base.en, roughly 142 MB. A hosted service can run something an order of magnitude larger and does not have to care that it would take thirty seconds to load.
Where that shows up is the hard cases: a noisy room, a strong accent, two people talking, domain jargon the model has barely seen. Where it mostly does not show up is one person speaking clearly into a laptop microphone, which is what dictation actually is almost all of the time. You can close part of the gap locally by switching to a larger variant, at the cost of speed — the technical walkthrough of the local pipeline lays out what each size costs you.
Spok3n also keeps a personal dictionary of terms and aliases, which is a narrower but effective fix for the specific failure most people hit: names, brands, and technical words the model gets wrong the same way every time.
This is the part that is genuinely not a matter of degree. A cloud service has to receive your audio to transcribe it. Whatever happens next — retained, deleted after thirty days, used for training, not used for training — is a policy commitment. It may well be an honest one. It is still a commitment rather than a property of the system.
On-device, there is no transmission step. The claim is checkable: turn off networking and dictate. If it still works, the audio was never going anywhere. That is a different kind of assurance than a policy document, and it is the main reason we build the product this way.
It also matters more in some contexts than others. Dictating a grocery list, it is close to irrelevant. Dictating patient notes, legal drafts, unreleased financials, or anything covered by an agreement you signed, it is the whole question.
Cloud transcription has a real marginal cost — somebody is paying for GPU time per minute of audio — so it tends to be priced per use or capped by tier. Local transcription has no marginal cost, because your Mac is doing the work and you already own it. Spok3n charges a flat subscription for unlimited use and offers a free tier of 5,000 dictated words a week, which is possible precisely because heavy use costs us nothing extra.
Choose cloud dictation if you regularly transcribe difficult audio, need many languages, work on hardware without a neural accelerator, or care about maximum accuracy more than anything else.
Choose on-device dictation if your audio is sensitive, you want it to work without a connection, you would rather not meter your own speech, or you simply prefer that the tools on your machine not phone home. If that is you, what Spok3n does is a shorter read than this page.
Usually yes, and the gap is widest on hard audio — strong accents, background noise, overlapping speakers, or specialist vocabulary. A hosted service can run a model far larger than a laptop can hold in memory. On clear speech from a single speaker, which is most dictation, the difference is small enough that many people never notice it.
Yes, once the speech model has been downloaded. The model is fetched once on first launch; after that transcription runs with networking disabled. Cloud dictation cannot work offline by definition.
That depends entirely on the provider and its retention policy. Audio has to be transmitted and processed on their servers, so the privacy outcome is a matter of what the policy says and whether it is followed. With on-device transcription there is no transmission step to have a policy about.
It depends on what you compare. On-device has no network round trip, so short dictations often finish faster. Longer dictations favor the cloud, because a server runs a bigger model on hardware far quicker than a laptop, and that advantage eventually outweighs the round trip.
Whisper base.en by default, an English-only model that is about 142 MB on disk, running through WhisperKit on the Apple Neural Engine. Larger variants can be selected in settings if you want more accuracy and can accept slower transcription.