How on-device transcription works
What actually runs on your Mac when you hold the Globe key, and why that matters.
“On-device” is a claim that is easy to make and hard to check. This page describes the specific pipeline Spok3n runs, the files it puts on disk, and the hardware it uses — so you can verify the claim rather than take it on faith.
The model is three models
Spok3n uses WhisperKit, a Swift framework from Argmax that runs OpenAI's Whisper speech-recognition model through Apple's CoreML. On first launch the app downloads a model directory from Hugging Face. Look inside it and you will find three compiled CoreML bundles, not one:
MelSpectrogram.mlmodelc
Raw microphone audio is resampled to 16 kHz and converted into a log-mel spectrogram — a time-by-frequency image of the sound. This is a fixed signal-processing step, not a learned one, but it ships as a compiled CoreML model so it runs on the same hardware path as the rest.
AudioEncoder.mlmodelc
The spectrogram goes through a transformer encoder, which turns it into a sequence of vectors representing what was said independent of who said it. This is the most compute-heavy stage and the one that benefits most from the Neural Engine.
TextDecoder.mlmodelc
A second transformer reads the encoder output and generates text one token at a time, each token conditioned on the ones before it. This autoregressive loop is why transcription time scales with how much you said, not just how long you spoke.
The three run in sequence: audio becomes a spectrogram, the spectrogram becomes encoded representations, and the decoder turns those into words. If any one of the three is missing, transcription cannot run at all — which is why an interrupted first-launch download produces a model folder that looks present but does not work.
Why CoreML and the Neural Engine
CoreML is Apple's on-device inference runtime. Given a compiled model it decides, per operation, whether to run on the Neural Engine, the GPU, or the CPU. The Neural Engine is a fixed-function block on Apple silicon built for exactly the matrix math a transformer spends its time on, and it does that work at substantially lower power than the GPU would.
That is the practical reason on-device dictation is viable on a laptop at all. The privacy property — no audio leaving the machine — follows from the architecture rather than from a policy promise, because there is no transcription server to send audio to in the first place.
Model size is the real tradeoff
Whisper comes in several sizes. Bigger models are more accurate, particularly on accented speech, background noise, and unusual vocabulary — and they are slower and take more memory. Spok3n ships base.en by default and lets you switch:
| Variant | On disk | Tradeoff |
|---|---|---|
| tiny.en | 75 MB | Fastest, noticeably lower accuracy. |
| base.en | 142 MB | The default. Balances speed and accuracy for dictation. |
| small.en | 466 MB | More accurate, slower to load and run. |
| medium.en | 1.5 GB | More accurate again, slow enough to notice. |
The .en variants are English-only. They take up the same space as the multilingual model of the same size but tend to be more accurate on English, because none of their capacity is spent on other languages.
What leaves your Mac, stated precisely
Audio never does. Neither does transcribed text. Both exist only in memory and in local storage on your machine. Three things do use the network, and it is worth being exact about them:
- The model download, once, on first launch — a one-way fetch from Hugging Face. After that, transcription works with networking off.
- Sign-in and billing, which handle your email address and subscription state. No audio or transcript is attached to any of it.
- Update checks against the release feed, to tell you a new version exists.
The full detail is in the privacy policy. If you want to check rather than read, run Spok3n with Wi-Fi off — dictation keeps working once the model is in place.
Where this approach loses
On-device transcription is not strictly better than the cloud alternative. A hosted service can run a far larger model than your laptop can, can improve it without you updating anything, and has no first-launch download. Those are real advantages. We went through the comparison honestly in on-device versus cloud dictation.
Try it
Spok3n runs on macOS 14 or later and is free for 5,000 dictated words a week. Setup, shortcuts, and model management are covered in the docs.
Download for Mac