Skip to content
TJ Miller

Where your voice actually goes when you dictate

4 min read

When you hold a key and talk to your computer, your voice goes somewhere.

It leaves your Mac, lands on a server, gets turned into text there, and the text comes back. Most dictation tools work this way. Most of the time you never think about it.

I’ve started thinking about it a lot.

Most of my work these days runs on LLMs. So I spend my days moving text and audio in and out of models that live on someone else’s hardware. That’s the normal shape of an AI feature right now: you send your data up, the model runs somewhere you can’t see, and an answer comes down. It works, and I use tools like that every day. It’s still worth being honest about what you hand over.

Dictation is the sharpest version of it. It’s your actual voice, mid-thought, often about something you’d never paste into a text box on purpose. Notes about a client. A text to someone you love. A draft you’re not ready to show anyone. The audio is you, and it goes to a server so a model can read it back.

The privacy policy will tell you it’s fine, and honestly, it might be. But “we don’t store your audio” is a decision a company makes, and decisions change with a funding round, an acquisition, or a new line in terms you won’t read.

Here’s the thing I kept coming back to. The safest audio is the audio that never leaves. If the model runs on my own machine, there’s no server in the path, no policy I’m asking you to trust, and no copy of your voice sitting in a data center waiting to leak. The question of what happens to your voice on someone else’s computer just stops existing.

So I’m building Sonari to work that way.

You press a shortcut, you talk, and it pastes clean text into whatever app you were already in. The speech model runs on your Mac. The cleanup that fixes the grammar and the filler words runs on your Mac. Your audio never gets uploaded, because there is nowhere for it to go.

I’ll say the limits up front, because running on your own chip instead of a data center costs something. Sonari is English only. It needs Apple Silicon and macOS 14 or newer. The optional rewrites, the ones that turn a rambling transcript into a clean email or a tidy list, use Apple’s own on-device models, so they need a Mac new enough for Apple Intelligence. Where that isn’t available your transcript still goes in unchanged, and it still never leaves the machine. Every one of those is a deliberate choice, and I’d rather you hear it from me than find it after you install.

Under the hood it runs Parakeet and Whisper on the Neural Engine, the part of your Mac’s chip built for exactly this. State of the art English transcription, on the silicon that’s already in your laptop. There’s a custom dictionary for the words a speech model gets wrong the same way every time. And every dictation is saved locally, so you can search it, replay the audio, and re-run it through a sharper model later. All of it stays on disk.

There’s something I genuinely love about watching a full speech model run with no network connection at all, no account, no key.

I’m not going to pretend I invented this. Local dictation exists, and Sotto in particular does it well and got there before me. What gets me is that the category default is still the cloud, and most people talking to their Mac right now have pretty much no idea their voice is making a round trip at all.

Sonari isn’t ready to buy yet. I’m starting with the list of people who want it. If you’d rather your voice stayed on your Mac, drop your email at sonari.audio and I’ll tell you the moment it’s ready.

TJ

Subscribe

More thoughts, links, projects, and personal updates. No cadence promises.