macOS 14+ / Apple silicon / English
Stop talking. The words are already there.
All of it happened on your Mac.
Voqela is dictation that runs on your own Apple silicon. The speech model is a file on your disk, so there is no upload, no queue, and no server to answer to. You stop speaking and the sentence is in front of you, wherever in the world you happen to be sitting and whether or not there is a network in the room.
Runs on this Mac No round trip to wait for Nothing to connect to Types where you type
Voqela
Recording Transcribing Ready 0:00 0:01 0:02
The field you were already in
- 01 You speak
- 02 The whole transcript lands
- 03 Correct Last
- 04 Right from now on
Measured, 31 July 2026
A retained fixture decoded in 0.071 seconds after model preparation; warm decodes were previously observed between 0.075 and 0.094 seconds. The fixture is generated synthetic English audio, so this is a receipt for the local decode path, not a natural-voice or stop-to-text latency claim.
00 All local
The model lives here, so there is nothing to wait for.
Dictation that transcribes on a server has to send your voice away, wait for a machine you do not own, and send words back. Voqela skips all three. The model is on your disk and your own chip runs it, which is why it feels the way it does.
Transcribed on a server
Two hops, one wait, and a copy you do not hold
Voqela
No hop to make, so there is nothing to wait for and nothing to send
A serious model, on your own hardware
The complete recording goes to the local model in one pass. Nothing is trimmed to keep a stream moving, and nothing is sent away to be finished somewhere else. This is the whole job, done here.
On the retained fixture, which is generated synthetic English audio scored at three of six protected spans exact, the beginning, the negation, and the ending were all present.
Privacy is what local gets you
Local is the fast choice. Staying private is what comes with it. Audio and transcript processing stay on this Mac, there is no cloud transcription path, there is no account, and there is no telemetry.
The one network operation in this build is the one-time download of the speech model from its public host. It carries none of your content, and it is logged where you can read it.
Some work cannot put a recording on somebody else's server
Health, law, finance, security, government, anything covered by an agreement you signed. And plenty of people simply do not want their voice on a stranger's machine, which is reason enough on its own. Local is the same answer to both.
That is a statement about where the work happens, not a compliance claim. Voqela makes no certification promise, and this page will not carry one.
01 Anywhere
Completely local, so it works no matter where you are.
Download the model once and there is nothing left for a dictation to reach out to. The model is a file on your disk, your own chip runs it, and no part of the job waits on a connection. A plane at altitude, a train in a tunnel, a cabin, a locked-down network: same shortcut, same speed, same sentence.
The developer with the network switched off
A high-end MacBook, local models running on it, and no internet by choice. It is the case where cloud dictation simply stops, and it is the case this is built for: a serious speech model already sitting on the disk, and dictation that keeps up with the way you actually talk.
Nothing about that setup is a special mode. It is the ordinary way Voqela runs.
Built that way, not promised that way
A check runs on every build and fails it if any connection-capable code appears anywhere in the sources outside the single model-download seam. There is no second place for a connection to be added, because the build would stop.
And Settings keeps an append-only log of every network operation the app has ever started, with no filter, no toggle and no clear button. It has one entry.
And watched doing it
On 3 August 2026 a transcription was run inside a process the kernel had denied the network. It succeeded: the words came back, and the app's own accounting reported zero bytes downloaded, on named hardware, with the model file untouched.
That run went through the command line against a saved recording, so it covers the decode rather than the microphone. Speaking into it with the network off is one attended test away, and it is on the list at the bottom of this page.
02 Any app
And anywhere you can put a cursor.
That is the other half of working everywhere. Voqela is pointed at the cursor rather than at an application, so there is no window to write in and no supported-app list to check: press the shortcut in whatever you were already doing and the words go into the field that had focus, the way a paste does.
You press
Control-Option-Space
Once, in whatever you were already doing. There is no window to go to first.
- A browser tab
- A terminal
- A chat box
- A browser tab
- A terminal
- Your notes
- A chat box
- A document
- Any standard field
Kinds of place, not a compatibility table. Voqela will name an application here when that application has actually been accepted in an attended test, and not one day earlier.
Tune it per application
One place behaves differently from another, so the rules do too. A per-application rule sets how text is rendered there and whether Voqela presses Return after it lands. Send straight away in the places you always send. Stay put in the ones you read first.
Automatic Return ships off. It can submit content, which is exactly why turning it on anywhere is your decision and not a default.
Guarded, never forced
The insertion is attempted once and then verified against the same field it started in. Voqela never posts a transcript twice on the chance that the first one missed, because a duplicated paragraph is worse than a missing one.
And if a field refuses it outright, nothing is lost: the transcript reached your History before the insertion was ever attempted, with Copy and Retry on the row.
One shortcut, and no interface to visit
Control-Option-Space on a fresh install, changeable to whatever you actually have free. Hold it down to talk, or press once to toggle, because holding a key for thirty seconds is the strain some people came to dictation to avoid. Known macOS conflicts are refused rather than quietly broken.
03 Your words
It learns your words, never your voice.
Everyone has a handful of words a general model has never met: the names, the jargon, the way your own corner of the world actually says things. Fix one with Correct Last and Voqela can offer to write the rule for you. Approve it and that is the spelling you get in every dictation after, dialect and all. Nothing about the model changes, which is exactly why the rule cannot quietly stop working next Tuesday.
- A colleague's first name Comes back as a stranger's.
- The product your team ships Arrives split into two ordinary words.
- A term only your field says Constant to you, unheard everywhere else.
- A word you made up Nothing has ever seen it before.
The same handful, every day. General transcription writes something close and moves on, and you are the one who finds it later. If a second editing pass is what dictation costs you today, that is the thing this is built against.
Raw transcript
start her on meth exis at the lower dose
file the response in web bold versus hartley
the sarn vale chapter is done
Four people, four sets of words no general model has ever seen
Your vocabulary applied
Start her on Methexis at the lower dose.
File the response in Vebrold versus Hartley.
The Sarnvale chapter is done.
Because you wrote each one down once, not because it guessed better this time
Vocabulary
Search Add wordYou write
Tommy
You say
tommitommieYou write
MethexisYou say
mythexismeth exisYou write
Vebrold
You say
veb rollweb boldWhole words only, so a spoken form of “vale” never rewrites “Valentine”. Longest exact match wins, then your priority.
“My Mac already has free dictation”
It does, and on a supported Mac it runs on the device. The difference is your words. On macOS a custom vocabulary belongs to Voice Control, which is an accessibility feature, rather than to Dictation, so built-in dictation has nowhere to put the spellings you need. Apple's newer on-device speech framework has no documented custom-vocabulary surface either, where the older one carried contextual strings.
That is a difference in what each tool exposes, read from published documentation. It is not a claim about which one hears better, and this page will not make one.
Learning suggestions off by default
Fix a word with Correct Last and Voqela can offer to add it for you, so the correction you already made becomes the rule. Nothing changes what gets written until you approve the suggestion yourself, and turning the whole thing off is one switch.
What it stores is bounded: the token span, an identifier, a timestamp, a status. Never the transcript, the audio, or the window you were in.
Not model training, and nothing decides what you meant
Your list is a deterministic text rule applied after the full decode. The model does not change, so nothing about your words can drift or be quietly retuned in a later build, and the raw transcript is kept beside the corrected one so you can always see what was actually heard.
No pass tidies your sentence into something more presentable. Correct Last and Improve Last are tiles you press, not things that happen to you.
Reliability, in four lines
- Written down while you talk The audio is journalled to disk as you speak, so an interruption leaves a file rather than a memory of one.
- Saved before it is typed The transcript reaches your History before any insertion is attempted, with Copy and Retry on the row.
- Seven honest states Ready, Recording, Finishing, Transcribing, Inserting, Ready to recover, Action needed. Nothing in between.
- Recovery waits for you Work found after an interruption does nothing at all until you press Recover Dictation.
Deterministic process-death tests recover the exact committed prefix. That is a strong statement about a tested path, not a promise about every crash on every Mac.
04 Your data
Privacy you can open in Finder.
A line on a marketing page is not evidence that what you dictated stayed put. Six marks, then the evidence for them.
-
Decoded on this Mac
The model runs on your own Apple silicon.
-
No cloud transcription path
Audio and transcript processing stay here.
-
Nothing to wait for
There is no trip to a server and back.
-
No account, no telemetry
Nothing to sign up for, nothing to opt out of.
-
One folder, real paths
Settings lists every file and reveals it in Finder.
-
Every connection recorded
Append-only, no clear button, one entry so far.
Network activity
Append-onlyNothing else has ever been recorded on this Mac.
No cap, no filter, no toggle, no clear button. A log you can empty is not evidence. The purposes are a closed list in the source, so a future update check could not happen without appearing here.
The build refuses to compile a leak
A check runs on every build and fails it if any connection-capable code appears anywhere in the sources outside the single model-download seam. Not a policy, not a promise in a document: the build stops.
One folder, and you can open it
Everything Voqela keeps is a file in one folder on this Mac. Settings lists every one of them with the real path on your install and a Show in Finder control on the row. The full inventory is on the privacy and data page.
A file leaves that folder only when you export it, and it goes where you choose, when you ask.
Nothing syncs, nothing uploads
No account to make, no telemetry to opt out of, and nothing that photographs your screen for context.
Retention and a delete-everything control are still open decisions, so this page promises neither.
Your records, your agent, your Mac.
Private is the floor. What sits on top of it is that everything you said is kept, as something you own: a local, searchable ledger of your own words that leaves whole, in a file, whenever you decide it should.
No account holds it, no server has a copy of it, and no export ever happens unless you ask for one.
Everything you have said, searchable
History is a local list of every dictation, with Copy, Retry and an explicit Delete on the row. It is the record of your own work, kept where your work is.
It leaves whole, in one file
One transcript exports as a plain text file. The entire history exports as one JSON archive, newest first, with stable bytes, so two exports of the same history can be compared line for line.
That file is yours to hand to whatever assistant you want reading your own words back to you.
And the app answers to a prompt
A local command line reports status, settings, per-application rules, diagnostics, and where every file lives, with JSON on every read. An assistant can understand your install without you narrating it.
What it will never print is a transcript. Your words leave by an export you performed, and by nothing else.
Six claims, and what earns each one.
The product keeps a written ledger of which claims are allowed and which are still being earned. These are the six biggest ones: what Voqela says today, and the named piece of evidence that would let it say the stronger version.
The point is not modesty. It is that everything left standing on this page is something you can hold us to, and everything below is dated work rather than a phrase somebody liked.
“100% private”
Today Audio and transcript processing stay on this Mac. There is no cloud transcription path, no account and no telemetry, and the one network operation in the build is the one-time model download, which carries none of your content.
Unlocks it The absolute needs that download gone: a build that ships the speech model inside it, so a fresh install never reaches a public host at all.
“Works offline”
Today Once the model is on your Mac, Voqela works with no internet, and that is now watched rather than argued: on 3 August 2026 transcription succeeded with the network denied at the kernel level, zero bytes downloaded, on named hardware. The model still needs one connection to arrive, and the run covered the decode rather than the microphone.
Unlocks it The same thing with a microphone: one attended dictation in airplane mode, shortcut to words on screen, which is what extends the receipt from the decode to the whole loop.
“Instant transcription”
Today A retained 10.25-second fixture decoded in 0.071 seconds after model preparation. The fixture is generated synthetic English audio, so it is a receipt for the local decode path rather than for the wait you feel.
Unlocks it The attended stop-to-text latency run. The bar is already set: under about a third of a second beyond reaction time is instant. Nobody has measured it yet, so nobody says it yet.
“Never loses a word”
Today Audio is journalled to disk while you speak, the transcript is saved before any insertion is attempted, and deterministic process-death tests recover the exact committed prefix.
Unlocks it Interruptions counted outside the test harness: real crashes on real Macs, in numbers, rather than the one failure mode a test can stage.
“Learns your voice”
Today It learns your words. A correction becomes a proposed rule, an approved rule writes that spelling every time after, and your own dialect goes in the list beside the jargon.
Unlocks it Nothing, and that is the answer rather than a gap. The speech model does not change, so this one is not a claim being earned. It is a claim being declined.
“Works everywhere”
Today Anywhere you can put a cursor. Insertion is aimed at the field that already had focus, attempted once and then verified rather than forced, and an application that refuses it costs you nothing because the transcript was saved first.
Unlocks it A published list of named applications, one row at a time, as each one is accepted in an attended test. Not one name earlier.
05 Built for macOS
A native utility, sized like one.
Swift, Apple silicon, a Core ML model, and a menu-bar surface that stays out of the way. Below are the parts of the design where a dictation tool is usually careless.
The rendered application has not been through its attended visual acceptance yet, which is why every product image on this site is drawn rather than captured, and labelled that way in its own caption. Real screenshots will replace them and say so.
06 Questions
The ones worth asking.
My Mac already has free dictation. Why would I want this?
Because of your words. On macOS, a custom vocabulary is part of Voice Control, which is an accessibility feature, rather than part of Dictation, so built-in dictation has nowhere to put the spellings you need. Apple's newer on-device speech framework has no documented custom-vocabulary surface either, where the older one carried contextual strings. That is a difference in what each tool exposes, not a claim about which one hears better. Voqela is built the other way round: your own word list first, per-application behaviour, and the recording kept so you can go back to it. If a general model already spells your world correctly, free is a good deal. If it does not, a name it has never seen stays wrong.
Why does running it on my own Mac make it feel fast?
Because there is no round trip. Dictation that transcribes on a server has to send your audio away, wait for a machine you do not own, and send words back. Voqela does none of that: the speech model is a file on your disk and your Apple silicon runs it, so the work starts the moment you stop talking. The measurement on record is a retained 10.25-second fixture decoding in 0.071 seconds after the model is prepared. That fixture is generated synthetic English audio, so it is a receipt for the local decode path rather than a natural-voice or stop-to-text latency claim.
It keeps getting a name wrong. Can I fix that for good?
Yes, and it is the thing this is built around. Voqela learns your words: you add one with the spelling you want and the ways you actually say it out loud, and from then on it is applied to every dictation as a deterministic text rule. It can do the writing down for you as well, because correcting a dictation with Correct Last lets it offer to add that word, though the suggestion sits there until you approve it and the whole feature is off until you turn it on. Your own names, your field's jargon and your own dialect all go in the same list. What never changes is the speech model, which is why the rule cannot quietly stop working later, and it is also why the honest phrase is that it learns your words and not your voice.
Where does the text actually go?
Into the field you were already in. Voqela does not have an application to switch to and it does not have a supported-app list: it puts the transcript into whatever field had focus when you pressed the shortcut, the way a paste does, and per-application rules let you set how it behaves in one place without changing anywhere else. It is designed for standard Mac text fields, and the insertion is attempted once and then verified rather than forced. If a field refuses the text, the dictation is already saved in your History with Copy and Retry on the row, so the paragraph is never the thing that gets lost. A published list of named applications waits until each one has been accepted in an attended test.
Does it change my words, or clean them up for me?
No pass runs that decides what you meant. Rendering is a set of exact, bounded text rules; none of them infers tone, emotion, or intent, and the Verbatim profile leaves output byte-identical to the decode. Correct Last and Improve Last are tiles you press, not something that happens on its own, and the raw transcript is kept beside the corrected one so you can always see what was actually heard.
Does my voice or my text ever leave my Mac?
No. Audio and transcript processing happen on your Mac, there is no cloud transcription path, and there is no account and no telemetry. The one network operation in the build is a one-time download of the speech model from its public host, and that request carries none of your content. Settings keeps an append-only log of every network operation the app has ever started, so you can check that rather than take it on faith.
How fast is it?
The measurement on record is a retained 10.25-second fixture decoding in 0.071 seconds after the model is prepared, with warm decodes previously observed between 0.075 and 0.094 seconds. That fixture is generated synthetic English audio. It is a receipt for the local decode path, not a natural-voice or perceived-latency claim, and Voqela will not publish one of those until it has been measured properly.
What happens if my Mac dies in the middle of a dictation?
Audio is journalled to disk while you are still speaking, and the transcript is saved locally before insertion is attempted. Work found after an interruption is shown as ready to recover and waits for you to press Recover Dictation; it never transcribes or inserts itself. Deterministic tests recover the exact committed prefix after process death, which is a different and smaller statement than a guarantee about every real-world crash.
Does it clip the first words, or drop the end of a long recording?
The complete recording is decoded in one pass rather than in chunks that can be discarded to keep a stream moving. On the retained fixture, which is generated synthetic English audio scored at three of six protected spans exact, the beginning, the negation, and the ending were all present. That is a measurement on one synthetic fixture, not an accuracy rate, and Voqela will not publish one of those until natural speech has been measured properly.
What happens when I am not connected to the internet?
Nothing changes. Voqela is completely local: the speech model is a file on your disk, your own Apple silicon runs it, and once that one-time download is done there is nothing left for a dictation to reach out to. A plane at altitude, a train in a tunnel, a cabin, a locked-down corporate network, all the same shortcut and the same sentence. That is by construction rather than by promise, because a check fails the build on any connection-capable code anywhere in the sources outside the single model-download seam, and Settings keeps an append-only log of every network operation the app has ever started so you can read the count yourself rather than take it on faith. It has also been watched: on 3 August 2026 a transcription ran inside a process the kernel had denied the network and it succeeded, zero bytes downloaded, on named hardware. That run went through the command line against a saved recording, so what is still not on the record is the same thing done through a microphone, which is one attended test in airplane mode and a named launch gate rather than an open question.
Do I get to keep what I said, and can my own AI read it?
Yes, and yes, with one line drawn deliberately. Every dictation is kept on this Mac in History, which is a searchable local list with Copy, Retry and an explicit Delete on the row. One transcript exports as a plain text file; the whole history exports as one JSON archive, newest first, with stable bytes so two exports can be compared. That file is yours, and handing it to whatever assistant you want reading your own words back to you is exactly what it is for. Separately, the app answers to a local command line that reports status, settings, per-application rules, diagnostics and where every file lives, with JSON on every read, so an assistant can understand your install without you narrating it. What that command line will never print is transcript text, audio, or history. Your words leave by an export you performed, and by nothing else.
What permissions does it need, and why?
Three. Microphone, to record what you say. Accessibility, to see which field has focus and put the text into it. Input Monitoring, to notice your shortcut. Each one is explained before the system prompt appears, each has its own button, and all three can be revoked in System Settings at any time.
What does it cost, and when can I get it?
Neither is decided. There is no price, no release date, and no distribution choice to announce, and this site will not carry one before it is real. Early access is a list, not a purchase.
What does it need to run?
An Apple silicon Mac running macOS 14 or newer, English speech, and a one-time model download. The downloaded model artifact occupied approximately 595 MB on the Mac it was measured on.
Is it a meeting recorder, or does it handle other languages?
No to both. Voqela is one person dictating on one Mac in English. It is not a meeting recorder, a transcription service, a mobile companion, a team workspace, or a multilingual suite, and there is no iPhone or iPad version.
07 Early access
Be there when it opens.
Voqela is in build. Leave an address and you get one message when there is something real to install, and nothing else. No price has been set, no release date has been promised, and nothing about this list is public.
One Mac, one person, English. macOS 14 or newer on Apple silicon.