Dictation for Coding: Talk to Claude Code, Cursor, ChatGPT
Forty minutes into a Claude Code session, the agent refactors the wrong layer. You know why: the billing module looks redundant but exists because of an old vendor quirk, and the fix belongs one level up. Explaining that takes a paragraph. Typing a paragraph mid-flow feels like a penalty, so you type eleven words, and the next attempt misses too. That gap is why dictation for coding has become one of the most discussed uses of voice input, and why it looks nothing like the voice coding of a decade ago.
The developers getting real value from it aren't reading out brackets and semicolons. They're dictating intent: the prompt, the spec, the commit message, the reason a review comment matters. Syntax is the agent's job now. The explaining is still yours, and most people explain faster out loud than through their fingers.
Voice coding meant something else before the agents arrived
For most of its history, coding by voice was a discipline rather than a convenience. Talon, often paired with the Cursorless extension for VS Code, lets people write and edit code hands-free through a compact vocabulary of spoken commands. Much of that work was built by and for developers with repetitive strain injuries, and if your hands are out of action it's still the serious option, after a few weeks of practice.
What changed is the unit of work. When an agent writes the function, the thing you produce is a description of the function: what it should do, what it mustn't break, how you'll know it's finished. That's prose. And prose is exactly what ordinary dictation has always handled well.
Read the Hacker News threads on this, with titles like "Ask HN: Anyone using dictation with coding agents?", and notice what's missing. The talk is about push-to-talk keys, custom dictionaries and which system-wide tool to pipe into a CLI agent. Almost nobody asks how to say a bracket.
What developers actually dictate (and what they still type)
Sort your day by who the words are for and the split becomes obvious. Anything addressed to a person, or to a model in plain language, is a candidate for voice. Anything a compiler or a shell has to parse character by character usually isn't.
| Task | Dictate or type? | Why |
|---|---|---|
| Prompts to Claude Code, Cursor, Codex or ChatGPT | Dictate | Long context is cheap to say, and models read straight through loose phrasing |
| Specs, plans and issue descriptions | Dictate | It's explanation; tidy the structure afterwards |
| Commit messages and PR descriptions | Dictate | Say what changed and why while it's fresh, then trim before pushing |
| Code review comments | Dictate, then reread | Tone lands on a colleague, so check it before posting |
| Docstrings and comments | Dictate | Prose that happens to live in a code file |
| Variable names, regex, config, one-line edits | Type | Casing, symbols and exact characters are slow and error-prone aloud |
| Shell commands | Type, or ask the agent | One misheard flag can do real damage |
On that last row: rather than dictating a shell command verbatim, describe the goal to the agent and read the command it proposes before approving it. A mishearing never reaches your shell.
The anatomy of a spoken prompt that works
Spoken prompts fail in a predictable way: they wander. You start with the goal, remember a constraint halfway through, backtrack, and finish on a tangent about the test suite. A model can usually untangle it, but a loose shape you can hold in your head gets better first attempts.
The four-beat voice prompt
Context: where we are and what already exists. Goal: the change you want, in one sentence. Constraints: what mustn't change, which patterns to follow, what to avoid. Done when: the test, behaviour or output that proves it worked.
Say them in order and you'll rarely need a follow-up for something you forgot.
A few habits separate a prompt that lands from one that needs three corrections:
- Read before you send. Claude Code's
/voicewaits for Enter by default, and that's the right setting to keep. A two-second scan catches the misheard file name before the agent goes hunting for it. Its hold mode also has a brief warmup, so if first words go missing, the docs suggest rebinding to a modifier combination, which records from the first keypress. - Mind the built-in limits. According to Claude Code's documentation, its tap-to-record mode stops after 15 seconds of silence or two minutes in total. For a long spec, dictate into a scratch file or note first and paste it in.
- Name things the way they're spelled. Say "the user service file in the auth folder" rather than hoping the engine produces the exact path, and give the agent the real path whenever it matters.
- Talk to it like a new teammate. "This looks redundant but it isn't, because..." is precisely the context typed prompts leave out.
One boundary: if you're dictating what was decided in yesterday's design review, that's a different job. Getting meeting context into Claude or ChatGPT is what MCP servers do, covered in our piece on how conversation apps are plugging into MCP.
Built-in voice input is often enough
Check what's already inside your tools first. The major agents have caught up, and for plenty of developers that's enough.
- Claude Code has a
/voicecommand. Hold Space to talk, or switch to tap mode. It's tuned for coding vocabulary, feeds your project and git branch names in as recognition hints, and streams audio to Anthropic for transcription. It needs a Claude.ai login, so it isn't available if you authenticate with an API key or through Bedrock, and it doesn't work in SSH sessions. - Cursor added Voice Mode in version 2.0 in October 2025: built-in speech-to-text for driving Agent, plus custom submit keywords so a spoken phrase can start the run.
- ChatGPT has a dictation button in the message box. The transcript arrives as editable text, so you can fix it before sending.
- VS Code Speech, Microsoft's free extension, adds voice chat for Copilot and dictation into the editor, and processes audio locally.
If your whole day happens inside one of those windows, start there. The gap appears when you leave it: the PR description lives in a browser tab, the review comment on GitHub, the standup update in Slack. Each built-in microphone works in its own box, so you either juggle several push-to-talk habits or go back to typing everything that isn't a prompt. That's the argument for a system-wide dictation tool: one hotkey that types wherever the cursor is, terminal included.
Jargon, identifiers and the terminal
The complaint that comes up most in developer threads isn't speed; it's vocabulary. General speech engines learned from general speech, so kubectl, useEffect, your internal service names and your colleague's surname come back as creative spelling. Tools tackle this three ways.
The first is a custom dictionary: add terms once and the engine stops guessing. In Laxis that's the Personal Dictionary; seed it with your stack, services and team names before judging accuracy. The second is reading context from the screen. Wispr Flow can recognise function, class and variable names visible in VS Code, Cursor and Windsurf, and can tag files by voice in Cursor's and Windsurf's chat panels; if your prompts are thick with camelCase, that's a genuine advantage. The third is a speech model trained on technical language, which is the pitch Aqua Voice makes for its Avalon model.
The terminal is the other trap. Many dictation apps insert text by pasting it, and terminals are fussier about paste than ordinary text boxes; Wispr Flow's own help pages describe splitting long dictations into parts for Claude Code and Codex on the Mac. Whatever you use, dictate a 150-word prompt into your real terminal and confirm every word arrives and nothing submits early.
For punctuation and formatting by voice, our guide to dictation commands lists what you can say. If recognition is the real problem, why voice to text keeps mishearing you walks through the usual causes and fixes.
Filler is harmless to a model, not to a reviewer
A language model reads straight through "um, so, wait, actually." A colleague reading your PR description doesn't enjoy it nearly as much. Cleanup that strips filler words and merges repeated phrases matters more for text humans read than for prompts, so judge any tool on your commit messages and review comments, not on a prompt.
Where your audio goes when the code is proprietary
Every option above, apart from the local ones, sends your voice to a server. Claude Code's /voice streams to Anthropic. ChatGPT's dictation goes to OpenAI. Laxis processes dictation in the cloud as well, and we'd rather say so plainly here than have you discover it in a policy page.
For most prompts this adds little new exposure, because the prompt text goes to the model vendor anyway. What changes is the number of companies in the chain: dictate through a separate app and your words pass through the dictation vendor before the agent vendor. For proprietary code, customer names or anything under NDA, clear that extra hop with whoever owns security where you work.
If audio can't leave the machine at all, there are real options: VS Code Speech runs locally, and Superwhisper can run its models on your own hardware. The trade-offs, and the questions worth asking any vendor, are laid out in our comparison of on-device and cloud transcription.
Three questions for your security team
Does spoken audio count as source code or customer data under our policy? Is the dictation tool on the approved-vendor list, and not just the coding agent? Do some repositories need a tool that processes locally while others don't?
What to look for in a dictation tool for coding
This isn't the place for a ranked list; our dictation software comparison does that job. For coding, the questions are narrower:
- Same hotkey everywhere? Terminal, IDE, browser and chat, or you're back to four habits.
- Can you teach it your vocabulary, and does the list follow you between machines?
- Hold or toggle? Holding a key suits short prompts; a toggle is kinder for long specs and for hands that don't want to grip a key.
- How much does it rewrite? You want filler gone and punctuation added, not your technical meaning paraphrased.
- Where is audio processed, and is that acceptable for the code you work on?
- What's metered? Words per week, minutes per month or unlimited, and whether that allowance is shared with anything else.
On that last point, here's ours. The free Laxis plan gives 300 minutes a month in a single pool: meetings and dictation draw from the same minutes. Prompts are short, but if you also record meetings with it, you'll feel the overlap. Premium raises the pool to 2,000 minutes, and every new account gets fourteen days of Premium free, without entering a card. The Laxis voice keyboard runs on Mac, Windows, iPhone and Android and types into any app, VS Code and the terminal included.
The bottom line
For years the case against coding by voice was that code isn't language. That's still true of the code. It's no longer true of the job. The part of programming that keeps growing is the part where you explain what you want to something capable of building it, and explaining is what people were doing out loud long before anyone typed. The developers who get the most out of agents may turn out to be the ones who are best at talking to them, not the fastest typists in the room.
Frequently asked questions
Can you code by voice?
Yes, in two different ways. Most developers now dictate natural-language prompts, specs and commit messages to AI agents such as Claude Code and Cursor, and let the agent write the syntax. Fully hands-free coding, where you speak and edit the code itself, is also possible with tools like Talon and Cursorless, but it takes a few weeks of practice.
Does Claude Code have voice input?
Yes. Claude Code has a /voice command: hold Space to record, or switch to tap mode, and your speech is transcribed into the prompt. It is tuned for coding terms and uses your project and branch names as hints. It requires a Claude.ai login, streams audio to Anthropic, and does not work over SSH or with API key authentication.
How do I use voice input in Cursor?
Cursor has had a built-in Voice Mode since version 2.0, released in October 2025, which lets you speak prompts to Agent and set custom submit keywords. It is designed for the Agent input. For commit messages, PR descriptions or the terminal, developers typically add a system-wide dictation app that types wherever the cursor is.
Is it safe to dictate prompts about proprietary code?
It depends on where the audio goes and what your employer allows. Most dictation, including Claude Code's /voice, ChatGPT dictation and Laxis, is processed in the cloud, which adds a vendor to the chain. If audio must stay on your machine, VS Code Speech and local-model tools like Superwhisper are the options. Check your security policy first.
How do I get dictation to spell technical terms and variable names correctly?
Add them to a custom dictionary before judging accuracy. Most dictation tools, including Laxis with its Personal Dictionary, let you save library names, internal services and teammates' names so the engine stops guessing. Some tools also read visible names from your editor. For identifiers with unusual casing, typing them or letting the agent fill them in is often quicker.
Can dictation software type into the terminal?
Yes, most system-wide dictation apps type into terminals as well as editors and browsers, though terminals handle pasted text differently from normal text boxes. Test with a long prompt in your actual terminal and check that every word arrives. Avoid dictating shell commands word for word; describe the goal to your agent and review the command it proposes.