
WhisperShortcut features
A macOS menu bar app for dictation, voice-to-AI, and an optional in-app chat experience. You bring your own provider API keys (BYOK) for cloud features; no backend and no WhisperShortcut account—see the Privacy Policy for what stays on your Mac.
Speech-to-Text (macOS dictation)
Transcribe in real time with Google Gemini, GPT, or Grok in the cloud, reach any audio-capable model on OpenRouter with a single key and a model slug, use a self-hosted OpenAI-compatible endpoint, or transcribe offline with Whisper when you want to keep audio on your machine. Tuned for everyday dictation, notes, and messages.
Cloud dictation is tunable for accuracy: temperature defaults to 0.0 (verbatim)so the model reproduces what you said instead of paraphrasing it, and Gemini's thinking effort can be raised for difficult audio, accents, or unusual vocabulary. A glossary fixes names and terms the model keeps getting wrong; select a correctly spelled term anywhere and one shortcut adds it.
Longer recordings start transcribing while you still speak—with Gemini, GPT, Grok, and on-device Whisper—so pressing Stop only waits for the last few seconds.
Shortcuts and clipboard
Every mode has a configurable global shortcut, and F1–F20 bind on their own, without a modifier—so a programmable (QMK/VIA) keyboard can dedicate a single key to dictation: press to start, press again to stop. Results are copied to the clipboard; the direct-download build can also paste them straight at the cursor, and with Restore clipboard whatever you had copied before comes back right after the paste. The menu bar keeps your last five transcriptions, so a dictation that landed in the wrong window is never lost.
Voice-Driven Text Editing (Dictate Prompt)
Speak instructions to edit or improve text using AI—handy for rewriting, summarizing, translating, or reformatting without typing. Works with the clipboard and your chosen model, including an on-device model offline. While you record, your most frequent instructions appear above the recording pill; one key runs them without speaking.
Writing Style: drafts of emails and messages come out the way you write, not in assistant prose. The app keeps a short style profile per context (email, messenger, work chat) and a pool of messages you actually wrote, imported from your Gmail sent mail or pasted in Settings, and stores both on your Mac.
Read Aloud
Listen to any selected text or chat reply with Gemini, GPT, or Grok text-to-speech voices, or on-device macOS voices offline—useful for proofreading, accessibility, or hands-free review. Pick a voice per provider, adjust playback speed, and optionally let AI rewrite text for more natural spoken output. Long text starts playing as soon as the first chunk is ready while the rest streams in behind it, and Markdown, links, and code fences are stripped before synthesis so chat replies read cleanly. A small player offers pause, ±10 s skip, a progress bar, and live speed changes.
AI chat: Gemini, GPT, Grok, Claude, and optional tools
Use an in-app chat with Gemini, GPT, Grok, or Anthropic Claude, run an on-device model with no server at all, or point it at a local OpenAI-compatible server such as Ollama or LM Studio. Paste a YouTube link and Gemini watches the video; link a timestamp to ask about one passage. Paste images into the composer or use /screenshot to attach screen captures to your next message. Grok models also search X.com, which makes them the pick for opinions, trends, and breaking social chatter—and /x @handle narrows that search to the accounts you trust. If you connect accounts, WhisperShortcut can use Google Calendar, Tasks, Gmail (read-only), and Trello in supported flows, only when you ask. Learn more in Terms and Privacy.
Workspace folders
Share a folder with the chat—in Settings, with the /folder command, or by dropping it onto the chat window—and the model can list, read, and search files inside it to answer questions about your notes, documents, or code. AGENTS.md / CLAUDE.md files and the rules and skills you already keep there are loaded automatically, so the chat works from your own context. Access is read-only by default and limited to the folders you pick. Turn on file editing and the chat can also create and change text files there; the previous version of every edited file is backed up first, and there is no delete.
Your own endpoint: Azure OpenAI, Vertex AI, or a proxy
Point chat—and dictation—at a URL you control: your own Azure OpenAI / Microsoft Foundry resource, a Google Vertex AI project, or an OpenAI-compatible proxy such as LiteLLM. One-click presets fill in the URL shape. Requests go straight from your Mac to that endpoint, in your region and under your own contract with Microsoft or Google.
Offline Mode
One switch makes the app device-local: dictation runs on an on-device Whisper model, voice editing and chat run on an on-device model, Read Aloud uses macOS voices, and requests the app builds to cloud services are blocked at the network layer. Nothing is written to the usage log. Built for regulated dictation—patient findings, case notes—where the recording must never leave the Mac. See Offline Whisper.
Live meeting transcription
Record with chunked, near-real-time transcription in the meeting flow. As the meeting runs, short notes are written into the chat itself, at the moment things were said—so you follow along in the same column you type in, and earlier notes are never rewritten.
The chat sees the whole transcript, so you can ask about anything said since the meeting started, with one-tap questions for Catch me up, Action items, Open questions, and Decisions. A shortcut flags an important moment while you are talking and can't type, and the final summary is written around what you flagged. Meeting transcripts are stored locally; Copy transcript + chat exports the transcript, notes, and every question and answer as one Markdown document; see the Privacy Policy for how local files and improvement features work.