
WhisperShortcut features
A macOS menu bar app for dictation, voice-to-AI, and an optional in-app chat experience. You bring your own provider API keys (BYOK) for cloud features; no backend and no WhisperShortcut account—see the Privacy Policy for what stays on your Mac.
Speech-to-Text (macOS dictation)
Transcribe in real time with Google Gemini, GPT, or Grok in the cloud, reach any audio-capable model on OpenRouter with a single key and a model slug, use a self-hosted OpenAI-compatible endpoint, or transcribe offline with Whisper when you want to keep audio on your machine. Tuned for everyday dictation, notes, and messages.
Cloud dictation is tunable for accuracy: temperature defaults to 0.0 (verbatim)so the model reproduces what you said instead of paraphrasing it, and Gemini's thinking effort can be raised for difficult audio, accents, or unusual vocabulary. A glossary fixes names and terms the model keeps getting wrong. Long recordings are split into chunks and transcribed in parallel.
Shortcuts and clipboard
Every mode has a configurable global shortcut, and F1–F20 bind on their own, without a modifier—so a programmable (QMK/VIA) keyboard can dedicate a single key to dictation: press to start, press again to stop. Results are copied to the clipboard and can optionally be pasted straight at the cursor; turn on Restore clipboard and whatever you had copied before comes back right after the paste. The menu bar keeps your last five transcriptions, so a dictation that landed in the wrong window is never lost.
Voice-Driven Text Editing (Speech-to-Prompt)
Speak instructions to edit or improve text using AI—handy for rewriting, summarizing, or reformatting without typing. Works with the clipboard and your chosen model.
Read Aloud
Listen to clipboard text or chat replies with Gemini, GPT, or Grok text-to-speech voices—useful for proofreading, accessibility, or hands-free review. Pick a voice per provider, adjust playback speed, and optionally let AI rewrite text for more natural spoken output. Long text starts playing as soon as the first chunk is ready while the rest streams in behind it, and Markdown, links, and code fences are stripped before synthesis so chat replies read cleanly. Trigger from the menu bar or under assistant messages in chat.
AI chat: Gemini, GPT, Grok, Claude, and optional tools
Use an in-app chat with Gemini, GPT, Grok, or Anthropic Claude—or point it at a local OpenAI-compatible server such as Ollama or LM Studio. Paste images into the composer or use /screenshot to attach screen captures to your next message. Grok models also search X.com, which makes them the pick for opinions, trends, and breaking social chatter—and /x @handle narrows that search to the accounts you trust. If you connect accounts, WhisperShortcut can use Google Calendar, Tasks, Gmail (read-only), and Trello in supported flows, only when you ask. Learn more in Terms and Privacy.
Workspace folders (read-only)
Share a folder with the chat—in Settings, with the /folder command, or by dropping it onto the chat window—and the model can list, read, and search files inside it to answer questions about your notes, documents, or code. Access is read-only and limited to the folders you pick; nothing else on your Mac is readable, and you can remove a folder at any time.
Live meeting transcription
Record with chunked, near-real-time transcription in the meeting flow. As the meeting runs, short notes are written into the chat itself, at the moment things were said—so you follow along in the same column you type in, and earlier notes are never rewritten.
The chat sees the whole transcript, so you can ask about anything said since the meeting started, with one-tap questions for Catch me up, Action items, Open questions, and Decisions. A shortcut flags an important moment while you are talking and can't type, and the final summary is written around what you flagged. Meeting transcripts are stored locally and can be copied out at any time; see the Privacy Policy for how local files and improvement features work.