Open source voice dictation for your whole desktop. Local whisper.cpp transcription, global hotkeys, optional LLM enhancement via Claude Code, Gemini CLI, Codex, or any API. Wayland-first.
| 1 | # mutterbox |
| 2 | |
| 3 | **Speak anywhere, type nowhere.** |
| 4 | |
| 5 | mutterbox is open source voice dictation for your whole desktop. Press a hotkey |
| 6 | in any application, speak, press it again, and your words land in whatever |
| 7 | input field has focus: your editor, your browser, your terminal, a chat box. |
| 8 | Transcription runs entirely on your machine through whisper.cpp. An optional |
| 9 | enhancement pass can clean up or transform the transcript before it is typed, |
| 10 | using either the AI CLI tools you are already logged into (Claude Code, Gemini |
| 11 | CLI, Codex) or direct API access. |
| 12 | |
| 13 | Think Wispr Flow, but open source, cross platform, and local-first. |
| 14 | |
| 15 | ## Mission |
| 16 | |
| 17 | Voice is the fastest way to get thoughts out of your head, yet system-wide |
| 18 | dictation is still locked behind closed source apps, subscriptions, and |
| 19 | single-platform support. mutterbox exists to make great dictation: |
| 20 | |
| 21 | - **Open**: MIT licensed, one Rust + web codebase anyone can read and extend. |
| 22 | - **Local-first**: audio is recorded, transcribed, and discarded on your |
| 23 | machine. Nothing is uploaded unless you explicitly enable enhancement, and |
| 24 | even then only the finished text transcript is sent, never audio. |
| 25 | - **Everywhere**: one codebase targeting Linux (X11 and Wayland), macOS, and |
| 26 | Windows, working in any app with a text field. |
| 27 | - **Yours**: your models, your prompts, your keybinds, your choice of AI |
| 28 | provider, or none at all. |
| 29 | |
| 30 | ## How it works |
| 31 | |
| 32 | ``` |
| 33 | hotkey ──▶ record mic (cpal) ──▶ whisper.cpp (local GGML model) |
| 34 | │ |
| 35 | ▼ |
| 36 | optional LLM cleanup (CLI tool or API) |
| 37 | │ |
| 38 | ▼ |
| 39 | inject into the focused field (paste or keystrokes) |
| 40 | ``` |
| 41 | |
| 42 | ## Features |
| 43 | |
| 44 | ### Dictation |
| 45 | |
| 46 | - **Global hotkey** with two activation styles: toggle (press to start, press |
| 47 | to stop) or hold (push to talk, records while the key is down). |
| 48 | - **Click-to-set hotkey capture**: click the hotkey field, press the combo you |
| 49 | want, done. |
| 50 | - **Local transcription** via whisper.cpp with automatic language detection or |
| 51 | a pinned language, resampling from any microphone format, and a model cache |
| 52 | so repeat dictations skip load time. |
| 53 | - **Recording overlay pill** showing listening and transcribing states, built |
| 54 | to never steal focus from the field you are dictating into. |
| 55 | - **System tray** with toggle, open, and quit; closing the settings window |
| 56 | keeps mutterbox running in the tray. |
| 57 | |
| 58 | ### One-click model manager |
| 59 | |
| 60 | Models download straight from Hugging Face into the app data directory with |
| 61 | live progress, and everything works offline afterwards. |
| 62 | |
| 63 | | Model | Size | Notes | |
| 64 | | --- | --- | --- | |
| 65 | | Tiny / Tiny (en) | 75 MB | fastest, lowest accuracy | |
| 66 | | Base / Base (en) | 142 MB | sensible default for quick notes | |
| 67 | | Small / Small (en) | 466 MB | good accuracy, still responsive | |
| 68 | | Medium | 1.5 GB | high accuracy, needs a strong CPU | |
| 69 | | Large v3 Turbo Q5 | 574 MB | quantized, the accuracy/speed sweet spot on CPU | |
| 70 | | Large v3 Turbo | 1.6 GB | near large-v3 accuracy, much faster | |
| 71 | | Large v3 | 3.1 GB | best accuracy whisper.cpp offers | |
| 72 | |
| 73 | ### Enhancement (optional) |
| 74 | |
| 75 | Pass the raw transcript through an LLM before it is typed. Off by default; |
| 76 | dictation is fully local without it. |
| 77 | |
| 78 | - **CLI backend**: shells out to agents you are already authenticated with, so |
| 79 | there are no API keys to manage: |
| 80 | - Claude Code (`claude -p`) |
| 81 | - Gemini CLI (`gemini -p`) |
| 82 | - Codex (`codex exec`, read-only sandbox) |
| 83 | - Custom: any shell command; the prompt arrives on stdin, the result is read |
| 84 | from stdout, so tools like `llm` or `aichat` drop right in |
| 85 | - **API backend**: direct HTTP with keys read from environment variables only, |
| 86 | never stored: |
| 87 | - Claude (`ANTHROPIC_API_KEY`) |
| 88 | - Gemini (`GEMINI_API_KEY`) |
| 89 | - Any OpenAI-compatible endpoint (`OPENAI_API_KEY` plus custom base URL: |
| 90 | OpenAI, OpenRouter, vLLM, LM Studio, llama.cpp server) |
| 91 | - Ollama for fully local enhancement, no key at all |
| 92 | - **Custom instructions**: a style prompt you control. Ask for cleaned-up |
| 93 | punctuation, bullet points, professional email tone, translation, summaries |
| 94 | of rambles, whatever fits how you dictate. |
| 95 | - **Transformation contract**: mutterbox always instructs the model to rewrite |
| 96 | the transcript and never respond to it. Dictate a question and you get the |
| 97 | polished question back, not an answer. Your instructions control style; they |
| 98 | cannot accidentally turn the enhancer into a chatbot. |
| 99 | - If enhancement fails for any reason, the raw transcript is inserted instead |
| 100 | and you get a notification. A dictation is never lost. |
| 101 | |
| 102 | ### Text injection that actually works |
| 103 | |
| 104 | - **Paste mode** (default): sets the clipboard, sends the paste chord, then |
| 105 | restores your previous clipboard. |
| 106 | - **Focus-aware pasting**: on Hyprland, mutterbox asks the compositor which |
| 107 | window is focused before pasting. Terminals (foot, kitty, alacritty, |
| 108 | wezterm, ghostty, konsole, and friends) automatically get Ctrl+Shift+V |
| 109 | instead of Ctrl+V. |
| 110 | - **Keystroke mode**: types the text character by character for fields that |
| 111 | block pasting. |
| 112 | - **Layered Wayland support**: wtype (virtual keyboard protocol), then ydotool |
| 113 | (uinput), then enigo (XWayland/X11), falling through automatically. |
| 114 | - **Never lose a dictation**: if every injection method fails, the transcript |
| 115 | stays on your clipboard and a notification tells you so. |
| 116 | |
| 117 | ### Built for Wayland, not just ported to it |
| 118 | |
| 119 | Wayland compositors do not let apps grab global keys or synthesize input, by |
| 120 | design. mutterbox embraces the compositor instead of fighting it: |
| 121 | |
| 122 | - `mutterbox --toggle`, `--start`, and `--stop` control the running instance |
| 123 | through single-instance forwarding, so any compositor keybind can drive |
| 124 | dictation. |
| 125 | - On Hyprland, a setup popup generates the exact `bind` lines for your chosen |
| 126 | hotkey and activation mode (including `bind` + `bindr` pairs for push to |
| 127 | talk), with a copy button. |
| 128 | - The recording overlay is made unfocusable automatically: mutterbox injects |
| 129 | the needed windowrules into Hyprland at runtime via `hyprctl`, supporting |
| 130 | both the 0.53+ rule engine syntax and older releases. Zero configuration. |
| 131 | - On X11, macOS, and Windows, the in-app global hotkey works directly. |
| 132 | |
| 133 | ## Installation |
| 134 | |
| 135 | ### Prerequisites |
| 136 | |
| 137 | - Rust (stable), Node 18+ |
| 138 | - Linux: Tauri system deps plus a C++ toolchain for whisper.cpp |
| 139 | |
| 140 | Arch: |
| 141 | |
| 142 | ```sh |
| 143 | sudo pacman -S --needed webkit2gtk-4.1 gtk3 cmake base-devel alsa-lib |
| 144 | # Wayland injection tools (either one works; wtype needs no daemon) |
| 145 | sudo pacman -S --needed wtype wl-clipboard |
| 146 | # or: sudo pacman -S ydotool && systemctl --user enable --now ydotool |
| 147 | ``` |
| 148 | |
| 149 | Debian/Ubuntu: |
| 150 | |
| 151 | ```sh |
| 152 | sudo apt install libwebkit2gtk-4.1-dev libgtk-3-dev cmake build-essential \ |
| 153 | libasound2-dev wtype wl-clipboard |
| 154 | ``` |
| 155 | |
| 156 | ### Build and run |
| 157 | |
| 158 | ```sh |
| 159 | git clone https://github.com/huncholane/mutterbox |
| 160 | cd mutterbox |
| 161 | npm install |
| 162 | npm run tauri dev # development |
| 163 | npm run tauri build # release bundles (deb, rpm, AppImage, dmg, msi) |
| 164 | ``` |
| 165 | |
| 166 | ## First run |
| 167 | |
| 168 | 1. Open the **models** page and install a model. Base (English) is a fast |
| 169 | starting point; Large v3 Turbo Q5 is the accuracy sweet spot. |
| 170 | 2. Set your hotkey on the **general** page (click the field, press keys), and |
| 171 | pick toggle or hold activation. |
| 172 | 3. On Hyprland, the setup popup appears with keybind lines to copy into |
| 173 | `~/.config/hypr/hyprland.conf`, then `hyprctl reload`. |
| 174 | 4. Focus any text field, hit the hotkey, speak, hit it again. Text appears. |
| 175 | 5. Optionally enable **enhance**, pick a backend, and tune the instructions. |
| 176 | |
| 177 | ## Configuration notes |
| 178 | |
| 179 | - Settings live in the platform config dir |
| 180 | (`~/.config/mutterbox/settings.json` on Linux) and models in the data dir |
| 181 | (`~/.local/share/mutterbox/models/`). |
| 182 | - API keys are read from environment variables at request time and never |
| 183 | written to disk by mutterbox. |
| 184 | - The spoken language setting needs a multilingual model for non-English |
| 185 | dictation. |
| 186 | - Too-short recordings (under half a second) are treated as accidental taps |
| 187 | and discarded. |
| 188 | |
| 189 | ## Architecture |
| 190 | |
| 191 | Tauri 2 app: Rust backend, React + TypeScript settings UI. |
| 192 | |
| 193 | | Module | Responsibility | |
| 194 | | --- | --- | |
| 195 | | `src-tauri/src/audio.rs` | cpal capture on a dedicated thread, resampling to 16 kHz mono | |
| 196 | | `src-tauri/src/transcribe.rs` | whisper-rs context cache and transcription | |
| 197 | | `src-tauri/src/models.rs` | model catalog, streaming downloads, install state | |
| 198 | | `src-tauri/src/llm.rs` | enhancement providers (CLI and API), transformation contract | |
| 199 | | `src-tauri/src/inject.rs` | focus-aware clipboard/keystroke injection with fallbacks | |
| 200 | | `src-tauri/src/pipeline.rs` | record → transcribe → enhance → inject state machine | |
| 201 | | `src-tauri/src/settings.rs` | settings persistence | |
| 202 | | `src-tauri/src/lib.rs` | tray, hotkeys, single-instance CLI triggers, windows | |
| 203 | | `src/` | settings UI (general, models, enhance) and the overlay pill | |
| 204 | |
| 205 | ## Roadmap |
| 206 | |
| 207 | - History of past dictations |
| 208 | - GPU inference feature flags (Vulkan, Metal, CUDA) |
| 209 | - Voice activity detection and streaming transcription |
| 210 | - Per-application vocabularies and text replacements |
| 211 | - Packaged releases and CI builds for all three platforms |
| 212 | - Sway and KDE setup popups like the Hyprland one |
| 213 | |
| 214 | ## Contributing |
| 215 | |
| 216 | Issues and PRs are welcome. The codebase is small and modular on purpose: |
| 217 | most features touch exactly one Rust module and one React page. Run |
| 218 | `cargo check` in `src-tauri/` and `npm run build` before sending a PR. |
| 219 | |
| 220 | ## License |
| 221 | |
| 222 | MIT |