irongit

Open source voice dictation for your whole desktop. Local whisper.cpp transcription, global hotkeys, optional LLM enhancement via Claude Code, Gemini CLI, Codex, or any API. Wayland-first.

mutterbox/README.md
222 lines9.2 KBMarkdown
1# mutterbox
2
3**Speak anywhere, type nowhere.**
4
5mutterbox is open source voice dictation for your whole desktop. Press a hotkey
6in any application, speak, press it again, and your words land in whatever
7input field has focus: your editor, your browser, your terminal, a chat box.
8Transcription runs entirely on your machine through whisper.cpp. An optional
9enhancement pass can clean up or transform the transcript before it is typed,
10using either the AI CLI tools you are already logged into (Claude Code, Gemini
11CLI, Codex) or direct API access.
12
13Think Wispr Flow, but open source, cross platform, and local-first.
14
15## Mission
16
17Voice is the fastest way to get thoughts out of your head, yet system-wide
18dictation is still locked behind closed source apps, subscriptions, and
19single-platform support. mutterbox exists to make great dictation:
20
21- **Open**: MIT licensed, one Rust + web codebase anyone can read and extend.
22- **Local-first**: audio is recorded, transcribed, and discarded on your
23 machine. Nothing is uploaded unless you explicitly enable enhancement, and
24 even then only the finished text transcript is sent, never audio.
25- **Everywhere**: one codebase targeting Linux (X11 and Wayland), macOS, and
26 Windows, working in any app with a text field.
27- **Yours**: your models, your prompts, your keybinds, your choice of AI
28 provider, or none at all.
29
30## How it works
31
32```
33hotkey ──▶ record mic (cpal) ──▶ whisper.cpp (local GGML model)
34 │
35 ▼
36 optional LLM cleanup (CLI tool or API)
37 │
38 ▼
39 inject into the focused field (paste or keystrokes)
40```
41
42## Features
43
44### Dictation
45
46- **Global hotkey** with two activation styles: toggle (press to start, press
47 to stop) or hold (push to talk, records while the key is down).
48- **Click-to-set hotkey capture**: click the hotkey field, press the combo you
49 want, done.
50- **Local transcription** via whisper.cpp with automatic language detection or
51 a pinned language, resampling from any microphone format, and a model cache
52 so repeat dictations skip load time.
53- **Recording overlay pill** showing listening and transcribing states, built
54 to never steal focus from the field you are dictating into.
55- **System tray** with toggle, open, and quit; closing the settings window
56 keeps mutterbox running in the tray.
57
58### One-click model manager
59
60Models download straight from Hugging Face into the app data directory with
61live progress, and everything works offline afterwards.
62
63| Model | Size | Notes |
64| --- | --- | --- |
65| Tiny / Tiny (en) | 75 MB | fastest, lowest accuracy |
66| Base / Base (en) | 142 MB | sensible default for quick notes |
67| Small / Small (en) | 466 MB | good accuracy, still responsive |
68| Medium | 1.5 GB | high accuracy, needs a strong CPU |
69| Large v3 Turbo Q5 | 574 MB | quantized, the accuracy/speed sweet spot on CPU |
70| Large v3 Turbo | 1.6 GB | near large-v3 accuracy, much faster |
71| Large v3 | 3.1 GB | best accuracy whisper.cpp offers |
72
73### Enhancement (optional)
74
75Pass the raw transcript through an LLM before it is typed. Off by default;
76dictation is fully local without it.
77
78- **CLI backend**: shells out to agents you are already authenticated with, so
79 there are no API keys to manage:
80 - Claude Code (`claude -p`)
81 - Gemini CLI (`gemini -p`)
82 - Codex (`codex exec`, read-only sandbox)
83 - Custom: any shell command; the prompt arrives on stdin, the result is read
84 from stdout, so tools like `llm` or `aichat` drop right in
85- **API backend**: direct HTTP with keys read from environment variables only,
86 never stored:
87 - Claude (`ANTHROPIC_API_KEY`)
88 - Gemini (`GEMINI_API_KEY`)
89 - Any OpenAI-compatible endpoint (`OPENAI_API_KEY` plus custom base URL:
90 OpenAI, OpenRouter, vLLM, LM Studio, llama.cpp server)
91 - Ollama for fully local enhancement, no key at all
92- **Custom instructions**: a style prompt you control. Ask for cleaned-up
93 punctuation, bullet points, professional email tone, translation, summaries
94 of rambles, whatever fits how you dictate.
95- **Transformation contract**: mutterbox always instructs the model to rewrite
96 the transcript and never respond to it. Dictate a question and you get the
97 polished question back, not an answer. Your instructions control style; they
98 cannot accidentally turn the enhancer into a chatbot.
99- If enhancement fails for any reason, the raw transcript is inserted instead
100 and you get a notification. A dictation is never lost.
101
102### Text injection that actually works
103
104- **Paste mode** (default): sets the clipboard, sends the paste chord, then
105 restores your previous clipboard.
106- **Focus-aware pasting**: on Hyprland, mutterbox asks the compositor which
107 window is focused before pasting. Terminals (foot, kitty, alacritty,
108 wezterm, ghostty, konsole, and friends) automatically get Ctrl+Shift+V
109 instead of Ctrl+V.
110- **Keystroke mode**: types the text character by character for fields that
111 block pasting.
112- **Layered Wayland support**: wtype (virtual keyboard protocol), then ydotool
113 (uinput), then enigo (XWayland/X11), falling through automatically.
114- **Never lose a dictation**: if every injection method fails, the transcript
115 stays on your clipboard and a notification tells you so.
116
117### Built for Wayland, not just ported to it
118
119Wayland compositors do not let apps grab global keys or synthesize input, by
120design. mutterbox embraces the compositor instead of fighting it:
121
122- `mutterbox --toggle`, `--start`, and `--stop` control the running instance
123 through single-instance forwarding, so any compositor keybind can drive
124 dictation.
125- On Hyprland, a setup popup generates the exact `bind` lines for your chosen
126 hotkey and activation mode (including `bind` + `bindr` pairs for push to
127 talk), with a copy button.
128- The recording overlay is made unfocusable automatically: mutterbox injects
129 the needed windowrules into Hyprland at runtime via `hyprctl`, supporting
130 both the 0.53+ rule engine syntax and older releases. Zero configuration.
131- On X11, macOS, and Windows, the in-app global hotkey works directly.
132
133## Installation
134
135### Prerequisites
136
137- Rust (stable), Node 18+
138- Linux: Tauri system deps plus a C++ toolchain for whisper.cpp
139
140Arch:
141
142```sh
143sudo pacman -S --needed webkit2gtk-4.1 gtk3 cmake base-devel alsa-lib
144# Wayland injection tools (either one works; wtype needs no daemon)
145sudo pacman -S --needed wtype wl-clipboard
146# or: sudo pacman -S ydotool && systemctl --user enable --now ydotool
147```
148
149Debian/Ubuntu:
150
151```sh
152sudo apt install libwebkit2gtk-4.1-dev libgtk-3-dev cmake build-essential \
153 libasound2-dev wtype wl-clipboard
154```
155
156### Build and run
157
158```sh
159git clone https://github.com/huncholane/mutterbox
160cd mutterbox
161npm install
162npm run tauri dev # development
163npm run tauri build # release bundles (deb, rpm, AppImage, dmg, msi)
164```
165
166## First run
167
1681. Open the **models** page and install a model. Base (English) is a fast
169 starting point; Large v3 Turbo Q5 is the accuracy sweet spot.
1702. Set your hotkey on the **general** page (click the field, press keys), and
171 pick toggle or hold activation.
1723. On Hyprland, the setup popup appears with keybind lines to copy into
173 `~/.config/hypr/hyprland.conf`, then `hyprctl reload`.
1744. Focus any text field, hit the hotkey, speak, hit it again. Text appears.
1755. Optionally enable **enhance**, pick a backend, and tune the instructions.
176
177## Configuration notes
178
179- Settings live in the platform config dir
180 (`~/.config/mutterbox/settings.json` on Linux) and models in the data dir
181 (`~/.local/share/mutterbox/models/`).
182- API keys are read from environment variables at request time and never
183 written to disk by mutterbox.
184- The spoken language setting needs a multilingual model for non-English
185 dictation.
186- Too-short recordings (under half a second) are treated as accidental taps
187 and discarded.
188
189## Architecture
190
191Tauri 2 app: Rust backend, React + TypeScript settings UI.
192
193| Module | Responsibility |
194| --- | --- |
195| `src-tauri/src/audio.rs` | cpal capture on a dedicated thread, resampling to 16 kHz mono |
196| `src-tauri/src/transcribe.rs` | whisper-rs context cache and transcription |
197| `src-tauri/src/models.rs` | model catalog, streaming downloads, install state |
198| `src-tauri/src/llm.rs` | enhancement providers (CLI and API), transformation contract |
199| `src-tauri/src/inject.rs` | focus-aware clipboard/keystroke injection with fallbacks |
200| `src-tauri/src/pipeline.rs` | record → transcribe → enhance → inject state machine |
201| `src-tauri/src/settings.rs` | settings persistence |
202| `src-tauri/src/lib.rs` | tray, hotkeys, single-instance CLI triggers, windows |
203| `src/` | settings UI (general, models, enhance) and the overlay pill |
204
205## Roadmap
206
207- History of past dictations
208- GPU inference feature flags (Vulkan, Metal, CUDA)
209- Voice activity detection and streaming transcription
210- Per-application vocabularies and text replacements
211- Packaged releases and CI builds for all three platforms
212- Sway and KDE setup popups like the Hyprland one
213
214## Contributing
215
216Issues and PRs are welcome. The codebase is small and modular on purpose:
217most features touch exactly one Rust module and one React page. Run
218`cargo check` in `src-tauri/` and `npm run build` before sending a PR.
219
220## License
221
222MIT