irongit

Open source voice dictation for your whole desktop. Local whisper.cpp transcription, global hotkeys, optional LLM enhancement via Claude Code, Gemini CLI, Codex, or any API. Wayland-first.

Add descriptive README and MIT license

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
huncholanehuncholaneauthored
parent 3d0809ccommit de30d9302e22779fd9c4eecd7aef215a91e007fbBrowse files

2 files changed, +213 -81

+21-0LICENSE
@@ -0,0 +1,21 @@
1+MIT License
2+
3+Copyright (c) 2026 mutterbox contributors
4+
5+Permission is hereby granted, free of charge, to any person obtaining a copy
6+of this software and associated documentation files (the "Software"), to deal
7+in the Software without restriction, including without limitation the rights
8+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9+copies of the Software, and to permit persons to whom the Software is
10+furnished to do so, subject to the following conditions:
11+
12+The above copyright notice and this permission notice shall be included in all
13+copies or substantial portions of the Software.
14+
15+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21+SOFTWARE.
+192-81README.md
@@ -1,110 +1,221 @@
11 # mutterbox
22
3-Open source voice dictation for your whole desktop. Press a hotkey anywhere,
4-speak, press it again, and the transcript lands in whatever input field has
5-focus. Transcription runs fully local through whisper.cpp; an optional
6-enhancement pass can route the raw transcript through Claude, Gemini, any
7-OpenAI-compatible endpoint, or a local Ollama model to clean it up or transform
8-it with your own instructions.
3+**Speak anywhere, type nowhere.**
4+
5+mutterbox is open source voice dictation for your whole desktop. Press a hotkey
6+in any application, speak, press it again, and your words land in whatever
7+input field has focus: your editor, your browser, your terminal, a chat box.
8+Transcription runs entirely on your machine through whisper.cpp. An optional
9+enhancement pass can clean up or transform the transcript before it is typed,
10+using either the AI CLI tools you are already logged into (Claude Code, Gemini
11+CLI, Codex) or direct API access.
912
1013 Think Wispr Flow, but open source, cross platform, and local-first.
1114
15+## Mission
16+
17+Voice is the fastest way to get thoughts out of your head, yet system-wide
18+dictation is still locked behind closed source apps, subscriptions, and
19+single-platform support. mutterbox exists to make great dictation:
20+
21+- **Open**: MIT licensed, one Rust + web codebase anyone can read and extend.
22+- **Local-first**: audio is recorded, transcribed, and discarded on your
23+ machine. Nothing is uploaded unless you explicitly enable enhancement, and
24+ even then only the finished text transcript is sent, never audio.
25+- **Everywhere**: one codebase targeting Linux (X11 and Wayland), macOS, and
26+ Windows, working in any app with a text field.
27+- **Yours**: your models, your prompts, your keybinds, your choice of AI
28+ provider, or none at all.
29+
1230 ## How it works
1331
1432 ```
15-hotkey ──▶ record mic (cpal) ──▶ whisper.cpp (local) ──▶ optional LLM cleanup ──▶ inject into focused field
33+hotkey ──▶ record mic (cpal) ──▶ whisper.cpp (local GGML model)
34+ │
35+ ▼
36+ optional LLM cleanup (CLI tool or API)
37+ │
38+ ▼
39+ inject into the focused field (paste or keystrokes)
1640 ```
1741
18-- One codebase for Linux, macOS, and Windows: Rust + Tauri 2, webview UI.
19-- Dictation models are whisper.cpp GGML files, installed with one click from the
20- Models page and stored in the app data dir. Everything works offline once a
21- model is downloaded.
22-- Enhancement is off by default and never required. API keys are read from
23- environment variables only and are never stored.
24-- Text insertion uses clipboard paste (with clipboard restore) or synthetic
25- keystrokes, selectable in settings.
42+## Features
2643
27-## Development
44+### Dictation
2845
29-Prerequisites: Rust, Node 18+, and on Linux the Tauri system deps
30-(`webkit2gtk-4.1`, `gtk3`), plus `cmake` and a C++ toolchain for whisper.cpp.
31-On Arch:
46+- **Global hotkey** with two activation styles: toggle (press to start, press
47+ to stop) or hold (push to talk, records while the key is down).
48+- **Click-to-set hotkey capture**: click the hotkey field, press the combo you
49+ want, done.
50+- **Local transcription** via whisper.cpp with automatic language detection or
51+ a pinned language, resampling from any microphone format, and a model cache
52+ so repeat dictations skip load time.
53+- **Recording overlay pill** showing listening and transcribing states, built
54+ to never steal focus from the field you are dictating into.
55+- **System tray** with toggle, open, and quit; closing the settings window
56+ keeps mutterbox running in the tray.
3257
33-```sh
34-sudo pacman -S --needed webkit2gtk-4.1 gtk3 cmake base-devel alsa-lib xdotool
35-```
58+### One-click model manager
3659
37-Run it:
60+Models download straight from Hugging Face into the app data directory with
61+live progress, and everything works offline afterwards.
62+
63+| Model | Size | Notes |
64+| --- | --- | --- |
65+| Tiny / Tiny (en) | 75 MB | fastest, lowest accuracy |
66+| Base / Base (en) | 142 MB | sensible default for quick notes |
67+| Small / Small (en) | 466 MB | good accuracy, still responsive |
68+| Medium | 1.5 GB | high accuracy, needs a strong CPU |
69+| Large v3 Turbo Q5 | 574 MB | quantized, the accuracy/speed sweet spot on CPU |
70+| Large v3 Turbo | 1.6 GB | near large-v3 accuracy, much faster |
71+| Large v3 | 3.1 GB | best accuracy whisper.cpp offers |
72+
73+### Enhancement (optional)
74+
75+Pass the raw transcript through an LLM before it is typed. Off by default;
76+dictation is fully local without it.
77+
78+- **CLI backend**: shells out to agents you are already authenticated with, so
79+ there are no API keys to manage:
80+ - Claude Code (`claude -p`)
81+ - Gemini CLI (`gemini -p`)
82+ - Codex (`codex exec`, read-only sandbox)
83+ - Custom: any shell command; the prompt arrives on stdin, the result is read
84+ from stdout, so tools like `llm` or `aichat` drop right in
85+- **API backend**: direct HTTP with keys read from environment variables only,
86+ never stored:
87+ - Claude (`ANTHROPIC_API_KEY`)
88+ - Gemini (`GEMINI_API_KEY`)
89+ - Any OpenAI-compatible endpoint (`OPENAI_API_KEY` plus custom base URL:
90+ OpenAI, OpenRouter, vLLM, LM Studio, llama.cpp server)
91+ - Ollama for fully local enhancement, no key at all
92+- **Custom instructions**: a style prompt you control. Ask for cleaned-up
93+ punctuation, bullet points, professional email tone, translation, summaries
94+ of rambles, whatever fits how you dictate.
95+- **Transformation contract**: mutterbox always instructs the model to rewrite
96+ the transcript and never respond to it. Dictate a question and you get the
97+ polished question back, not an answer. Your instructions control style; they
98+ cannot accidentally turn the enhancer into a chatbot.
99+- If enhancement fails for any reason, the raw transcript is inserted instead
100+ and you get a notification. A dictation is never lost.
101+
102+### Text injection that actually works
103+
104+- **Paste mode** (default): sets the clipboard, sends the paste chord, then
105+ restores your previous clipboard.
106+- **Focus-aware pasting**: on Hyprland, mutterbox asks the compositor which
107+ window is focused before pasting. Terminals (foot, kitty, alacritty,
108+ wezterm, ghostty, konsole, and friends) automatically get Ctrl+Shift+V
109+ instead of Ctrl+V.
110+- **Keystroke mode**: types the text character by character for fields that
111+ block pasting.
112+- **Layered Wayland support**: wtype (virtual keyboard protocol), then ydotool
113+ (uinput), then enigo (XWayland/X11), falling through automatically.
114+- **Never lose a dictation**: if every injection method fails, the transcript
115+ stays on your clipboard and a notification tells you so.
116+
117+### Built for Wayland, not just ported to it
118+
119+Wayland compositors do not let apps grab global keys or synthesize input, by
120+design. mutterbox embraces the compositor instead of fighting it:
121+
122+- `mutterbox --toggle`, `--start`, and `--stop` control the running instance
123+ through single-instance forwarding, so any compositor keybind can drive
124+ dictation.
125+- On Hyprland, a setup popup generates the exact `bind` lines for your chosen
126+ hotkey and activation mode (including `bind` + `bindr` pairs for push to
127+ talk), with a copy button.
128+- The recording overlay is made unfocusable automatically: mutterbox injects
129+ the needed windowrules into Hyprland at runtime via `hyprctl`, supporting
130+ both the 0.53+ rule engine syntax and older releases. Zero configuration.
131+- On X11, macOS, and Windows, the in-app global hotkey works directly.
132+
133+## Installation
134+
135+### Prerequisites
136+
137+- Rust (stable), Node 18+
138+- Linux: Tauri system deps plus a C++ toolchain for whisper.cpp
139+
140+Arch:
38141
39142 ```sh
40-npm install
41-npm run tauri dev
143+sudo pacman -S --needed webkit2gtk-4.1 gtk3 cmake base-devel alsa-lib
144+# Wayland injection tools (either one works; wtype needs no daemon)
145+sudo pacman -S --needed wtype wl-clipboard
146+# or: sudo pacman -S ydotool && systemctl --user enable --now ydotool
42147 ```
43148
44-Build a release bundle:
149+Debian/Ubuntu:
45150
46151 ```sh
47-npm run tauri build
152+sudo apt install libwebkit2gtk-4.1-dev libgtk-3-dev cmake build-essential \
153+ libasound2-dev wtype wl-clipboard
48154 ```
49155
50-## The hotkey on Wayland
51-
52-Global hotkeys are an X11 concept; Wayland compositors do not let apps grab
53-keys. mutterbox handles this with single-instance triggers: `mutterbox
54---toggle`, `--start`, and `--stop` control dictation in the running instance.
55-When mutterbox detects Hyprland it shows a popup with the exact lines to copy
56-into your config, generated from your chosen hotkey. For example:
57-
58-- Toggle mode: `bind = CTRL ALT, SPACE, exec, mutterbox --toggle`
59-- Hold (push to talk): `bind = CTRL ALT, SPACE, exec, mutterbox --start` plus
60- `bindr = CTRL ALT, SPACE, exec, mutterbox --stop`
61-- Sway: `bindsym $mod+d exec mutterbox --toggle`
62-- KDE / GNOME: add a custom shortcut running `mutterbox --toggle`
63-
64-On X11, macOS, and Windows the in-app hotkey (default `Ctrl+Alt+Space`, set by
65-clicking the hotkey field and pressing keys) works directly, in both toggle and
66-hold mode.
67-
68-For text injection on Wayland, install `ydotool` and enable the `ydotoold`
69-service; mutterbox falls back to it automatically when synthetic input via
70-enigo is unavailable:
156+### Build and run
71157
72158 ```sh
73-sudo pacman -S ydotool
74-systemctl --user enable --now ydotool
159+git clone https://github.com/huncholane/mutterbox
160+cd mutterbox
161+npm install
162+npm run tauri dev # development
163+npm run tauri build # release bundles (deb, rpm, AppImage, dmg, msi)
75164 ```
76165
77-## Enhancement providers
78-
79-Two backends, chosen in the enhance page:
80-
81-**CLI tools** run an installed agent with your existing login, no API keys:
82-
83-| Provider | Binary | Notes |
84-| --- | --- | --- |
85-| Claude Code | `claude` | `claude -p` with the transformation contract as system prompt |
86-| Gemini CLI | `gemini` | |
87-| Codex | `codex` | `codex exec` in its read-only sandbox |
88-| Custom | any | your command via `sh -c`; prompt on stdin, result on stdout |
89-
90-**API keys** call the provider HTTP APIs directly:
91-
92-| Provider | Key variable | Notes |
93-| --- | --- | --- |
94-| Claude | `ANTHROPIC_API_KEY` | default model `claude-opus-4-8` |
95-| Gemini | `GEMINI_API_KEY` | |
96-| OpenAI-compatible | `OPENAI_API_KEY` | custom base URL supported (OpenRouter, vLLM, LM Studio, llama.cpp server) |
97-| Ollama | none | local, defaults to `http://localhost:11434/v1` |
98-
99-The instructions field is a system prompt, so enhancement can do more than
100-cleanup: summarize rambles, force bullet points, translate, match email tone.
101-
102-## Status
103-
104-Early scaffold. Working: recording, local transcription, model manager with
105-one-click downloads, LLM enhancement, paste/type injection, tray, overlay
106-indicator, settings persistence. Not yet: push-to-talk mode, history, GPU
107-inference feature flags (Vulkan/Metal/CUDA), packaged releases.
166+## First run
167+
168+1. Open the **models** page and install a model. Base (English) is a fast
169+ starting point; Large v3 Turbo Q5 is the accuracy sweet spot.
170+2. Set your hotkey on the **general** page (click the field, press keys), and
171+ pick toggle or hold activation.
172+3. On Hyprland, the setup popup appears with keybind lines to copy into
173+ `~/.config/hypr/hyprland.conf`, then `hyprctl reload`.
174+4. Focus any text field, hit the hotkey, speak, hit it again. Text appears.
175+5. Optionally enable **enhance**, pick a backend, and tune the instructions.
176+
177+## Configuration notes
178+
179+- Settings live in the platform config dir
180+ (`~/.config/mutterbox/settings.json` on Linux) and models in the data dir
181+ (`~/.local/share/mutterbox/models/`).
182+- API keys are read from environment variables at request time and never
183+ written to disk by mutterbox.
184+- The spoken language setting needs a multilingual model for non-English
185+ dictation.
186+- Too-short recordings (under half a second) are treated as accidental taps
187+ and discarded.
188+
189+## Architecture
190+
191+Tauri 2 app: Rust backend, React + TypeScript settings UI.
192+
193+| Module | Responsibility |
194+| --- | --- |
195+| `src-tauri/src/audio.rs` | cpal capture on a dedicated thread, resampling to 16 kHz mono |
196+| `src-tauri/src/transcribe.rs` | whisper-rs context cache and transcription |
197+| `src-tauri/src/models.rs` | model catalog, streaming downloads, install state |
198+| `src-tauri/src/llm.rs` | enhancement providers (CLI and API), transformation contract |
199+| `src-tauri/src/inject.rs` | focus-aware clipboard/keystroke injection with fallbacks |
200+| `src-tauri/src/pipeline.rs` | record → transcribe → enhance → inject state machine |
201+| `src-tauri/src/settings.rs` | settings persistence |
202+| `src-tauri/src/lib.rs` | tray, hotkeys, single-instance CLI triggers, windows |
203+| `src/` | settings UI (general, models, enhance) and the overlay pill |
204+
205+## Roadmap
206+
207+- History of past dictations
208+- GPU inference feature flags (Vulkan, Metal, CUDA)
209+- Voice activity detection and streaming transcription
210+- Per-application vocabularies and text replacements
211+- Packaged releases and CI builds for all three platforms
212+- Sway and KDE setup popups like the Hyprland one
213+
214+## Contributing
215+
216+Issues and PRs are welcome. The codebase is small and modular on purpose:
217+most features touch exactly one Rust module and one React page. Run
218+`cargo check` in `src-tauri/` and `npm run build` before sending a PR.
108219
109220 ## License
110221