irongit

Open source voice dictation for your whole desktop. Local whisper.cpp transcription, global hotkeys, optional LLM enhancement via Claude Code, Gemini CLI, Codex, or any API. Wayland-first.

4 commits
Clone ▾
HTTPS
SSH
CLI
README.md

mutterbox

Speak anywhere, type nowhere.

mutterbox is open source voice dictation for your whole desktop. Press a hotkey in any application, speak, press it again, and your words land in whatever input field has focus: your editor, your browser, your terminal, a chat box. Transcription runs entirely on your machine through whisper.cpp. An optional enhancement pass can clean up or transform the transcript before it is typed, using either the AI CLI tools you are already logged into (Claude Code, Gemini CLI, Codex) or direct API access.

Think Wispr Flow, but open source, cross platform, and local-first.

Mission

Voice is the fastest way to get thoughts out of your head, yet system-wide dictation is still locked behind closed source apps, subscriptions, and single-platform support. mutterbox exists to make great dictation:

  • Open: MIT licensed, one Rust + web codebase anyone can read and extend.
  • Local-first: audio is recorded, transcribed, and discarded on your machine. Nothing is uploaded unless you explicitly enable enhancement, and even then only the finished text transcript is sent, never audio.
  • Everywhere: one codebase targeting Linux (X11 and Wayland), macOS, and Windows, working in any app with a text field.
  • Yours: your models, your prompts, your keybinds, your choice of AI provider, or none at all.

How it works

hotkey ──▶ record mic (cpal) ──▶ whisper.cpp (local GGML model)
                                        │
                                        ▼
                       optional LLM cleanup (CLI tool or API)
                                        │
                                        ▼
                    inject into the focused field (paste or keystrokes)

Features

Dictation

  • Global hotkey with two activation styles: toggle (press to start, press to stop) or hold (push to talk, records while the key is down).
  • Click-to-set hotkey capture: click the hotkey field, press the combo you want, done.
  • Local transcription via whisper.cpp with automatic language detection or a pinned language, resampling from any microphone format, and a model cache so repeat dictations skip load time.
  • Recording overlay pill showing listening and transcribing states, built to never steal focus from the field you are dictating into.
  • System tray with toggle, open, and quit; closing the settings window keeps mutterbox running in the tray.

One-click model manager

Models download straight from Hugging Face into the app data directory with live progress, and everything works offline afterwards.

Model Size Notes
Tiny / Tiny (en) 75 MB fastest, lowest accuracy
Base / Base (en) 142 MB sensible default for quick notes
Small / Small (en) 466 MB good accuracy, still responsive
Medium 1.5 GB high accuracy, needs a strong CPU
Large v3 Turbo Q5 574 MB quantized, the accuracy/speed sweet spot on CPU
Large v3 Turbo 1.6 GB near large-v3 accuracy, much faster
Large v3 3.1 GB best accuracy whisper.cpp offers

Enhancement (optional)

Pass the raw transcript through an LLM before it is typed. Off by default; dictation is fully local without it.

  • CLI backend: shells out to agents you are already authenticated with, so there are no API keys to manage:
    • Claude Code (claude -p)
    • Gemini CLI (gemini -p)
    • Codex (codex exec, read-only sandbox)
    • Custom: any shell command; the prompt arrives on stdin, the result is read from stdout, so tools like llm or aichat drop right in
  • API backend: direct HTTP with keys read from environment variables only, never stored:
    • Claude (ANTHROPIC_API_KEY)
    • Gemini (GEMINI_API_KEY)
    • Any OpenAI-compatible endpoint (OPENAI_API_KEY plus custom base URL: OpenAI, OpenRouter, vLLM, LM Studio, llama.cpp server)
    • Ollama for fully local enhancement, no key at all
  • Custom instructions: a style prompt you control. Ask for cleaned-up punctuation, bullet points, professional email tone, translation, summaries of rambles, whatever fits how you dictate.
  • Transformation contract: mutterbox always instructs the model to rewrite the transcript and never respond to it. Dictate a question and you get the polished question back, not an answer. Your instructions control style; they cannot accidentally turn the enhancer into a chatbot.
  • If enhancement fails for any reason, the raw transcript is inserted instead and you get a notification. A dictation is never lost.

Text injection that actually works

  • Paste mode (default): sets the clipboard, sends the paste chord, then restores your previous clipboard.
  • Focus-aware pasting: on Hyprland, mutterbox asks the compositor which window is focused before pasting. Terminals (foot, kitty, alacritty, wezterm, ghostty, konsole, and friends) automatically get Ctrl+Shift+V instead of Ctrl+V.
  • Keystroke mode: types the text character by character for fields that block pasting.
  • Layered Wayland support: wtype (virtual keyboard protocol), then ydotool (uinput), then enigo (XWayland/X11), falling through automatically.
  • Never lose a dictation: if every injection method fails, the transcript stays on your clipboard and a notification tells you so.

Built for Wayland, not just ported to it

Wayland compositors do not let apps grab global keys or synthesize input, by design. mutterbox embraces the compositor instead of fighting it:

  • mutterbox --toggle, --start, and --stop control the running instance through single-instance forwarding, so any compositor keybind can drive dictation.
  • On Hyprland, a setup popup generates the exact bind lines for your chosen hotkey and activation mode (including bind + bindr pairs for push to talk), with a copy button.
  • The recording overlay is made unfocusable automatically: mutterbox injects the needed windowrules into Hyprland at runtime via hyprctl, supporting both the 0.53+ rule engine syntax and older releases. Zero configuration.
  • On X11, macOS, and Windows, the in-app global hotkey works directly.

Installation

Prerequisites

  • Rust (stable), Node 18+
  • Linux: Tauri system deps plus a C++ toolchain for whisper.cpp

Arch:

sudo pacman -S --needed webkit2gtk-4.1 gtk3 cmake base-devel alsa-lib
# Wayland injection tools (either one works; wtype needs no daemon)
sudo pacman -S --needed wtype wl-clipboard
# or: sudo pacman -S ydotool && systemctl --user enable --now ydotool

Debian/Ubuntu:

sudo apt install libwebkit2gtk-4.1-dev libgtk-3-dev cmake build-essential \
  libasound2-dev wtype wl-clipboard

Build and run

git clone https://github.com/huncholane/mutterbox
cd mutterbox
npm install
npm run tauri dev      # development
npm run tauri build    # release bundles (deb, rpm, AppImage, dmg, msi)

First run

  1. Open the models page and install a model. Base (English) is a fast starting point; Large v3 Turbo Q5 is the accuracy sweet spot.
  2. Set your hotkey on the general page (click the field, press keys), and pick toggle or hold activation.
  3. On Hyprland, the setup popup appears with keybind lines to copy into ~/.config/hypr/hyprland.conf, then hyprctl reload.
  4. Focus any text field, hit the hotkey, speak, hit it again. Text appears.
  5. Optionally enable enhance, pick a backend, and tune the instructions.

Configuration notes

  • Settings live in the platform config dir (~/.config/mutterbox/settings.json on Linux) and models in the data dir (~/.local/share/mutterbox/models/).
  • API keys are read from environment variables at request time and never written to disk by mutterbox.
  • The spoken language setting needs a multilingual model for non-English dictation.
  • Too-short recordings (under half a second) are treated as accidental taps and discarded.

Architecture

Tauri 2 app: Rust backend, React + TypeScript settings UI.

Module Responsibility
src-tauri/src/audio.rs cpal capture on a dedicated thread, resampling to 16 kHz mono
src-tauri/src/transcribe.rs whisper-rs context cache and transcription
src-tauri/src/models.rs model catalog, streaming downloads, install state
src-tauri/src/llm.rs enhancement providers (CLI and API), transformation contract
src-tauri/src/inject.rs focus-aware clipboard/keystroke injection with fallbacks
src-tauri/src/pipeline.rs record → transcribe → enhance → inject state machine
src-tauri/src/settings.rs settings persistence
src-tauri/src/lib.rs tray, hotkeys, single-instance CLI triggers, windows
src/ settings UI (general, models, enhance) and the overlay pill

Roadmap

  • History of past dictations
  • GPU inference feature flags (Vulkan, Metal, CUDA)
  • Voice activity detection and streaming transcription
  • Per-application vocabularies and text replacements
  • Packaged releases and CI builds for all three platforms
  • Sway and KDE setup popups like the Hyprland one

Contributing

Issues and PRs are welcome. The codebase is small and modular on purpose: most features touch exactly one Rust module and one React page. Run cargo check in src-tauri/ and npm run build before sending a PR.

License

MIT