irongit

Open source voice dictation for your whole desktop. Local whisper.cpp transcription, global hotkeys, optional LLM enhancement via Claude Code, Gemini CLI, Codex, or any API. Wayland-first.

mutterbox/README.md
111 lines3.9 KBMarkdown

mutterbox

Open source voice dictation for your whole desktop. Press a hotkey anywhere, speak, press it again, and the transcript lands in whatever input field has focus. Transcription runs fully local through whisper.cpp; an optional enhancement pass can route the raw transcript through Claude, Gemini, any OpenAI-compatible endpoint, or a local Ollama model to clean it up or transform it with your own instructions.

Think Wispr Flow, but open source, cross platform, and local-first.

How it works

hotkey ──▶ record mic (cpal) ──▶ whisper.cpp (local) ──▶ optional LLM cleanup ──▶ inject into focused field
  • One codebase for Linux, macOS, and Windows: Rust + Tauri 2, webview UI.
  • Dictation models are whisper.cpp GGML files, installed with one click from the Models page and stored in the app data dir. Everything works offline once a model is downloaded.
  • Enhancement is off by default and never required. API keys are read from environment variables only and are never stored.
  • Text insertion uses clipboard paste (with clipboard restore) or synthetic keystrokes, selectable in settings.

Development

Prerequisites: Rust, Node 18+, and on Linux the Tauri system deps (webkit2gtk-4.1, gtk3), plus cmake and a C++ toolchain for whisper.cpp. On Arch:

sudo pacman -S --needed webkit2gtk-4.1 gtk3 cmake base-devel alsa-lib xdotool

Run it:

npm install
npm run tauri dev

Build a release bundle:

npm run tauri build

The hotkey on Wayland

Global hotkeys are an X11 concept; Wayland compositors do not let apps grab keys. mutterbox handles this with single-instance triggers: mutterbox --toggle, --start, and --stop control dictation in the running instance. When mutterbox detects Hyprland it shows a popup with the exact lines to copy into your config, generated from your chosen hotkey. For example:

  • Toggle mode: bind = CTRL ALT, SPACE, exec, mutterbox --toggle
  • Hold (push to talk): bind = CTRL ALT, SPACE, exec, mutterbox --start plus bindr = CTRL ALT, SPACE, exec, mutterbox --stop
  • Sway: bindsym $mod+d exec mutterbox --toggle
  • KDE / GNOME: add a custom shortcut running mutterbox --toggle

On X11, macOS, and Windows the in-app hotkey (default Ctrl+Alt+Space, set by clicking the hotkey field and pressing keys) works directly, in both toggle and hold mode.

For text injection on Wayland, install ydotool and enable the ydotoold service; mutterbox falls back to it automatically when synthetic input via enigo is unavailable:

sudo pacman -S ydotool
systemctl --user enable --now ydotool

Enhancement providers

Two backends, chosen in the enhance page:

CLI tools run an installed agent with your existing login, no API keys:

Provider Binary Notes
Claude Code claude claude -p with the transformation contract as system prompt
Gemini CLI gemini
Codex codex codex exec in its read-only sandbox
Custom any your command via sh -c; prompt on stdin, result on stdout

API keys call the provider HTTP APIs directly:

Provider Key variable Notes
Claude ANTHROPIC_API_KEY default model claude-opus-4-8
Gemini GEMINI_API_KEY
OpenAI-compatible OPENAI_API_KEY custom base URL supported (OpenRouter, vLLM, LM Studio, llama.cpp server)
Ollama none local, defaults to http://localhost:11434/v1

The instructions field is a system prompt, so enhancement can do more than cleanup: summarize rambles, force bullet points, translate, match email tone.

Status

Early scaffold. Working: recording, local transcription, model manager with one-click downloads, LLM enhancement, paste/type injection, tray, overlay indicator, settings persistence. Not yet: push-to-talk mode, history, GPU inference feature flags (Vulkan/Metal/CUDA), packaged releases.

License

MIT