Voicebox clones a voice, dictates into any field, and gives any MCP-aware agent a voice via a single voicebox.speak tool call, all running locally.
Yes, Voicebox is open source under the MIT license.
Voicebox is free to use.
Yes, Voicebox can be self-hosted (the source is available under the MIT license).
Open-source, fully-local AI voice studio (MIT, Python + React) from Jamie Pine: clone a voice from a short sample, generate speech across 7 TTS engines and 23 languages, and dictate into any text field with a global hotkey (a Wispr Flow alternative). The GTM hook is one MCP tool call, voicebox.speak, that lets any MCP-aware agent talk back in a voice you cloned. Models, voice data and captures never leave the machine; runs on Apple Silicon (MLX) or CUDA/ROCm/CPU (PyTorch), macOS and Windows.
Two things make this more than another voice toy for a GTM operator. First, MIT plus fully local means narration and dictation stop being a metered line item: docs, demos and voice notes become a build step you run as often as the product changes, on hardware you own. Second, voicebox.speak turns an agent's output audible with one tool call, so a research or monitoring agent can brief you in a voice instead of a wall of text. The honest part: it shares the voice lane with VoiceStudio, so pick on license and dictation, not hype; quality tracks your hardware and the chosen engine; and the consent to clone a voice is on you, every time.