# Voicebox

> Updated 2026-09-14 · type: tool · category: content-copywriting · status: active · rev 1

Voicebox clones a voice, dictates into any field, and gives any MCP-aware agent a voice via a single voicebox.speak tool call, all running locally.

- Open source: yes (MIT)
- Self-hostable: yes
- Pricing model: free
- Best for: Founders and operators who want narration, dictation and an agent's spoken output on their own hardware with no per-character meter, and agent builders who want to give an MCP agent a real voice with one tool call.
- Last verified: 2026-09-14

- **Canonical:** https://gtmstacker.com/registry/tool/voicebox/
- **Source:** [github · jamiepine/voicebox](https://github.com/jamiepine/voicebox)
- **Tags:** content-copywriting, local-first, voice-cloning, mcp-agents, dictation
- **Repository:** https://github.com/jamiepine/voicebox

## Is Voicebox open source?

Yes, Voicebox is open source under the MIT license.

## How much does Voicebox cost?

Voicebox is free to use.

## Can I self-host Voicebox?

Yes, Voicebox can be self-hosted (the source is available under the MIT license).

## Alternatives & related

- [VoiceStudio](https://gtmstacker.com/registry/tool/voicestudio/)


---

Open-source, fully-local AI voice studio (MIT, Python + React) from Jamie Pine: clone a voice from a short sample, generate speech across 7 TTS engines and 23 languages, and dictate into any text field with a global hotkey (a Wispr Flow alternative). The GTM hook is one MCP tool call, voicebox.speak, that lets any MCP-aware agent talk back in a voice you cloned. Models, voice data and captures never leave the machine; runs on Apple Silicon (MLX) or CUDA/ROCm/CPU (PyTorch), macOS and Windows.

## Provenance

- MIT license, local-first design, the 7-engine / 23-language range, dictation, and the voicebox.speak MCP tool independently WebFetch-verified on the repo 2026-09-14 (github.com/jamiepine/voicebox, ~53.2k stars, 638 commits, published by Jamie Pine, the Spacedrive founder). Engine and transcription details are from the project docs.
- Surfaced via the 2026-09-14 viral-posts brief ("52,000+ stars, open-source alternative to ElevenLabs + Wispr Flow, it's called Voicebox"). Note: many mirror repos of the name exist; the canonical one verified here is jamiepine/voicebox.
- Distinct from VoiceStudio (upd0912-voicestudio, AGPL-3.0). Both are local ElevenLabs alternatives; Voicebox adds Wispr-Flow-style dictation and the MCP agent voice, under a more permissive MIT license.
- Curated from the GTM Stacker signal registry (2026-09-14 pass: daily pull + viral-posts brief); license independently verified 2026-09-14.

## Why it matters for a GTM stack

Two things make this more than another voice toy for a GTM operator. First, MIT plus fully local means narration and dictation stop being a metered line item: docs, demos and voice notes become a build step you run as often as the product changes, on hardware you own. Second, voicebox.speak turns an agent's output audible with one tool call, so a research or monitoring agent can brief you in a voice instead of a wall of text. The honest part: it shares the voice lane with VoiceStudio, so pick on license and dictation, not hype; quality tracks your hardware and the chosen engine; and the consent to clone a voice is on you, every time.
