VoiceStudio
Open-source, fully local ElevenLabs alternative for voice cloning, voice design, video dubbing, dictation, transcription and audiobooks, with a local API and MCP server for agents.
Repository Health
Technical Analysis
Dependency Health
VoiceStudio is a desktop app and local backend for working with voices on your own machine. You can clone a voice from a short reference recording, design a new one from a description, dub videos with timed speech, dictate into any app, transcribe audio, and produce audiobooks, in 646 languages. Its default engine is OmniVoice, and you can switch to others from a model catalogue.
It runs as an Electron desktop app for macOS, Windows and Linux, with a Python FastAPI service behind it, and a Docker image is available too. Models download on first use from Hugging Face, and the compute device (CUDA, Apple MPS or CPU) is detected automatically. Remote GPU workers are optional.
For developers and agents, it exposes a local HTTP API and an MCP server that can generate speech, clone voices, transcribe audio and list voices. Everything is local by default, with analytics off unless you consent.
What You Get
- Voice cloning from a clean reference clip, and voice design from a text description
- Video dubbing with timed speech, word timestamps and speaker detection
- Dictation with a floating widget, plus transcription across many engines
- Audiobook and batch jobs from text, EPUB and PDF
- A local API and an MCP server for agents
- An interface in 21 languages
Common Use Cases
- Cloning a voice for narration
- Dubbing videos into another language
- Dictating in any app without a cloud
- Transcribing recordings
- Turning documents into audiobooks
Under The Hood
Architecture A local-first desktop shell over a Python service. The Electron app supervises a FastAPI backend that serves loopback HTTP, and it is also the only place user-chosen file destinations are authorised, never through HTTP parameters. GPU-bound work (speech generation, transcription, source separation, muxing) goes through a single serialized job queue with cancellation and queue-position reporting, so concurrent requests do not exhaust VRAM. Speech engines sit behind a plugin registry with an availability probe and a shared contract, and heavy engines can run in isolated sidecar environments. Local state lives in SQLite with migrations, and an event bus streams progress to the interface.
Tech Stack Python with FastAPI and Uvicorn, PyTorch and Transformers, WhisperX, faster-whisper, pyannote and Demucs, and ONNX-based engines such as sherpa-onnx. Apple Silicon gets MLX variants, and a native Rust bridge handles desktop integration. The desktop shell is Electron with React, TypeScript, Vite, Tailwind and shadcn components. Packaging uses electron-builder and PyInstaller, Docker is supported, dependency management is uv and bun, and the MCP server uses the official Python SDK. Models come from Hugging Face.
Code Quality Pull requests are gated by CI with backend and frontend tests, installer contract tests, Electron typecheck and production builds, a Linux native-package check, and a smoke matrix across macOS, Windows and Linux. Separate workflows cover security scans, evaluations and documentation drift. Mechanical rules live in deterministic tests (locale parity, changelog style, version lockstep, hardcoded-text checks), and tests run with Hugging Face offline so cache state cannot hide failures. The codebase has regression-test discipline, strict path handling, and detailed comments that explain why. A very large main application file is the main maintainability drag.
What Makes It Unique It treats voice work as a local workstation: cloning, design, dubbing, dictation and audiobooks share one engine catalogue and one hardware-detection layer. The engine bar is a named-job policy rather than a list, with documented acceptance rules. The MCP server is built for agent context budgets, with output modes that return file URLs instead of inline audio and a confined base path for file inputs. All generated speech is watermarked, and anything that leaves the machine is opt-in.
Self-Hosting
The application is AGPL-3.0. You can use and modify it freely, including commercially, but if you offer a modified version to others over a network you must offer them the source. The maintainer also offers a commercial licence for embedding it in closed-source products. The notice says pricing tiers are coming soon, and you contact the maintainer to arrange one.
Model weights keep their own licences. The default OmniVoice code is Apache-2.0, but its pretrained weights are CC-BY-NC, so commercial use of the default model needs separate review. Its tokenizer carries extra Boson Higgs Audio 2 and Meta Llama community terms, and some engines (Supertonic-3, PocketTTS) require accepting a model licence before first use. The README tells commercial users to review each model’s licence and says to clone voices only with permission.
It runs as a desktop app or in Docker, and there is no hosted version. The backend supports macOS on Apple Silicon, Windows and Linux. Intel Macs can only act as a client to a remote backend. Hardware needs vary by engine: the default engine uses CUDA, MPS or CPU, and some engines are CPU-only. You carry disk space for models, since the default multilingual model alone is a couple of gigabytes, plus drivers and backups. There is no licence key and no gated feature, and the sponsorship page says there is no paid tier and no cloud. Optional analytics are off unless you consent. Support is community-based through Discord and GitHub Issues, and the maintainer describes it as a one-developer project with a large contributor base.
Related Apps
Voicebox
AI Development · Productivity · Voice AI
Clone voices, dictate anywhere, and give AI agents your voice — all locally.
GitNexus
AI Code Assistants · Developer Tools · Mcp
Index any codebase into an interactive knowledge graph and give your AI agents deep architectural context via MCP — with zero servers required.
fish-speech
AI Development · Developer Tools · Music Audio
SOTA open-source dual-autoregressive text-to-speech model with rapid voice cloning, inline emotion tags, and real-time streaming inference across 80+ languages.