VoiceStudio

Open-source, fully local ElevenLabs alternative for voice cloning, voice design, video dubbing, dictation, transcription and audiobooks, with a local API and MCP server for agents.

52K stars
5.8K forks
GNU AGPLv3
Python

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
85 /100 Excellent
Development Activity 100
Maintenance 100
Community 76
Maturity 24
Momentum 40

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
83 /100 Excellent
Architecture 85
Code Quality 85
Innovation 78
Learning Curve 85

Dependency Health

Score based on the health, technical quality, freshness, and vulnerability profile of runtime dependencies. How we score it →
70 /100 Good
Library Repo Health 80
Library Technical Quality 83
Version Staleness 81
Vulnerabilities 33
Dependency Footprint 40

VoiceStudio is a desktop app and local backend for working with voices on your own machine. You can clone a voice from a short reference recording, design a new one from a description, dub videos with timed speech, dictate into any app, transcribe audio, and produce audiobooks, in 646 languages. Its default engine is OmniVoice, and you can switch to others from a model catalogue.

It runs as an Electron desktop app for macOS, Windows and Linux, with a Python FastAPI service behind it, and a Docker image is available too. Models download on first use from Hugging Face, and the compute device (CUDA, Apple MPS or CPU) is detected automatically. Remote GPU workers are optional.

For developers and agents, it exposes a local HTTP API and an MCP server that can generate speech, clone voices, transcribe audio and list voices. Everything is local by default, with analytics off unless you consent.

What You Get

  • Voice cloning from a clean reference clip, and voice design from a text description
  • Video dubbing with timed speech, word timestamps and speaker detection
  • Dictation with a floating widget, plus transcription across many engines
  • Audiobook and batch jobs from text, EPUB and PDF
  • A local API and an MCP server for agents
  • An interface in 21 languages

Common Use Cases

  • Cloning a voice for narration
  • Dubbing videos into another language
  • Dictating in any app without a cloud
  • Transcribing recordings
  • Turning documents into audiobooks

Under The Hood

Architecture A local-first desktop shell over a Python service. The Electron app supervises a FastAPI backend that serves loopback HTTP, and it is also the only place user-chosen file destinations are authorised, never through HTTP parameters. GPU-bound work (speech generation, transcription, source separation, muxing) goes through a single serialized job queue with cancellation and queue-position reporting, so concurrent requests do not exhaust VRAM. Speech engines sit behind a plugin registry with an availability probe and a shared contract, and heavy engines can run in isolated sidecar environments. Local state lives in SQLite with migrations, and an event bus streams progress to the interface.

Tech Stack Python with FastAPI and Uvicorn, PyTorch and Transformers, WhisperX, faster-whisper, pyannote and Demucs, and ONNX-based engines such as sherpa-onnx. Apple Silicon gets MLX variants, and a native Rust bridge handles desktop integration. The desktop shell is Electron with React, TypeScript, Vite, Tailwind and shadcn components. Packaging uses electron-builder and PyInstaller, Docker is supported, dependency management is uv and bun, and the MCP server uses the official Python SDK. Models come from Hugging Face.

Code Quality Pull requests are gated by CI with backend and frontend tests, installer contract tests, Electron typecheck and production builds, a Linux native-package check, and a smoke matrix across macOS, Windows and Linux. Separate workflows cover security scans, evaluations and documentation drift. Mechanical rules live in deterministic tests (locale parity, changelog style, version lockstep, hardcoded-text checks), and tests run with Hugging Face offline so cache state cannot hide failures. The codebase has regression-test discipline, strict path handling, and detailed comments that explain why. A very large main application file is the main maintainability drag.

What Makes It Unique It treats voice work as a local workstation: cloning, design, dubbing, dictation and audiobooks share one engine catalogue and one hardware-detection layer. The engine bar is a named-job policy rather than a list, with documented acceptance rules. The MCP server is built for agent context budgets, with output modes that return file URLs instead of inline audio and a confined base path for file inputs. All generated speech is watermarked, and anything that leaves the machine is opt-in.

Self-Hosting

The application is AGPL-3.0. You can use and modify it freely, including commercially, but if you offer a modified version to others over a network you must offer them the source. The maintainer also offers a commercial licence for embedding it in closed-source products. The notice says pricing tiers are coming soon, and you contact the maintainer to arrange one.

Model weights keep their own licences. The default OmniVoice code is Apache-2.0, but its pretrained weights are CC-BY-NC, so commercial use of the default model needs separate review. Its tokenizer carries extra Boson Higgs Audio 2 and Meta Llama community terms, and some engines (Supertonic-3, PocketTTS) require accepting a model licence before first use. The README tells commercial users to review each model’s licence and says to clone voices only with permission.

It runs as a desktop app or in Docker, and there is no hosted version. The backend supports macOS on Apple Silicon, Windows and Linux. Intel Macs can only act as a client to a remote backend. Hardware needs vary by engine: the default engine uses CUDA, MPS or CPU, and some engines are CPU-only. You carry disk space for models, since the default multilingual model alone is a couple of gigabytes, plus drivers and backups. There is no licence key and no gated feature, and the sponsorship page says there is no paid tier and no cloud. Optional analytics are off unless you consent. Support is community-based through Discord and GitHub Issues, and the maintainer describes it as a one-developer project with a large contributor base.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers