YuE2 Studio

Windows desktop studio that runs the YuE2 song model on your own NVIDIA GPU: it writes an editable sheet-music score first, then sings it as a full song with vocals, offline, with covers, LoRA training and an MCP server for agents.

320 stars
MIT License

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
60 /100 Good
Development Activity 100
Maintenance 100
Community 40
Maturity 0
Momentum 0

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
83 /100 Excellent
Architecture 84
Code Quality 78
Innovation 85
Learning Curve 85

Dependency Health

Score based on the health, technical quality, freshness, and vulnerability profile of runtime dependencies. How we score it →
76 /100 Good
Library Repo Health 73
Library Technical Quality 83
Version Staleness 88
Vulnerabilities 52
Dependency Footprint 81

YuE2 Studio is a native Windows desktop app for YuE2, the open song model from M-A-P that writes a score before it sings. You describe a style and write lyrics. The model composes a melody with chords in ABC notation, the studio engraves it as sheet music, and then the model performs it as a full song with vocals of up to six minutes, entirely on your own graphics card.

Because the score comes back with the track, you can edit notes, tempo or key and render the same composition again with a different sound. You can also re-render from saved audio codes with other steps or seeds. Cover mode uses SheetSage2 to transcribe a recording’s melody, and YuE2 then sings it in your style. Around generation sits a full studio: a library, stem separation, audio-to-MIDI, word-level karaoke, audio processing with VST3 and mastering, video clips, and a player with an equalizer, MilkDrop visualiser and a skinned Winamp mode.

It ships as one Tauri executable with a Rust service inside and a bundled C++/CUDA engine from yue2.cpp, so there is no Python or Node.js in the runtime path. NVIDIA cards run on CUDA, AMD and Intel are experimental on Vulkan, and models download on first use. While it is open, it also serves an MCP endpoint so agents such as Claude Code can drive everything the interface does.

What You Get

  • Full songs with vocals, up to six minutes, from a style and lyrics
  • An editable sheet-music score that comes back with every track
  • A Windows installer with auto-update, or a portable folder that holds all its data
  • A local library with playlists, cover art and MP3 export with ID3 tags
  • An interface in English, Russian, Chinese, Japanese and Korean
  • A local MCP endpoint for agents

Common Use Cases

  • Sketching songs from lyrics and a style description
  • Covering existing melodies in new styles
  • Producing stems and MIDI from any track
  • Training a LoRA on a particular artist’s style

Under The Hood

Architecture A layered, process-supervised design. A React interface talks over local HTTP to a Rust service compiled into the desktop shell. The service owns the library, the job queue, downloads and result import, and supervises a separate native inference engine bound only to loopback. Finished songs are imported by the service itself, so a reload or closed window never loses a result. Progress is derived from the engine’s own log. Engines are treated as adapters behind capability-level provider selection (local or OpenRouter), so no interface code hardcodes a model. Child processes are grouped so none outlive the studio. The same service code backs the UI and the MCP endpoint, plus a bridge that lets an agent see and operate the window.

Tech Stack Rust with Axum and Tokio for the service, Tauri for the desktop shell, and an NSIS installer with a signed auto-updater. The UI is React and TypeScript on Vite with Tailwind, abcjs for engraving, Webamp, Butterchurn, audioMotion and an in-browser FFmpeg. Storage is bundled SQLite plus plain files. Audio work uses Symphonia for decoding, LAME for MP3 and a dedicated Rust DSP crate. ONNX Runtime is loaded dynamically for recognition models. The engine is C++ on ggml with GGUF weights, with CUDA, Vulkan and CPU backends loaded at run time. Models come from Hugging Face.

Code Quality Unit tests cover most modules of the service and audio crates. They include a test that reads the staged engine bundle’s import tables to catch missing DLLs, a check that the release can actually load. The frontend uses Vitest for its service layer. Errors are explicit and contextual, and input handling is strict: a loopback-only host check, and a multipart parser that rejects malformed results. Comments explain intent, and a guidelines file sets architectural rules. The only CI workflow in the repo renders the star-history chart, so test automation appears to run locally or at release time.

What Makes It Unique Generation is score-first, and the ABC score is a user-editable artifact rather than a hidden intermediate. Each track stores its request and audio codes, so replay is exact. Adapter slots map to the model’s two independent halves (composition and sound) with separate strengths. The MCP server goes beyond tools: it lets an agent screenshot and click the window, and the studio can use the connected agent as its own writing assistant. A LoRA training pipeline runs from a folder of songs to a usable adapter.

Self-Hosting

The studio is MIT, as is the yue2.cpp engine it bundles, so you can use, modify and redistribute it commercially. The model weights are separate. YuE2-3B and its VAE are CC BY-NC 4.0 with an individual-creator permission: personal users, creators and musicians can monetize their songs, but companies need a licence from the authors. SheetSage2 and the MuScriptor MIDI weights are CC BY-NC 4.0 without that permission. The visualiser’s spectrum looks come from audioMotion-analyzer, which is AGPL-3.0, so anyone redistributing modified builds should account for that.

It is a desktop app, not a server. It needs Windows 10 or 11 on x64 and a GPU with 6 GB of VRAM or more. NVIDIA is the fast path, AMD and Intel on Vulkan are experimental, and CPU-only works but is far slower. You carry the disk space, drivers and backups, and NVIDIA users download NVIDIA’s cuBLAS libraries once on first start. Training a LoRA needs a stronger card and extra downloads. Updates come through the built-in updater.

No licence key is required, and I found no gated features: there are no enterprise, pro or cloud directories, and a code search for licence keys only matched model-licence notices. There is no hosted or paid tier. The trade-off is community-level support with no SLA, a very small maintainer team, and Windows-only releases. Optional OpenRouter calls for the writing assistant, cover art and transcription use your own key.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers