YuE2 Studio
Windows desktop studio that runs the YuE2 song model on your own NVIDIA GPU: it writes an editable sheet-music score first, then sings it as a full song with vocals, offline, with covers, LoRA training and an MCP server for agents.
Repository Health
Technical Analysis
Dependency Health
YuE2 Studio is a native Windows desktop app for YuE2, the open song model from M-A-P that writes a score before it sings. You describe a style and write lyrics. The model composes a melody with chords in ABC notation, the studio engraves it as sheet music, and then the model performs it as a full song with vocals of up to six minutes, entirely on your own graphics card.
Because the score comes back with the track, you can edit notes, tempo or key and render the same composition again with a different sound. You can also re-render from saved audio codes with other steps or seeds. Cover mode uses SheetSage2 to transcribe a recording’s melody, and YuE2 then sings it in your style. Around generation sits a full studio: a library, stem separation, audio-to-MIDI, word-level karaoke, audio processing with VST3 and mastering, video clips, and a player with an equalizer, MilkDrop visualiser and a skinned Winamp mode.
It ships as one Tauri executable with a Rust service inside and a bundled C++/CUDA engine from yue2.cpp, so there is no Python or Node.js in the runtime path. NVIDIA cards run on CUDA, AMD and Intel are experimental on Vulkan, and models download on first use. While it is open, it also serves an MCP endpoint so agents such as Claude Code can drive everything the interface does.
What You Get
- Full songs with vocals, up to six minutes, from a style and lyrics
- An editable sheet-music score that comes back with every track
- A Windows installer with auto-update, or a portable folder that holds all its data
- A local library with playlists, cover art and MP3 export with ID3 tags
- An interface in English, Russian, Chinese, Japanese and Korean
- A local MCP endpoint for agents
Common Use Cases
- Sketching songs from lyrics and a style description
- Covering existing melodies in new styles
- Producing stems and MIDI from any track
- Training a LoRA on a particular artist’s style
Under The Hood
Architecture A layered, process-supervised design. A React interface talks over local HTTP to a Rust service compiled into the desktop shell. The service owns the library, the job queue, downloads and result import, and supervises a separate native inference engine bound only to loopback. Finished songs are imported by the service itself, so a reload or closed window never loses a result. Progress is derived from the engine’s own log. Engines are treated as adapters behind capability-level provider selection (local or OpenRouter), so no interface code hardcodes a model. Child processes are grouped so none outlive the studio. The same service code backs the UI and the MCP endpoint, plus a bridge that lets an agent see and operate the window.
Tech Stack Rust with Axum and Tokio for the service, Tauri for the desktop shell, and an NSIS installer with a signed auto-updater. The UI is React and TypeScript on Vite with Tailwind, abcjs for engraving, Webamp, Butterchurn, audioMotion and an in-browser FFmpeg. Storage is bundled SQLite plus plain files. Audio work uses Symphonia for decoding, LAME for MP3 and a dedicated Rust DSP crate. ONNX Runtime is loaded dynamically for recognition models. The engine is C++ on ggml with GGUF weights, with CUDA, Vulkan and CPU backends loaded at run time. Models come from Hugging Face.
Code Quality Unit tests cover most modules of the service and audio crates. They include a test that reads the staged engine bundle’s import tables to catch missing DLLs, a check that the release can actually load. The frontend uses Vitest for its service layer. Errors are explicit and contextual, and input handling is strict: a loopback-only host check, and a multipart parser that rejects malformed results. Comments explain intent, and a guidelines file sets architectural rules. The only CI workflow in the repo renders the star-history chart, so test automation appears to run locally or at release time.
What Makes It Unique Generation is score-first, and the ABC score is a user-editable artifact rather than a hidden intermediate. Each track stores its request and audio codes, so replay is exact. Adapter slots map to the model’s two independent halves (composition and sound) with separate strengths. The MCP server goes beyond tools: it lets an agent screenshot and click the window, and the studio can use the connected agent as its own writing assistant. A LoRA training pipeline runs from a folder of songs to a usable adapter.
Self-Hosting
The studio is MIT, as is the yue2.cpp engine it bundles, so you can use, modify and redistribute it commercially. The model weights are separate. YuE2-3B and its VAE are CC BY-NC 4.0 with an individual-creator permission: personal users, creators and musicians can monetize their songs, but companies need a licence from the authors. SheetSage2 and the MuScriptor MIDI weights are CC BY-NC 4.0 without that permission. The visualiser’s spectrum looks come from audioMotion-analyzer, which is AGPL-3.0, so anyone redistributing modified builds should account for that.
It is a desktop app, not a server. It needs Windows 10 or 11 on x64 and a GPU with 6 GB of VRAM or more. NVIDIA is the fast path, AMD and Intel on Vulkan are experimental, and CPU-only works but is far slower. You carry the disk space, drivers and backups, and NVIDIA users download NVIDIA’s cuBLAS libraries once on first start. Training a LoRA needs a stronger card and extra downloads. Updates come through the built-in updater.
No licence key is required, and I found no gated features: there are no enterprise, pro or cloud directories, and a code search for licence keys only matched model-licence notices. There is no hosted or paid tier. The trade-off is community-level support with no SLA, a very small maintainer team, and Windows-only releases. Optional OpenRouter calls for the writing assistant, cover art and transcription use your own key.
Related Apps
GitNexus
AI Code Assistants · Developer Tools · Mcp
Index any codebase into an interactive knowledge graph and give your AI agents deep architectural context via MCP — with zero servers required.
fish-speech
AI Development · Developer Tools · Music Audio
SOTA open-source dual-autoregressive text-to-speech model with rapid voice cloning, inline emotion tags, and real-time streaming inference across 80+ languages.
agentmemory
AI Agents · AI Memory · Mcp
Persistent memory for AI coding agents — built on the iii engine, it works across Claude Code, GitHub Copilot CLI, Cursor, Gemini CLI, Codex CLI, and any MCP client, so agents stop needing everything re-explained every session.