Profile
Back to NewsBack
GitHub Trending 12 min
Reader Mode
multimodal-art-projection/YuE: YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.

multimodal-art-projection/YuE: YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.

Looking for the original YuE? Its code, documentation, and license are preserved on the YuE-v1 branch.

YuE

HKUST, M·A·P, Tokenwave.AI, NYU, Stanford, MBZUAI, NOIZ, and ACE Studio

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Compose in symbols. Create in sound.

📄 Technical report · 🎧 Demos · 🚀 Try online (free) · 🗳️ Music Arena · 📰 News · 🤗 YuE2 · 🚀 Quick start · 🤖 Agent skill · 📊 Benchmarks · 🤗 MERT2 · 🤗 SheetSage2 · 🤗 WSB · 📦 Release · Join us on Discord

YuE — GitHub Trending #1 Repository of the Day
All languages · September 14, 2026

Hugging Face Global Model Trending: reached #3 on September 17, 2026 Hugging Face Text-to-Audio Trending: reached #1 on September 20, 2026
Global: September 17, 2026 · Text-to-Audio: September 20, 2026

YuE2 brings frontier song quality to music generation with an editable composition. Give it lyrics and a style prompt: it writes a melody-and-chord plan, then realizes that plan as a complete song with vocals and accompaniment.

🎧 YuE2 needs your ears
> We're running a public blind listening study comparing YuE2 with leading proprietary music generation systems. Hear anonymous clips and choose A, B, or a tie.
> Listen & vote → · No account needed; headphones recommended.
  • Frontier quality. YuE2 is competitive with Suno v5/v6 on WildSongBench. YuE2 (best-of-8) achieves 6.9632 SongBench Avg, the highest observed mean among all evaluated settings.
  • White-box music generation through symbolic planning. Read, play, and change the composition before rendering it. Melody and chords become explicit controls that a person or an agent can inspect and edit.
  • Zero-shot covers and agentic editing. Reimagine a transcribed song in a new style, or refine a song through a conversation about its score, arrangement, and lyrics—all with the same generation checkpoint.
YuE2 song quality and text alignment on WildSongBench</a>

192 WildSongBench prompts. Both YuE2 settings use symbolic planning. Bo8 = best-of-8. The axes are normalized comparison indices; bubble area represents AudioBox production quality. Scores and evaluation protocol. Vector PDF · SVG.

News

  • 📄 September 26, 2026 — YuE2 technical report. Our technical report is now available, with the full model and training methods, automatic and expert listening results, and evaluations of score editing and zero-shot covers.
  • 🎹 September 25, 2026 — Instrumental generation and covers. The yue2-music agent skill now turns a simple description, ABC score, or reference recording into instrumental music. YuE2 writes the score by default; the skill moves the vocal melody into the instrumental part before rendering. Recommended agent: GPT-6 Astra. Download the updated skill →
  • 🚀 Try YuE2 online for free. Create a song in your browser with the NOIZ-hosted demo →. No installation required.
  • ⚡ YuE2-Turbo. NOIZ's inference and serving toolkit → accelerates YuE2 and supports concurrent requests.
  • 🎛️ ComfyUI. YuE2 has native nodes and an official text-to-music workflow →.
  • 🎬 Maestro. Maestro → includes YuE2 for local song generation, composition planning, and covers. Creator's post →

Hear what you can make

| Create | Cover | Edit with an agent | |---|---|---| | Lyrics + style → score → full song | Source recording → melody score → a new interpretation | Musical feedback → score, style, or lyric revisions → a new recording | | Listen and inspect the score | Hear zero-shot covers | Follow an editing conversation |

The agentic demo follows The Last Train through 9 steps and 14 versions, from Mandarin pop to English jazz with new harmony and a saxophone solo. Listen to each version and inspect its conversation, score, prompt, and lyrics.

How it works

!YuE2 architecture: style and lyrics become an editable score, semantic music tokens, acoustic latents, and audio

One AR–NAR Mixture-of-Transformers backbone predicts the score and semantic tokens autoregressively, then generates acoustic latents with flow matching. A VAE decodes those latents into stereo audio. Creation, covering, and editing differ in where the score comes from: YuE2, a transcribed recording, or an edited composition.

The staged Python API exposes plan() → generate_semantic() → synthesize() → decode(). See the generation guide for exact-plan reuse and decoder selection.

Quick start

Linux · Python 3.12 · NVIDIA GPU with BF16 support and 24 GB VRAM. YuE2 produces 48 kHz stereo audio without quantization. Model files download from Hugging Face on first use.

git clone https://github.com/multimodal-art-projection/YuE.git
cd YuE
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .
python examples/generate.py --output outputs/first-song

Open outputs/first-song/audio.flac. The output directory also retains the score, semantic tokens, acoustic latents, generation settings, and model identities.

The Python interface is equally short:

import json
from pathlib import Path
from yue2 import YuE2Pipeline

request = json.loads(Path("examples/song.json").read_text(encoding="utf-8")) with YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda") as pipe: song = pipe(**request) song.save_artifacts("outputs/my-song") print(song.truncated)

| Setting | Behavior | |---|---| | cot="full" | Generate an editable melody-and-chord plan; the default for new songs | | cot="melody" | Use a melody plan with free accompaniment; recommended for covers | | cot="off" | Generate directly from lyrics and style | | abc=... | Supply your own score in full or melody mode |

Generation guide · Original example inputs · v0.1.6 wheel archive

Cover a song

Transcribe a source recording with 🤗 SheetSage2, review its melody ABC, and provide new lyrics or a target style. For covers, use cot="melody" and a score without chord symbols so the accompaniment can adapt to the new style.

from pathlib import Path
from yue2 import YuE2Pipeline

with YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda") as pipe: cover = pipe( style="English, jazz-funk, warm lead vocal, Rhodes, bass and drums", lyrics=Path("cover-lyrics.txt").read_text(encoding="utf-8"), abc=Path("cover-score/score.abc").read_text(encoding="utf-8"), cot="melody", seed=42, ) cover.save_artifacts("outputs/cover")

SheetSage2 runs in a separate environment and loads its MERT2 encoder automatically. The cover guide gives the complete transcription and generation commands. An included original melody example also lets you try score-conditioned generation immediately.

Edit a composition

Export a plan, revise the musical details, and render the edited score:

import json
from pathlib import Path
from yue2 import YuE2Pipeline

request = json.loads(Path("examples/song.json").read_text(encoding="utf-8")) with YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda") as pipe: plan = pipe.plan(**request) plan.save("outputs/plan")

Copy outputs/plan/score.abc to edited.abc, then ask an agent to change its harmony, melody, tempo, or form. Supply the edited file as a new score:

python examples/generate.py --request examples/song.json \
  --abc-file edited.abc --cot full --output outputs/edited

The editable score is the white-box interface: you can inspect the intended composition and intervene on it. Editing generates a new complete recording; it does not preserve the original waveform outside an edit. Editing guide and a reproducible harmony example.

Agent skill

The yue2-music skill teaches an agent how to generate songs and instrumental music, transcribe and cover recordings, edit ABC scores, check musical invariants, and organize listening comparisons. We recommend GPT-6 Astra as the agent. YuE2 remains the model that composes the default score and generates the audio.

Use skills/yue2-music/ from this repository or download the updated skill ZIP. Install it using your agent's skill-directory or import mechanism. The song workflow uses the Python runtime installed with pip install .; the instrumental workflow includes its own pinned setup recipe for the agent to follow. The earlier v0.1.6 release remains unchanged.

For instrumental music, give the agent a simple request:

Use the yue2-music skill to create gentle piano instrumental music for reading, with no vocals. Give me the playable audio and the full prompt.

For an instrumental cover, attach a reference recording or ABC score:

Turn this melody into an acoustic-guitar instrumental cover. Keep the melody and give me the audio and full prompt.

The default flow is YuE2 score → move Vocal notes to Ins → render. Audio covers first use SheetSage2 to transcribe the reference. The agent composes a new score only when explicitly asked. The helpers preserve vocal-note pitches, onsets, and durations in the converted score and record overlapping parts; listening is still needed to check for vocal leakage and audible melody fidelity.

Try a concrete request:

Use the yue2-music skill to create an English piano-pop song. Keep the original audio and score. Make a second version with jazz harmony, preserve the vocal melody and lyric order, and give me both versions to compare.

Benchmarks

WildSongBench: 192 prompts, automatic evaluation, September 12, 2026.

| System / setting | SongBench Avg ↑ | AudioBox PQ ↑ | MuLan ↑ | PER ↓ | |---|---:|---:|---:|---:| | YuE2 (best-of-8) † | 6.9632 | 8.2714 | 0.5051 | 9.79% | | Mureka 9 | 6.9377 | 8.0226 | 0.4394 | 11.69% | | Suno v5 | 6.8721 | 8.1698 | 0.5428 | 8.10% | | YuE2 † | 6.7316 | 8.2598 | 0.5068 | 8.44% | | Suno v5.5 | 6.7150 | 8.1955 | 0.5089 | 5.96% | | Suno v4.5 | 6.6995 | 8.2541 | 0.5022 | 5.80% | | Suno v6 | 6.5562 | 8.1296 | 0.4916 | 7.58% | | Suno v6 Wild | 6.4195 | 8.1785 | 0.4999 | 7.45% | | LeVo 2 † | 6.3247 | 8.3966 | 0.3542 | 26.12% | | MiniMax Music 2.6 | 6.3222 | 8.1711 | 0.4251 | 24.55% | | MiniMax Music 3 † | 6.2830 | 8.2825 | 0.3928 | 6.27% | | HeartMuLa † | 6.2483 | 8.2933 | 0.3823 | 10.71% | | Muse † | 6.0349 | 8.0517 | 0.3937 | 33.42% | | ACE-Step 1.5 † | 6.0118 | 8.0518 | 0.4372 | 7.46% | | DiffRhythm 2 † | 5.2428 | 7.9782 | 0.3782 | 18.41% | | YuE 1 † | 4.9165 | 7.8683 | 0.2623 | 36.38% | | SongBloom † | 4.2350 | 8.1539 | 0.2697 | 19.19% |

† Publicly available model weights. All 17 evaluated settings are shown, sorted by SongBench Avg; bold values mark the best result in each column.

Both YuE2 settings use symbolic planning and the benchmark decoder, YuE2-Vae-legacy. Standard YuE2 selects from two candidates; best-of-8 selects from eight. Rankings vary by metric; the small gap between the highest means does not establish statistical significance. Full results and selection protocols.

Zero-shot covers. On 948 works, full-score YuE2 reaches 0.647 CLEWS mAP, compared with 0.006 without a score, while using the general generator without cover-specific fine-tuning. Source-identity preservation and target-style quality are measured separately; melody-only covers offer more freedom to change the arrangement. Cover evaluation.

Reproduce the benchmarks

To reproduce the reported benchmark scores, follow the instructions on 🤗 WildSongBench (WSB).

MERT2

State-of-the-art music understanding: MERT2-30s leads on 14 of 15 MARBLE metrics, and MERT2-FS (full-song) leads on 13 of 15, each against the listed external baselines. MERT2-30s reaches 91.72% genre accuracy on GTZAN.

Demo and results · 🤗 MERT2-30s · 🤗 MERT2-FS

SheetSage2

State-of-the-art audio-to-score transcription: SheetSage2-AR leads on 12 of 15 benchmark metrics in the reported comparison, including JAAH chord recognition and Rock Corpus vocal melody transcription. Vocal melody pitch-class F1 is 82.51% on RWC-Pop and 67.08% on Rock Corpus.

Demo and results · 🤗 Model and inference

Models and resources

| Resource | Purpose | |---|---| | 📄 Technical report | Model, training, and evaluation details | | 🤗 YuE2-3B | Song generation, symbolic planning, covering, and editing | | 🤗 YuE2-Vae | Default generation and listening decoder | | 🤗 YuE2-Vae-legacy | Decoder for the reported benchmark protocol | | 🤗 SheetSage2 | Audio-to-score transcription for covers and editing | | 🤗 MERT-v2-FullSong | Full-song music representations; SheetSage2's encoder | | 🤗 MERT-v2-30s | Music representations for short recordings | | 🤗 WildSongBench | Evaluation prompts and benchmark resources |

MERT2 feature extraction is optional for generation. YuE2's pipeline does not require a separate MERT2 model download. Demos and interactive results · Release downloads.

License

| Use | Terms | | --- | --- | | Personal users, content creators, and musicians | Free to use YuE2 and monetize generated outputs, with no fees or royalties payable to us. | | Academic research and education | Free for non-commercial use. | | Commercial use by companies | Contact us to discuss a commercial license for the model weights. |

We strongly encourage crediting YuE2 or using #YuE2 when sharing generated work; attribution is optional.

Responsible use. The additional creator permission prohibits illegal, harmful, deceptive, or unethical use. YuE2 is provided as is, without warranties. Users are responsible for their inputs, outputs, and use; liability limits are set out in the full terms.

Code, agent skill, and documentation: Apache 2.0. Model weights: CC BY-NC 4.0 with additional creator permission.

Copyright (c) 2026 the YuE2 authors. Third-party components and earlier releases retain their respective licenses.

Citation

Please cite the YuE2 paper when using YuE2, MERT2, SheetSage2, or WildSongBench:

@article{yuan2026yue2,
  title = {{YuE2}: Unifying Symbolic and Audio Music Generation at Frontier Quality},
  author = {Yuan, Ruibin and Pan, Jiahao and Jiang, Junyan and Wu, Zhiyue and Zhou, Ziya and Sun, Jiankai and Li, Yizhi and Zhang, Ge and Gu, Yicheng and Tian, Zeyue and Dai, Junyu and Lin, Hanfeng and Li, Kai and Wu, Shangda and Liu, Xuanjie and Wang, Jiaming and Liu, Zihan and Wang, Yue and Ma, Yinghao and Yin, Hanzhi and Chen, Kangrui and Zhang, Xinyue and Ma, Ziyang and Liao, Mengqi and Zhao, Hejia and Huang, Guowei and Yan, Chao and Ke, Lei and Yu, Jianwei and Liu, Bei and Guo, Joe and Xue, Liumeng and Xia, Gus and Xue, Wei and Guo, Yike},
  journal = {arXiv preprint arXiv:2609.33757},
  year = {2026},
  eprint = {2609.33757},
  archivePrefix = {arXiv},
  primaryClass = {eess.AS},
  url = {https://arxiv.org/abs/2609.33757}
}

For the original MERT and YuE models, please cite:

@article{li2023mert,
  title = {{MERT}: Acoustic Music Understanding Model with Large-Scale Self-supervised Training},
  author = {Li, Yizhi and Yuan, Ruibin and Zhang, Ge and Ma, Yinghao and Chen, Xingran and Yin, Hanzhi and Xiao, Chenghao and Lin, Chenghua and Ragni, Anton and Benetos, Emmanouil and Gyenge, Norbert and Dannenberg, Roger and Liu, Ruibo and Chen, Wenhu and Xia, Gus and Shi, Yemin and Huang, Wenhao and Wang, Zili and Guo, Yike and Fu, Jie},
  journal = {arXiv preprint arXiv:2306.00107},
  year = {2023},
  eprint = {2306.00107},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2306.00107}
}

@article{yuan2025yue, title = {{YuE}: Scaling Open Foundation Models for Long-Form Music Generation}, author = {Yuan, Ruibin and Lin, Hanfeng and Guo, Shuyue and Zhang, Ge and Pan, Jiahao and Zang, Yongyi and Liu, Haohe and Liang, Yiming and Ma, Wenye and Du, Xingjian and Du, Xinrun and Ye, Zhen and Zheng, Tianyu and Jiang, Zhengxuan and Ma, Yinghao and Liu, Minghao and Tian, Zeyue and Zhou, Ziya and Xue, Liumeng and Qu, Xingwei and Li, Yizhi and Wu, Shangda and Shen, Tianhao and Ma, Ziyang and Zhan, Jun and Wang, Chunhui and Wang, Yatian and Chi, Xiaowei and Zhang, Xinyue and Yang, Zhenzhu and Wang, Xiangzhou and Liu, Shansong and Mei, Lingrui and Li, Peng and Wang, Junjie and Yu, Jianwei and Pang, Guojian and Li, Xu and Wang, Zihao and Zhou, Xiaohuan and Yu, Lijun and Benetos, Emmanouil and Chen, Yong and Lin, Chenghua and Chen, Xie and Xia, Gus and Zhang, Zhaoxiang and Zhang, Chao and Chen, Wenhu and Zhou, Xinyu and Qiu, Xipeng and Dannenberg, Roger and Liu, Jiaheng and Yang, Jian and Huang, Wenhao and Xue, Wei and Tan, Xu and Guo, Yike}, journal = {arXiv preprint arXiv:2503.08638}, year = {2025}, eprint = {2503.08638}, archivePrefix = {arXiv}, url = {https://arxiv.org/abs/2503.08638} }

Contact

 WeChat
Chinese-speaking users
 Join Discord
Global users
Show WeChat QR code
YuE2 WeChat group QR code
Click to enlarge
Valid until Oct 5, 2026
Chat with me