All articles

YuE2 vs Suno: What the Benchmarks Really Show

A benchmark chart has been bouncing around AI music communities for the past week, and it makes two claims that sound almost too neat. The first is that YuE2, a free model you can run on a gaming PC, now beats Suno v6 on song quality. The second sits quietly in the same table: Suno v6 scores below Suno v5, the model it just replaced.

Here’s my short answer to the YuE2 vs Suno question. The chart is real, but it was built and published by the same team that made YuE2, using a selection method that flatters their own model. YuE2 is still the most interesting open release this field has ever seen, for reasons the chart barely captures. And that ugly Suno v6 number happens to match what thousands of paying Suno users are saying independently, which makes it the most believable line on the whole graph.

I’ve spent the past couple of years rebuilding AI tracks in a DAW, a digital audio workstation, registering AI assisted songs, arguing with distributors, and testing every major generator as new versions land. Let me walk you through what this benchmark actually measures, what it hides, and what it means if you plan to release music rather than just generate it.

The Chart Everyone Keeps Sharing

The numbers come from WildSongBench, an evaluation the YuE2 team ran on 192 prompts across 17 model settings, dated September 12, 2026. Every song is scored automatically, with no human listeners involved, across nine metrics. The headline column is called SongBench Avg, a blend of seven judged dimensions rolled into a single number.

On that column, the top of the table reads like a shock result. YuE2 in its best of eight setting scores 6.9632, the highest of any system tested, open or commercial. Mureka 9 follows at 6.9377, then Suno v5 at 6.8721. A standard single pass of YuE2 lands at 6.7316.

Suno v5.5 sits at 6.7150 and Suno v4.5 at 6.6995. Then comes Suno v6, the brand new flagship, at 6.5562, with its looser v6 Wild variant at 6.4195. Down at the bottom, the original YuE 1 scores 4.9165, which tells you how far this line of models has traveled in under two years.

Read that middle section again slowly. On this table, every older Suno model beats the new one, and a free download you run at home beats them all.

The other columns tell a messier story, which is your first clue that one number can’t summarize a song. Suno v4.5 posts the lowest PER, short for phoneme error rate, a measure of how clearly the sung words match the written lyrics, at 5.80 percent. LeVo 2 wins AudioBox PQ, an automated estimate of production quality, at 8.3966, and Suno v5 takes the best MuLan score, which gauges how closely the audio matches the style you asked for. Sort by a different column and a different system claims the crown.

The scatter plot version that spreads on social media plots song quality against text alignment, with bubble size showing AudioBox PQ and black outlines marking the Pareto optimal points, the systems you can’t beat on both axes at once. It looks authoritative. It’s also just a picture of the same self-run test.

Who Built the Benchmark, and How the Winner Got Picked

Now for the part most reposts leave out: WildSongBench was created and published by the YuE2 team itself. That doesn’t make the numbers fake. It makes them marketing, the way nearly every vendor benchmark in AI has been marketing since the beginning.

Then there’s the winner’s asterisk. The chart topping 6.9632 comes from a best of eight setting, labeled Bo8 on the chart: the system generates eight complete candidates for each prompt and keeps the one that scores best on musicality, then prompt adherence, then lyric clarity. The single pass number, 6.7316, is the honest like-for-like figure, and it lands above Suno v6 but below Suno v5. As Mervin Praison put it in his breakdown of the release, the headline of beating Suno v6 holds for a single cheap generation, while matching Suno v5 takes eight times the compute.

Even the baselines weren’t sampled equally. YuE2’s own project page states that YuE2, the open baselines, and the two Suno v6 variants were each run with two candidates and the clearer one kept, while the older commercial systems retained whatever candidates they originally delivered. Different selection budgets, same table.

To the team’s credit, they say this openly, and the notes on GitHub even concede that the small gap between the top scores doesn’t establish statistical significance. That sentence should be printed across the chart. It never travels with the screenshot.

And there’s one thing automatic judges cannot do, which is listen like a person. The team seems to know it, because the model card links a blind listening arena where anyone can vote on anonymous clips. Six months of that data will be worth more than any launch week table.

So when someone in a forum calls this a trust me bro benchmark, they’re roughly right. And when the same table gets shared as proof that Suno is finished, that’s wrong in the other direction. Both reactions mistake a marketing document for a measurement.

What YuE2 Actually Is

Strip away the chart drama and YuE2 is still a genuinely new kind of music model. It comes from Multimodal Art Projection, the research collective behind the original YuE, and it’s a roughly 3.6 billion parameter model with open weights you download from Hugging Face and run on your own machine.

The design choice that matters is this: YuE2 writes the song before it renders the song. Given a style prompt and lyrics, it first composes a symbolic plan, melody and chords written out in ABC notation, a plain text format musicians have used for decades to jot down tunes. You can open that plan, read it, change a chord, rewrite a phrase of melody, and only then render the audio.

There are three planning modes, full melody and chords, melody only, or no plan at all when you want chaos. Every other mainstream generator hands you a finished waveform and a reroll button. Here the composition is a document you can edit, a different philosophy entirely.

The same trick powers its covers. Feed it a reference recording and a companion model called SheetSage2 transcribes the melody into that same editable score, which YuE2 then re-renders in whatever style you describe. On the team’s own cover benchmark, conditioning on the full score preserves the original song’s identity dramatically better than generating without one. This is the feature early adopters rave about most, and I’ll come back to why it’s also the feature with the sharpest legal edges.

The practical picture, drawn from the official documentation and early hands-on reports: output is 48kHz stereo, saved as FLAC, a lossless audio format, rather than the compressed files most cloud tools hand you by default. The ComfyUI integration supports songs up to fifteen minutes. The documented setup is Linux, Python, and a 24GB NVIDIA graphics card, though the team’s own measurements show the full pipeline peaking around 11GB of VRAM, the working memory on a graphics card, and the tester at MindStudio measured usage under 8GB, so a 16GB card is comfortable in practice. That same MindStudio test clocked a roughly two minute song rendering in about two minutes on a high end card.

None of this is one click software yet. You’ll need patience with a terminal or with ComfyUI’s node graphs, and that filter alone will keep most casual users on cloud services for now.

What It Sounds Like, Honestly

Here’s my reality check, because the benchmark argument dissolves the moment you press play. The listening reaction across the communities splits into two camps, and both are telling the truth.

Camp one says YuE2 sounds clean but simple. Arrangements default to safe, sparse choices, closer to a competent demo than a finished record, and plenty of experienced Suno users who went to audition it came back unimpressed, with comparisons to elevator music making the rounds. Camp two says the simplicity is a default, not a ceiling, that the model rewards score editing and careful prompting, and that it’s weeks old and nobody has learned its tricks yet.

My own take after living with these tools: the fidelity is real, the vocals are better than an open model has any right to have, and the arrangement intelligence still trails what Suno v5 could do on a good day. Automatic metrics don’t hear boring. People do.

Covers are the exception. Handing it a reference song and hearing the melody survive into a completely different production is genuinely startling, and it’s the one capability no cloud service will ever match, because no platform lawyer would allow it. There’s also a longer game starting: since the weights are open, communities are already discussing fine tuning lightweight adapters on their own catalogs so the model learns a specific artist’s sound. Subscription services offer custom model features too, but you’ll never hold those weights in your hand.

The License Is Stranger and Better Than It Looks

Everyone calls YuE2 open source, and the pedants pushing back have a point. The model weights ship under CC BY-NC 4.0, a Creative Commons license that forbids commercial use, which by the strict definition means the model is not open source at all. But the team layered a specific permission on top, and it changes everything for musicians.

According to the terms published on the project’s GitHub page, individual creators are free to use YuE2 and monetize the songs it generates, with no fees or royalties owed to the authors. Companies that want to build the model weights into a commercial product must contact the team for a license. The surrounding code ships separately under Apache 2.0, a permissive software license.

In plain words: you, a musician, can generate a song tonight and send it to Spotify tomorrow, and the YuE2 team asks for nothing. A startup, meanwhile, can’t quietly wrap the model in a paid app. As a working producer I find that the most creator friendly reading of noncommercial I’ve seen shipped with any serious model.

The training data story matters just as much. The team states the model was trained primarily on public domain CC0 music plus synthetic data licensed from a provider called Tokenwave.AI, around 346,000 hours for the main model. Whether that claim survives outside scrutiny is a fair question, but notice the shape of it: this is a training recipe designed to be defensible, announced in the middle of an industry still suing over scraped catalogs.

Now the part that trips up almost everyone, on every tool: permission to monetize is not ownership. Under the US Copyright Office’s January 2025 report on AI and copyrightability, material generated entirely by an AI system is not eligible for copyright protection, and prompts alone don’t count as authorship no matter how detailed they get. What can be protected is the human layer: lyrics you wrote, real edits and arrangements you made, the creative selection and shaping of AI material into a larger work. YuE2’s editable score is quietly interesting here, because rewriting the melody and harmony yourself inside the ABC plan is a far stronger authorship story than pressing generate five times, though none of this has been fully tested in court and none of it is legal advice.

The uncomfortable flip side stays true as well. A raw AI output you release is something anyone else can copy without infringing your rights, because you don’t hold rights in it. If a track starts to matter commercially, the cleanest fix is still the old fashioned one: have real musicians re-record and produce the song, so the finished master is a human performance you actually own and can register. It costs more than a download, and it’s the only version of ownership that doesn’t arrive with an asterisk.

Suno’s v6 Problem Is the Real Story Here

While the open model crowd argues about benchmarks, Suno is having the roughest launch in its history, and that context is what makes the WildSongBench table land so hard.

Suno shipped the v6 family on September 9, 2026: v6, a looser v6-wild, and a faster v6-mini available on every plan, with generations up to eight minutes. As Jack Righteous documents in his hands-on guide to the v6 family, all older models were retired for new creation the same day, so v4.5, v5, and v5.5 are gone as options and your old songs simply remain playable in your library. On paper the feature set is genuinely impressive, with plain language section edits, mashups built from multiple sources, and a sample and isolate workflow, per Suno’s own release notes.

v6 is also the first generation built in Suno’s settlement era. Warner settled its lawsuit and signed a licensing partnership in November 2025, a deal Rolling Stone described as a pact for next generation licensed AI music, and BMG followed with a global agreement in August 2026. MusicRadar reports the v6 suite was developed with Warner, BMG, and Believe, trained on licensed catalogs and user data rather than, in Suno’s earlier description of its sources, “all music files of reasonable quality that are accessible on the open internet”. Universal and Sony are still litigating, and they’ve moved to expand the case to more than 61,000 recordings, which pushes the theoretical damages past nine billion dollars.

The user response to v6 has been brutal. Digital Music News rounded up subscribers describing the output as “muffled, dull, and strangely lifeless”, and the recurring technical complaints are smeared or missing high end, buried vocals, heavy compression, and synths with a metallic, artificial edge. MusicRadar covered a heavily upvoted cancellation thread and notes the platform serves more than two million paid subscribers, which makes the scale of the revolt commercially meaningful.

In fairness, the reception is genuinely split: some producers, especially in dense, distorted genres where the model’s rough edges hide, report keepers and even better covers than v5.5 gave them. But when a self-interested benchmark from a rival team and Suno’s own paying customers independently agree that v6 sits below v5, I stop treating either signal as noise.

Nobody outside Suno knows why. The popular theory is training data: a licensed only dataset is far smaller and less diverse than the open internet scrape behind earlier models, and people believe they can hear the shrinkage. That remains speculation, and I’ll flag it as exactly that, but it’s the kind of speculation the evidence keeps agreeing with.

The launch also came bundled with policy changes that matter more than the model itself. In August 2026 Suno rolled out audio watermarking and fingerprinting of its outputs plus caps on downloads, moves Music Business Worldwide connects directly to the label partnerships and to reining in mass distribution. Pricing itself is unchanged, Pro at $10 a month and Premier at $30, or $8 and $24 on annual billing, with a standard generation producing two songs for 10 credits.

What changed is the deal around the price. You now rent access to a model that can be swapped out from under you, delivering watermarked files through metered downloads. That, more than any benchmark score, is what radicalized so many Suno power users toward local models this month.

Where Running Locally Actually Wins

Set the quality debate aside for a moment, because the structural advantages of a local model are hard to argue with. Nothing gets retired. The exact weights on your drive today will make the same music in ten years, which sounds trivial right up until the week your favorite cloud model vanishes mid project, which is precisely what just happened to every v5.5 workflow in existence.

Nothing is metered. Once the GPU is paid for, generation four thousand costs the same as generation one, a little electricity. Files come out at 48kHz, lossless, unwatermarked, and unlimited, which is real headroom when you’re pulling a track into a DAW for proper mixing.

Your unreleased demos never leave your machine, which matters the moment you start feeding reference audio into a cover workflow. And there’s no cloud side content filter rejecting an upload for reasons nobody will explain.

Be honest about the other column too. Suno works in a browser and on a phone within thirty seconds of signing up. It wraps the model in genuinely useful tools, stems, section replacement, personas, a full production suite on the top plan. Its best outputs are still more radio shaped than anything I’ve coaxed out of an open model.

And a capable GPU is a real cost: if you don’t already own one, several hundred dollars up front against $10 a month is not a subtle comparison. Local wins on freedom, privacy, and permanence. Suno still wins on convenience and, for now, on polish.

The Mass Upload Dream Is Already Dead

Every conversation about free unlimited generation ends at the same tempting math: if songs cost nothing, why not upload thousands and farm streaming pennies. Because the pipes are closing, fast, and this is the section I’d staple to the top of every AI music forum.

The scale is real. According to figures Deezer published in July 2026, fully AI generated tracks now make up more than half of all daily uploads to the platform, about 90,000 songs a day, up from roughly ten percent at the start of 2025. Now the number that should end the fantasy: all of that flood earns just 1 to 3 percent of total streams, Deezer demonetizes the 85 percent of detected AI streams it classifies as fraudulent, strips AI tagged tracks out of its recommendations, and has begun deleting AI tracks nobody streams.

Spotify moved in September 2025 with an AI policy update that pairs a spam filter aimed at mass uploads and duplicates with standardized AI disclosures in credits through a system called DDEX, plus a ban on unauthorized voice clones, and TechCrunch reports the platform had already removed some 75 million spam tracks. Suno itself now watermarks, fingerprints, and caps what leaves its servers.

Then, five days before I wrote this, Universal Music Group sued DistroKid, the biggest independent distributor in the US market, accusing it of running an “AI slop pipeline” and seeking statutory damages of up to $150,000 per infringed work. Read the complaint’s framing carefully, because it tells you where the line is being drawn: UMG says the case is not about AI music that’s clearly disclosed as AI, it’s about undisclosed machine output masquerading as artist backed releases. Every distributor watched that filing land. Expect tighter screening at DistroKid, CD Baby, and TuneCore, not looser.

So the honest strategic read is this: volume is a dying strategy operating in a shrinking window, and the companies you’d depend on to execute it are the ones getting sued. What remains is the boring opportunity, fewer songs, made better, with rights you can actually document.

Five Questions Before You Pick a Side

Ask whether you want an appliance or an instrument. If you want finished sounding songs with zero setup, Suno remains the appliance, and a free account plus one $10 Pro month will tell you quickly whether v6 suits your genre better than the angriest threads suggest. If you want a model you can open up, edit at the score level, and keep forever, YuE2 is the first open release worth the friction.

Ask what your hardware situation is. A 16GB or 24GB NVIDIA card and some tolerance for Linux puts YuE2 fully within reach, and the ComfyUI templates soften the setup considerably. No GPU means cloud tools by default, and there Mureka deserves a real look alongside Suno: the version this benchmark tested, Mureka 9, was the closest commercial score to the top, and it has already been superseded by version 9.5, which shipped at the end of August 2026 with a push toward more natural, human sounding arrangements.

Ask whether covers are the point. If yes, YuE2’s score based covers are the strongest tool I’ve seen, and the legal work stays yours. Releasing a cover in the US still requires a mechanical license, the permission to reproduce someone else’s composition in your own recording, which distributors arrange cheaply, DistroKid handles it for about $12 a year per song. That license doesn’t let you rewrite lyrics or alter the fundamental character of the composition, and it does nothing to authorize cloning a recognizable singer’s voice, so cover the song, never the singer.

Ask what happens if a track starts to matter. If a release is headed for sync pitches, a label conversation, or anywhere ownership gets audited, remember that a raw AI master from any tool carries no copyright of its own. The human contributions are what count, and rebuilding the track with real performances, whether you replay every part yourself or bring in players to re-record it, is what turns a generated demo into a master someone can actually own. I said it earlier and it bears one repeat, because it’s the question clients ask me most.

Ask, finally, whether you’d bet against the trend line. Not long ago the open models were toys, and the 4.9 the original YuE scores on this very benchmark is the receipt. Today a single free generation outscores the newest Suno on a friendly test and lands within earshot of everything else.

Model weights don’t regress, they accumulate, and they can’t be retired by a settlement. Whatever the YuE2 vs Suno scoreboard says next quarter, the direction of travel only points one way, and for the first time the most interesting music model available is one that nobody can take away from you.

Sources

Common questions

Is YuE2 actually better than Suno?

On the benchmark its own team published, a single YuE2 generation scores above Suno v6 and below Suno v5, and only the best of eight setting tops the whole table. Most listeners still describe Suno's output as more polished and YuE2's as cleaner but simpler. Call it frontier level, not a clear winner.

Who created the WildSongBench benchmark?

The same team that built YuE2, the research collective Multimodal Art Projection. The scores are automatic, the selection method favors their model, and the team itself notes the small gaps at the top are not statistically significant. Treat it as a strong marketing document rather than an independent test.

Can I sell or stream music I make with YuE2?

Yes. The weights carry a Creative Commons noncommercial license, but the authors added explicit permission for individual creators and musicians to monetize generated songs with no fees or royalties. Companies that want to build products on the model weights must negotiate a commercial license.

Do I own the copyright to a song made with YuE2 or Suno?

Not the raw output. The US Copyright Office says purely AI generated material is not copyrightable and that prompts alone do not count as authorship. Your own lyrics, real edits and arrangements, and human performances can be protected, which is why rebuilding a track with real players creates ownership the raw file never has.

What computer do I need to run YuE2?

The official setup is Linux, Python, and a 24 GB NVIDIA graphics card, though the team's measurements show the full pipeline peaking around 11 GB of video memory. In practice people run it comfortably on 16 GB cards. There is also native ComfyUI support if you prefer a visual workflow.

Why does Suno v6 score lower than Suno v5?

Nobody outside Suno knows for certain. The v6 family is the first built after the label deals, reportedly trained on licensed catalogs and user data instead of a broad internet scrape, and many users hear a duller, more compressed sound. The benchmark result matches those complaints, but the training data theory remains speculation.

Can YuE2 make covers of existing songs?

Yes, and it is the standout feature. A companion model transcribes a reference recording into an editable score, and YuE2 renders a new version that keeps the melody intact. Releasing a cover in the US still requires a mechanical license, which distributors arrange for a small yearly fee, and cloning a recognizable voice is never covered.

Is it worth mass uploading AI songs to streaming platforms?

No. AI tracks are now more than half of daily uploads on Deezer yet earn only 1 to 3 percent of streams, and most detected AI streams get flagged as fraud and demonetized. Spotify runs a spam filter, Suno watermarks and caps downloads, and Universal is suing DistroKid over exactly this kind of pipeline.

What audio quality does YuE2 produce?

It renders 48 kHz stereo audio and saves lossless FLAC files with no watermark and no download limits. That gives you a little more headroom than typical cloud exports when you take a track into mixing and mastering. The quality of the arrangement itself still depends on your prompting and score edits.

Is YuE2 open source?

Not by the strict definition, because the model weights forbid commercial use of the weights themselves. The surrounding code is under a permissive license, and individual musicians are explicitly allowed to monetize the songs they generate. Open weights with a creator carve out is the accurate description.