Suno v6: generation results, and what Variety and Max Mode actually do … v6 is weak. Is this a copyright move?

2026-09-25 04:43 (2 hours ago)

Suno v6: generation results, and what Variety and Max Mode actually do … v6 is weak. Is this a copyright move?

Suno shipped v6 on 2026-09-09. At the same time every older model, v5 included, disappeared from the model picker.

When v5 came out it was genuinely striking, and although I came to feel the limits of how much information a tag could carry, I expected v6 to be something like "the AI that makes all of modern popular music obsolete."

I even wrote a song about that frustration and that expectation (More Tag).

Then v6 actually shipped, and my first impression from using it was that the musical interest had broken down. It felt below my expectations, or rather below v5. There is no sense of a major-label sound.

Looking around online, the same reaction showed up worldwide at the same time: the sound is muffled, like a blanket over the speakers. MusicRadar ran "muffled, generic, soulless garbage" and MusicTech ran "The death of Suno". On the other hand, an engineer with 30 years of experience who holds a Recording Academy vote compared v5.5 and v6 on the same lyrics and said the sound goes to v6.

This blog, ytyng.com, writes lyrics for each article and has Suno make the background music. Which means I still have songs generated before v6 on hand. Now that the old models are gone, this is a comparison only people who kept their output can make.

The songs I generated

Style prompts

Three Drawers

100BPM, french electro pop, hyperpulse, artcore, glitch, sci-fi understones, lo-fi beat, ruspy tone, 808 sub base, atomospheric synths, haunting, stacked vocals on chorus part

The distinctive tags here are french electro pop, hyperpulse and glitch; the intent is electronic pop that carries a kind of digital chaos.

Among the songs I made with v5, I was aiming for something close to Frozen Opacity (article), Turtle Badge (article), Antimatter (Suno), Undersong (Suno).

Credential Cereal

160BPM, alternative rock, math rock, solid guitar cutting, guittar arrpeggio, warm saturation, garage band energy, stacked vocals on chorus part, digital effects

Here math rock is the distinctive tag; the intent is a solid rock sound with intricate guitar cutting.

Among the songs I made with v5, I was aiming for something close to Terabyte Whispers (article) and Byoushin (Suno).

Generation patterns

Everything was fixed at Weirdness = 20% and Style Influence = 80%. On top of that I generated with Variety both Off and Max, and with Max Mode both On and Off.

What Variety is

A feature where, at generation time, the AI reinterprets and rewrites your style prompt before generating.

This is not my guess; the official v6 FAQ says so.

The Variety slider is designed to introduce variety in your outputs by adjusting and updating your style prompts.

If you'd like to retain full control of your style tags, reduce the Variety slider to 0.

So Variety is not a feature that widens the randomness of generation; it rewrites the style prompt you wrote. Only at 0 do your own tags pass through as written.

Suno produces two songs per generation, and the prompt differs slightly between the two.

For the Variety On cases I generated several times, listened to the results myself, and picked one per pattern — the one that sounded most acceptable to me — to publish. That choice is subjective.

Rewrite examples

100BPM, french electro pop, hyperpulse, artcore, glitch, sci-fi understones, lo-fi beat, ruspy tone, 808 sub base, atomospheric synths, haunting, stacked vocals on chorus part

Feeding in that style and generating with Variety = Max turns it into styles like these:

French electro pop, artcore: clipped electric guitar and ticket-stamp percussion over 808 sub base and atmospheric synths, 100BPM hyperpulse lo-fi beat, female ruspy-tone vocals with two-part harmony and stacked chorus layers, glitch processing, sci-fi undertones, distorted tape dropouts
French electro pop, female ruspy two-part vocal harmony with stacked chorus vocals, glitchy lo-fi production and sci-fi undertones, 808 sub base, atmospheric synths, clipped electric guitar, ticket-stamp percussion, haunting 100BPM hyperpulse groove
French electro pop, artcore at 100BPM with glitch hyperpulse and lo-fi tape distortion; female vocals in a raspy tone, close-mic verses with stacked harmonies and two-part chorus doubling; 808 sub bass, atmospheric synths, clipped electric guitar, ticket-stamp percussion; haunting half-time pulse with clipped electronic transitions
French electro pop, female vocal with a ruspy tone, close-mic verses and stacked harmonies in the chorus, hyperpulse glitch production with lo-fi tape grain, 100BPM steady midtempo pulse, 808 sub base, atmospheric synths, clipped electric guitar, ticket-stamp percussion, sci-fi undertones and haunting reverb tails

And for this style:

160BPM, alternative rock, math rock, solid guitar cutting, guittar arrpeggio, warm saturation, garage band energy, stacked vocals on chorus part, digital effects

it becomes styles like these:

Alternative rock, math rock at 160BPM with solid cutting guitar and interlocking arpeggios, male vocal driving tight verses and stacked harmonized chorus vocals in two-part harmony, warm saturation, digital effects, garage-band punch, handclaps, cash-register chimes, brass glissando
Alternative rock with math rock precision: solid cutting guitar and interlocking guitar arpeggios, male two-part vocal harmony with stacked harmonized vocals, warm saturation and digital effects, driving 160BPM garage-band pulse
Alternative rock, male lead with clipped talk-sung verses and stacked two-part chorus harmonies; solid cutting guitars interlocking with arpeggios, handclaps, cash-register chimes, and brass glissando; driving 160BPM math-rock pulse with garage-band punch; warm saturation, digital effects, tight vocal doubles, and a wide final-chorus lift
Alternative rock, math rock; male English vocals with clipped verse delivery and stacked two-part harmonized choruses; warm saturation, digital effects, cash-register chimes, handclaps, brass glissando; driving 160BPM garage-band pulse; solid cutting guitar interlocked with precise arpeggios, syncopated math-rock riffs, tight drums, and a full-power final chorus

How the rewriting behaves

With these prompts, the rewriting consistently made things longer. A 10-to-12-word tag list swells into a 30-to-50-word prose instruction. And the part that got added contains elements I never wrote.

In the Three Drawers case, clipped electric guitar, ticket-stamp percussion, distorted tape dropouts, close-mic verses and half-time pulse are not in my input. A vocal gender, female, was added on its own. On the Credential Cereal side, handclaps, cash-register chimes, brass glissando and male were added.

It appears to infer the arrangement from the lyrics (Three Drawers is a song about drawers and a ticket booth, Credential Cereal is a song about a supermarket), so I can see the intent. But it also means control over the style specification moves away from me.

On a Korean Suno board there is a test claiming roughly 1,000 attempts at this rewriting (DCinside). It is an anonymous post so I cannot verify it, but the reported tendencies line up with the rewrites above.

  • The rewriting priority is genre names > instrument names > general words. The more general the word, the more likely it gets replaced
  • Writing Rock with Variety at Max drops the share of electric guitar by nearly half and mixes in synths and harmonica
  • Feeding in Japanese lyrics produces a Japanese style prompt, which can lower quality
  • The recommendation is: Normal to Max while exploring, OFF for the final take

English-language sources also report that the rewritten prompt comes back shorter and weaker than what you wrote (roo.beehiiv.com). That is a blog compiling Reddit posts, and it does not show the original post URLs. In my cases it got longer instead, so it does not seem to shorten consistently.

What Max Mode is

An option that spends more compute on a single generation. The official v6 FAQ describes it like this.

Max Mode is an option you can turn on for any generation when you want v6 to spend more on getting it right. It costs more credits and it's best for: songs longer than two minutes, covers where you want the result to stay close to the original, transferring the style of one song onto another, and keeping vocals and style consistent through the whole track.

What stands out is that every use case Suno lists is about consistency and fidelity.

  • Songs longer than two minutes
  • Covers you want to stay close to the original
  • Transferring the style of one song onto another
  • Keeping vocals and style consistent through the whole track

It does not say this is a feature that improves sound quality. It is positioned as a feature for not missing, not a feature for being better.

Credit consumption doubles. My subjective impression, with v6-wild, was that Max Mode = Off gave results even weaker than v5, while Max Mode = On gave results close to v5.

Differences between the models

All three models go up to 8 minutes per generation. v6 and v6-wild are Pro / Premier only; v6-mini is available on every plan including the free tier (official help).

v6

Suno describes it in two places, with slightly different wording.

Official blog:

reliable, precise and consistently delivers polished music across every genre and style

Official help:

Balanced, expressive, and built for a wide range of styles and prompts. Great for everyday creation.

reliable, precise, consistently, balanced. Every one of these points at not missing, and there is not a single word pointing at "you can make an amazing song."

Other people's assessments concentrate on the same blandness. MusicRadar collected the voices of users announcing cancellations under the headline "muffled, generic, soulless garbage". What the article repeats is that the sound is muffled and dull and strangely lifeless, the vocals get buried, and the high end is almost gone. MusicTech's headline is "The death of Suno". On prompt adherence, reports that tempo, genre, arrangement and duet role assignments get missed have been organized as multiple corroborating accounts (jackrighteous.com).

There is a positive side too. An engineer with 30 years of experience who holds a vote in the Recording Academy's Producers & Engineers Wing ran the same lyrics and prompt through v5.5 and the three v6 models and concluded that the sound itself goes to v6 (warmer, wider, stronger low end, more depth in the mix). He also said that people who hear a smear on the vocals are not wrong, and that as an engineer he hears it too (test video). It is one non-blind review, so no general law follows from it.

German-language media rated the technical side positively, writing that it understands musical concepts such as instrumentation, vocals and atmosphere far better than the previous version, and calling it polished, radio-compatible (basic-tutorials).

I think what is happening here is not that opinion is split, but that people are looking at different axes. On "is this finished as a recording," v6 can win. On "is this interesting to listen to," it loses.

My own assessment is the latter. Overall it feels like it deviates from pleasing music theory in a bad way. I do not get much sense of commercial viability, of a major-label sound, of the well-trodden path; it keeps disappointing me in the bad direction.

It feels like it has been trained on a mass of casually made songs by people who did not stake their lives on music.

It does not feel usable at all.

Suno does not publish the composition of its training data. More broadly, a 2024 audit of AI music training data reported that about 86% of the data hours concentrate on music of the Global North, and that guitar, piano and drums appear in 52-67% of training clips while sitar, tabla, accordion and bagpipes together account for under 3%. Niche genres being weak is not specific to v6; it is a structural tendency across this whole field.

v6-wild

Official blog:

less predictable and more varied, producing unexpected, textured and ambitious results

Official help:

A more experimental take on v6. Leans into unexpected choices, genre-blending, and results that push further from the prompt.

unexpected, genre-blending, push further from the prompt. It is positioned as the model for chasing creative accidents.

Other people's assessments are split. A blog compiling Reddit posts relays a launch-day tester's report that roughly half of their v6-wild generations came out broken and unusable (roo.beehiiv.com). A Taiwanese personal blog writes that wild distorts partway through, so you should listen to the whole thing (jiuliyue.com). There are posts reporting good results with avant-garde and glitch jazz, while criticism of artifacts and repetitive generation is also on record (jackrighteous.com).

Incidentally, in the WildSongBench scores discussed below, v6-wild is lower than v6 (6.4195 against 6.5562).

Subjectively I feel v6-wild is the model closest to v5, and it is the easier one to work with. v5.6, maybe?

I could not find anyone else making this assessment. Suno explains it as a separate model, and I have no material that overturns that. This is only my subjective impression.

Officially the line is that it suits emergent results, but that is not my image of it; coming from v5, it is the one that feels reassuring.

Or rather, it feels slightly weaker than the v5 line. So: v5 minus.

For what it is worth, v5 and v5.5 can no longer be selected, so perhaps they retrained on safe material only, taking out anything problematic, out of consideration for commercial music licensing. Thinking of it that way, quite a lot falls into place.

I put in tags that were distinctive on v5, like math rock, hyperpulse and glitch, and they do not fire well, so it feels like there is less information — fewer tags that actually manifest — than v5.

I find v6-wild closer to my taste than v6, but it is weaker than v5, and I feel it will take more time to use it well. Or rather, I do not feel like using it.

v6-mini

Official blog:

faster, more efficient version available to everyone … delivers better, faster results than any free model

Official help:

A lighter, faster variant. Ideal for quick ideas, high-volume generation, and exploring concepts before committing to a full generation.

It is positioned as the only v6-line model available on the free tier, and Suno itself writes that it is for exploring ideas before the real take. It is not sold as the model that wins on quality.

The assessments roughly follow that line. A Taiwanese personal blog calls it bland trans-pop (jiuliyue.com). On a Korean Suno board there was an assessment that it has a plain lo-fi quality worth using deliberately (DCinside). On the other hand there are reports of more artifacts than the paid flagship, and also reports that it unexpectedly produces keepers and works well for remixes (jackrighteous.com).

My own impression is that it makes sense as a lower-tier model. Against v6 and v6-wild it is fast and cheap in exchange for being coarse, and there are situations where that coarseness works as plain lo-fi.

Spectrum measurement

As an objective assessment of sound quality, I also measured the spectrum.

Spectrum does not have much to do with whether a song is good, but it was the only objective, machine-measurable difference I could think of, so I did it.

v6 is down 2 to 3.4 dB from 2 to 12 kHz, and the spectral centroid is 36% lower. On the other hand, the "over-compressed" verdict was not supported. What narrowed is the bandwidth, not the dynamics.

What I measured

In the 40 minutes from 23:03 to 23:43 on 2026-09-22 I generated 24 tracks from two sets of lyrics.

  • Songs: Credential Cereal and Three Drawers (both lyrics written from articles on this blog)
  • Models: v6, v6-wild, v6-mini
  • Max mode: On / Off
  • Variety slider: Off / enabled

Two songs x three models x 2 x 2 = 24 tracks. Same account, same 40 minutes, same lyrics, so everything except the model and the settings is roughly held constant.

For comparison I used 12 pre-v6 songs published on this blog. They were generated between 2026-07-26 and 2026-08-31, more than 10 days before the v6 release on the 9th, so they are from the v5 line. Their generation prompts differ from the ones used for v6 here.

I decoded with ffmpeg and ran an FFT. Band energy is reported in dB relative to full-band RMS; with absolute values, loudness differences alone would reorder the results. I used only the middle 80% of each track to avoid the fades at the start and end, and excluded silent frames.

The codec difference

The two sets use different codecs. Suno has become able to download in m4a, so I used that; previously it was only mp3 and wav.

That means this measurement also includes changes in character caused by the codec, so please note that it is not a strict v6 vs v5 comparison.

The 24 v6 tracks are Opus 132kbps / 48kHz in an mp4 container. The 12 pre-v6 tracks are MP3 about 180kbps VBR / 48kHz. Both are 48kHz, so the codec's own bandwidth ceiling only starts to matter above 16 kHz.

I did not re-encode to align them. The sound normally published on this blog is mp3, and the sound I was listening to for v6 is m4a, so comparing these two as they are matches what I actually hear. Aligning them would make both sets of numbers describe audio nobody actually listened to.

Results

Averages for 12 pre-v6 tracks (mp3) and 24 v6 tracks (m4a).

Metric Pre-v6 (12) v6 (24) Difference
Integrated LUFS -14.84 -13.72 +1.11 dB
True peak -2.35 dBTP -0.60 dBTP +1.75 dB
LRA 3.74 5.65 +1.91
Crest factor 14.53 dB 14.70 dB +0.17 dB
200-400Hz -10.28 -10.73 -0.45 dB
2-5kHz -13.07 -15.17 -2.10 dB
5-8kHz -17.98 -20.17 -2.19 dB
8-12kHz -17.00 -20.35 -3.36 dB
12-16kHz -23.48 -25.75 -2.27 dB
Spectral centroid 866 Hz 557 Hz -36%
95% rolloff 4774 Hz 2228 Hz -53%

Averages can be dragged by outliers, so I also looked at all 288 pairings across the 36 tracks.

  • For 8-12kHz, pre-v6 is higher in 260 of the pairings (90.3%)
  • 95% rolloff in 250 (86.8%), spectral centroid in 244 (84.7%)
  • 2-5kHz in 206 (71.5%)
  • True peak goes the other way: v6 is higher in 276 pairings (95.8%)

There are only two songs on the v6 side, so I split it by song as well.

LUFS True peak LRA 2-5kHz 8-12kHz Centroid
Pre-v6 (6 songs, 12 tracks) -14.84 -2.35 3.74 -13.07 -17.00 866 Hz
v6 Credential Cereal (12) -14.25 -0.58 5.39 -12.88 -19.35 685 Hz
v6 Three Drawers (12) -13.19 -0.63 5.92 -17.46 -21.36 429 Hz

Both songs are independently darker than pre-v6. 8-12kHz is down by at least 2.3 dB in both.

I believe this is what the words "muffled" and "a blanket over the speakers" are pointing at.

Changes in dynamics and compression

One complaint that keeps coming up online is "overproduced, heavily compressed" — the claim that it has been squashed flat. The tracks I have do not show that.

  • Crest factor goes from 14.53 to 14.70 dB, essentially unchanged (+0.17dB). The peak-to-RMS ratio has not moved, which means the degree of waveform squashing has not moved. Across the pairings, pre-v6 is higher in 50.7% and v6 is higher in 49.3%, so there is no difference.
  • LRA is actually 1.91 wider on v6 (3.74 to 5.65). The swing in loudness within a track has increased. Across the pairings, v6 is wider in 92.0%.

If compression had been increased, the crest factor would go down and LRA would narrow; both came out the opposite way.

So what narrowed is the bandwidth, not the dynamics. Audio with no high end is easy to perceive as clogged or squashed, so the impression makes sense to the ear, but it points at the wrong cause. Suspecting the compressor while misreading this will not fix anything.

Changes in peak level

There is one more clear change, separate from compression.

True peak has risen from -2.35 dBTP to -0.60 dBTP. Across the pairings, v6 is higher in 95.8%. Integrated LUFS is also up by 1.11 dB (87.5%).

Since the crest factor has not changed while the ceiling alone rose by 1.75 dB, the waveform shape was left alone and the whole thing was lifted — in other words, the output's mastering / limiter settings changed. Leaving only 0.6 dB of headroom to 0 dBTP is on the better side by streaming standards.

Limits of this measurement

I have put numbers out, so let me write down the limits too.

The biggest one is that song and genre are not controlled. There are two songs on the v6 side and six on the pre-v6 side, each built from different prompts. Spectral centroid moves a lot with genre, so I cannot rule out that part — or all — of this difference comes from the songs being different. Running the same lyrics and the same prompt through v5 and v6 would settle it, but v5 can no longer be selected.

The next biggest is the codec. As written above, the v6 side is m4a (Opus 132kbps) and the pre-v6 side is mp3 (about 180kbps VBR), compared as they are. That is a deliberate choice to match what I actually hear, but the possibility remains that part of the difference comes from the codec. That said, 8-12kHz is a band both codecs pass through, and the difference there is 3.36 dB, so explaining all of it with the codec seems like a stretch to me.

Beyond that, whether the 12 pre-v6 tracks are v5 or v5.5 is only inferred from their generation dates; I have not confirmed it.

Suno's internal master must have more high end than either format; what I am measuring is the audio after export.

What is going on with the shift to v6

Reading Suno's own descriptions makes v6's positioning clear.

The official blog describes v6 as "reliable, precise and consistently delivers polished music across every genre and style," putting not missing front and center. v6-wild is "less predictable and more varied," v6-mini is "faster, more efficient." And it states the collaboration with labels: "developed with our industry partners, including Warner Music Group, BMG and Believe."

On the older models: "As v6 rolls out, we will retire our previous models and move Suno entirely onto the v6 generation."

Laid out on a timeline, this looks less like a musical update than the execution of rights clearance.

Date Event
2024-06 UMG / Sony sue Suno
2025-11-25 Settlement with Warner. The published terms include "release a licensed model in 2026 and deprecate the current models" and "monthly download caps, including on paid accounts"
2026-07-31 GEMA wins at the Munich Regional Court (LG München I, 42 O 763/25, not final)
2026-09-03 Download caps take effect (Pro 20 songs/month, Premier 60 songs/month)
2026-09-09 v6 released, older models retired
2026-09-18 UMG / Sony file a second suit (No. 1:26-cv-14275, D. Mass.)

The download caps on September 3 and the retirement of the older models on September 9 did not happen to coincide; they are the execution of terms that had been announced in a settlement 10 months earlier. That is written in Warner's press release.

What is interesting is the substance of the second suit: UMG and Sony call v6 "the fruit of the same poisoned tree." Their claim is that because v6 was built using the older models' output, user preference data and knowledge distillation, taking a license does not launder the infringement. The complaint lists 60,202 recordings matched by fingerprinting Suno's training data against Audible Magic.

To be clear, this is the plaintiffs' claim, not a court's finding. Suno has responded that it is fundamentally flawed on both the facts and the law.

Still, there is one thing that can be stated as fact. Suno signed licensing deals and moved to a new model, and the litigation risk with the two majors has not gone away.

For what it is worth, two weeks after release, Suno has said nothing publicly about the criticism of quality. The blog and X are promotion only. Meanwhile it issued a rebuttal to the second suit the next day.

How generation quality actually changed

The reason "it got worse" and "it got better" can both hold at once is, I think, that the evaluation conditions are not aligned.

Resolving it would just take running the same prompt through v5 and v6, but the moment the older models disappeared, same-condition comparison on new generations became impossible for everyone. People who kept their earlier output can measure after the fact, as I did here. But nobody can align the conditions anew.

And Suno publishes neither listening tests nor benchmark scores.

The only quantitative data I found is the WildSongBench scores that the open model YuE2 publishes on its official site (192 prompts, 9 automatic metrics, 17 system settings compared).

Model SongBench average
YuE2 (best-of-8) 6.9632
Suno v5 6.8721
Suno v6 6.5562
Suno v6-wild 6.4195

Within Suno's own lineage, the score goes down from v5 to v6.

That said, these are numbers published by the developers of a competing open model, and I found no independent verification. Their own model is scored best-of-8 (generate 8 times and pick using SongBench Musicality) while the other systems are picked differently. The metrics are automatic, not blind human listening tests. The site itself notes that because the selection procedures are not aligned, direct comparison is methodologically complex. It is hard to adopt as decisive evidence.

So "v6 got worse" and "v6 got better" are both beyond anyone's ability to prove now. Personally I find that the most troublesome part.

If I were to switch

If the goal is high-quality AI music generation, what should I do from here? I do not feel Suno is in the running.

Among commercial services, the alternatives are not encouraging from what I looked into. Udio stopped downloads from October 2025, so you cannot get anything out of it. ElevenLabs Music v2's terms limit film, TV and game use to Enterprise.

The realistic direction is running an open model locally, and after reading the licenses, ACE-Step 1.5 came out as the first candidate.

License VRAM Japanese
ACE-Step 1.5 MIT. The model card explicitly states the output is usable commercially Runs from under 4GB. Under 10 seconds per song on a 3090 Over 50 languages
YuE2-3B Code is Apache-2.0, but the weights are CC BY-NC 4.0. Individuals can monetize the output; corporate commercial use needs a separate agreement 24GB class (14GiB peak) Official tags are zh / en only
MiniMax Music 3 Commercial use allowed, but "MiniMax-Music3" must be displayed in the UI. Over $20M revenue requires prior permission Under 24GB (8GB with layer streaming) Unconfirmed

"Open weights" means neither "open source" nor "unconditionally usable commercially." Going to YuE2 on benchmark scores alone will get you stuck on the license if you are using it at a company.

That said, it is not as if the open models carry no rights risk. Only YuE2 (mostly CC0 and synthetic data) and ACE-Step (licensed tracks plus public domain plus MIDI synthesis) disclose their training data; the rest do not.

The storyline that licensing cleanup cost commercial models their expressiveness and thereby opened room for open models to catch up is interesting as a story, but there is no material showing causation. It just looks that way.

Please rate this article (No signup or login required)
Currently unrated
The author runs the application development company Cyberneura.
We look forward to discussing your development needs.

Categories

Archive