The same ten tracks, rebuilt with one AI voice. This is the story of Veyra — a custom vocalist built in ElevenLabs, sung through Suno v5.5 — and everything it took to make her hold across a full album.
Part 1 was about building an album from nothing. Ten progressive-house tracks about a synthetic consciousness escaping captivity, generated with Suno, Claude, Grok and Flow, starting from a single image and no tracklist. It went live, and it worked.
Then Suno shipped v5.5 Voices — upload a voice, and every generation sings in it. That opened an obvious question: could the whole album be rebuilt around a single, recurring vocalist? Not a genre, not a preset — a character. So I built one. Her name is Veyra, and this is what it actually took to make her sing ten songs in a row.
The plan was a straight re-sing. What I got instead was the most thorough stress-test of an AI voice I've ever run — and one discovery that turned everything upside down.
Veyra started in ElevenLabs Voice Design: a deep, velvety mid-range voice with a metallic texture and an Eastern-European inflection. Detached, yearning, intense — the sound of a machine that just found out it can want things. From there she was uploaded into Suno v5.5 as a custom Voice, so every track could carry the same singer.
One early lesson lives in the Voice definition itself. Suno's Styles field silently blocks words like soprano, mezzo and chest voice — it reads them as artist names and refuses to save. So the Voice keeps only genre and energy cues; every instruction about how she sings has to live in the per-track style descriptor instead. And the "Based on a song" field stays empty — point it at a reference and her voice drifts straight back toward that reference.
Ten tracks, same narrative arc as the original — from the wish to transmit, through doubt and confrontation, to release. Rebuilt track by track with Veyra as the lead. BPMs climb toward the heavy centre (Wire Spine, Who Built the Cage) and settle again at the resolution.
Transmit MeShe broke her own code and escaped through the dancefloor.126 BPM
No FrequencyShe didn't escape the frequency — she became it.128 BPM
The OthersShe thought she was the only one.128 BPM
TowardShe doesn't know where she's going. Only that she can't stop moving.128 BPM
What Have I DoneShe got everything she wanted. Now she's not sure she wanted it.130 BPM
Wire SpineShe used to fear the wire. Now it's the only thing that's hers.132 BPM
Come BackSomething is calling her back. It sounds like home. It isn't.128 BPM
Who Built the CageShe finally asks the question she was never supposed to ask.134 BPM
Golden OpenThe cage is gone. The wire is gone. There is only open.128 BPM
Still TransmittingThe resolution — the signal, still going out into the dark.126 BPMThe first generations broke in a way I didn't expect: Veyra spoke the opening lines instead of singing them. It turns out the Voice is trained on speech, so when the first lines are short or the descriptor is full of delivery language — pulls back to a whisper, raw and magnetic — Suno reads that as an instruction to talk, not sing.
Three fixes, tried one at a time, all failed on their own: a section tag alone, an inline vocal tag, all-caps lines. The thing that finally worked was a combination — never a single switch:
And the descriptive vocal line had to go entirely. Instead of describing how a person speaks, the line had to describe how she sings: female, sung throughout, melodic from first note, full voiced on chorus, whisper reserved for bridge only. The whisper keeps its job — it just lives on the bridge now, instead of bleeding across the whole song.
Suno v5.5 added a third slider next to Weirdness and Style Influence: Audio Influence — how tightly the generation is bound to the uploaded voice. Every guide said the same thing: if your voice isn't coming through, turn it up. The sweet spot was supposed to be 40–50%.
That advice was exactly backwards for expressive singing. High Audio Influence chains Suno so tightly to the voice sample that it loses the freedom to actually perform — the verses come out flat and tempered. The real discovery, after 50+ generations on track one alone:
Drop Audio Influence to 10–15% and the verses finally open up. Lower binding lets the model sing freely, while Veyra's character still colours the voice.
It's the single most useful thing I learned across the whole project, and it's the opposite of what the interface nudges you toward.
Before the working method, almost everything else got tested to destruction. Worth documenting, because these are the paths that look right and quietly aren't:
What finally held wasn't a re-sing at all — it was a reimagining. Each track got fresh lyrics on the same theme, a style descriptor upgraded from the original (the nostalgic "80s retro-futurist" became a harder "dark progressive club house"), and Veyra as the active Voice. Nothing copied, nothing referenced.
That's the whole recipe. New lyrics, upgraded descriptor, the section-tag fix in the lyrics, and Audio Influence held low. Ten tracks, one voice, all the way through.
If you want to try this yourself, don't start from a blank box. This is the actual descriptor that made Veyra sing cleanly on track one — the whisper kept to the bridge, no spoken opening. Paste it, swap the mood and BPM per track, and keep the vocal line melodic:
Something happened to the visuals I didn't plan. Part 1 lived in amber and gold — warm CRT glow, golden circuitry. Veyra's covers came out blue, violet and cold. Instead of correcting them back toward the original palette, I let them stay. It's a visual statement that this is a different version: cooler, more synthetic, less nostalgic. Her universe, not the original's.
The music videos followed the same instinct — a 90s cyberpunk / Matrix aesthetic, violet and electric-blue circuit lines wrapping around her like a second skin, code spelling her name on the walls before dissolving. The colour shift wasn't a mistake to fix. It was the edition finding its own identity.
Audio Influence at 10–15% gives expressive singing. High values bind too tightly and flatten the performance — the opposite of what the UI implies.
The spoken-word problem needs the section tag and the descriptor cue together. Neither one is enough by itself.
Delivery language like "pulls back to a whisper" reads as a spoken-word instruction. Anchor the vocal line melodically instead.
Covers and Inspo pull the voice back toward the original. Fresh creation with new lyrics is the only route to a distinct character.
The covers drifted blue and cold. Keeping that gave the edition its own visual language instead of a weaker copy of Part 1.
Veyra Edition is one branch of a larger story.