YuE2-style lyric sheet (as generated)
[Verse 1] [Male Vocals] (saxophone) City lights fading in the rearview (Forever!) [Chorus] We run all night
AI song subtitle guide
By Yana Li
If you generated a song with YuE2, you already hold the two inputs a lyric video needs: the lyric sheet you wrote and the audio the model produced. The remaining work is timing — turning that known lyric sheet into subtitle lines that land on the sung words. This guide shows the whole path: clean the generator's performance markers out of the lyric sheet, run a script-plus-audio alignment with the song option enabled, review the quality checks, and export SRT, VTT, SBV, ASS, or JSON files for your editor or upload.
YuE2-style lyric sheet (as generated)
[Verse 1] [Male Vocals] (saxophone) City lights fading in the rearview (Forever!) [Chorus] We run all night
Cleaned script and timed lines
1 00:00:03,120 --> 00:00:07,480 City lights fading in the rearview 2 00:00:07,900 --> 00:00:11,240 We run all night
Removed vs. kept
"[Male Vocals] (saxophone)" is a performance marker nobody sings, so it is removed before timing. "(Forever!)" is sung, so it stays in the lyrics. Nothing is removed silently: the pre-submit preview lists every removed line, and you can submit the original text instead.
Decision points
YuE2 takes a structured lyric sheet — verses, choruses, and performance hints — plus style settings, and returns finished audio. Your lyric sheet is the source of truth for the words: you wrote them, and the model sang them. So the task is not transcription (guessing words from audio) but alignment (putting known words on the audio's timeline). That is exactly the script-plus-audio workflow, and it keeps your approved wording intact instead of replacing it with whatever a recognizer thinks it heard.
Lyric sheets written for song-generation models are full of whole-line performance and instrument markers: "[Male Vocals]", "(saxophone) (brass) (guitar)", "[Vocal ad-libs]", "(电吉他)", "【副歌】", "[主歌1]". Across the 99 public YuE2 example lyric sheets, marker lines like these appear constantly — and they are never sung. Left in the script, they become subtitle lines that can never match the audio. The trap has a second half: parenthetical lines that ARE sung, like "(I'm so good)", "(Forever!)" or backing-vocal callouts, must stay. A cleanup that only looks at brackets can delete real lyrics; a cleanup that removes nothing leaves unsingable lines in every file you export.
Prepare a plain-text file with your final lyrics, in the exact lines you want sung. Save the generated audio as MP3 or WAV. In a Script + Audio project, paste or upload the lyrics, upload the song audio, and switch on the "This is a song" option. The pre-submit preview then lists the lyric markers it will tidy — section labels, performance markers, and trailing repeat markers — line by line, before anything is submitted. After processing, review the quality summary and any reported issues, then export SRT for editors, VTT or SBV for video platforms, ASS for styled lyric videos, or JSON for custom tooling.
The cleanup follows a simple rule: a whole line is removed only when every bracketed group on it is a recognized performance, instrument, or section marker — and a line with any unrecognized word is kept whole. Removals are counted and shown line by line before you submit, the cleaned script is what gets stored and aligned, and one click submits your original text instead. Section labels like "[Verse 1]" are removed; sung lines like "(Forever!)" are not touched.
YuE2 prompting is its own craft, and the official project documentation covers it far better than a subtitle tool should: the official YuE2 project page hosts the examples this guide references, and the YuE2 repository documents the lyric format and generation settings. Three habits help both the model and the aligner: keep section labels on their own lines, write repetitions out line by line rather than as "×2" shorthand (the shorthand is tidied, not expanded — if the chorus is sung twice, write it twice), and keep parenthetical backing vocals that are actually sung.
Exports cover SRT, VTT, SBV, ASS, JSON, and a TXT transcript, bundled together in a single ZIP with a manifest and readme. SRT and VTT cover most editors and platforms; ASS carries fonts and positioning for styled lyric videos; JSON gives you word-level structure for custom generators. The product does not export LRC; if your player needs LRC, convert from the exported SRT with any standard converter.
Sung vocals are harder to time than spoken audio, and alignment confidence varies by mix and language. When the checks cannot establish reliable timing, the project reports the problem and releases the attempt instead of handing you silently wrong timing — treat any reported issue list as part of the deliverable, not as noise. Chinese songs now anchor well when the vocal carries the words: a Chinese YuE2 example song was delivered as a precise-timing file in a production check, with about three quarters of its lines anchored to the sung words. Heavily produced mixes where instruments mask the vocal can still fail the checks — those songs are reported back with the reason instead of delivered with unreliable timing. Any delivery is a workflow check, not a per-song accuracy guarantee.
Import the exported SRT into your editor or lyric-video tool, style the lines once, and every timing lands where the song sings it. CapCut, Premiere, Resolve, and most platforms accept SRT directly; the CapCut workflow guide covers the round trip in detail. Keep the exported JSON alongside the SRT — word-level structure makes future restyling or translation cheap.
Practical workflow
Export your final lyric sheet as plain text: one sung line per line, section labels on their own lines, repetitions written out rather than as "×2".
Save the generated song audio as MP3 or WAV.
Start a Script + Audio project, paste or upload the lyrics, upload the audio, and switch on the "This is a song" option.
Open the pre-submit cleanup preview: check the listed lyric markers line by line, and choose "Use original text" if anything you sing is on the removal list.
When processing finishes, read the quality summary and issue list before trusting the timing.
Export SRT, VTT, SBV, ASS, or JSON and import it into your editor or lyric-video tool.
Product boundary
This guide covers turning an existing lyric sheet plus generated song audio into subtitle files. It does not cover generating music, prompting YuE2 in depth, or distributing generated songs, and it links to the official YuE2 resources for those topics. Timing accuracy for sung vocals varies; quality checks and issue lists are part of the deliverable, not a guarantee.
Official references checked for workflow posture
Official reference review: 2026-09-28
FAQ
No. TimedSubs works after generation: it takes the lyric sheet and audio you already have and produces checked subtitle files. For model usage, prompting, and downloads, use the official YuE2 project page and repository.
SRT, VTT, SBV, ASS, JSON, and TXT. SRT and VTT are accepted by nearly every editor and platform; ASS supports styled lyric videos; JSON keeps word-level structure for custom tooling. LRC is not an export format — convert from SRT if your player requires it.
No. Repeat shorthand is tidied from the script, not expanded into more lines. If a chorus is sung twice, write it twice in your lyric sheet so every sung line has its own subtitle entry.
Only if the whole line consists of recognized performance, instrument, or section markers. Sung backing vocals such as "(ooh)" or "(Forever!)" stay, because no word on the line matches a known marker vocabulary. Every removal is shown in the pre-submit preview, and you can submit the original text instead.
Yes for vocal-forward mixes: a Chinese YuE2 example song was delivered as a precise-timing file in a production check, with about three quarters of its lines anchored to the sung words. Heavily produced mixes where instruments mask the vocal can still be rejected by the delivery checks rather than delivered with unreliable timing.
The cleaned script — with section labels and performance markers removed — is what gets stored and timed, and the pre-submit preview shows every removal in advance. Your original text stays in the editor and can be submitted instead with one click.
Related guides
View all guidesWorkflow guide
Prepare YouTube SRT or VTT subtitle files from an approved script and final voiceover without relying on auto-caption wording.
Text to SRT guide
Learn how to turn text or TXT into a valid SRT draft with estimated timestamps, and when matching audio is required for accurate synchronization.
Workflow guide
Prepare SRT/VTT subtitle assets with timing and reading-speed checks before importing them into CapCut or Jianying for short-video editing.