Fix wrong auto captions in this order: words, line breaks, timing, then style. If an image to video AI draft already has a caption track, start with the spoken audio and a reference script when available, then record each repair before moving on. This keeps the caption text aligned with the message from the first correction through the final review.
This guide focuses on correcting an existing caption track. Use a separate workflow for voiceover or timed subtitle-file handoff.
Check the caption track before editing
Start with an editable project that contains the spoken audio and caption blocks. CapCut's caption tools support manual changes to caption text, timing, and appearance after generation.
Before changing anything, confirm three things:
- The audio you hear is the version the captions should represent.
- A written script or transcript is available for names, numbers, units, and technical terms.
- The caption blocks can be edited individually or as a group without starting a new recognition run.
Choose the repair location before editing:
If the project is only a flattened video, this article's correction workflow does not apply. Keep the original project or use a subtitle-file workflow instead.
Correct words before formatting
Play each caption block with the surrounding audio. Fix meaning-changing errors first, especially names, numbers, units, punctuation, and terms that could change the instruction. When a word repeats incorrectly, check every occurrence against the same reference instead of correcting only the first one.
Use this short pass for each error:
- 1
- Note the timecode of the block. 2
- Write what the speaker says or what the reference script specifies. 3
- Replace only the incorrect text. 4
- Replay the block and the neighboring phrase.
An image to video AI workflow is easiest to review when you make one controlled correction at a time. Keep the source audio and the corrected text together so you can replay the changed phrase before moving to the next cue.
Split long lines without changing the message
After the words are correct, check whether each block can be read comfortably at normal playback. Split a block when it contains two separate ideas or becomes crowded. Keep a name with its title, a number with its unit, and a short phrase with the word that completes its meaning.
Do not split solely because a line looks short in the editor. Replay that spoken section and ask whether the break follows the spoken idea. If a split changes the intended emphasis, undo it and choose a break that preserves the sentence's meaning.
Align timing with speech and scene cuts
Set a block's start where the spoken section begins and its end where it finishes. Then inspect the nearby visual cut. A caption can match the audio yet appear over the wrong shot if the scene changes during that section.
Review at least the first, middle, and last block of each scene. In an image to video AI workflow, this sampling keeps a shared timing pattern separate from one local cue. If several neighboring blocks share the same offset, note the shared pattern before dragging them one by one. If only one block is late or early, correct that block and replay both adjacent boundaries.
Use the timing controls to change only the affected cue, then replay the words immediately before and after it. CapCut's subtitle guidance supports manual timing edits after recognition.
Apply one consistent style after timing is stable
Keep font, color, size, position, and animation consistent unless emphasis has a clear purpose. Keep the caption inside safe margins and away from faces, diagrams, or other details the viewer must inspect.
Watch the corrected section twice:
- 1
- With sound on, confirm that words and timing follow the speech. 2
- With sound off, confirm that the caption sequence still communicates the message and remains legible.
If a style change makes a line wrap differently, return to the line-break pass and check it again. Style is the last adjustment because it can hide a text or timing problem.
Keep a correction log you can recheck
Use a small log for the first three errors and continue it for the rest of the track:
Timecode: Heard or written reference: Caption currently shown: Repair: word / line break / timing / style Replayed with sound: yes / no Checked without sound: yes / no
The final check is complete when every logged repair has been replayed, the first and last caption around each scene cut are aligned, and no new line wrap or overlap appeared after styling.
Know when to stop and choose the next workflow
Stop instead of guessing when there is no reliable spoken reference, the caption track is not editable, or the only available correction would start a new recognition or generation action. Preserve the original text and timecodes before making a broader change.
If you need to add voiceover, continue with the caption and voiceover workflow. If corrected text needs to move as a timed file, use the SRT caption workflow.
Once the caption track passes the word, line-break, timing, and silent-review checks, finish the caption pass in the AI video generator before moving to the CapCut video workflow. For a different focused task, browse the caption and export guides.