Captions
Why word-level, script-accurate captions are different from ordinary auto-captions.
Captions are word-level and snapped to your own script — not to what the speech-to-text model thinks it heard.
Why that distinction matters
Timing comes from transcription, but the caption text is your script itself, which is already known to be correct. A name or technical term the transcription would have misheard still displays correctly, because it was never generated from the transcription's guess at the word.
Styling
Font, size, highlight colour and position are project settings on the timeline, independent of the word-matching step above. The karaoke-style word highlight follows the same timing.
When captions generate
Automatically once a scene has voiceover — nothing to trigger by hand.