Bring your own voiceover

Already recorded? Turn your voiceover into a video.

Add a narration, a podcast clip or a voice actor's take instead of a script. Vid Optimus writes out what you said, cuts it into scenes that line up with your recording, then builds the visuals and captions around it. Your voice is what plays, at your own pace.

  • MP3, WAV, M4A and more
  • Fix any misheard word
  • No AI voice to pay for
  • 8audio formats
  • 0AI voices needed
  • 1recording, in one piece
  • 2sizes, 9:16 and 16:9

Audio to video, built around your voice

Text-to-speech is quick, but nothing sounds like you. On the Script screen, choose Use a voiceover and pick your audio file. Vid Optimus transcribes it with Groq or OpenAI using your own key, so you get the words and exactly when each one was said.

You check the transcript, fix any word it misheard, and join or split scenes where you like. The recording is then cut into scenes on its own timings, and AI Director plans the visuals over your real pacing.

Transcript

Fix the words, then shape the scenes

The transcript is yours to correct before anything is built. Change a misheard name, then join two scenes or split one in two where the story turns.

Transcript
  • MP3, WAV, M4A, AAC, OGG, FLAC, Opus or WMA
  • Transcribed with your Groq or OpenAI key
  • Join or split scenes before the cut
Timing

Captions and shots timed to how you speak

Because the timing comes from your recording, every caption word and every cut follows your real delivery, pauses and all. Nothing is stretched to fit a robot voice.

Timing
  • Word-by-word captions on your own voice
  • Shots land on the words you say
  • Music dips under your voice and returns in the pauses
Have the pictures too?

Use SceneSync instead

If you also have your own images and clips, SceneSync lines them up with your recording: changing picture in your pauses, at exact times you set, or from a scene sheet it writes for you.

See SceneSync
Have the pictures too?
  • JPG or PNG images and MP4 clips
  • Four ways to sync
  • Then the same timeline and export
How it works

Your own voiceover, step by step

  1. 1 Pick your audio

    On the Script screen, choose Use a voiceover and select the file.

  2. 2 Check the transcript

    Fix any misheard word, then join or split scenes.

  3. 3 Build the visuals

    AI Director plans shots on your timing, or pick visuals by hand.

  4. 4 Export

    Your recording plays as it is, with captions, music and your visuals.

Try it free for 7 days Every feature included. No card needed.
Who it's for

Made for creators like you

Podcasters

Turn an episode clip into a video for YouTube or Shorts.

Voice actors and narrators

Your performance stays exactly as recorded.

Teachers

Record a lesson once and get a video with visuals and captions.

FAQ

Your own voiceover: common questions

Straight answers about how it works and what it costs.

Still have a question? Send us a message and we'll help you out. Contact us
Which audio files can I use?

MP3, WAV, M4A, AAC, OGG, FLAC, Opus or WMA.

Do I need an API key?

Yes, a Groq key (free tier available) or an OpenAI key to transcribe the recording. No voice key is needed, because no AI voice is generated.

What if the transcription mishears a word?

Fix it in the transcript before the scenes are cut. Captions follow your corrected text.

Is my recording changed?

No. It is cut into scenes for planning, but what plays in the video is your own recording, at your own pace.

What is the difference from SceneSync?

With Use a voiceover, Vid Optimus builds the visuals for you. SceneSync is for when you already have your own images and clips and want them lined up with your audio.

Can I still add music and a hook?

Yes. Music ducks under your voice, and you can add an opening hook, a colour look and everything else on the timeline.

Your own voiceover

Try it on your next video.

Free for 7 days with every feature. No card needed. Works on Windows and macOS.