Scene-accurate sync
Scene changes land on the word that starts them, read from the narration itself. The same script always produces the same timing.
SceneSync matches every image to the exact word in the narration that introduces it — then renders the video with narration and animated captions already in place.
You bring the writing and the images. SceneSync does the timing, the speaking, the captions and the render.
Scene changes land on the word that starts them, read from the narration itself. The same script always produces the same timing.
Upload the table you already write in — scene number, style, narration and image prompt. It becomes the timeline.
Give different characters or sections their own voice, so a single script can be read by more than one speaker.
Add a voice from a sample you provide, then keep using it across every video you make.
Choose a caption look and it stays consistent for the whole video, with the words timed to the narration.
Scenes, narration and animated captions are composited into one downloadable file.
Nothing to assemble by hand. The alignment is computed from the narration, not guessed from a number you typed.
Paste or upload the script table. Each row is one scene and one image.
Supply the pictures for each scene and pick who speaks them.
Get a finished video with animated captions already composed.
Most tools ask a model to guess a timestamp for each scene and hope the guess survives. Small errors compound, and by the end of a long script the pictures are drifting away from the words.
SceneSync instead reads the narration audio for word-level timing, then matches your script against it as one sequence. Scene boundaries come from where the words actually are — so a 60-scene video stays aligned from the first frame to the last.
Re-render the same project and you get the same scene timings. That is what makes reviewing a long video practical — you fix one scene without the rest of the timeline moving underneath you.
Our free Script Preparation Skill shows you how to split an existing document into clear, uniquely introduced scenes without changing the narration. Use it yourself or paste it into an AI assistant before uploading.
Scene 1: The first exact narration passage. Scene 2: The next passage, beginning with distinct words. Scene 3: One primary visual for each scene.
Scene 1.png Scene 2.png Scene 3.png Scene 4.png
Use this short skill in Google Flow so every generated image is exported as Scene 1, Scene 2, and so on. The filename then matches the corresponding scene heading in your prepared script.
Watch step-by-step masterclasses on syncing narration to scenes, or explore viral videos created entirely with SceneSync.
Download ready-to-install Markdown skills for ChatGPT or Claude. Generate voice-over documents, exact SceneSync scene maps, consistent image prompts, and complete publishing packs.
Bring a script and a folder of images. Leave with a finished file.