Your script, your images, one finished video.

SceneSync matches every image to the exact word in the narration that introduces it — then renders the video with narration and animated captions already in place.

What it does

Everything between your script and a finished file

You bring the writing and the images. SceneSync does the timing, the speaking, the captions and the render.

#1

Scene-accurate sync

Scene changes land on the word that starts them, read from the narration itself. The same script always produces the same timing.

#2

Shot-list import

Upload the table you already write in — scene number, style, narration and image prompt. It becomes the timeline.

#3

Multi-voice narration

Give different characters or sections their own voice, so a single script can be read by more than one speaker.

#4

Voice cloning

Add a voice from a sample you provide, then keep using it across every video you make.

#5

Styled captions

Choose a caption look and it stays consistent for the whole video, with the words timed to the narration.

#6

HD render

Scenes, narration and animated captions are composited into one downloadable file.

How it works

Three steps, then wait

Nothing to assemble by hand. The alignment is computed from the narration, not guessed from a number you typed.

1

Bring your shot-list

Paste or upload the script table. Each row is one scene and one image.

2

Add images and voices

Supply the pictures for each scene and pick who speaks them.

3

Render and download

Get a finished video with animated captions already composed.

Scenes matched 64 / 64
Narration voices 2
Captions Timed
Animated captions Mixed
Export HD
Why it stays in sync

Timing is measured, not estimated

Most tools ask a model to guess a timestamp for each scene and hope the guess survives. Small errors compound, and by the end of a long script the pictures are drifting away from the words.

SceneSync instead reads the narration audio for word-level timing, then matches your script against it as one sequence. Scene boundaries come from where the words actually are — so a 60-scene video stays aligned from the first frame to the last.

Same script, same result

Re-render the same project and you get the same scene timings. That is what makes reviewing a long video practical — you fix one scene without the rest of the timeline moving underneath you.

Better input, better timing

Prepare any script for accurate SceneSync timestamps

Our free Script Preparation Skill shows you how to split an existing document into clear, uniquely introduced scenes without changing the narration. Use it yourself or paste it into an AI assistant before uploading.

Download SKILL.md

SceneSync-ready structure
Scene 1: The first exact narration passage.

Scene 2: The next passage, beginning with distinct words.

Scene 3: One primary visual for each scene.
Google Flow image sequence
Scene 1.png
Scene 2.png
Scene 3.png
Scene 4.png
Correct image order

Name Google Flow images by scene number

Use this short skill in Google Flow so every generated image is exported as Scene 1, Scene 2, and so on. The filename then matches the corresponding scene heading in your prepared script.

Masterclasses & Community Showcase

See SceneSync in action

Watch step-by-step masterclasses on syncing narration to scenes, or explore viral videos created entirely with SceneSync.

Loading video showcase…
Masterclass Playlist
Creator automation skills

Build your faceless YouTube pipeline

Download ready-to-install Markdown skills for ChatGPT or Claude. Generate voice-over documents, exact SceneSync scene maps, consistent image prompts, and complete publishing packs.

Creator reviews

What creators say about SceneSync

Share your experience

Make your first video

Bring a script and a folder of images. Leave with a finished file.