AUDIO-DRIVEN LIP-SYNC

AI Talking Photo

Upload one photo with a visible face and one audio track you recorded or own. The tool builds a lip-sync MP4 from that pair. It is not a still-plus-audio overlay and not generic image-to-video motion.

Photo · required(0/1)

Audio · required

A clearly visible face in the still gives the lip-sync a better chance. This is guidance, not a hard gate. Sign in before uploading.

0 / 300
Resolution
Account required32 credits/s at 720 · 48 credits/s at 1080
PreviewReady

Add a photo and an audio track

Use a still with a visible face and a clean uploaded track. The generated MP4 replaces this preview when it is ready.

32 credits/s at 720 · 48 credits/s at 1080 · billed from measured audio seconds
Open My Creations
Direct answer

What is an AI talking photo?

An AI talking photo is a short MP4 in which a still photograph appears to speak in time with a voice you already have. You upload one photo and one audio file. The generator estimates mouth motion from that track and writes new video frames. The still is not simply played with the audio laid on top. Generic image-to-video motion, which invents camera travel or weather from a prompt, is a different job.

The useful input is a photograph where a face is clearly visible, well lit, and not cropped through the mouth. Sunglasses, heavy occlusion, extreme profile, or a tiny face in a crowd make lip-sync harder. That is a quality hint, not a hard gate. The workbench will still accept a file that meets size and type rules. Results stay probabilistic. Teeth, tongue, and the edges of the lips can drift. Identity around the jaw can soften. A second person in the frame may pick up leftover motion. Inspect the download before you publish it.

Banana Pro AI charges from the audio duration the server measures during upload, not from the seconds a browser player reports. 32 credits per second at 720 and 48 credits per second at 1080. Billed length is one second through 60 minus one. Tracks at 60 seconds or longer are rejected. A variable-bitrate MP3 with no readable duration metadata can fail even when the player preview looks fine. Optional hint text may run to 300 characters. v1 does not add text-to-speech, voice clone, a seed, or a mask.

Source stills for commercial talking photo jobs

Three source still posters, not generated output

These frames are source stills and still posters for a talking photo brief. They are not generated lip-sync clips and not live recordings. Use them to judge whether your own still is a similar kind of input.

Source still poster of a founder holding a product, used as talking photo input, not a generated clip
Source still
Still poster

Product pitch from a founder still

A catalog or landing still of a person holding a product, recorded later as a spoken pitch. The talking photo path is for a face that already exists in the photograph plus a track you recorded in a quiet room. Keep the product label readable in the still. Do not expect the generator to invent a second camera angle or a new pair of hands.

Source still poster of a formal portrait intended as talking photo input, not a live memorial video
Source still
Still poster

Memorial greeting from a family still

A formal portrait used with a short greeting you recorded yourself. A talking photo can carry a remembered voice onto a still you have the right to use. It cannot restore a person who is not in the photograph, and it cannot promise a perfect match to every family memory. Keep the upload lawful. Confirm rights before you spend credits.

Source still poster of a presenter facing camera, used as talking photo input rather than a finished explainer video
Source still
Still poster

Explainer from a presenter still

A presenter still for a course, FAQ, or support clip. The talking photo output should keep the same clothing and background while the mouth follows your uploaded track. If you need a new room, a walking shot, or a cut between slides, use image-to-video or a filmed session instead of this workbench.

A practical workflow

How to make a talking photo from a photo and audio

Prepare the still and the track first. Credits are quoted only after the server measures the audio. Generation is a tracked task you can reopen from My Creations.

1. Upload one still with a visible face

Choose a JPG, PNG, or WebP up to 10 MB, with each side at least 256 pixels. A face that fills a reasonable part of the frame, looking toward the camera, is the honest starting point for a talking photo. Crop before you upload if the subject is tiny.

2. Upload the track that should drive the mouth

Upload MP3, WAV, AAC, OGG, or M4A up to 10 MB. Prefer a single speaker and a readable file. The server returns billed seconds. The on-page player is a preview only. A warning appears when measured length is over 15 seconds. Hard rejection is 60 seconds or longer.

3. Pick 720 or 1080, then confirm rights

Default output is 1080 at 48 credits per billed second. 720 costs 32 credits per billed second. Confirm that you own the photo and the audio, or that you have permission to use them. The generator will not add a second voice or clone a voice from a prompt.

4. Generate the lip-sync MP4

Sign in, spend the quoted credits, and wait. Status checks run every few seconds and pause after fifteen minutes. The same request id can resume without a second charge. If the task fails, reserved credits return. Retry from My Creations restores only the hint and resolution. You must upload the photo and audio again.

5. Inspect mouth sync, then download

Watch the whole clip. Check the first and last phonemes, teeth, and whether a second face moved. Download through My Creations. Treat the file as a draft until you have watched it. A talking photo is not a substitute for a filmed interview when legal or archival accuracy matters.

Choose the right video path

Talking Photo vs Talking Avatar vs Image-to-Video vs Traditional filming

Pick the path that matches the assets you already have. A talking photo needs one still and one uploaded track. A talking avatar is a persistent character. Image-to-video is motion from a still. Traditional filming captures a real performance on a timeline.

Talking Photo vs Talking Avatar vs Image-to-Video vs Traditional filming
Talking Photo vs Talking Avatar vs Image-to-Video vs Traditional filmingBest forControlTradeoff
Talking photoA real photograph of one person plus a recorded track you can uploadYou choose the still, the audio file, and 720 or 1080. Duration comes from the trackLip-sync can drift. No TTS, no voice clone, no extra camera move beyond what the model invents around the mouth
Talking avatarA reusable digital presenter that should appear across many scriptsCharacter setup and later takes share an identity you maintain over timeYou are not locking output to one existing photograph. Avatar pages are a different product
Image-to-videoA still that should gain camera motion, weather, or a slow product orbit from a written promptStart frame, optional end frame, duration, and a motion promptThe mouth is not driven by your audio. Do not use it when the job is a talking photo
Traditional filmingInterviews, legal statements, and any take where the real performance must exist on a camera cardLighting, lenses, retakes, and an editor on a local timelineHighest fidelity and highest production cost. No model will replace a shoot you still need

What this talking photo tool will not claim

The output is a generated lip-sync MP4. It is not a recovered recording of the day the still was taken. It is not a lossless overlay of audio onto the JPEG. If the face is small, turned, or covered, sync quality usually drops. If the audio is noisy, clipped, or full of overlapping speakers, the mouth motion will show it. Credits are billed from server-measured audio seconds. A browser preview duration is not the bill.

v1 requires a user-uploaded audio file. There is no text-to-speech box, no voice clone from a sample, and no duration slider on generate. Tracks at 60 seconds or longer are rejected. Length above 15 seconds is allowed under that cap and may look less stable. Variable-bitrate MP3 files without duration metadata may be rejected even when they play in the preview.

Limits and storage

A signed-in run creates a tracked task. You can reopen the result from My Creations after a dropped connection without another charge for that request id. The current storage backend has no automatic expiry for uploaded or generated objects. Delete in My Creations soft-deletes the task record from account history. It does not delete the underlying object or a provider copy. Review the Privacy Policy and the Terms of Service before you upload sensitive faces or private speech.

My Creations · Privacy Policy · Terms of Service

U.S. Copyright Office, Copyright and Artificial IntelligencePublic materials on copyright questions involving artificial intelligence. Not legal advice.

Related video tools

Other Banana Pro AI video paths

Stay on this talking photo page when you have one still and one uploaded track. Use the related tools when the job is motion from a still, a broader video generator, or an avatar identity.

Questions

AI Talking Photo FAQ

Straight answers about audio-driven lip-sync, credits, length limits, and what this talking photo page will not do.

What does an AI talking photo generate?

A lip-sync MP4 from one still photo and one uploaded audio track. The mouth is driven by that track. The still is not played as a slideshow with sound on top.

Do I need both a photo and audio?

Yes in v1. The talking photo workbench requires a photo and a user-uploaded MP3, WAV, AAC, OGG, or M4A file. There is no TTS field and no voice clone.

How are talking photo credits billed?

From billed seconds the server measures on audio upload. 32 credits per second at 720. 48 credits per second at 1080. Default resolution is 1080. The player duration is a preview only.

How long can the audio be?

Shorter than 60 seconds. A warning appears over 15 seconds. 60 seconds or longer is rejected. Files above 10 MB are rejected. A VBR MP3 without readable duration metadata may fail.

Does the face have to be visible?

A visible face is strongly recommended for a talking photo. It is not a hard software gate. Occluded or tiny faces usually produce weaker lip-sync. Inspect the result.

Is this the same as image-to-video?

No. Image-to-video invents motion from a prompt. A talking photo locks the performance to your uploaded audio. Use image-to-video when you want camera travel without a voice track.

Generate an AI talking photo from a still and a track

Upload a photo you may use, upload the audio, pick 720 or 1080, and inspect the lip-sync MP4 before you share it.

Open the workbench