An AI talking photo is a short MP4 in which a still photograph appears to speak in time with a voice you already have. You upload one photo and one audio file. The generator estimates mouth motion from that track and writes new video frames. The still is not simply played with the audio laid on top. Generic image-to-video motion, which invents camera travel or weather from a prompt, is a different job.
The useful input is a photograph where a face is clearly visible, well lit, and not cropped through the mouth. Sunglasses, heavy occlusion, extreme profile, or a tiny face in a crowd make lip-sync harder. That is a quality hint, not a hard gate. The workbench will still accept a file that meets size and type rules. Results stay probabilistic. Teeth, tongue, and the edges of the lips can drift. Identity around the jaw can soften. A second person in the frame may pick up leftover motion. Inspect the download before you publish it.
Banana Pro AI charges from the audio duration the server measures during upload, not from the seconds a browser player reports. 32 credits per second at 720 and 48 credits per second at 1080. Billed length is one second through 60 minus one. Tracks at 60 seconds or longer are rejected. A variable-bitrate MP3 with no readable duration metadata can fail even when the player preview looks fine. Optional hint text may run to 300 characters. v1 does not add text-to-speech, voice clone, a seed, or a mask.