How to Use AI Lip Sync
Text or audio to lip-sync video
Paste a script or drop in a voice file (MP3 or WAV). The text to video engine aligns every word to natural mouth movements and renders a finished talking video. No timeline, no keyframes, nothing to install.

Preserve voice across languages
Clone your voice from a 15-second sample with AI voice cloning and keep the same tone in every lip-synced video. Translate the script into 177+ languages and dialects without re-recording a single line.

Talking photo from a single image
Turn any portrait into a speaking presenter from a single still. Upload a photo and an audio file, or type a script, and the AI photo avatar animates the face with matched lip movement, natural head motion and expressive delivery, so a headshot becomes a talking or singing video without a camera.

Multi-speaker and multi-face sync
Lip sync every face in a scene, not only the main speaker. Precision mode tracks speaker switches, camera-angle changes and side profiles, and holds sync through hands and microphones in front of the mouth. Dialogue, group videos and duet scenes stay aligned throughout.

Realistic AI avatar lip sync with Custom Motion
Direct the performance, not only the lips. Custom Motion takes prompts for gesture, posture and expression, so the same script can read as a calm explainer or a punchy social hook. Avatar V for real people, Avatar IV for cartoon, 3D and stylized characters. Delivered by a lifelike AI spokesperson.


Manual dubbing takes weeks per language. Upload a video, pick a language, and the AI video translator re-voices it and re-syncs the lips to the new audio. Ship across 30 markets in an afternoon.

Record one take, then swap the name, company or offer per viewer. Variables in the script pull from a Google Sheet or HubSpot, and each version renders with matched lip movement. Sales teams send hundreds of one-to-one videos without recording one at a time.

Repurpose one English upload into native lip-synced versions for every market. Paste the YouTube link, pick up to 10 languages, and the video translator handles voice cloning, lip sync and captions in one pass, so a single channel serves audiences who never watched the original.

Policy changed? Edit the script, regenerate with lip sync, and republish. The AI video editor lets you swap a paragraph or a language without booking talent again.

Take one UGC hook and localise it across regions rather than reshooting per market. Dub each variant into the local language with matched lip movement, or generate the creator with Veo 3.1 or Seedance 2.0 inside HeyGen and lip sync that. A/B test in days instead of months.

Pick one of 1100+ stock avatars, paste today's script, and generate. The avatar delivers it with matched lip movement, no filming. Queue a week of clips in Batch Mode with different scripts, avatars or voices, then swap a line when the product changes.
How to lip sync a video or photo with AI
Four steps, from a video or just a photo and audio. Short clips render in 2–5 minutes.
Add an MP4, MOV or WEBM (up to 5 GB), a YouTube or Google Drive link, or a photo or avatar. HD by default, 4K for final cuts.
Type a script for AI narration, upload an MP3 or WAV, or clone your voice.
Choose Speed for front-facing clips or Precision for side profiles, multiple speakers and occlusions. The AI matches every sound to natural lip movement.
Preview, adjust timing and export as MP4, ready for social, your LMS or your team.
AI lip sync matches the mouth movements in a video to any audio track. The model breaks the audio into phonemes, maps each one to a mouth shape, and re-renders the face frame by frame so the speech looks natural even when the audio was generated from text or translated.
Yes. The Free plan includes 1 to 3 videos (varies by region), up to one minute each, with 720p sharing and a HeyGen watermark. No credit card required. Paid plans remove the watermark, raise the limit to 30 or 60 minutes per video and add 1080p and 4K export.
Paste your script into the AI lip sync generator, pick a voice, choose a video or photo for the face, and click generate. The AI lip sync tool writes the audio, syncs the lips, and renders the finished video. Short videos finish in under five minutes.
One that does both in a single pass, in the browser. HeyGen takes a YouTube link or upload, translates into up to 10 languages, clones your voice and re-syncs the lips in one job, then exports 9:16 for Shorts, Reels and TikTok. For talking-avatar clips, HeyGen's Avatar V handles upper-body movement, jaw and tongue dynamics, and Custom Motion to produce realistic AI talking avatars that lip sync perfectly.
Modern AI lip syncing matches or beats hand-animated dubbing for most footage. For the cleanest result, keep the speaker within about 45 degrees of the camera, within 10 feet, in even lighting, with one person speaking at a time. Footage with side profiles, hands or microphones near the mouth, or frequent cuts should use Precision mode.
Both run the same lip sync engine. Speed is faster and cheaper (6 credits per minute) and is built for front-facing footage with a single speaker. Precision (10 credits per minute) handles side profiles, camera-angle changes, speaker switches and objects in front of the mouth. Start with Speed; switch to Precision if sync drifts.
Yes. Translation and AI voices cover 177+ languages and dialects. If you upload your own audio, the lip sync engine is language-agnostic: it maps mouth shapes to sound, so any language, dialect or accent works. Würth Group shipped a 65-minute presentation in eight languages in four days and cut translation costs by 80%.
Final ready-to-share video files export as MP4 in 16:9, 9:16, and 1:1 ratios, with optional captions from the subtitle generator. Vertical 9:16 fits TikTok, Reels, and Shorts; 16:9 is ready for YouTube, LinkedIn, or an LMS.
One pipeline, four stages. HeyGen transcribes the speech, translates it, and regenerates it in the speaker's own voice. Dynamic Duration stretches or compresses each line by up to 20% so the new audio lands on the original timing, then the lip sync engine re-renders the mouth to the translated phonemes frame by frame. Pick Speed for front-facing footage or Precision for side profiles and speaker switches. Use dubbing and video translation to reach 30 markets without splitting tools or workflows.
Both work. Upload an audio file and the AI syncs the lips to it directly to synchronize lip movements with your track. Or clone your voice from a 15-second sample and have the model speak any script in your own voice across every language.
MP4, MOV or WEBM, up to 5 GB and from 2 seconds long. Per-video duration caps depend on plan: 30 minutes on Creator and Pro, 60 minutes on Business, 5 hours on Enterprise. You can also paste a YouTube or Google Drive link instead of uploading. Output keeps the source resolution up to 4K on Pro and above.
Yes. Combine lip sync with face swap to replace the on-screen talking head and re-sync the audio in the same pass. Useful for UGC, paid ads, and marketing videos when you want to test a new face without reshooting.
Yes. A clip from Sora, Veo, Kling or any other generator is treated like any upload: if a face is visible, HeyGen re-syncs the mouth to your audio. For still characters, Avatar IV animates cartoon, 3D and stylized faces from a single image. Avatar V is tuned for real people. Heavily stylized features can reduce mouth accuracy, so keep faces close to human proportions.
ChatGPT itself does not generate video or lip sync output. HeyGen runs natively inside the ChatGPT App Store, so you can prompt ChatGPT to create talking videos with lip sync and get the finished MP4 back. The video model does the lip sync work; ChatGPT is the front door.
Short lipsync clips (under one minute) render in two to five minutes. Longer videos and high-res exports take 10 to 20 minutes. Multilingual batch jobs (one script, ten languages) finish in roughly the time of a single render plus translation overhead.
Yes. The Lip Sync API (POST /v3/lipsyncs) takes your video and your audio and returns the synced MP4. Pay-as-you-go from $0.84 per minute (Speed) or $1.74 per minute (Precision), no subscription. Batch up to 100 videos per call. Use it for personalization at scale, in-app video generation, or batch dubbing pipelines. See the Lip Sync API docs (developers.heygen.com/lipsync-speed) for request fields and batch endpoints.
Yes, on any paid plan. Generated videos can be used in video ads, client deliverables, paid social, and published content. Voice cloning requires consent from the voice owner. Your uploads are hosted on AWS in the US, never shared with third-party vendors, and covered by HeyGen's GDPR commitment and Data Processing Addendum. Avatar use follows the standard usage license shipped with your plan.
Yes. Upload the dubbed track (MP3 or WAV) as the avatar's audio and Avatar IV or Avatar V build the mouth shapes from the sound itself, so a line dubbed into Spanish or Japanese syncs as cleanly as the original recording. For real footage, Video Translation dubs and re-syncs in the same job. For your own dubbed audio on any video, the Lip Sync API swaps the track and redraws the mouth.
No download. HeyGen runs in the browser on desktop, and there is a HeyGen mobile app. A free account is required so your projects, voices and avatars are saved to your workspace.
Yes. If your slides and speaker notes are the source, PPT to video can turn them into narrated scenes with an avatar, script, captions, and synchronized speech.
Explore more AI powered tools
Bring any photo to life with hyper‑realistic voice and movement using Avatar IV.
Transform your ideas into professional videos with AI.
