AI Talking Head

Commercial use OK 567 models No watermark No sign-up needed
Model:
+ GPT-5, Claude, Gemini
Animate a portrait photo to speak. Drop a face image and an audio file (or paste a script), and AI generates a video of the face talking with synchronized lip movements.
Animating a photo is a premium feature and uses purchased tokens. There is no free option: no self-hosted talking-head model with a commercial license is installed. You see the price before anything is charged.
The free self-hosted model (SadTalker) is no longer offered: part of it was trained on the Basel Face Model, which is licensed for non-commercial research only. Premium lip sync uses purchased tokens.

JPG, PNG or WebP - front-facing portrait, one clear face

Have a video of the person instead? Use AI Lip Sync

MP3/WAV - or leave empty + use TTS below

If you provide audio above, this text is ignored. Keep it under about 800 characters (under 60 seconds of speech).
Premium, priced by audio length
Download
Advanced options
Result
Tokens running low. Get More Tokens
Want better results? Premium models (GPT-5, Claude, Gemini) deliver higher quality. View Plans

❤️ Love Free.ai? Tell your friends!

Sign up to get a referral link and earn 30,000 tokens per friend.

Want more? Sign up free: 30,000 tokens/day
Sign Up Free

Processing your request...

Animate any portrait photo to speak with premium lip sync - drop a face image + audio, get a lip-synced talking-head video back. Ideal for explainers, avatars, voice-over to video.

How to Use AI Talking Head

1
Enter your input

Type text, upload a file, or describe what you want. No account needed.

2
Click generate

Our AI processes your request in seconds using the best open-source models.

3
Download & share

Download, copy, or share your result. Free for personal and commercial use.

Use this tool via API

Automate this tool from your own code. OpenAI-compatible REST endpoint, Bearer-token auth, no extra SDK required. Token costs match the web interface.

curl -X POST https://api.free.ai/v1/video/generate/ \
  -H "Authorization: Bearer sk-free-..." \
  -H "Content-Type: application/json" \
  -d '{"prompt": "A cat playing piano", "duration": 4}'

AI Talking Head - FAQ

Upload a portrait photo and an audio clip (or type a script), and AI animates the face to lip-sync the audio. Output is an MP4 video of the photo "speaking" with synchronized mouth movements. Animating a photo is a premium feature: it uses purchased tokens and the price is shown before you pay.

No. The free self-hosted model, SadTalker, was withdrawn on September 15, 2026. Its code is Apache-2.0, but its face-reconstruction weights were trained on the Basel Face Model, whose license allows non-commercial research only, and Free.ai is a commercial service. No commercially licensed replacement is installed yet.

An open-source code license does not cover the model weights or the data they were trained on. SadTalker bundles a 3D face reconstruction network trained on the Basel Face Model 2009, and the Basel license forbids using the model or its contents in for-profit products. Photo animation is now a premium feature that uses purchased tokens.

Front-facing portrait, clear face, even lighting, neutral expression. The face should fill at least 30% of the frame. Avoid heavy sunglasses (they break eye tracking), profile shots (the model needs both eyes visible), and extreme expressions. Studio headshots and good selfies work great.

MP3, WAV, OGG, M4A or AAC of clear speech, under 60 seconds per clip. For best lip-sync, use a single speaker, low background noise, and clearly enunciated speech. Generate the audio first via /tts/ if you want to script the talking head.

A few minutes for most clips; longer audio takes longer. You can close the tab and the result lands in your dashboard.

D-ID charges $5.99/month for 5 minutes of video. HeyGen is $24/month. Synthesia is $30/month. On Free.ai you pay per clip with tokens instead of a subscription, and the premium lip sync model is comparable to D-ID Studio quality for explainer and avatar videos.

Yes - generate a face via /image/avatar/ or /image/generate/, then feed it here. The model treats any front-facing portrait the same way. Common chain: prompt, SDXL portrait, /tts/ for the voice, then animate it here.

Lip sync animates the face region. The shoulders, clothing, and background stay nearly static. For a talking head with body movement, start from a short video of the person instead of a still photo.

Yes - POST to /v1/video/lip-sync/ with multipart `file` (the photo) + `audio_file` and `consent_given=1`, once per clip, or use /scheduled/ to queue many runs.

Yes - POST multipart `file` (a JPG, PNG or WebP photo, or a video), `audio_file` (or `text`) and, for a photo, `consent_given=1` to /v1/video/lip-sync/ on api.free.ai with Bearer auth. A photo is animated by the premium photo model and a video is lip-synced; both use purchased tokens and scale with audio duration. GET /v1/video/lip-sync-quote/?input=photo&duration=20 returns the price first. The old free /v1/video/talking-head/ endpoint returns 410. /api/ has the curl example.

Photos and audio are deleted within 24 hours of generation. Output videos sit on our CDN for 24 hours (7 days for paid users) so you can re-download from /account/?tab=history. Never used for training. Privacy policy in full at /privacy/.

Sign up free: 30,000 tokens/day

Create Free Account

No credit card required

How would you rate this tool?

Love Free.ai? Tell your friends!