OmniHuman 1.5 AI Talking Avatar Generator

Give OmniHuman 1.5 one portrait and a voice track, and it returns a video of that person saying your exact words — lips, expressions and gestures follow the sound.

0 / 2,000
130 cr/s · ≤35s

What is OmniHuman 1.5?

OmniHuman 1.5 is ByteDance's audio-driven avatar model: give it one portrait and a voice track, and it returns a video of that person speaking your audio — lips, expressions, head motion and gestures all driven by the sound. Unlike prompt-driven video models, you supply the exact words. The research paper pairs a multimodal language model with a diffusion transformer, so the performance reacts to what is being said rather than just tracking the mouth.

The honest part: it needs a frontal portrait. Our three-quarter-profile test below shows lips that barely articulate and a head that drifts away. Audio is capped at 35 seconds per run, so longer scripts mean splitting and stitching, and multi-person dialogue — a headline feature of the paper — is not available here. It is speech video, not scene video: for camera movement or action, use Seedance 2.5 or Kling 3.0 Turbo.

Architecture, positioning and user-study numbers are from the ByteDance research paper: OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation (arXiv:2508.19209)

What testers say

Version 1.5 demonstrates improved robustness for challenging inputs.
Brad Rose, fal — OmniHuman 1.5 vs OmniHuman
A waist-up portrait with clear hands, shoulders, and natural lighting usually gives better body movement.
Z.Tools — OmniHuman-1.5 deep dive
Heavy shadows, sunglasses, strange occlusions, and tiny faces make the job harder.
Z.Tools — OmniHuman-1.5 deep dive

Why run OmniHuman 1.5 on CreateVision AI?

The model is ByteDance's; the reason to use it here is that the other two inputs — the face and the voice — come from the same workspace.

GPT Image 2 portrait used as the input face for an OmniHuman 1.5 avatar clip

The whole input chain lives on one page

A talking-avatar clip needs three things: a face, a voice, and the model. Here all three come from the same workspace — generate the portrait with GPT Image 2 or Seedream, type the script and synthesize it with one of 60 built-in voices, then feed both straight into OmniHuman 1.5. Every clip on this page was made exactly that way, end to end, without recording a second of audio or uploading a single photo of a real person.

Input portrait for the Mandarin-speaking OmniHuman 1.5 tea shop case

A voice library instead of a recording booth

The text mode is not an afterthought: 60 voices across English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Russian and Portuguese, each with a preview you can play before spending anything. Type up to 600 characters, pick a voice, and the platform synthesizes the speech and drives the avatar with it in one run. The Mandarin tea-shop clip above is this exact path — no microphone involved.

Input portrait for the news anchor OmniHuman 1.5 case

One subscription, and the cost is visible up front

OmniHuman 1.5 bills by the second of audio — 130 credits per second, rounded up — and the workspace shows the exact cost of your clip before you press generate. The same plan covers GPT Image 2, Nano Banana Pro, Seedream, Seedance, Kling and Veo, so the sensible workflow is cheap: draft the portrait for a handful of credits, lock the script, and only then spend on avatar seconds. Trim the silence off your audio — you are paying for every second of it.

Three-quarter profile input portrait used for the OmniHuman 1.5 stress test

Measured, not promised — including the failure

Our seven-clip test set was generated on August 19, 2026 — a product pitch, a news brief, a gym pep talk, a Mandarin tea-shop welcome, a storybook character, a deliberate three-quarter-profile failure, and a wide stage shot. Every accepted job rendered successfully on its first take, in 3m 20s to 4m 42s for around ten seconds of audio. We also publish what went wrong: on a busy Tuesday evening the upstream rejected 13 of our 20 submissions with capacity errors before accepting them (rejected runs are refunded automatically), one accepted job queued for 14 minutes, and the profile clip on this page shows exactly what happens when you ignore the frontal-portrait advice.

Real OmniHuman 1.5 clips from this workspace

All 7 clips below were generated here on August 19, 2026 — each one is the first render its inputs produced, no retakes and no cherry-picking. Portraits are GPT Image 2 outputs; every voice was typed as text and synthesized with the built-in library. Accepted jobs returned in 3m 20s to 4m 42s for roughly 10 seconds of audio (one outlier queued 14 minutes); on a busy evening the upstream rejected several submissions with capacity errors before accepting — rejected runs are refunded. Sound on.

Prompt9.2 s audio · Trustworthy Man voice · 1,300 credits · 4m 42s

Input: GPT Image 2 portrait of a creator holding a portable speaker + the built-in Trustworthy Man voice. Script: "This little gadget honestly surprised me. One charge lasts the whole week, it pairs instantly, and the sound is way bigger than the size suggests. I've been using it every single day."

Prompt10.2 s audio · News Anchor voice · 1,430 credits · 4m 09s

Input: GPT Image 2 portrait of an anchor at a news desk + the built-in News Anchor voice. Script: "Good evening, and welcome to the briefing. Tonight we look at how small studios are using digital presenters to publish in ten languages at once — without booking a single camera crew."

Prompt10.7 s audio · Energetic Guy voice · 1,430 credits · 4m 00s

Input: GPT Image 2 waist-up photo of a coach mid-gesture + the built-in Energetic Guy voice. Script: "Alright team, listen up! Thirty seconds, that's all I'm asking for. Three rounds, full effort, no excuses. When that timer starts, you give me everything you've got. Let's go!"

Prompt11.8 s audio · Mandarin Sweet Girl voice · 1,560 credits · 4m 19s

Input: GPT Image 2 portrait of a tea-shop owner + the built-in Mandarin Sweet Girl voice. Script: "大家好,欢迎来到我们的小店。这一款云雾绿茶是今年春天刚采的,用八十度的水慢慢泡,入口特别清甜。喜欢的话,记得点个关注哦。"

Prompt9.7 s audio · Sweet Girl voice · 1,300 credits · 4m 01s

Input: GPT Image 2 storybook illustration of a character called Pip + the built-in Sweet Girl voice. Script: "Hi there! I'm Pip, and I live inside a storybook. Every page you turn is a brand new adventure — so, shall we see what's on the next one?"

Prompt8.6 s audio · Wise Lady voice · 1,170 credits · lip-sync visibly degrades at this angle

Stress test — input deliberately breaks the guidance: a three-quarter profile turned well away from the camera, + the built-in Wise Lady voice. Script: "People ask why I still write my letters by hand. Because slowness is not a flaw. Some things are only worth saying at the speed of ink."

Prompt5.4 s audio · Movie Narrator voice · 780 credits · this frame is the page backdrop

Input: GPT Image 2 wide 16:9 shot — a small spotlit figure on a dark stage, testing how the model handles a face that is far from the camera. + the built-in Movie Narrator voice. Script: "Every story begins in the dark, long before the first spotlight finds it."

Who OmniHuman 1.5 is actually for

Each of these maps to a clip demonstrated above — not a wish list.

  • Course creators and explainers

    Type the narration and get a presenter on screen — no camera, lighting or reshoots. The 35-second cap per clip maps neatly onto one concept per clip, which is how most course editors cut anyway.

  • UGC and product marketers

    The speaker-pitch case above is the whole workflow: generate a creator-style portrait holding your product, type the pitch, pick a voice. No talent fees, and no likeness release to chase — the face never existed.

  • Localization teams

    The same face can deliver the same message in ten languages: keep the portrait, swap the voice. Lip-sync follows whatever audio you provide — the Mandarin case above is the proof.

  • Faceless channels and virtual hosts

    A persona you generate is a persona you own. Design the face once, give it a consistent library voice, and it can front every episode — the illustrated storybook character above shows it does not even have to be photoreal.

  • Podcasters and audio-first creators

    You already have the hard part: the audio. Trim a highlight to 35 seconds, add a portrait, and you have a visual clip for social without opening an editor.

OmniHuman 1.5 vs Seedance 2.5 vs Veo 3.1 Pro vs Kling 3.0 Turbo

Speech video versus scene video — what each model needs from you, and what it hands back in this workspace.

OmniHuman 1.5Seedance 2.5Veo 3.1 ProKling 3.0 Turbo
What you provide1 portrait photo + your audio, or typed text and a library voiceA prompt, plus optional image, video and audio referencesA prompt, plus an optional reference imageA prompt or a start image
Who speaks the wordsYou do — the avatar lip-syncs to your exact soundtrackThe model generates voices from the promptThe model generates voices from the promptAudio is generated with the clip
Clip length hereLength of your audio, up to 35 s4–30 s4–8 s3–15 s
Credits130 per audio second110/s at 480p · 230/s at 720p320–1,920 per clip (4–8 s · 720p–4K · ± audio)55/s at 720p · 70/s at 1080p
Measured here~4 min for ~10 s of audio (7 clips, August 2026)Not measured on this taskNot measured on this taskNot measured on this task

Credits, duration and input options are read from this workspace's own model registry on August 19, 2026 — they describe what you get here, not the vendors' published specifications. Timings are our own measurements on the clips shown above; we have not run the same scripts through the other three models, so those cells say so rather than guessing.

More Video Models Like OmniHuman 1.5

Compare OmniHuman 1.5 with other video models — same workspace, one click to switch.

OmniHuman 1.5 specs on CreateVision AI

Modelomnihuman-1.5, by ByteDance (research paper published August 2025)
ModeTalking avatar: one photo + one audio track → speaking video, in the avatar workspace on this page
InputsExactly 1 portrait image, plus audio up to 35 seconds (mp3/wav upload) — or typed text up to 600 characters, synthesized to speech
Voice library60 built-in voices across 10 languages: English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Russian, Portuguese
Optional guidanceA short performance note for actions, emotion and camera — useful, never required
OutputOne continuous shot that keeps your photo's framing; clip length equals audio length. Measured output: 1248×1664 from a 2:3 photo, 1920×1088 from 16:9, 25 fps, H.264 with mono AAC audio
Credits130 per second of audio, rounded up — 130 minimum, 4,550 at the full 35 seconds; the exact cost is shown before you generate
Measured speed3m 20s–4m 42s per accepted job for ~10 s of audio in our 7-clip test; one outlier queued 14 minutes (August 2026, a busy Tuesday evening)

How to make a talking avatar from a photo

Start Generating Videos
  1. Add the face

    Upload a clear, front-facing, waist-up portrait — or generate one with GPT Image 2 or Seedream in this same workspace, which is how every face on this page was made. One person per photo; keep the face large and unobstructed.

  2. Give it a voice

    Upload an audio clip (mp3 or wav, up to 35 seconds), or switch to text mode: type up to 600 characters and pick one of 60 voices in 10 languages, each with a preview you can play first.

  3. Optionally direct the performance

    A short note like "calm, small nods, subtle smile" steers gesture and emotion. Leave it empty and the model reads the rhythm and meaning of the audio on its own.

  4. Generate and check the sync

    The exact credit cost appears before you run — it scales with audio length. Generation is asynchronous; the clip lands in your history with every input attached, so a working setup is reproducible next week.

Beyond OmniHuman 1.5

One account, the full image and video generation workflow.

Related reading

OmniHuman 1.5 Pricing and Free Daily Credits

Run OmniHuman 1.5 on the free tier's daily credits, or upgrade for premium models, faster generation, and 4K output.

Free

$0

Perfect for getting started

  • 200 welcome credits, up to 80/day
  • Z Image Turbo drafts at 0 credits
  • Standard generation speed
  • Generation history included
Get Started

Premium

Most popular
$29/month

Billed annually · $240/year

  • 8,000 credits/month, no daily limits
  • All premium models — GPT Image 2, Nano Banana, Seedream, Seedance, Veo, Kling
  • 5x generation speed, priority queue
  • AI prompt enhancement
Upgrade to Premium

Ultimate

$49/month

Billed annually · $380/year

  • 18,000 credits/month, no daily limits
  • HD generation up to 4K resolution
  • Priority support
  • Fastest queue and early access to new models
Upgrade to Ultimate

OmniHuman 1.5 — Frequently Asked Questions

1

What is OmniHuman 1.5?

ByteDance's audio-driven avatar model. It takes one portrait photo and one audio track and returns a video of that person speaking the audio — lips, expressions, head motion and hand gestures all follow the sound. The research paper behind it was published in August 2025.

2

Can the avatar say my exact words?

Yes — that is the difference from prompt-driven video models. You supply the soundtrack, either as an mp3/wav upload or by typing the script and picking a built-in voice, and the avatar lip-syncs to it word for word. Prompt-driven models generate their own voices and rarely say precisely what you wrote.

3

Is OmniHuman 1.5 free to use?

No — it bills at 130 credits per second of audio, rounded up, so a 10-second clip is 1,300 credits. The exact cost is shown before you generate, and failed runs are refunded automatically. Keep scripts tight and trim silence: every second of audio is a second you pay for.

4

What makes a good source photo?

Front-facing, waist-up, one person, with the face large, evenly lit and unobstructed — reviewers and our own tests agree. Our deliberate three-quarter-profile test above shows what happens otherwise: the lips barely articulate, the model slowly rotates the head further away, and objects appear on the desk that were never in the photo. Generated portraits work exactly as well as camera photos.

5

Does it work with illustrations or anime characters?

Yes — our storybook test above is a flat 2D illustration, and OmniHuman 1.5 animated it convincingly while keeping the art style: eyes, brows and mouth move like a hand-animated character. One quirk we observed: when the mouth opens wide, the model draws surprisingly realistic teeth for a flat illustration. ByteDance's paper also claims non-human subjects; we have only tested this illustrated character.

6

How long can a clip be, and how long does generation take?

Up to 35 seconds of audio per clip on this platform — longer scripts need to be split and stitched. In our August 2026 test, accepted jobs returned in 3m 20s to 4m 42s for roughly 10 seconds of audio. One caveat from the same evening: under heavy load the upstream rejected several submissions with capacity errors before accepting the job. Rejected runs cost nothing, but plan for retries at peak hours rather than a deadline five minutes away.

7

What is OmniHuman 1.5 bad at?

Non-frontal faces, above all: our three-quarter-profile test shows near-static lips, a head that drifts further from the camera, and desk objects that were never in the photo. Small identity details can wander mid-clip — an earring changed shape in that test, and lip colour reads slightly stronger than the input in two others. Multi-person dialogue is not available here, output resolution is modest (1248×1664 in our portrait tests — no 4K), and it only makes speech video: for action, camera movement or scene changes, use Seedance 2.5 or Kling 3.0 Turbo.

8

Can I use the videos commercially — and whose face can I use?

Videos you generate can be used in personal and commercial projects under our terms of service. The face is the part to take seriously: animating a real person's photo without their consent can violate likeness and publicity rights, and several jurisdictions now require disclosure of synthetic presenters. Generated faces — like every face on this page — avoid the consent problem entirely.

Give a photo something to say

Start Generating Videos

Pricing & Access — The FactsOmniHuman 1.5

Model
OmniHuman 1.5
Developer
ByteDance
Access
Runs in the CreateVision AI workspace via API — sign in and generate. No API key, no cloud project, no software install.
Price per generation
From 650 credits ≈ $2.36 per 5-second clip on the Premium plan ($29/month for 8,000 credits). The free plan includes 80 credits every day.
Privacy
Private by default — no other user can see your generations, on any plan. We don't sell or misuse member data. Privacy Policy

What's newOmniHuman 1.5

  1. New

    OmniHuman 1.5 avatar model integrated with a 60-voice TTS library (6 voice archetypes × 10 languages) and ready-made avatar gallery.