How I Built a Local AI Talking Head That Keeps the Real Person
Instead of regenerating the person, the pipeline keeps the real recording and synthesizes only the pixels inside a mouth mask. SyncNet holds the normal track at a zero-frame offset, and a deliberate 400 ms shift moves it by 10 frames — the evaluator is tracking timing, not faces.
Watch (9:47)
Overview
Instead of regenerating the person, the pipeline keeps the real recording and synthesizes only the pixels inside a mouth mask. SyncNet holds the normal track at a zero-frame offset, and a deliberate 400 ms shift moves it by 10 frames — the evaluator is tracking timing, not faces.
Full transcript (from the video)
Most AI avatars regenerate the whole person. I wanted the opposite. Keep the real performance, replace only the mouth region, and make every new voice earn its place through measured proof. Start with the contract.
A local source video model creates the picture, but its output is deliberately silent. The final speech is the same fish narration used across the project. SyncNet checks the mouth against that audio. In the current reviewed evidence, the normal track lands at zero frames.
A 400-ms shift moves the detected offset by 10 frames, and unrelated cyclic audio collapses confidence. Those controls show that the evaluator responds to timing instead of merely recognizing a face. Photo-driven avatars begin with almost no physical performance. One model must invent gaze, blink timing, head motion, shoulders, lighting response, and the mouth together.
A smooth clip can still read as a warped portrait when eyes drift, teeth change shape, the lower face softens, or the mouth interior seems detached. My reversal was to stop recreating everything. If a real recording already contains convincing human behavior, preserve that performance and generate only the new motion required by the new words. The architecture is easier to reason about when every layer has one owner.
The source recording owns the eyes, blinks, gaze, head motion, shoulders, skin, lighting, and camera texture. Fish owns the spoken performance and exact waveform. Latency Inc. receives a copy only to create mouth motion.
Deck Smith owns the final composition. That division narrows the failure surface. I can reject a weak mouth without touching accepted narration, or regenerate one short insert without asking a video model to redefine the person. This is the strongest invariant.
The source recording never becomes the final voice. This generates the canonical narration and a copy conditions the lip sync provider. Provider audio does not survive. Every promoted presenter clip is stripped to zero audio streams while the editor and deterministic renderer keep visual video muted.
Dex Smith then places the original fish waveform on the shared timeline exactly once. This prevents doubled speech, changes to accepted voice data, hidden provider audio, and drift between approved narration and the final video. The key change is starting from real video. The source already contains the eyes, blinks, posture, lighting, and camera texture.
So, the model has far less of Michael to invent. The private studio profile is a performance bank lasting just over 62 and 1/2 seconds, not one repeated clip. For a previous 12 and 1/2 second segment, the selector evaluated 113 windows and chose a take beginning at 11.6 seconds. It matches duration, motion, and performance energy without borrowing the recorded words.
A hook can use one gesture, a transition another, and a conclusion a calmer hold. Better source coverage raises the ceiling because every generated insert begins with more suitable human acting. Each provider exposed a different constraint. Joy Veza proved the photo plus audio route and the review architecture, but a portrait model invents too much of the person.
Muse Talk became the fast source video baseline. In inspected frames, it preserved source motion, although its central and lower face appeared softer. Latent Sync is the current quality baseline because the source upper face stays sharper and mouth generation starts from higher resolution video. Mouth and teeth quality still limits Latent Sync.
Echo Mimic remains compatibility blocked here, so it is a bounded experiment rather than a hidden production option. Latent Sync can generate a face sized result, but Deep Sync Myth accepts only the pixels inside the current mouth mask. Hair, forehead, brows, eyes, blinks, upper cheeks, head, neck, shoulders, lighting, and background still come from the source frame. The synthesized region covers lips, inner mouth, teeth, visible tongue, necessary jaw motion, and a narrow transition band.
Every extra generated pixel risks losing identity or camera texture. Generating less is not a shortcut. It is how the compositor keeps more of the real performance intact. This is the important part.
The generator does not get ownership of the face. It gets a narrow mouth mask while the real upper face and physical performance stay in control. SyncNet creates no pixels and does not decide whether a face looks human. It is a specialist alignment model.
Deep Sync Myth gives it moving windows of fish audio and corresponding mouth frames. SyncNet searches temporal offsets, then reports the strongest match and a confidence signal. At 25 frames per second, a zero frame best offset means the tested clip is not measurably early or late. One score is not enough, so the result must respond predictably when I deliberately move or replace the audio.
These are the current measurements I am willing to show. The normal Latent Sync track reports a zero frame offset with the exact confidence displayed beside it. Moving the audio by 400 milliseconds moves the detected offset to minus 10 frames, the expected scale at 25 frames per second. Cyclically mismatched audio drops confidence to the displayed value just above one.
These are conditional results for this evidence, not a universal benchmark. Their value is causal. Offset follows the shift and confidence collapses on the wrong speech. No judge answers every question.
Deterministic gates prove that the clip decodes, timestamps, and duration are sane. No black or frozen run appears. No audio stream survives. And the fish hash is unchanged.
SyncNet measures lip timing, temporal and identity proxies track face coverage, mouth flow, eye continuity, and source preservation. A multimodal model then watches continuous realism. Its judgement is advisory. It cannot dismiss a broken file, approve hidden audio, or authorize release.
A human still decides whether the result serves the story honestly. Perceptual review asks what timing metrics cannot. Does this read as a believable speaker over time? The current explicit presenter judge gives Quinn blinded validation facts while withholding provider names and numeric scores.
Quinn watches the continuous editor proof twice, checking mouth and teeth detail, eyes and blinks, mask seams, head and neck coherence, and overall speaker realism. Some older stored reviews used an earlier fact format, so I do not claim they were score independent. Current claims use the explicit path and its opinion remains subordinate to deterministic facts. Generation is a bounded tournament, not a success flag.
Each provider and parameter set keeps separate cache lineage, so evidence cannot leak. A candidate is stripped of audio and passes hard validation first. Deck Smith renders it with the canonical fish track, measures timing and continuity, and runs blinded visual review. Only then does it face the incumbent.
Weighted and Pareto comparisons prevent one metric from hiding a regression, while candidate budgets keep the search finite. Incumbent media bites remain unchanged until evidence supports promotion and human explicitly approves it. This video uses the architecture it explains. Fish generated the narration once.
Each presenter insert comes explicitly from the registered source profile and that slide's accepted narration, while the generated clip stays silent. Deck Smith promotes only the reviewed picture into project media, and the editor plays the canonical narration track. Remotion renders the surrounding diagrams, controls, comparisons, and evidence boards deterministically. There is no universal automatic presenter schedule in every build.
Today, presenter generation is an explicit production decision wired into the normal visual contracts. The limits remain visible. Mouth and teeth quality sets the visual ceiling, especially in long takes. Even aligned speech can feel wrong when the selected source performance has the wrong energy.
Larger pose changes, accessories, and difficult lighting also challenge tracking and blending. The response is to keep presenter inserts purposeful. Hook, transition, opinion, conclusion. The selector finds compatible acting, and every result is watched.
The next major gain may come from a richer studio library with varied gestures, energy levels, and clean holds, not from generating more pixels. The goal was never to manufacture a fake person. It was to preserve a real creator, give the final voice to fish, and reject any result the system cannot measure and watch.