KM-Speaker: Keypoint-Based Style Control for High-Quality
Speech-Driven 3D Facial Animation and Dialogue Localization

Conditionally Accepted at ACM SIGGRAPH Asia 2026

1Ecole de Technologie Supérieure, 2Ubisoft La Forge

KM-Speaker is a keypoint-conditioned flow-based generative framework that generates high-quality animations based on global style guidance and frame-level temporal control.

Example-based generation and dialogue localization

SVG Example

Abstract

Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality. Existing controllable models typically learn global style control by relying on large-scale, low-quality in-the-wild datasets that compromise overall animation realism. Furthermore, these frameworks often lack the fine-grained temporal precision required for demanding tasks such as dialogue localization (e.g., dubbing), where matching specific facial expressions is as critical as lip synchronization. We present KM-Speaker (Keypoint-Matching Speaker), a novel keypoint-conditioned flow-based generative framework that provides both global style guidance and frame-level temporal control from reference performances. We propose a disentanglement strategy that encourages the separation of audio-driven lip motion from keypoint-driven upper-face dynamics, together with a global style context preservation mechanism to ensure coherent full-face expressiveness. KM-Speaker advances example-based 3D facial animation by achieving high-fidelity motion and flexible controllability in a data-constrained setting, consistently outperforming state-of-the-art methods in lip-sync accuracy, style adherence, and expressive temporal control.

Video

Framework

SVG Example

KM‑Speaker architecture and applications. A source audio signal and two sets of target keypoints are processed independently. Full‑face keypoints provide global style features, while upper‑face keypoints provide temporal style cues. Conditioning the flow model Ψ on all inputs enables dialogue localization, where the target upper‑face motion (green boxes) is matched to a new audio clip. Conditioning the same model only on audio and global style enables example‑based generation, where the generated animation's overall style matches the target.

Example-Based Results

The following animations are driven by arbitrary audio and conditioned on the global style of a short target animation.

For each video: Target animation (left); Generated animation using KM-Speaker (right).


Example-based SOTA comparison

For each video: Target style animation with no audio (upper-right); Mimic [1], MSMD [2], Ours with no context preservation, and Ours generation (lower).

Dialogue Localization Results

The following animations are driven by translated speech and conditioned on both the global and temporal styles of a short target animation.

For each video: Target animation to dub (left); Generated animation using KM-Speaker (right).


Dialogue localization SOTA comparison

For each video: Target animation to dub (upper); MeshTalk [3], Ours_meshTalk_masks, Ours_no_disentangle, and Ours generation (lower).

Generalization Results

The following animations are driven by arbitrary audio while conditioned on the global style of a short target video.

For each video: Target animation (left); Generated animation (right).

BibTeX

@article{josi2026km,
  title={KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization},
  author={Josi, Arthur and Got, Emeline and Dib, Abdallah and Hafemann, Luiz Gustavo and Cruz, Rafael MO},
  journal={arXiv preprint arXiv:2606.28568},
  year={2026}
}