Making MetaHuman speech feel alive beyond the lips — tips for blending A2F lipsync with facial/head motion?

Hi all!

I run an experiment that might be fun to dissect: a fully autonomous, interactive live talk show hosted by two MetaHuman twins in UE 5.6 — on air 12 hours a day, 7 days a week.

Topics come from real news feeds, dialogue and voices are generated locally, NVIDIA Audio2Face drives the lipsync via LiveLink, and viewers can talk to the twins in chat: they greet you and answer questions live on air.

The part I want to improve: when the twins speak, the lipsync is accurate but the rest of the face stays too quiet. Real speech carries into cheeks, brows, micro head motion, emphasis nods. Right now it reads as “a talking mouth on a calm face.”

What I’m doing: A2F blendshapes via LiveLink, blended with an idle body anim; I add procedural blinks, gaze shifts and small head motion in the idle layer. My question for people who’ve fought this: what are your favorite techniques to make speech propagate naturally into the whole face and head? Additive animation layers on top of A2F? Control Rig post-processing driven by audio amplitude? Curve remapping of A2F output onto brow/cheek shapes? Something else entirely?

And honestly, tips on anything else you notice are just as welcome — fresh eyes catch things I’ve long stopped seeing.

It runs live every day noon–midnight ET, so you can see the current state (and its limits) for yourself: https://www.twitch.tv/thelunaandnovashow Happy to share details about the setup — contact@lunanovashow.com

Thanks!