Skip to main content
-20% on StoryFlow EditorGet Now
MetaHuman Support Is Here: Better Lip Sync in Unreal Engine - Featured image
Announcements Unreal Engine October 6, 2026 • 14 min read

MetaHuman Support Is Here: Better Lip Sync in Unreal Engine

MetaHumans are now supported in StoryFlow Editor's Unreal Engine plugin. This update brings a dedicated MetaHuman lip sync component, optional dialogue baking with MetaHuman Animator, and smoother automatic lip sync for Synty Sidekick characters. It arrives alongside the new StoryFlow Editor and plugin updates.

Connect a MetaHuman to its dialogue character and it can follow the voice as the conversation plays. Use live procedural lip sync while working on your dialogue, or bake facial animation from recorded lines for more detailed articulation. The plugin handles matching those animations to the voices, so you can concentrate on writing and testing the conversation.

MetaHuman lip sync support in StoryFlow Editor's Unreal Engine plugin

Two ways to bring your MetaHuman dialogue to life

MetaHumans have their own facial controls and RigLogic, so they need a different animation path from the morph targets used by a Sidekick. The new MetaHuman component takes care of that connection and comes with its own mouth poses and response settings. It follows the selected speaker through the dialogue system, including voice playback and replay.

Live lip sync from the playing voice

The live option analyzes the voice's frequency spectrum and blends between mouth shapes as the audio plays. There is no facial animation to prepare for each line. Replace a recording, sync your dialogue and the mouth follows the new audio. Strength, sensitivity, jaw bias and smoothing let you adjust the response for the face.

This is a lightweight way to give a character automatic mouth movement. It responds to acoustic cues rather than recognizing words or phonemes, so it cannot promise an exact lip shape for every consonant. That makes it useful during iteration and for dialogue where a simple, responsive performance is enough.

Baked facial animation for recorded dialogue

For recorded lines, the plugin also integrates with MetaHuman Animator's audio-driven animation workflow. The audio-to-animation processing happens in the Unreal Editor. During gameplay, the MetaHuman plays the generated animation alongside its voice, with playback position following the audio.

I preferred the baked performance on the MetaHuman and voice used in our comparison: the mouth shapes and transitions looked more convincing. The plugin makes that option practical for dialogue by keeping a shared library of voice-to-animation associations. A line with no usable bake can use live procedural lip sync, so you can work with a mix of prepared and newly recorded dialogue.

Baking adds a preparation step. When a recording changes or you add another language, its animation needs to be generated too. The setup panel tracks ready, missing, changed and failed bakes, helping you keep the facial animation in step with the dialogue as the project grows.

Setting up a MetaHuman

The workflow starts on your assembled MetaHuman character in Unreal Engine. You connect the character once; the plugin then follows that speaker's lines without a separate animation trigger on every dialogue node.

  1. Add the StoryFlow MetaHuman Lipsync component to the MetaHuman actor or Blueprint.
  2. Set Source Actor to the actor running your dialogue. In the component's StoryFlow Editor Lip Sync section, choose the imported project under Baking Project and select the Dialogue character this face belongs to.
  3. Check the setup status for the detected face, then play a voiced conversation to try live lip sync.
  4. To prepare facial animations, enable MetaHuman Animator and click Bake Dialogue Voices. The plugin finds the character's voice assets and prepares missing or changed animations in its library.
  5. Choose a baked voice and use Preview on Character to check it. Enable Bake new and changed voices after sync if you want baking to follow your dialogue updates automatically.

Creating new bakes through this workflow requires Windows, Unreal Engine 5.6 or newer, and the MetaHuman Animator plugin. That dependency is optional: live lip sync and playback of compatible existing bakes do not require the editor's baking solver. The plugin also accepts compatible MetaHuman Animator exports through its voice-to-animation assignments.

Use a face whose animation setup the component can drive. If another system already controls it, check the setup status before proceeding; the component checks animation ownership to avoid replacing an incompatible setup. The built-in workflow handles the facial connection for a supported assembled MetaHuman.

Smoother lip sync for Synty Sidekicks

The same update improves the existing procedural lip sync component used by Sidekicks. The main problems I worked on were small, rapid lip movements and a mouth that could stay too wide open through softer parts of a line. The new response is calmer, with a smaller opening range and more room to close as the voice softens.

  • Steadier movement: improved spectrum processing and smoothing reduce small fluctuations in the mouth poses.
  • More restrained opening: revised poses and lower jaw bias keep the mouth from opening too far.
  • A gentler response to quiet sounds: reduced sensitivity helps softer portions sit closer to a resting mouth.
  • Coordinated facial parts: the head, teeth and tongue receive their matching weights across a modular Sidekick.
  • Dialogue-focused analysis: the default voice analysis happens before output volume, keeping music out of the signal and preserving movement when dialogue volume is lowered.

Sidekicks continue to use the existing lip sync component. The updated defaults are Strength 0.46, Sensitivity 1.05, Jaw Bias 0.62 and Smoothing 20. If your Blueprint or placed character has saved overrides, check those values when updating; an existing override can keep the previous tuning. Custom pose maps can also change the result, so use these settings as a starting point for your own character and voice.

These improvements keep the procedural workflow lightweight and automatic. They improve how acoustic cues translate into movement, while the underlying method remains audio-driven approximation rather than phoneme recognition.

See the updated lip sync in action

The MetaHuman video compares baked playback and live procedural lip sync using the updated plugin. A third character uses a custom integration of Epic's realtime audio solver as a visual reference. Each character follows the same recording.

Left to right: baked MetaHuman animation, StoryFlow Editor's live procedural lip sync, and Epic realtime. The first two use the updated plugin. Epic realtime is a separate comparison adapter and is not included in this release. Play with sound to compare the mouth movement with the voice.

Sidekick mouth movement with the latest tuning

This recording includes the final reduction in Sidekick jaw opening and sensitivity. The comparison on the right uses a custom mapping of Epic's realtime output to the Sidekick's morph targets.

Left: the updated StoryFlow Editor procedural lip sync, including the latest Sidekick mouth tuning. Right: Epic realtime mapped to the same Sidekick as a comparison only. Both follow the same recording.

Both videos were captured at 30 FPS in standalone editor playback after installing the plugin in the comparison project. They show the updated animation; the performance figures below come from the separate packaged benchmark captured during development.

Where Epic's realtime solver fits into the comparison

The supported MetaHuman choices in this release are procedural lip sync and baked playback. The Epic realtime adapter shown in the videos and charts is a separate benchmark implementation. It is not a third mode included in the plugin, and the Sidekick mapping shown beside our component is comparison-only too.

I included it to understand the tradeoff between generating detailed facial controls during playback and preparing them ahead of time. In this UE 5.7.4 comparison, the adapter runs Epic's neural model through DirectML, synchronously on the game thread. It processes mono 16 kHz audio in 20 ms steps with an 80 ms lookahead. The packaged test uses audio samples prepared beforehand, so input conversion is outside its runtime measurements. It does not measure microphone-to-face latency or a general live-input workflow.

That implementation detail matters when reading the charts. A synchronous solver can make the game thread wait for the GPU even when overall CPU utilization stays low. The results describe this adapter and these scenes; a different integration could distribute the work differently.

The performance measurements behind the implementation

The development benchmark contains 42 accepted runs of a packaged UE 5.7.4 Development build: seven cases, tested capped at 60 FPS and uncapped, with three process runs per case and frame-rate setting. Each portrait scene contains one active speaker. The MetaHuman cases cover audio only, baked playback, procedural lip sync and Epic realtime; the Sidekick cases cover audio only and the two live methods.

These captures were made on October 4, 2026, before the final plugin port and Sidekick sensitivity and jaw-bias adjustment. They document the measured implementations and settings in the downloadable dataset. They have not been rerun as a benchmark of the release defaults.

  • CPU: AMD Ryzen 9 7900X, 12 cores and 24 logical processors.
  • GPU: AMD Radeon RX 7900 XTX, 24 GB VRAM; driver 32.0.22029.9039.
  • System: 32 GB RAM, Windows 11 Pro build 26200, Balanced power plan.
  • Build: Windows x64 Development, UE 5.7.4, engine changelist 51494982.
  • Rendering: DirectX 12, 1080 × 1080 offscreen, Epic scalability, native-resolution TAA at 100% screen percentage, VSync and dynamic resolution disabled.

Within each rig, the camera, lighting, character position and body pose stay the same. Each method is compared with its own character's audio-only baseline. A MetaHuman and a Sidekick have different geometry, materials and rig costs, so comparing their total frame times would not isolate lip sync.

What the numbers measure

Component time is elapsed game-thread time. It includes synchronous inference waits in the Epic realtime case and excludes some asynchronous animation and RigLogic work. Baked playback includes the measured component's retained audio analysis and playback bookkeeping; generating the animation in the editor is outside the runtime measurement. Whole-scene frame timings provide the wider context.

CPU utilization covers the whole process, normalized across 24 logical processors. GPU engine activity is reported separately because those engines can run concurrently, as Microsoft's GPU counter documentation explains. Their percentages should not be added together. Unreal's rendering GPU timer is also not an isolated measurement of neural inference.

Runtime cost at a 60 FPS target

Every case averaged close to 60 FPS, but the component timings reveal different amounts of work within that frame budget. Procedural lip sync averaged 0.038 ms on the MetaHuman and 0.013 ms on the Sidekick. Baked MetaHuman playback averaged 0.036 ms. The separate Epic realtime adapter averaged 1.30–1.31 ms, including its waits.

Benchmark / 01Lower is better

The cost of moving the mouth

Component time per rendered frame · 60 FPS target

Component time (ms)

MetaHuman 3 runs

Synty Sidekick 3 runs

The procedural and baked ticks are small; the realtime neural adapter spends more of the frame in its synchronous call.

Component time per rendered frame · 60 FPS target · Mean
Mean
MetaHuman Epic baked 0.036 ms
MetaHuman Procedural 0.038 ms
MetaHuman Epic realtime 1.301 ms
Synty Sidekick Procedural 0.013 ms
Synty Sidekick Epic realtime 1.309 ms

MetaHuman: baked playback, procedural and Epic realtime

One MetaHuman, 60 FPS target. Values average three process results.
MeasurementAudio onlyEpic baked playbackProceduralEpic realtime
Component game-thread time, mean—0.036 ms0.038 ms1.30 ms
Component game-thread time, p95—0.042 ms0.044 ms2.36 ms
Whole-process CPU usage5.63%5.88%5.88%5.74%
Rendering GPU frame time, mean3.72 ms3.61 ms3.73 ms3.52 ms
GPU 3D-engine activity24.39%23.65%24.53%26.02%
GPU compute-engine activity11.49%10.93%11.44%10.77%
Dedicated graphics memory2,980 MiB2,971 MiB2,961 MiB3,024 MiB
Frame time, p9516.82 ms16.85 ms16.83 ms18.43 ms
Average frame rate59.99 FPS60.00 FPS59.99 FPS59.99 FPS

Sidekick: procedural and Epic realtime

One Sidekick, 60 FPS target. Values average three process results.
MeasurementAudio onlyProceduralEpic realtime
Component game-thread time, mean—0.013 ms1.31 ms
Component game-thread time, p95—0.016 ms2.38 ms
Whole-process CPU usage4.42%4.61%4.50%
Rendering GPU frame time, mean2.27 ms2.28 ms2.10 ms
GPU 3D-engine activity15.67%15.77%17.62%
GPU compute-engine activity10.88%10.88%10.12%
Dedicated graphics memory2,278 MiB2,279 MiB2,347 MiB
Frame time, p9516.77 ms16.78 ms18.65 ms
Average frame rate60.00 FPS59.99 FPS60.00 FPS

The Epic adapter's mean used about 7.8% of a 16.67 ms frame budget. Its p95 component time was 2.36 ms on the MetaHuman and 2.38 ms on the Sidekick. The similar whole-process CPU percentages do not make those costs equivalent: time spent waiting can delay a frame without keeping a processor busy.

Compared with procedural lip sync, Epic realtime added 1.49 percentage points of GPU 3D-engine activity on the MetaHuman and 1.86 on the Sidekick. Dedicated graphics memory rose by about 63 MiB and 68 MiB. Those are whole-process differences, not model sizes. The slightly lower compute-engine activity and rendering times in those runs should not be read as evidence that inference has no GPU cost.

Benchmark / 02Lower means less usage

CPU and GPU tell different stories

Whole-process resource use · 60 FPS target

Process CPU usage (%) Observed run range

MetaHuman 3 runs

Synty Sidekick 3 runs

CPU usage is normalized across 24 logical processors. A GPU wait can block the game thread without keeping a CPU core busy.

Whole-process resource use · 60 FPS target · CPU
CPU Observed range
MetaHuman Audio only 5.63% 5.42% – 5.80%
MetaHuman Epic baked 5.88% 5.35% – 6.63%
MetaHuman Procedural 5.88% 5.66% – 6.16%
MetaHuman Epic realtime 5.74% 5.20% – 6.19%
Synty Sidekick Audio only 4.42% 4.33% – 4.55%
Synty Sidekick Procedural 4.61% 4.51% – 4.79%
Synty Sidekick Epic realtime 4.50% 4.43% – 4.54%

Uncapped playback and frame-time spikes

Benchmark / 03Higher is better

What changes without a frame cap?

Whole-scene frame rate · uncapped

Frames per second Observed run range

MetaHuman 3 runs

Synty Sidekick 3 runs

Epic realtime averaged 11.9% fewer FPS on the MetaHuman and 4.8% fewer on the Sidekick than procedural. The Sidekick baseline also varied substantially between runs.

Whole-scene frame rate · uncapped · Average FPS
Average FPS Observed range
MetaHuman Audio only 229.5 FPS 229.0 FPS – 230.3 FPS
MetaHuman Epic baked 227.8 FPS 226.7 FPS – 228.7 FPS
MetaHuman Procedural 219.7 FPS 208.6 FPS – 225.4 FPS
MetaHuman Epic realtime 193.5 FPS 191.5 FPS – 197.2 FPS
Synty Sidekick Audio only 309.9 FPS 291.0 FPS – 321.2 FPS
Synty Sidekick Procedural 317.0 FPS 307.9 FPS – 321.7 FPS
Synty Sidekick Epic realtime 301.7 FPS 299.1 FPS – 304.0 FPS
Uncapped playback. Range is the minimum and maximum of three process averages.
Character and methodAverage FPSRun-average rangeMean frame timeFrame time, p95
MetaHuman — audio only229.5 FPS229.0–230.3 FPS4.36 ms5.03 ms
MetaHuman — baked227.8 FPS226.7–228.7 FPS4.39 ms5.13 ms
MetaHuman — procedural219.7 FPS208.6–225.4 FPS4.56 ms5.50 ms
MetaHuman — Epic realtime193.5 FPS191.5–197.2 FPS5.17 ms13.31 ms
Sidekick — audio only309.9 FPS291.0–321.2 FPS3.23 ms3.99 ms
Sidekick — procedural317.0 FPS307.9–321.7 FPS3.16 ms3.89 ms
Sidekick — Epic realtime301.7 FPS299.1–304.0 FPS3.31 ms7.54 ms

Baked MetaHuman playback averaged 227.8 FPS against 229.5 FPS for audio only, a difference of about 0.7% in this scene. That supports the appeal of doing the audio-to-animation work before gameplay when recordings are already available. The character still has animation, rig-evaluation and rendering costs.

Epic realtime averaged 11.9% fewer frames per second on the MetaHuman and 4.8% fewer on the Sidekick than the matching procedural cases. Its effect on slower frames was more visible: p95 frame time rose from 5.50 to 13.31 ms for the MetaHuman, and from 3.89 to 7.54 ms for the Sidekick. The solver advances at 50 steps per second, so its work is not spread evenly across rendered frames.

Benchmark / 04Lower is better

The averages hide the slower frames

Whole-frame time · uncapped

Whole-frame time (ms) 60 FPS budget · 16.67 ms

MetaHuman 3 runs

Synty Sidekick 3 runs

The realtime adapter has longer frame-time tails. Each value averages the corresponding playback-window percentiles; these are not worst-frame times.

Whole-frame time · uncapped · 95th percentile
95th percentile
MetaHuman Audio only 5.03 ms
MetaHuman Epic baked 5.13 ms
MetaHuman Procedural 5.50 ms
MetaHuman Epic realtime 13.31 ms
Synty Sidekick Audio only 3.99 ms
Synty Sidekick Procedural 3.89 ms
Synty Sidekick Epic realtime 7.54 ms

Small differences need care. The Sidekick audio-only baseline ranged from 291.0 to 321.2 FPS across its three runs; the higher procedural mean does not show that lip sync makes rendering faster. These are results from one powerful desktop, one voice per rig and one active speaker at a time. They do not establish performance for crowds, weaker hardware, a Shipping build or a scene busy with gameplay and streaming.

Capture method and downloadable results

Each process settles for ten seconds, plays one complete warm-up line and measures two playbacks of an approximately 14-second recording. Each measurement window covers 0.5–13.5 seconds, excluding partial frames at its boundaries. The two windows are averaged within each process, then all three process results receive equal weight. Reported p95 values average the window percentiles; they are not pooled percentiles. Ranges show the three process averages.

Validation checks voice playback, mouth movement, baked playback without procedural fallback, solver progression and competing Unreal processes. Audio processing stays active while application output is muted. Case order rotates between repetitions. A round affected by browser video decoding and two captures with invalid GPU-counter jumps were replaced using external-load and counter-validity checks. The accepted set retains three runs for every case and cap, with the capture-selection history preserved.

Playback startup is tracked separately from steady playback. All published measurements use the corrected executable after a replay-position bug was found during validation; affected captures are excluded. The dataset includes the tested component settings, individual playback windows and package hashes, so the measurements remain tied to the build that produced them.

Download all 42 accepted benchmark runs as JSON.

Start with your dialogue

Update StoryFlow Editor and the Unreal Engine plugin, sync your project and add the MetaHuman lip sync component to your character. Start with live movement to test the conversation, then bake recorded lines when you want the more detailed performance shown here. For Sidekicks, check the updated settings on your existing component and try a line with both open vowels and quiet pauses.

MetaHuman support now belongs to the same dialogue workflow as your other characters: select the speaker, play the voice and let the component follow the conversation. The StoryFlow Editor Unreal Engine integration connects that character setup to the stories you build in the editor.

Enjoyed this post?

Subscribe to get notified about new articles and updates.