Skip to main content

Lip Sync

StoryFlowLipsync moves a character's mouth while their StoryFlow Editor dialogue audio plays. Add it to the character's scene as a separate node. Projects that only display dialogue text do not need it.

Supported Faces

The node supports ARKit-style facial blendshapes. This widely used facial-animation convention gives individual movements names such as jawOpen, mouthPucker and mouthFunnel. Each blendshape acts like a slider that deforms the face. The built-in mapping uses the mouth-related shapes from this set; it does not require the full set of 52. See Apple's blendshape documentation for the convention.

Your character's mesh must already contain the facial blendshapes you want to animate. The node drives those existing shapes from audio; it does not generate a facial rig. Use a Viseme Map to connect differently named blendshapes. No iPhone or face-tracking hardware is needed for this audio-driven animation.

Set Up a Face

  1. Add a StoryFlowLipsync node as a child of the character node containing the face meshes.
  2. Assign the StoryFlowComponent that plays dialogue to Source (source). If left empty, the node searches the scene for a dialogue component.
  3. Set Character Id (character_id) to this character's stable ID from StoryFlow Editor. An empty ID makes the face react to every line.
  4. Set Face Root (face_root) to the node containing all facial MeshInstance3D nodes. If left empty, it uses the lipsync node's parent. Include the head, teeth and tongue if they are separate meshes.
  5. Play a voiced line and adjust Strength, Sensitivity, Jaw Bias and Smoothing to suit the face and voice.

The face must have blendshapes that match the active viseme mapping. Live audio analysis works without a timing track for each line, but it follows sound energy rather than recognizing phonemes, so it does not guarantee exact consonant timing.

The node follows dialogue playback automatically, closes the mouth when audio pauses or finishes, and can follow audio that continues after the dialogue ends. Turning enabled off releases analysis and immediately clears the blendshapes it owns.

Character binding

Use the character ID rather than its display name. Your game owns the character scene and meshes; the dialogue component supplies the speaking character's identity.

Audio Routing

Leave Analysis Bus (analysis_bus) empty for normal dialogue playback. The node creates a temporary analysis bus for the audio player and sends its sound to the player's original bus. Faces following the same player share the analyzer, and the original route is restored when the last listener releases it.

If an Area2D or Area3D audio bus override, or your own routing, sends the voice elsewhere, set Analysis Bus to the name of the existing bus that actually receives it. This observes that bus without changing the player's route. Every sound on the selected bus can affect the mouth, so use a dedicated dialogue bus.

For audio started by your game, call start_lipsync_for(player) on the lipsync node with an AudioStreamPlayer, AudioStreamPlayer2D or AudioStreamPlayer3D. Call stop_lipsync() when done; it releases analysis and lets the mouth close without stopping the audio.

Mapping a Rig

The built-in table uses ARKit-style mouth blendshape names. It drives matching shapes on every mesh under Face Root, while leaving unrelated expressions alone. Rigs without the extra tongueUp and tongueRaise shapes can still use the available jaw and lip shapes.

For different shape names or weights, create a StoryFlowVisemeMap resource in Viseme Map (viseme_map). Its poses array contains StoryFlowVisemePose resources; each pose's morphs array contains StoryFlowVisemeMorph resources with a shape name and weight.

Use Reset To Builtin in the resource's Inspector, or call reset_to_default(), to fill the ten editable poses: rest, AA, EE, IH, OH, OO, MM, FF, TH and L. Live speech blends along the OO, OH, AA, EE vowel axis; text-only idle motion can use the other poses too.

An unassigned map, or a map whose poses array is empty, uses the built-in table. Custom maps default to Exact names; Mesh Blends Prefix supports names such as MESHBlends.jawOpen. Save the resource as a .tres file to reuse it across character scenes.

Tuning

Start with the defaults below and check a few loud and quiet lines on your character. Adjust the controls to suit the voice recordings and face rig.

Property Default Purpose
strength0.5Scales the pose's blendshape weights.
sensitivity1.0Adjusts how strongly the mouth responds to speech.
jaw_bias1.12Scales the jawOpen shape.
smoothing40.0Controls how quickly weights approach the current pose. Higher values respond faster.
idle_mouth_without_audiotrueAdds mouth movement to text-only lines.
analysis_full_scale32.0Sets the spectrum reference used to interpret audio levels.

Unvoiced lines

Idle Mouth Without Audio affects text-only lines. A voiced line stays still if its audio cannot be found or analyzed, and does not switch to idle motion when its audio finishes. Turn this option off if you prefer a still face on unvoiced lines.

Troubleshooting

  • No audio response: log get_level() and get_raw_peak() while a matching voiced line plays. If both stay near zero, check that the player is playing, its volume is audible and Analysis Bus is empty or names the bus receiving that voice. A missing explicit bus produces a warning.
  • Audio responds but the face stays still: check that Face Root includes the meshes and that the mapping names match their blendshapes. Read missing-shape warnings in Godot's Output panel.
  • Only the head moves: use the common parent of the head, teeth and tongue as Face Root, and check that each part has the required blendshapes.
  • The wrong character moves: check Source and Character Id. Leaving the ID empty deliberately listens to every speaker.
  • The mouth barely opens or stays open: lower Analysis Full Scale for quiet audio that barely responds, or raise it if the mouth stays open between words. Then adjust Strength and Jaw Bias for the rig. get_centroid() shows the current position from OO to EE on the vowel axis.

Audio setup

See Audio for dialogue playback settings and API Reference for the lipsync node's public methods.

Need Help?

Join our Discord community to ask questions, share your projects, report bugs, and get support from the team and other users.

Join Discord