GNM + Three.js Talking Avatar
GNM HEAD v3.0 MEAN IDENTITY 67 MORPHS
INTERACTIVE DEMO ON DEVICE

02 / LIVE AVATAR

Meet your local talking avatar.

Speak naturally and watch a local agent listen, think, answer, and articulate in real time. Choose a brain and voice below—the full conversation runs in your browser.

DEMO STATUS ready to load
MICROPHONE INPUT MICROPHONE PAUSED

Input is off. Start listening when you’re ready.

  1. 01LISTENMoonshine Tiny q8
  2. 02HEARwaiting
  3. 03THINKwaiting
  4. 04SPEAKwaiting
FIRST TOKEN
FIRST CLAUSE
TTS SYNTH
TURN → AUDIO
YOU

Start the demo, then speak naturally.

AVATAR

Your avatar’s reply will appear here. First run downloads the selected local models once.

LISTEN Moonshine Tiny34 MB
THINK Gemma 4 E2B2.0 GB
SPEAK Kokoro Timestamped fp32326 MB
Local model manager checking browser…
01
BRAINreply intelligence
COLD
Gemma 4 E2B

Browser-optimized LiteRT-LM intelligence.

2.0 GB download · additional GPU working memory FIRST TOKEN —
02
VOICEspeech synthesis
COLD
Kokoro Timestamped fp32

WebGPU voice synthesis with native phone durations.

326 MB download · WebGPU · synthesis-native timestamps SYNTH —
03
LIP SYNCGNM timing policy
ACTIVE
Acoustic refinement

IPA windows refined by local energy, voicing, and transient cues.

No model download · deterministic browser DSP AUDIO START —
04
MIC INPUTspeech recognition
COLD
Moonshine Tiny q8

Low-latency English speech recognition for local conversation.

34 MB download · CPU/WASM recognition TRANSCRIBE —

Models load only when you ask.

Model inference stays on this device. Starting the demo confirms each model’s one-time download; switching a model unloads its previous runtime first.

WebGPU
checking
Cache
checking
Local loop
zero cloud calls
Try any written phrase ElevenLabs voice + timed lip sync
OPTIONAL CLOUD VOICE Connect ElevenLabs

Add your own API key to generate a one-off phrase. The local conversation demo above does not need this key.

API key required not configured

Stored in this origin’s local storage and sent directly to ElevenLabs when you generate. Site scripts can read it—use a restricted, low-quota key and never save one on a shared device.

⌘ / Ctrl + Enter to speak 0 / 1000
Ready for a phrase no audio loaded
ACTIVE GESTURE articulatory rest
JAW
LIPS
TONGUE
IPA / DETERMINISTIC PHONE WINDOWS
awaiting speech
How speech becomes motion

The outer lips and mouth interior belong to the same native GNM mesh. Teeth and tongue are the model’s own internal anatomy; no detached facial shell or replacement mouth is used.

Loading native GNM anatomy

ABOUT / RUNTIME ARCHITECTURE

One face.
One deformation surface.

This app bakes Google GNM Head v3 offline with the population-mean identity and 51 named facial targets. Browser playback maps waveform-refined speech gestures directly into that expanded native basis every frame—without a learned animation model.

Identity
Mean / 0
Targets
51
Surface
Native
local microphone local STT + LLM local PCM + IPA coarticulation GNM morphs

Voice options

Conversational generation can use Supertonic, Kokoro, or KittenTTS locally. Optional direct-phrase playback calls ElevenLabs from this browser with a locally stored key; this static demo has no application server.