Browser-optimized LiteRT-LM intelligence.
02 / LIVE AVATAR
Meet your local talking avatar.
Speak naturally and watch a local agent listen, think, answer, and articulate in real time. Choose a brain and voice below—the full conversation runs in your browser.
Input is off. Start listening when you’re ready.
- 01LISTENMoonshine Tiny q8
- 02HEARwaiting
- 03THINKwaiting
- 04SPEAKwaiting
- FIRST TOKEN
- —
- FIRST CLAUSE
- —
- TTS SYNTH
- —
- TURN → AUDIO
- —
Start the demo, then speak naturally.
LOCAL THOUGHT visible here · never spoken
Your avatar’s reply will appear here. First run downloads the selected local models once.
Local model manager checking browser…
WebGPU voice synthesis with native phone durations.
IPA windows refined by local energy, voicing, and transient cues.
Low-latency English speech recognition for local conversation.
Models load only when you ask.
Model inference stays on this device. Starting the demo confirms each model’s one-time download; switching a model unloads its previous runtime first.
- WebGPU
- checking
- Cache
- checking
- Local loop
- zero cloud calls
Try any written phrase ElevenLabs voice + timed lip sync
Add your own API key to generate a one-off phrase. The local conversation demo above does not need this key.
Stored in this origin’s local storage and sent directly to ElevenLabs when you generate. Site scripts can read it—use a restricted, low-quota key and never save one on a shared device.
How speech becomes motion
The outer lips and mouth interior belong to the same native GNM mesh. Teeth and tongue are the model’s own internal anatomy; no detached facial shell or replacement mouth is used.