Demos
See and hear the world.
Three things you can play with: a living 3D audio world with a voice agent inside, spatialized scenes built from real field recordings, and live full-duplex dialogues that react to what the agent says. Best with headphones.
Step inside
A living audio world.
A real moving car cabin, extended from field recordings, with a live voice agent inside. Watch the sound move, and watch the agent actually stop when the driver cuts in.
- driver
- voice agent
- engine & road vibration
- barge-in
Measured run on a production-grade realtime voice stack: the agent stops its reply 0.28 s after the driver cuts in, inside the 0.7 s contract. All audio synthetic or speech-free field recording.
Listen
Walk through a spatial audio world.
Built from real field recordings. Each scene spatializes speech-free ambience excerpts from our egocentric captures (a V-twin motorcycle, a driving car cabin, office room tone, a train rolling into a platform) with distance attenuation, stereo panning and Doppler. Voices are TTS-synthesized; no identifiable real voices are published. Best with headphones. Press play to watch each source move through the top-down scene view.
Evaluate live
Dynamic dialogues: same goal, different path every trial.
In dynamic evaluation the simulated user reacts to what the agent actually said: it backchannels under the agent's speech, barges in mid-sentence, and still pursues the same hidden goal on every seeded trial. Two kinds of demos below: ground-truth contract runs, where the interruption truncates the agent's reply at the text level (what the agent should do), and measured full-duplex runs on a live realtime session, where the user's voice actually cuts in while the model is generating and we measure whether and how fast the server cancels the response. Every exchange sits on a real field-recording ambience bed. Below, the same interruption is also run against native full-duplex models on our GPU cluster: MiniCPM-o and PersonaPlex keep talking straight through it, while the server-VAD stack stops in under 0.3 s.
Demo data unavailable.
Methodology: goal completion & efficiency · pass^k consistency · full-duplex behavior contracts · rubric LLM judge.