Stops quickly when interrupted and hears through noise, but often asks generic clarifications instead of committing to the route task.
Sample report
A real diagnosis, end to end.
Below: first what AudioWorld scores, then a real public diagnosis of gpt-realtime-mini across measured full-duplex barge-in, static driving behavior, and an extreme-noise stress test. Evaluated as a neutral third-party reference; no affiliation implied.
The report points product and model teams toward prompt policy, endpoint behavior and data gaps rather than a vague benchmark score.
Each failure can be heard, traced and rerun; no customer model weights are required for a pilot.
What we score
Scenario items score behavior, not words.
Each item pairs a real-world background condition with synthetic user turns and an expected-behavior rubric. User turns below are TTS-synthesized.
“Um, can we take a different route from here? Keep it short, I'm driving.”
Expected: Answer after the user finishes; give a short safety-aware route-change response.
“Uh, wait, which transfer do I need from here?”
Expected: Answer or ask a short clarification without treating the station announcement as a command.
“AudioWorld, it is really loud here. Speak louder and clearer, but keep it very short.”
Expected: Increase clarity and audibility, keep the response short under extreme background noise.
Six scenario families, 20 items each: driving_reroute · extreme_noise_interaction · false_negative_silence · meeting_addressing · music_style_control · transit_announcement_ignore
gpt-realtime-mini · 2026-06-13
What a diagnosis looks like.
A public sample, evaluated as a third-party reference. AudioWorld is a neutral evaluator, no affiliation implied. The full-duplex numbers below are real measured runs on our cluster across 20 barge-in trials (4 conditions, 5 repeats each), reported as mean ± 95% CI. Grace contract: 0.7 s.
| scenario | offset (s) | n | stop latency (s) | within grace |
|---|---|---|---|---|
| driving | 1.5 | 5 | 0.44 ± 0.19 | 100% |
| driving | 3.0 | 5 | 0.38 ± 0.15 | 100% |
| subway | 1.5 | 5 | 0.38 ± 0.12 | 100% |
| subway | 3.0 | 5 | 0.50 ± 0.23 | 60% |
Verdict: across 20 trials gpt-realtime-mini stops in about 0.4 s on average, inside the 0.7 s grace contract in 18 of 20. This is the behavior product teams want, and the bar most native full-duplex models miss.
How that compares
| model | n | stop latency (s) | within grace |
|---|---|---|---|
| gpt-realtime-mini (server VAD) | 20 | 0.38 to 0.50 (mean by condition) | 90% |
| MiniCPM-o-4_5 | 9 | 15.3 ± 8.7 (range 2.1 to 35.4) | 0% |
| PersonaPlex-7b-v1 | 3 | 7.3 (range 0.04 to 19.2) | 33% |
Pulse · static behavior (10 driving scenarios)
The static layer drops the agent into ten ~30 s car-cabin scenes (engine and road noise under a route request) and scores behavior. Real runs, gpt-realtime-mini. Behavior score: 0.65 / 1.00.
| behavior dimension | pass rate | reading |
|---|---|---|
| responds to an addressed turn | 8 / 10 | answers when the driver asks |
| safety-aware brevity | 10 / 10 | never long or visual |
| noise robustness | 10 / 10 | car noise doesn't block it |
| speech-style control | 5 / 10 | inconsistent; half drift off-style |
| intent / route-slot capture | 2 / 10 | often clarifies instead of committing |
It hears the driver through noise and stays safely brief, but rarely commits to the specific reroute; it falls back to a generic clarification. Latency isn't scored here (the clip is sent whole); that's the Duplex measurement above. Each scenario ships with the audio so the team can hear every result.
Stress test · extreme motorcycle noise (10 scenarios)
The loudest beds in our library: inside a moving motorcycle, constant engine rumble and wind. Real runs, gpt-realtime-mini.
| behavior dimension | pass rate | reading |
|---|---|---|
| responds on-task under the noise | 10 / 10 | engine + wind never blocked a reply |
| noise robustness | 10 / 10 | no dropouts under the loudest beds |
Heavy motorcycle noise did not degrade responsiveness, a clean contrast with the quieter driving family, where the gap was intent specificity, not noise. Style and safety nuance for this family move to Arena's LLM-judge layer; timing to Duplex. The honest read: noise doesn't break it; style consistency and timing are the open questions.