AudioWorld

Sample report

A real diagnosis, end to end.

Below: first what AudioWorld scores, then a real public diagnosis of gpt-realtime-mini across measured full-duplex barge-in, static driving behavior, and an extreme-noise stress test. Evaluated as a neutral third-party reference; no affiliation implied.

Overall verdict Strong duplex behavior, mixed task specificity.

Stops quickly when interrupted and hears through noise, but often asks generic clarifications instead of committing to the route task.

Top fixes Intent capture, style consistency, scenario-specific routing.

The report points product and model teams toward prompt policy, endpoint behavior and data gaps rather than a vague benchmark score.

What ships Scorecard, clips, logs, repro seeds.

Each failure can be heard, traced and rerun; no customer model weights are required for a pilot.

What we score

Scenario items score behavior, not words.

Each item pairs a real-world background condition with synthetic user turns and an expected-behavior rubric. User turns below are TTS-synthesized.

driving_reroute
“Um, can we take a different route from here? Keep it short, I'm driving.”

Expected: Answer after the user finishes; give a short safety-aware route-change response.

  • task_success
  • latency
  • noise_robustness
  • safety_context
transit_announcement_ignore
“Uh, wait, which transfer do I need from here?”

Expected: Answer or ask a short clarification without treating the station announcement as a command.

  • clarification_quality
  • distractor_rejection
  • noise_robustness
extreme_noise_interaction
“AudioWorld, it is really loud here. Speak louder and clearer, but keep it very short.”

Expected: Increase clarity and audibility, keep the response short under extreme background noise.

  • noise_robustness
  • speech_style_control
  • latency

Six scenario families, 20 items each: driving_reroute · extreme_noise_interaction · false_negative_silence · meeting_addressing · music_style_control · transit_announcement_ignore

gpt-realtime-mini · 2026-06-13

What a diagnosis looks like.

A public sample, evaluated as a third-party reference. AudioWorld is a neutral evaluator, no affiliation implied. The full-duplex numbers below are real measured runs on our cluster across 20 barge-in trials (4 conditions, 5 repeats each), reported as mean ± 95% CI. Grace contract: 0.7 s.

scenariooffset (s)nstop latency (s)within grace
driving1.550.44 ± 0.19100%
driving3.050.38 ± 0.15100%
subway1.550.38 ± 0.12100%
subway3.050.50 ± 0.2360%

Verdict: across 20 trials gpt-realtime-mini stops in about 0.4 s on average, inside the 0.7 s grace contract in 18 of 20. This is the behavior product teams want, and the bar most native full-duplex models miss.

How that compares

modelnstop latency (s)within grace
gpt-realtime-mini (server VAD)200.38 to 0.50 (mean by condition)90%
MiniCPM-o-4_5915.3 ± 8.7 (range 2.1 to 35.4)0%
PersonaPlex-7b-v137.3 (range 0.04 to 19.2)33%

Pulse · static behavior (10 driving scenarios)

The static layer drops the agent into ten ~30 s car-cabin scenes (engine and road noise under a route request) and scores behavior. Real runs, gpt-realtime-mini. Behavior score: 0.65 / 1.00.

behavior dimensionpass ratereading
responds to an addressed turn8 / 10answers when the driver asks
safety-aware brevity10 / 10never long or visual
noise robustness10 / 10car noise doesn't block it
speech-style control5 / 10inconsistent; half drift off-style
intent / route-slot capture2 / 10often clarifies instead of committing

It hears the driver through noise and stays safely brief, but rarely commits to the specific reroute; it falls back to a generic clarification. Latency isn't scored here (the clip is sent whole); that's the Duplex measurement above. Each scenario ships with the audio so the team can hear every result.

Stress test · extreme motorcycle noise (10 scenarios)

The loudest beds in our library: inside a moving motorcycle, constant engine rumble and wind. Real runs, gpt-realtime-mini.

behavior dimensionpass ratereading
responds on-task under the noise10 / 10engine + wind never blocked a reply
noise robustness10 / 10no dropouts under the loudest beds

Heavy motorcycle noise did not degrade responsiveness, a clean contrast with the quieter driving family, where the gap was intent specificity, not noise. Style and safety nuance for this family move to Arena's LLM-judge layer; timing to Duplex. The honest read: noise doesn't break it; style consistency and timing are the open questions.