Send an endpoint or logs
No model weights required. Start with a realtime endpoint, a webhook adapter, or recorded response logs from your stack.
Voice-agent QA for real audio
We find the failures your scripted tests miss: barge-in, false activation, noise robustness, addressing and recovery when a realtime voice agent is placed in cars, transit, meetings and other messy acoustic worlds.
What we sell
A customer gives us a realtime voice endpoint or response logs. AudioWorld replays controlled real-world scenarios, scores behavior, and returns a scorecard, hearable failures, reproducible logs and concrete fixes.
What the world measures
How it works
No model weights required. Start with a realtime endpoint, a webhook adapter, or recorded response logs from your stack.
We replay seeded audio worlds: interruptions, bystander speech, car cabins, transit noise, meetings and long-session memory pressure.
You receive pass/fail scorecards, failure clips, JSON/CSV logs, repro seeds and specific recommendations for model or product fixes.
What you receive
Overall verdict, failure families, confidence, scenario coverage and the top fixes to prioritize.
Short audio clips that show exactly where the agent kept talking, activated falsely or missed the task.
Scenario IDs, seeds, timing traces, transcripts, expected behavior and machine-readable score outputs.
Recommendations for endpoint behavior, prompt policy, VAD/cancel settings, data collection or fine-tuning.
Products
Pulse is the intake report. Duplex measures live interruption behavior. Arena, Boardroom and Forge extend the same evaluation lineage into reliability, enterprise scenes and training data.
Ten short scenarios in one environment. A fast scorecard for noise robustness, false activation, silence compliance and basic task behavior.
Live realtime barge-in measurement: stop latency, cancel verdict and grace-contract pass rate, benchmarked against reference systems.
Agent-conditioned user simulation over multiple seeds for pass^k reliability, interruption recovery and goal completion.
Multi-party and long-form scenarios for addressing, silence compliance, false activations and custom acoustic environments.
RL-ready supervised, preference and duplex-timing data generated from evaluation failures and reproducible simulation runs.
Evidence
The public site shows synthetic speech and speech-free ambience demos. The private pipeline already produces measurable failure cases across static items, duplex trials, long-session cards and Boardroom scenes.
Security and data handling
Customers can provide endpoints or logs without sending model weights. Pilot scope defines retention, deletion and allowed artifacts before runs begin.
Published demos use TTS voices and speech-free ambience. Identifiable raw voices and customer outputs are not published without explicit approval.
Scenario assets track source, processing version, consent scope and export policy. Deletion requests are handled at the pilot artifact level.
Why this becomes infrastructure
Text benchmarks do not catch the failures that make voice agents feel broken: late cancellation, accidental activation, missed addressee, noise collapse and long-session drift. AudioWorld turns those failures into repeatable QA, procurement evidence and training data.
Teams shipping realtime agents into support, mobility, meetings, wearables and embedded environments.
Scenario generator, field-recording-derived ambience, timing instrumentation and failure-to-data pipeline.
Seeking design partners for Pulse and Duplex pilots; deeper enterprise packs scoped from real customer failures.