Veris AI Launches VAmoS Bench: Evaluating 11 Voice Agents Across 100 Scenarios
rdesh26 · x · 2026-08-01
Veris AI has introduced the Voice Agent Simulation Bench (VAmoS). It evaluates 11 mainstream voice agent frameworks and platforms through 100 simulated credit card support scenarios, totaling 3,300 calls.
Key Leaderboard Insights:
- Task Completion: Pipecat (71.0%) and Livekit (70.3%) top the chart, followed closely by Vapi and ElevenLabs. OpenAI Realtime Mini lags at 51.0%.
- Latency & Cost: Gemini 3.1 Live offers the lowest latency (1.38s) and cost ($0.016/call) but has a mediocre completion rate (62.3%). Vapi and Cartesia are among the most expensive ($0.135 and $0.146).
- Tech Stack: Most top-performing implementations utilize a combination of Deepgram (STT), GPT-4.1-mini (LLM), and ElevenLabs (TTS).
The benchmark is actively maintained and open for developers to submit their own implementations for evaluation.
More from coding & agent
- Minor Harness Setting Tweaks Radically Alter ARC-AGI Scores — teortaxesTex · 2026-08-01
- Google's Gemini Enterprise Agent Platform Hits GA with Robust Agent Evaluation Tools — rseroter · 2026-08-01
- Astounding AI Coding Efficiency: Feature Shipped in Under 2 Hours — charliedeets · 2026-08-01
- Practicing Long-Running Async AI Workflows: Automating Sales and Conversion — edgarpavlovsky · 2026-08-01
- SkillsGate: Open-Source Visual Skill Manager for 20+ AI Agents — tom_doerr · 2026-08-01
- Sleepwalker: Export Web Pages to AI-Readable OKF Markdown — spicemelange13 · 2026-08-01