Roark is a testing and evaluation platform for voice and chat AI agents. Before launch, it simulates your agent against synthetic callers across customer flows, personas, accents, and edge cases; after launch, it ingests every production call, transcribes it, and scores it against your quality metrics.
The workflow forms a loop: an issue caught in a production call becomes a simulation scenario, a fix is validated against that scenario, and the agent is redeployed with evidence that the metric now passes. Human reviewers label calls and set ground truth so that automated metrics stay aligned with what the team considers correct.
Roark connects to the platform your agent already runs on and syncs calls automatically. Metrics combine built-in checks with custom LLM-based evaluators and a purpose-built evaluation model, and results are searchable with transcripts, tool calls, and recordings attached.
Features
- Simulation testing: synthetic callers exercise flows, personas, accents, and edge cases in staging
- Post-call analysis: every production call is transcribed, traced, and scored on more than 60 audio-native metrics, including sentiment, emotion, interruptions, and speech patterns
- Custom metrics: LLM-based evaluators with pass/fail thresholds alongside the built-in set
- Human review: label calls and set ground truth to calibrate automated scoring
- Fix loop: turn a caught issue into a scenario, prove the fix in simulation, then deploy
- Load tests and health checks: capacity and uptime checks for the agent endpoint
- Platform integrations: Vapi, Retell, ElevenLabs, LiveKit, Pipecat, Leaping, and custom stacks
- SDKs and API: Node.js and Python SDKs plus a REST API for calls, metrics, and simulations
- Compliance: SOC 2 Type II and a HIPAA BAA