liberate research

Liberate Research

Catching It First: Evaluating Insurance AI Agents at Scale
Liberate built domain-specific, automatic evaluation infrastructure that scores every agent conversation, clusters and ranks failures by severity, and traces each one to the responsible agent, prompt, or tool, catching the insurance-specific failures that generic quality scores miss, and often correcting them before the caller ever notices.

.png)

Simulating Insurance Calls to Test Voice AI Agents at Scale
Agent Arena is a testing system that runs hundreds of synthetic insurance calls against Liberate's voice AI in parallel, complete with the misheard policy numbers, interruptions, and noisy transcripts of a real call. It turns a full day of manual QA into minutes, catching regressions before any policyholder ever hears them.