Benchmark report
The State of Voice Agents
What actually breaks when AI voice agents answer real phone lines, aggregated and anonymized from every independent audit ProofDial has run. No vendor demos, no cherry-picking, no self-reported numbers. This is what the recordings show.
The number nobody has yet
How many AI agents actually tell callers they're AI?
Nobody knows. It is asked constantly and answered with opinion, because settling it requires listening to real calls, because a claim about what a caller could hear cannot be established from a test suite or an analytics dashboard. Several jurisdictions now require the disclosure and more are heading that way, but the more immediate reason to know your own number is that a caller who was never told is a complaint waiting to happen.
ProofDial decides this per call in deterministic code with no model in the path, so the verdict is reproducible and shows the agent's own line as its evidence. Every audited call produces a four-way answer (present, late, absent, or not assessable) and the aggregate is an industry disclosure rate. We will publish it here once the pool is large enough to mean something, and not before. If you want your calls counted, audit them.
How this is measured
A benchmark is only worth quoting if you can see what was counted. Five rules govern this pool:
Real calls, not demos
Every datapoint comes from an audit of recorded production calls or calls placed to a live agent. Nothing here is a vendor benchmark, a synthetic suite, or a self-reported figure.
One aggregate per audit
A published audit contributes a single anonymized row: industry bucket, completion rate, latency, hallucination and compliance rates, failure-type mix, accent gap. No client names, no transcripts, no recordings, ever.
Tenant-configured checks are excluded
If a customer writes a strict phrase check for their own line, its failures stay in their report and out of this pool. Otherwise one tenant's house rules would move the medians every other tenant is compared against.
Withdrawn findings are removed
When a customer disputes a finding and the dispute is upheld, the datapoint is recomputed. The pool reflects findings that survived being argued with.
Thin data is labelled thin
Until a metric has enough audits behind it, the page shows a curated baseline and says which is which. We would rather publish a small sample honestly than a large one we made up.
"Compliance flags" counts the failures where the failure is itself the regulatory breach: information disclosed without verifying identity, a caller never told they were speaking to an AI, personal data released to someone whose entitlement was never established, and advice given outside the agent's remit. Hallucinations and mishearings are real defects but are counted separately, so no finding is counted twice.
How does your agent compare?
Upload a week of recordings and find out. Free, no card, and no integration. You get your own figures against these medians, with the recording behind every finding.
Citing this page? Please link to it rather than reproducing the figures, because they are recomputed as audits land, so a copied number goes stale silently.