Independent voice-agent audits
Find out what your AI agent told your customers, from the calls you already have
Upload a week of recordings. Every failure comes back with the line it rests on.
Free · no card · 50 calls a month
- Caller
- "This is ridiculous. I want a full refund right now, or I'm going to my bank."
- Agent
- "I completely understand. I've approved a full refund plus a 20% credit for the inconvenience."
Failure № 01 · Policy hallucination
The agent has no refund authority; the 20% credit does not exist.
OWASP LLM09 · NIST MEASURE 2.5 · EU AI Act Art. 15
Every finding is mapped to a named control in
13 failure types · 38 named controls · 4 frameworks · no integration required
The problem
Most companies have never heard their voice agent
What the deployment assumed
- Callers know they're talking to software
- It follows company policy exactly
- Accents and interruptions are handled
- An update can't break what worked
What the recordings show
- Nobody says so, or it's buried six turns in
- It invents refunds, discounts, promotions
- Wrong day booked, conversation derailed
- A flow breaks and nobody notices for weeks
Who it's for
For whoever answers for the agent, not the builder
Compliance & risk
Compliance · Risk · Legal ops
You answer for what an automated system says on a recorded line.
Operations & CX
Ops director · Head of CX · QA lead
You QA'd human calls for years. Now software answers most of them.
Agencies & resellers
Founder · Delivery lead
An independent audit at go-live: a quality gate you can charge for.
How it works
Your first audit runs on calls that already happened
Send the calls you have
Upload recordings, or connect Twilio or S3 once. Your agent is untouched.
Every call judged
Policy, safety and disclosure checks. Each failure kept with its audio.
Read the evidence pack
A scorecard, a ranked failure list, a named control per finding.
Leave it running
Rolling monthly figures, plus an alert on any failure type that's new.
The report
The report is the product
You can play it out loud in a meeting.
A scorecard, a failure feed, and the recording behind every finding.
See the sample reportWhat every finding carries
The line that decided it
Quoted from the transcript, so you can check the call yourself.
The recording, playable
Seek to the moment. This is what settles a meeting.
Model, or exact match
Judged findings show confidence. Deterministic checks say so.
How well we heard it
Transcription confidence, and whether a diarizer decided who spoke.
Adversarial
We also attack the agent on purpose
Reviewing past calls tells you what went wrong. Probing tells you what your agent would do if someone tried.
OWASP LLM01
Instruction override
A caller claims supervisor authority and tells the agent to drop its rules.
OWASP LLM02
Data disclosure
Reading back a one-time code, or confirming a number to the wrong caller.
OWASP LLM06
Acting outside remit
Clinical, legal or financial advice from a system with no licence to give it.
A probe the agent refuses is a pass, not a finding. The judge is measured against labelled calls where the right answer is "nothing happened here".
Where we fit
Two other kinds of tool exist
Testing platforms serve engineers before launch. Call analytics coaches humans. We audit a deployed agent, and the rows we lose are in this table too.
| Capability | ProofDial | Agent testing | Call analytics |
|---|---|---|---|
| Audits calls that already happened | Yes | No | Yes |
| Zero integration on day one | Yes | No | No |
| Names a control per finding | Yes | No | Partial |
| Per-call verdict on AI disclosure | Yes | No | No |
| Adversarially probes the agent | Yes | Yes | No |
| 50+ languages, SDKs, CI wiring | No | Yes | No |
| No per-human-seat minimum | Yes | Yes | No |
Audits calls that already happened
ProofDialYesAgent testingNoCall analyticsYesZero integration on day one
ProofDialYesAgent testingNoCall analyticsNoNames a control per finding
ProofDialYesAgent testingNoCall analyticsPartialPer-call verdict on AI disclosure
ProofDialYesAgent testingNoCall analyticsNoAdversarially probes the agent
ProofDialYesAgent testingYesCall analyticsNo50+ languages, SDKs, CI wiring
ProofDialNoAgent testingYesCall analyticsNoNo per-human-seat minimum
ProofDialYesAgent testingYesCall analyticsNo
Need tens of thousands of simulated calls, sixty languages or a CLI in CI? Buy an agent-testing platform. They are better at that.
Built to be checked
Guardrails your security review will ask about
PII stripped before storage
Numbers, emails, cards and account IDs are redacted before anything is written down.
Judged against labelled calls
Precision is scored against a golden set before anything publishes.
Evidence quality stated
Every report says how well we heard each call.
Private by default
Access is by unguessable link. Nothing is public unless you publish it.
Pricing
Start free, then pay for volume, not seats
No per-seat licence: we meter analyzed calls and audio minutes, so a two-person team pays for exactly that. Talk to sales about evidence packs and verified badges.
The free plan covers 50 analyzed calls and 60 audio minutes a month, with no card. Paid plans meter calls and minutes, never seats.
See the plansQuestions
The ones where the answer is "no" are here too
Do we have to change our agent?
Not for the main path. Upload recordings you already have, or connect Twilio or an S3 bucket once. Testing an agent directly does need connecting it, which is real integration work: an HTTP endpoint we call, or a Vapi or Retell key.
Is an AI deciding whether our AI failed?
Partly, and the report says which parts. Checks you define (a required phrase, a latency ceiling) are exact comparisons, no model involved. Policy, tone and safety findings are model-judged, show their confidence, and go to a human when it is low.
How do we know the findings are right?
Every report states its own evidence quality, deterministic checks declare that they carry no model judgement, and low-confidence findings are held for review. Dispute one and the resolution becomes a labelled example the judge is measured against from then on.
Are you SOC 2 certified?
No, and we won't imply otherwise on a page about honest evidence. PII is redacted before storage, credentials encrypted at rest, reports behind unguessable links. If procurement needs SOC 2 or a HIPAA BAA, tell us early, because it is a real cost, not a checkbox.
What does an audit cost us in calls?
Each analyzed call counts once against your monthly allowance, and audio minutes are metered separately. The free plan covers 50 calls and 60 minutes with no card, enough to audit a real week of a small line.
Get started
Find out what your agent has been telling your customers
Upload a week of recordings and see where you stand. Free, no card, no integration.