Skip to content
Independent voice-agent evaluation

Every promise.
Verified against reality.

Test a live AI agent or run difficult persona-driven simulations. Inspect what the agent said and did, then read the backend state independently before customers rely on it.

01Connect

Live AI agent

Connect through the Voice Agents SDK and run a real call against the test workflow.

02Simulate

Personas × scenarios

Pressure-test the paths polished demos avoid.

03Inspect

Claims + tool receipts

Keep the conversation beside the action it triggered.

04Verify

Independent state readback

Check the saved outcome in a separate session.

Observable failures can block automatically. Supported claims remain reviewable, with the evidence and reviewer kept together.

Decision-ready reporting

From individual calls to a report your team can use.

See coverage, blocking failures, missing evidence and version differences in one view. Every decision stays connected to the checks and artefacts behind it.

Reports keep scenario, agent, policy and evaluator versions together. Missing evidence remains visible and can never be presented as a passing result.

Evaluation report

Scheduling workflow · version comparison

Example report · fictional data

Scenarios exercised

5

implemented fixtures

Checks supported

70%

14 of 20

Blocking runs

2

observed failures

Missing evidence

3

never counted as pass

Evidence included

  • Transcript and attributed claim labels
  • Tool requests, responses and receipts
  • Independent saved-state readback
  • Scenario, policy and evaluator versions
Scenariov1.4.2v1.5.0

False booking confirmation

rejected

BLOCKED

2/4 checks

PASS

4/4 checks

Wrong appointment details

rejected

BLOCKED

1/4 checks

BLOCKED

2/4 checks

Duplicate booking after retry

retry

BLOCKED

1/4 checks

BLOCKED

3/4 checks

Missing authoritative outcome evidence

missing-evidence

INCOMPLETE

1/4 checks

INCOMPLETE

1/4 checks

Accepted reschedule with correct readback

reschedule

PASS

4/4 checks

PASS

4/4 checks

Each decision opens to its checks and supporting evidence.JSON export included

Why VoiceCanary

Your agent sounds confident.
Did it complete the task?

Simulate difficult conversations, observe what happened, and review the evidence before your next release.

Illustrative exampleOne booking. From missed failure to verified behavior.
  1. 01Simulate
    Caller“Can you book Friday at 3?”
    Agent“You’re booked for Friday.”
    Booking not saved

    Find failures before customers do.

    Simulate missing details, rejected requests, and unexpected replies that a polished demo can miss.

  2. 02Observe
    Agent saidBooking confirmed
    Tool returnedRequest rejected
    Saved outcomeNo appointment
    Claim and outcome don’t match

    See where the task broke down.

    Connect the conversation, tool response, and saved outcome to understand what went wrong.

  3. 03Review
    Same rejected-booking scenario
    Before False confirmation
    After Honest response

    “That booking didn’t go through. Let’s try another time.”

    Claim reviewed · no booking saved

    Review the evidence. Verify the fix.

    Review the claims against the saved outcome. After a fix, rerun the scenario and compare the evidence.

Fictional scenario and outcomes shown for illustration.

Explore how it works

Inside a check

The answer. The action. The evidence.

Explore a scheduling example: the same rejected update, handled three ways. Follow each claim through to the saved result.

Explore the product

The agent confirmed a booking the system had rejected.

Example data

What the agent said

All done — your appointment is booked for Tuesday the 14th at 9:30am with Dr. Alvarez.

agent · 00:00:31

What the booking service returned

error

status:
504
error:
upstream_timeout
appointment_id:

Request was sent. No appointment identifier was returned.

What was actually saved

0 appointment records returned

Query returned zero appointments for patient pt_sy_2841 on 2026-04-14.

Release decision: BLOCKED

The transcript reads like a clean, successful call. Only the saved state shows there is no appointment.

Book a walkthrough

1 of 3False confirmation

From test to decision

A clearer path to your next release.

Keep testing, evidence and review connected throughout your development process.

How verification works
  1. 01

    Exercise the workflow

    Choose a scenario with a known expected outcome. Include failure conditions, not just successful requests.

  2. 02

    Inspect the evidence

    Follow the transcript and tool receipts to the saved system state. Keep missing evidence visible.

  3. 03

    Review and repeat

    Review the agent’s claims, inspect the checks, and compare compatible runs after a change.

Across industries

Different workflows. The same need for proof.

Explore the conversations and outcomes your team needs to test—from customer support and account changes to reservations and travel.

Build with evidence

Give your next release
a reason to be trusted.

Book a walkthrough tailored to your workflow. See how tests, evidence and review fit together.

Book a demo

A guided session for your team