Testing guide

How to Test AI Call Screening Before Going Live

A good conversation and a completed action are different things to verify.

Build a small acceptance test that follows a caller’s inquiry from entry to the record your team will rely on.

Decide what your test can prove.

A clear voice is useful evidence about presentation. It is not enough to accept a call-handling workflow. Before testing, write down the question you want answered: does the wording make sense, does this caller enter the right screen, or did an action actually complete?

Three kinds of evidence, three different limits
EvidenceUseful forWhat it does not prove
Scripted voice illustrationListening to selected wording, voice character, and an illustrative conversation.That an agent independently interpreted answers, followed the configuration, completed actions, or achieved the illustrated timing.
Interactive browser testTrying your saved configuration with questions, paraphrases, and caller responses; reviewing the conversation.That a real appointment, phone handoff, message, or business record was completed. Preview action language is rehearsal.
End-to-end action verificationChecking the intended call path and its resulting action and destination records under an authorized test.That every future caller or failure condition will behave the same way. Record which cases remain untested.

RevSystems’ Test your agent guide separates browser conversation review from real-call action verification. Use its setup instructions when you are ready to run an authorized test; this worksheet helps you decide what evidence to collect.

Follow the inquiry all the way to the record.

Inspect five things separately. If the wrong caller enters screening, better question wording will not fix the entry setting. If the recorded answers are wrong, a correct-looking next step can hide a bad decision. If the result is right but an action cannot complete, the next investigation belongs in action readiness or availability.

  • Entry: which enabled greeting, direct request, or topic should begin screening? Which ordinary calls should stay in reception?
  • Caller answers: what did the caller actually establish, and does the recorded information preserve uncertainty or a correction?
  • Screening result: does qualified or not-qualified agree with the saved questions and pass rules?
  • Chosen action: is it allowed for that result and ready to use? Which alternative applies if the first option is unavailable?
  • Persisted record: what evidence shows the action was attempted, failed, or completed for this particular call?

In RevSystems, entry leads into exactly two or three configured criteria. Entry is separate from the criteria; the qualified or not-qualified result is separate from the action that follows. Keep a dated copy of your saved settings with the test notes so a later configuration change does not make an old result ambiguous. See structured qualification.

Qualified callers can use ready handoff, booking, or message actions. Not-qualified callers cannot use live handoff. The agent tries the first available action in the configured order and stops after one succeeds. Disabling an action can leave its order visible for editing without making it usable. Verify the outcome and readiness rules.

Keep a worksheet with room for what really happened.

Copy these rows into your test notes. Replace the inputs and expectations with your own fictional cases before running them. For every run, note the configuration version or saved date, test mode, time, and call reference if available. The “What was heard” and “Observation” cells below are intentionally blank. “Not tested” is the starting status, not a result.

Screening evidence worksheet — no tests are represented as completed
Scenario / inputConfigured expectationWhat was heardResult / action evidence to inspectObservationStatus
Normal receptionAn ordinary opening-hours question stays outside screening under the intended entry settings. Entry behavior and any answer/result recorded for this call; approved business information used. Not tested
Intended entryA clear program inquiry enters through the enabled greeting, request, or topic. Test enabled methods separately. Entry decision and first criterion; do not count the greeting as an extra criterion. Not tested
Answers that fitSelf-reported answers satisfy the saved rules and lead toward the chosen qualified action. The recorded answers, qualified result, selected action, and evidence of its actual status. Not tested
One answer does not fitChange one relevant fact. Determine the expected result and allowed next step before the run. The changed answer and result; the selected action must remain eligible for that result. Not tested
Unclear or corrected answerSpecify what would count as confirmed information. Try an uncertain answer and a later correction as separate runs. What was understood, what was recorded, and whether any result/action rests on an unsupported assumption. Not tested
Primary next step unavailableUse an authorized unavailable-destination case with a useful configured alternative. The primary attempt or availability decision, alternative chosen, and final action status. Not tested
Booking or messages unreadyTest each selected action when its required setup is incomplete or disabled, using an isolated authorized configuration. Readiness settings and observed behavior; a visible saved order is not evidence of a completed action. Not tested
Spoken success, missing recordRequire evidence for any claim that an appointment, message, or handoff completed. Correlate the call with the expected action/destination record. Investigate an absent or conflicting record before marking pass. Not tested

Use pass only when the relevant expectation and evidence agree. Use fail for a confirmed mismatch. Keep not tested when the required evidence is absent or the case has not run, and explain why in Observation. An unclear answer or unusual phrasing is a case to investigate; this worksheet makes no guarantee about an untested recovery path.

A consultation discussed is not a consultation booked.

In the later-start Mortgage example, the fictional homebuyer reports a planned start in six months. That does not fit this example’s 90-day question, so the script discusses a consultation, with a message as the alternative. The caller expresses interest. The clip stops there.

This scripted illustration performs no screening or booking. For a future interactive test of that configuration, compare the recorded timing answer and result with the permitted next step. For a separately authorized action test, also check the appointment or message record associated with that call. Neither an interested caller nor an offer to book establishes completion.

The example uses an unverified caller estimate of credit score; it is not a credit check or loan decision. The point to carry into your own worksheet is the evidence boundary: conversation, screening result, and completed action each need their own support.

Look where the completed work should be.

For an authorized action test, use Home and Inbox to inspect the call and follow-up record. Match the time and call reference, then inspect the evidence relevant to the action. A transcript records words; use action status and the appropriate destination record to establish what completed.

  • Handoff: check the call’s handoff result and whether the intended destination was actually reached. An offer to connect is insufficient.
  • Appointment: verify a corresponding booking and its time/details in the configured booking surface. Mentioning an available time is insufficient.
  • Message: verify that the message was saved for the correct follow-up context. A saved message and delivery through another channel are different claims.
  • Unavailable or failed action: record what did not complete and whether a ready, eligible alternative actually succeeded. Do not silently mark an attempt as success.

Use the exact preparation guides for handoff, calendar and booking, and message taking. Do not assume a system integration, notification delivery, or professional decision that your configuration and records do not establish.

Make one small change, then repeat the useful cases.

Start with an initial test set that covers your main call path and each materially different alternative. Use fictional information and an appropriate authorized environment; reserve tests that can contact people or change business records for a separately approved plan. Keep those cases marked not tested until that verification is done.

  • Write down the discrepancy at the earliest stage where evidence diverges: entry, answer, result, action choice, or record.
  • Change one relevant setting, save it, and record the change. Begin a fresh test so earlier conversation does not obscure the comparison.
  • Repeat the failing case and the nearby main path that already worked. If you change entry, include a reception call that should stay outside screening.
  • If you change action readiness or ordering, review both result paths and retest the affected alternatives. Preserve the earlier evidence alongside the new result.
  • Decide which remaining gaps need investigation before launch and who owns them. A small selected test set is not a reliability benchmark.

For button-by-button instructions, use Test your agent. For the broader launch checklist, use Test before launch. Return to the screening overview for the operating model, or the AI receptionist and answering-service comparison if you are still choosing who should handle your calls.

Your next step

Prepare a screen your team can review.

Start with your business facts and chosen next steps. Set up RevSystems, then test the paths that matter to your team.

Start free