Skip to content

Testing an IVR without calling it

· ivrloom team

#testing #ivr-design #five9

The default way to test an IVR change is to call it. Dial in, press the keys, listen. It works, and it does not scale — because the paths most likely to be broken are the ones least likely to be dialled.

What dial-testing is good at

Genuinely useful things, and worth keeping:

  • Audio quality and whether the right recording plays
  • Timing and pacing — whether a prompt feels rushed
  • The end-to-end experience of the main path

What it systematically misses

Branches nobody dials. The holiday path, the after-hours path, the third retry, the “customer not found” branch. You cannot test the holiday path in August without changing the system clock or the flow.

Paths that need specific state. Anything conditioned on a CRM lookup, a variable set upstream, or a queue depth. Reproducing that by calling means arranging the state first.

Regressions in what you did not touch. You changed the billing branch. Did anything else move? Dial-testing answers this only if you re-dial everything, which nobody does.

What is verifiable without a call

More than most teams expect, because a call flow is a graph and a lot of questions about graphs are structural. These are the checks ivrloom’s Issues panel runs on every edit, and each one is a question you could answer by hand with enough patience:

CheckThe question it answers
Unreachable node / branchIs there any path from incomingCall to this module, or to this whole region of them?
Dead-end nodeDoes this non-terminal module have somewhere to go?
No mis-entry handlingDoes this menu have somewhere to send a caller who presses nothing, or a key it doesn’t offer?
Transfer has nowhere to go if it failsIf this transfer fails, is there an error branch or a next module for the call to continue to?
Possible infinite loopIs there a cycle with no exit?
Branch points to a missing nodeDoes this desc reference a module that exists?
Variable read but never setCan a script variable this condition tests still be NULL when it runs?
Prompt not in catalog / Unused promptDo the prompt references and the prompt library agree?
Long prompting before reaching an agentHow many prompts must a caller hear in full, unable to skip, before a transfer?

None of these needs a phone, and several of them — the missing branch, the variable nothing sets — cannot be found by phone at all, because you cannot dial your way to a node no path reaches.

Deterministic replay is the other half

Structure tells you the graph is well-formed. It does not tell you the call goes where you meant. For that you run it.

Given a fixed set of inputs — these digits, this time of day, these variables — the path through the flow is fully determined. That means you can record the expected path once and re-check it after every edit. In ivrloom that is a saved scenario: the DTMF you queued, any variables you set, the clock, and the outcome the run produced. Pin it, and it becomes an assertion: from then on the panel replays it against the current flow and reports pass or fail.

A few properties of that replay matter more than they sound:

  • The time is an input. The simulator never reads the real clock. __DAY__, __TIME__ and __DATE__ come from a clock you set in the panel, so the holiday branch is testable in August by typing a date.
  • Silence is an input. At any digit collection you can say “the caller stayed silent” and watch the no-input path, which is the path most often broken and least often dialled.
  • A dead end is always a failure. If a replayed call ends with no path forward, the scenario fails even if that dead end was what you pinned. The panel will not let you bless a broken flow as the new baseline; it names the node the call could not leave.
  • Coverage is reported. With a handful of scenarios saved, the panel lists every branch none of them takes — typically the ELSE of the hours check, the No Match option, the error exit — so you know what is still untested rather than assuming.

Six or eight recorded scenarios, replayed after each edit, catch the class of problem that dial-testing structurally cannot: something you did not touch, breaking.

The paths worth recording first

If you are starting from nothing, these are the scenarios with the best return:

  1. The main path, one per top-level menu option.
  2. Silence at the first collection. What does a caller who says nothing hear, and where do they end up?
  3. An unlisted digit at the main menu. Do they get a “not a valid choice” or a disconnect?
  4. After hours, with the clock set to a weekend evening.
  5. Each holiday on your list, with the clock set to that date — and the day after each, to confirm normal service resumes.
  6. Lookup failed, for any CRM-driven branch, by leaving the looked-up variable empty.

That is roughly ten scenarios for a typical main flow, and each takes under a minute to set up.

What still needs a phone

Audio. Whether the recording is the right recording, whether the voice matches, whether there is a click at the start. Simulation tells you which prompt plays, not whether it sounds right. (ivrloom will play the recording from your bundle in the browser, which covers “is this the file I think it is” — but not how it sounds over a real line.)

So the split is: structure and routing before you deploy, audio after. The point is not to stop calling — it is to stop using calls to answer questions a graph can answer faster and more completely.