Test with real accents
The most common cause of a bad launch is an agent validated entirely by colleagues who all speak the same way. Recruit testers who match your actual caller base — regional accents, older speakers, people calling from a car or a factory floor. On a Saudi line that means Najdi speakers alongside Egyptian, Levantine, and South Asian expatriate speakers, because those are the people who will call.The scenarios worth running
The three most common requests
The three most common requests
Whatever your line exists for. These should work flawlessly before anything else is considered.
The caller who changes their mind
The caller who changes their mind
“Actually, no, it’s the other building.” Agents that have collected information often fail to update it, and confidently log the original.
The caller who volunteers everything at once
The caller who volunteers everything at once
Real callers open with a paragraph containing every detail out of order. Agents built around one-question-at-a-time often ask for what they were just told.
Mixed Arabic and English
Mixed Arabic and English
Product names and numbers in English inside an Arabic sentence. Check the agent stays in the right reply language rather than drifting.
The angry caller
The angry caller
Interruptions, raised voice, no patience for questions. Check the agent escalates rather than persisting.
Silence and noise
Silence and noise
Dead air, background conversation, a bad line. The agent should handle these without looping.
Out of scope
Out of scope
Something the line does not handle. The agent should route it, not improvise.
Read the transcripts, not the summaries
Summaries tell you what the agent thought happened. Transcripts tell you what happened. A call can produce a clean summary and still have frustrated the caller for two minutes. Read turn by turn and find the exact turn where each bad call went wrong. It is almost always one turn, and the fix is almost always one line of instruction.Keep a regression set
Save the calls that went wrong. After each instruction change, run them again. This takes minutes and catches the most common failure in agent development: fixing one case while breaking another. Without it you will make the same fix three times.Before you point real traffic at it
- Tested by people who sound like your customers, not just colleagues
- The three core requests work end to end
- Escalation works, and lands somewhere staffed
- Action failures produce a sensible reply, not an invented answer
- After-hours behaviour is defined
- Someone is assigned to read transcripts daily for the first week

