How to assess customer support agents — before and after you hire them
Every support leader has made this hire: great interview, warm manner, said all the right things about empathy — then three months in, they escalate anything that isn't in the macro list. The reverse happens too: the quiet candidate who turns out to be the one everyone routes hard tickets to. Both are the same failure. We assess support agents on how they talk about the work, and the job is actually about how they think through it.
Why the standard methods mislead
- Interviews measure articulacy. "Tell me about a difficult customer" rewards people who tell good stories. It cannot reveal whether someone can navigate an unfamiliar system and find the cause of a problem.
- CVs measure tenure. Five years in support can mean five years of investigating — or five years of pasting macros and escalating the rest. The CV looks identical.
- Knowledge quizzes measure memory. Product knowledge is learnable in weeks. The ability to reason through a ticket you've never seen before is the scarce skill, and quizzes don't touch it.
- QA reviews only see production. Post-hoc ticket reviews catch what went wrong after customers already felt it — a lagging metric, not a predictor — and they say nothing about capability you haven't hired yet.
What actually predicts performance: work samples
Decades of hiring research keep landing on the same conclusion: the best predictor of job performance is a sample of the job. For support work, that means: give the person a realistic ticket, a system that holds the evidence, and let them investigate. Watch what they do — not what they say they'd do.
The reason this works is that real tickets aren't black and white. Most candidates can find an answer. The difference between an average agent and a great one is who digs one layer deeper — who notices the order shipped to a different address than the customer is watching, who checks whether the discount actually applied rather than just validated. That extra layer is what prevents the follow-up ticket, and it is completely invisible in an interview.
What to measure
"Did they get the right answer?" is the wrong scoring model — it hides everything useful. Assess the thinking in dimensions. PRISM scores eight:
- Problem identification — did they find the real issue, or answer the surface question?
- Investigation quality — did they look in the right places and capture specific, usable detail?
- Logical rationale — does the reasoning connect the evidence to the resolution?
- Observation awareness — did they notice anything beyond the obvious?
- Resolution clarity — could the customer act on the response immediately?
- Calmness under pressure — measured and confident, or hedging?
- Self-sufficiency — solved without hand-holding?
- Creative thinking — workarounds, reduced customer effort, prevented the next ticket?
Two candidates who both "solve" the same ticket can score completely differently across these — and that difference is precisely what you're hiring for.
How to run a work sample yourself — no tools required
You can do this manually, today, and you should at least once — it will change how you interview. Here's the whole method:
- 1. Pick a real ticket from your queue — one where the obvious answer was wrong and the real answer needed one extra check. Strip customer details. The best tickets have a discoverable twist: the refund that went to a different card, the "missing" delivery sent to the address on the order.
- 2. Give the candidate the evidence, not the answer. Screenshots of the relevant system screens — order record, audit log, account history — including two or three screens that are irrelevant. Part of the skill is knowing what to ignore.
- 3. Ask for three things in writing: what they found, what they'd say to the customer, and what they'd do so it doesn't come back. Thirty minutes, open book. You're not testing recall.
- 4. Score against a rubric you wrote before you read any answers. Steal this one: found the actual cause (0–3) · response the customer could act on immediately (0–3) · noticed anything beyond the question asked (0–2) · calm, confident wording, no hedging (0–2). Ten points. Decide your pass line first.
- 5. Same ticket, every candidate. The moment you vary the ticket, the comparison is worthless — the entire value is scoring different people against identical evidence.
Common traps: don't grade tone over substance (interviewers reliably overweight polish); don't mark down bullet-point answers if the content is right; and don't let whoever will manage the hire score the papers alone — hopes contaminate rubrics.
Where PRISM comes in — the honest bit
The manual version works. Its cost is your time: building a sanitised scenario, screenshotting evidence, scoring every paper against the rubric, for every candidate, consistently, without drift. That consistency is exactly what PRISM Ready automates: candidates work a live ticket in a real system (not screenshots), the PRISM scoring engine scores all eight dimensions against pre-calibrated criteria, and you get ranked results with a recommendation per candidate — proceed, interview to clarify, or reject. Fifteen minutes of candidate time, none of yours.
For teams you already have: the same engine runs as scenario-based training — agents work through a library of scenarios across different sandboxed systems, get screen-by-screen coaching on their own investigation after every one, and managers see dimension-level strengths and gaps per agent on one dashboard. Assessment and development become the same motion: every scenario both measures and teaches.
Play one scenario as the agent — a real ticket in a live sandbox, scored across all 8 dimensions by the PRISM scoring engine. 15 minutes, no signup. Fair warning: nobody has hit a 10 yet.
Try the live scenario →