The agent opens a QA review and sees Empathy: 2/5. No excerpt. No explanation. Their last three customers wanted refunds the policy did not allow.
What should the agent do differently on the next ticket? The score cannot tell them. It can only tell them that somebody—or some model—judged them.
Supportman grounds feedback in specific conversations scored against a consistent rubric — the kind of evidence agents actually trust.
This is how feedback backfires in support: a manager intends to improve the work, but the intervention makes the agent think about fairness, status, or self-protection instead. The next conversation becomes more cautious, more scripted, and sometimes worse.
Feedback can make good agents worse
Kluger and DeNisi’s landmark meta-analysis examined 131 studies and 607 effect sizes. Feedback interventions improved performance on average, yet more than one-third reduced it. Feedback itself was not a reliable cure; its effect depended on where it directed the recipient’s attention. Read the Feedback Intervention Theory research.
Their Feedback Intervention Theory describes three broad levels of attention:
- Task learning: How do I perform this action better?
- Task motivation: How much effort am I applying, and how does my result compare with the standard?
- Meta-task or self: What does this judgment say about me?
Feedback is more useful when it keeps attention close to the task. It becomes less reliable as attention climbs toward the self. “Add the troubleshooting already completed to the transfer note” gives the agent a move to make. “You need to take more ownership” invites them to interpret a verdict about their character.
This does not mean feedback must be gentle. It means it must be usable. A clear account of a consequential mistake can be direct without turning the person into the mistake.
The four attention traps
| What the agent receives | Where attention goes | What to deliver instead |
|---|---|---|
| “You are careless.” | Identity and defence | The observable action, its effect, and the next action |
| “Resolution: 12/20.” | Guessing at the evaluator’s logic | Criterion, excerpt, rationale, and a stronger alternative |
| A low score caused by policy or product failure | Arguing about fairness | Separate agent-controlled behaviour from company-controlled conditions |
| A “coaching” meeting tied to formal consequences | Status and job security | Name the purpose and process before discussing the work |
The third trap is especially common in customer support. A customer can be unhappy because a refund window is restrictive, a feature is missing, the queue is understaffed, or an automated handoff lost context. Those are real service failures. They are not automatically agent failures.
Score what the agent controlled: accuracy, explanation, expectation-setting, ownership, and correct escalation. Tag the policy, product, or workflow defect separately and send it to the owner who can change it. Otherwise QA trains agents to perform contrition for company decisions they cannot alter.
The fourth trap needs honesty. Developmental feedback asks what to learn and practise; administrative evaluation determines consequences or rewards. If the same conversation serves both purposes, do not call it a safe coaching chat. Explain what is being decided, what evidence counts, and how the agent can challenge an error.
Triage the feedback before you deliver it
Before putting a low score into a one-to-one, run this five-minute check:
- Verify the evidence. Open the conversation and confirm the scored criterion actually applies. Check previous messages, internal notes, channel, customer history, and any required policy.
- Classify control. Mark the cause as agent-controlled, company-controlled, or shared. Coach only the controllable part; route the rest as operational feedback.
- Check recurrence. Pull five comparable conversations, not five random ones. As a practical trigger, coach a pattern found in at least three of the five. Treat one serious incident through the appropriate incident or performance process rather than pretending it is a trend.
- Choose one behaviour. If the review finds six gaps, select the one with the clearest customer impact and highest recurrence. Do not hand the agent a personality diagnosis disguised as a backlog.
- Define the follow-up sample. Decide now which five comparable conversations you will review and when. A vague promise to “keep an eye on it” is not a feedback loop.
The three-of-five trigger is an operating rule, not a scientific threshold. High-risk work may require action after one case; low-volume teams may need a longer window. Write down your rule so two managers looking at the same pattern do not invent different standards.
Once a pattern is verified, turn it into one specific goal, practice method, and review date. The detailed evidence and planning workflow belongs in the support QA coaching-plan guide.
A worked example: from score to next action
Imagine a subscription support team. An agent handles five requests from customers who missed the refund window. All five receive the correct policy answer. Three conversations, however, end immediately after “we’re unable to refund this charge.” The QA rubric marks empathy and ownership down.
The lazy feedback is:
You need to be more empathetic when you deny refunds.
That phrasing drags attention toward personality and leaves the agent to guess whether “more empathy” means a longer apology, a warmer tone, or granting an exception they cannot authorize.
The manager checks the evidence and separates two issues:
- Company-controlled: the refund window and lack of a self-serve renewal reminder.
- Agent-controlled: explaining the charge, acknowledging the unwanted outcome, and offering the available next step.
The revised feedback is:
In three of the five refund-window conversations, the reply gave the correct denial but ended without acknowledging the impact or offering the available cancellation and exception-review options. On the next refund denial, name the charge, acknowledge that the outcome is disappointing, and give the customer the next option you are authorized to offer.
A strong alternative might read:
The charge is for the annual renewal on 8 July. I understand you did not intend to renew, and I’m sorry this was unexpected. It falls outside our standard refund window, so I cannot issue the refund directly. I can cancel the next renewal now and submit this case for an exception review; that review normally follows the timeframe shown in our policy.
The example does not ask the agent to sound sorrowful on command or promise an outcome outside their authority. It gives them three observable actions. The manager can review the next five eligible refund conversations for those actions and separately raise the renewal-notification gap with the product owner.
Use a seven-minute feedback sequence
- State the scope: “I reviewed five refund-window conversations from the past two weeks.”
- Show the evidence: Put the three relevant excerpts on screen. Do not paraphrase from memory.
- Describe the impact: Explain what the missing action leaves the customer unable to understand or do.
- Ask for context: “What were you trying to achieve?” and “Is there a policy, tool, or queue constraint I have missed?”
- Agree on one target: Write the observable behaviour in language both people can recognize in a transcript.
- Test it: Have the agent rewrite one response or talk through the next case. If the instruction cannot survive a realistic example, it is still too vague.
- Book the follow-up: Name the date, sample, and success evidence before the meeting ends.
Seven minutes is a forcing function, not a limit for sensitive or high-stakes conversations. It prevents a routine coaching point from becoming a 40-minute monologue. If new context reveals a policy conflict, rubric defect, or capability issue, stop and investigate that problem on its own terms.
Use AI evaluation responsibly
AI evaluation can search a much larger body of conversations than a manager can sample manually. That makes it useful for finding candidates for review, not for skipping review.
Use it to:
- Surface the exact excerpts behind a score
- Group repeated misses on the same criterion
- Build a comparable before-and-after sample
- Flag disagreement between evaluators for calibration
Do not let it:
- Infer traits such as carelessness, attitude, or motivation
- Turn an unreviewed score into a disciplinary decision
- Score a conversation without the policy and preceding context
- Remove the agent’s ability to challenge evidence or rationale
Calibration matters most where the score carries the most weight. Review a sample of AI-flagged cases with managers and agents, record why judgments differ, and revise the rubric or evaluator instructions when the same ambiguity recurs.
Make the feedback about the work
A useful feedback record should let a third person answer five questions: Which task was reviewed? What happened? Why did it matter? What will change? When will the change be checked?
Supportman connects rubric scores to per-attribute reasons and conversation evidence, then surfaces repeated patterns for coaching. The score becomes the start of an inquiry instead of the final word.