Skip to content
Supportman
Support Operations

How to Build a Customer Support QA Rubric Agents Can Actually Use

“Deliver excellent support” is a worthwhile ambition. It is a terrible scoring criterion.

One reviewer may interpret excellence as speed. Another may reward warmth. A third may focus on technical accuracy. The final QA score looks precise, but the judgment underneath it changes from person to person.

Supportman scores every Intercom conversation against your own QA rubric automatically — no sampling, no spreadsheet.

A useful rubric turns broad expectations into distinct, observable criteria with a defined scale, explicit weights, and rules for the awkward cases. This article builds one end to end. To keep it concrete, we will follow an illustrative team throughout: imagine a 12-person support team at a B2B scheduling product, handling email and chat in Intercom, whose biggest complaint from customers is “I had to contact you twice for the same thing.”

Start with the purpose

Before choosing attributes, define what the rubric is intended to improve: more complete resolutions, fewer repeat contacts, clearer product guidance, safer policy compliance, a more consistent voice. A rubric used for agent development needs richer explanations than a compliance audit; a rubric spanning chat and phone may need channel-specific criteria. Do not build one scorecard for every purpose.

Write one sentence:

This rubric helps us evaluate whether a support conversation delivers ___ for ___.

The scheduling team writes: “This rubric helps us evaluate whether a support conversation delivers a complete, first-contact resolution for business customers.” That sentence becomes the filter for every attribute that follows — an attribute that does not serve first-contact resolution has to argue its way in.

Pick and weight a short list of attributes

Four to six attributes is a workable starting heuristic, not a law. Fewer than four and the score hides why a conversation succeeded or failed; many more and reviewers start pattern-matching instead of judging. A regulated team might justifiably run eight with a compliance block; a two-person team might run three.

We break down a full starting set — technical knowledge, problem resolution, communication, brand voice and tone, ownership — in The Five Dimensions of Customer Support Quality, along with the argument for weighting them unevenly. The short version: unweighted attributes just mean reviewers apply their own hidden weights.

The scheduling team picks four, folding ownership into problem resolution because their purpose sentence is about first-contact resolution:

AttributeWeightWhy this weight
Technical knowledge30%Wrong answers cause the repeat contacts they are trying to eliminate
Problem resolution30%The purpose sentence is literally about resolution
Communication20%Matters, but a clear non-answer is still a non-answer
Brand voice and tone20%Table stakes for B2B customers; rarely the reason someone contacts twice

Scoring dimensions separately has research behind it. The LLM-Rubric framework evaluates distinct questions about qualities such as naturalness, concision, and citation quality, then combines them into an overall prediction; in its study setting it reported substantially better agreement with human judgments than an uncalibrated baseline. Read that as evidence for multidimensional scoring generally — it does not validate any particular attribute list or weighting for support QA. Yours has to come from your purpose sentence.

Define observable criteria

Each attribute should describe evidence a reviewer can find in the conversation.

Weak criterion:

Agent showed empathy.

Stronger criteria:

  • Acknowledged the customer’s stated impact when the situation called for it
  • Did not use generic sympathy in place of action
  • Maintained a calm, respectful tone during frustration
  • Avoided blaming the customer or another team

Weak criterion:

Agent resolved the issue.

Stronger criteria:

  • Identified the customer’s actual goal, not just the stated symptom
  • Supplied a complete answer or appropriate action
  • Verified assumptions before closing
  • Made ownership and the next step explicit when immediate resolution was impossible

Observable criteria make scores explainable, and they follow the logic of behaviourally anchored rating systems: attach evaluation to examples of work rather than personality labels.

Choose a rating scale and an N/A rule

A rubric without a defined scale is a list of opinions. Three decisions turn criteria into a scorecard:

Scale. A 0–2 scale per attribute is the easiest to calibrate: 0 = missed the standard, 1 = partially met it, 2 = fully met it. Five- and seven-point scales feel more precise but mostly manufacture disagreement — two reviewers who agree a reply was “good” will still split between 5 and 6. Start narrow; widen only if a real coaching decision hinges on the distinction.

Anchors. Every score level needs an observable description, written per attribute, so “1” means the same thing to every reviewer. The worked scorecard below shows a full set.

N/A handling. Some attributes will not apply to some conversations — technical knowledge on a billing-address change, for instance. Decide the rule in advance: mark the attribute N/A and redistribute its weight proportionally across the remaining attributes. Never score an inapplicable attribute 2 “to be fair” — that silently inflates scores on easy tickets.

The worked scorecard

Here is the scheduling team’s complete rubric — the thing a reviewer or an AI evaluator actually holds while scoring. Copy the structure, replace the content.

Attribute (weight)0 — missed1 — partial2 — met
Technical knowledge (30%)Gave incorrect or outdated information, or guessed without flagging uncertaintyAccurate but incomplete; missed a relevant limitation, setting, or known issueAccurate, complete, and relevant to this customer’s configuration
Problem resolution (30%)Ignored the underlying goal, repeated an already-failed suggestion, or closed prematurelyAddressed the goal but left a loose end: no verification, vague next step, or unowned handoffResolved the actual goal, or escalated with a named owner, reference, and timeframe
Communication (20%)Confusing structure, jargon, or key information buried or missingUnderstandable but inefficient; customer must reread or infer the stepsClear, ordered, skimmable; the customer knows exactly what to do next
Brand voice and tone (20%)Cold, defensive, blaming, or wildly off-registerPolite but templated; ignored stated frustration or effortWarm, professional, acknowledged the customer’s specific situation

Score calculation. Each attribute contributes (score ÷ 2) × weight. A conversation scored 2 / 1 / 2 / 1 works out to 30 + 15 + 20 + 10 = 75%.

N/A recalculation. On a billing-address change, technical knowledge is N/A. The remaining weights (30/20/20) are normalised to 100 by dividing by 0.7: resolution becomes ~43%, communication and tone ~28.5% each. The conversation is scored out of what could actually be observed.

Critical failure. One rule sits outside the arithmetic: if the agent gives account access or account data to an unverified requester, the conversation scores 0 regardless of the attribute scores. A 95% conversation with a verification failure is not a 95% conversation. The next section makes that rule precise.

Write the critical-failure rule down

Critical-failure rules exist so that catastrophic mistakes cannot be averaged away by good manners. Use them sparingly — one or two, for genuinely unacceptable outcomes — but write them with the precision of a policy, because they will be contested. A vague rule (“security issues cap the score”) creates exactly the reviewer-by-reviewer variance the rubric was meant to remove.

The scheduling team’s rule, in full:

  • Trigger: the agent shares account data, grants access, or changes account ownership without completing identity verification as defined in the verification policy.
  • Effect: the conversation’s total score is set to 0 (not capped, not discounted) and the conversation is flagged as an auto-fail.
  • Correction edge case: if the agent catches and corrects the error in the same conversation before the customer acts on it, it is not an auto-fail — but technical knowledge scores 0 and the conversation is flagged for lead review anyway.
  • Confirmation: no auto-fail lands on an agent’s record until a second reviewer or team lead confirms it. False auto-fails destroy trust in the whole rubric.
  • Missing evidence: if the reviewer cannot see whether verification happened — say it occurred by phone and the notes are silent — the conversation is marked “insufficient evidence” and excluded from auto-fail counts. Nobody gets zeroed on a guess.

Whatever your equivalents are — regulated financial advice, medical claims, data-deletion promises — give each one this same four-part treatment: trigger, effect, edge cases, confirmation.

Score real conversations, not hypotheticals

Criteria only become usable when paired with scored examples. For each attribute, collect an example that exceeds, one that meets, a borderline case, a clear miss, and an N/A case — identifiers removed, reviewed with the team. Here is what that looks like in practice for the scheduling team.

The customer writes:

Your Outlook integration keeps dropping my calendar sync. I’ve reconnected it twice this week and I’m honestly done troubleshooting this myself.

Reply A: “So sorry for the trouble! Reconnecting the integration usually fixes this — here’s our step-by-step guide. Let us know if you need anything else!” Polite, accurate link, tidy formatting. It scores: technical knowledge 1 (the article is accurate but ignores that reconnecting has already failed twice), problem resolution 0 (repeats an already-failed fix and closes prematurely), communication 2, tone 1 (friendly but ignores the stated frustration). Total: 15 + 0 + 20 + 10 = 45%.

Reply B: “Reconnecting twice and losing sync again means something deeper is wrong — you shouldn’t have to keep doing that. I’ve checked and there’s a known token-refresh issue affecting some Outlook accounts; I’ve added yours to the engineering ticket (SYNC-482) and I’ll update you by Thursday either way. Until then, here’s a one-click workaround that holds the connection.” It scores 2 / 2 / 1 / 2 — communication drops a point for a dense paragraph that buries the workaround — for 30 + 30 + 10 + 20 = 90%.

Two superficially similar replies — both prompt, both polite, both “helpful” — separated by 45 points, for reasons a reviewer can point at. That is the whole value of anchors.

Include at least one genuinely difficult case in your example set. The scheduling team’s: a customer asks for two-way sync of recurring events, the agent correctly explains the product does not support it, offers the closest workaround, and logs the feature request — and the customer leaves a 1-star CSAT. The rubric scores it 2 / 2 / 2 / 2, with attribution marked “product or policy.” If your examples never show a high QA score alongside a bad CSAT, reviewers will quietly learn to score the customer’s mood instead of the agent’s work.

Separate agent performance from system failure

A fair rubric does not punish an agent for a missing feature, an approved policy, an engineering delay, a misrouted conversation, or a tool outage — it evaluates what the agent controlled: clarity, expectation-setting, ownership, escalation quality, accuracy. We cover the reasoning in the quality dimensions article; what the rubric itself needs is one extra field on every scorecard:

  • Agent-controlled
  • Product or policy
  • Workflow or tooling
  • Shared responsibility
  • Insufficient evidence

The attribution field is what turns QA from an agent leaderboard into an organisational instrument. When the scheduling team’s scores dip every time the Outlook issue flares up, attribution data shows the problem is SYNC-482, not the agents answering for it.

Test before launch

Before the rubric touches anyone’s performance record, pilot it: have at least two reviewers independently score the same set of representative conversations. For a team the size of our example, 30–50 conversations across channels and issue types is a reasonable starting sample; a high-volume, multi-language operation needs more before the results mean anything.

Then make an actual launch decision, against thresholds you set in advance. A workable starting rule — treat these numbers as defaults to adjust, not standards to cite:

  • On each attribute, reviewers give the identical score on at least 80% of conversations (on a 0–2 scale, adjacent scores are still disagreement — the scale is too coarse for “close enough”).
  • On critical-failure calls, aim for 100% agreement and investigate every single miss. Reviewers who disagree about auto-fails do not yet share a rubric.
  • Every disagreement gets adjudicated in discussion, and the criterion that caused it gets rewritten or gains a new anchored example. Do not solve disagreement by telling reviewers to “be consistent.”
  • Score a fresh batch of 15–20 conversations with the revised rubric. Launch when the second round clears the bar; repeat if it does not.

Also compare scores by channel and issue type, and ask reviewers which scores felt unfair — perceived unfairness in the pilot becomes open resistance after launch. The value of making the rubric this specific to your domain has empirical support: Microsoft’s RUBICON research found that domain-specific rubrics differentiated conversation quality better than generic baselines in its evaluation of developer-tool conversations. Take it as evidence for specificity, not as a template for support QA.

Roll out with a cadence, not a memo

A rubric agents can “actually use” is mostly a matter of how it operates after launch. The loop that works:

  • Involve agents before launch. Have the team review draft anchors and argue with them. Agents who helped write the standard defend it; agents who received it by memo appeal every score.
  • Publish scored examples. The example library from the pilot — including the difficult cases — lives somewhere every agent can read it. New anchors get added as new patterns appear.
  • Hold a weekly calibration session during launch. Thirty minutes, one contested conversation, everyone scores it independently, then discussion. After six to eight weeks, drop to monthly.
  • Allow score challenges. A defined window (say, five business days) and a defined adjudicator. Challenges are free QA on the rubric itself: repeated challenges on one criterion mean the criterion is broken.
  • Coach on patterns, not incidents. One low score is noise; the same criterion missed across three or more conversations is a coaching topic. We cover turning those patterns into a plan in How to Turn QA Scores Into a Coaching Plan.
  • Version the rubric. Every change gets a version number, an effective date, and a line in a changelog. Conversations are always scored against the version in force when they happened — rescoring history against new rules is how trust dies.

If an AI evaluator will apply the rubric at full coverage, it joins this same loop: it gets the identical anchors, weights, N/A rules, and critical-failure definitions as human reviewers, and its agreement with your calibration set is checked at the attribute level before you rely on it. That process is its own discipline, covered in How to Calibrate the Judge.

Rubric checklist

  • One-sentence purpose statement
  • Roughly four to six distinct attributes, chosen against that purpose
  • Observable criteria per attribute
  • A defined scale with written anchors for every score level
  • Explicit weights and a documented score calculation
  • An N/A rule with weight redistribution
  • Critical-failure rules with trigger, effect, edge cases, and confirmation
  • Agent-control attribution field
  • Scored example library, including hard cases
  • Two-reviewer pilot with pre-set agreement thresholds
  • Launch cadence: calibration sessions, challenge window, pattern-based coaching
  • Version numbers, effective dates, changelog

From draft to version one

The sequence is short even if the work is not: write the purpose sentence, pick and weight the attributes, anchor every score level, add the awkward rules — N/A, critical failures, attribution — then pilot on real conversations, rewrite whatever caused disagreement, and launch it as version 1.0 with a date on it. A rubric that has survived that process will score your hardest conversations defensibly; one that has not is a mood with a spreadsheet.

Once version 1.0 exists, applying it to every conversation instead of a 2% sample is the part software should do. Supportman scores every eligible Intercom conversation against your attributes, weights, and critical rules — not a generic definition of “good support.”

Build your rubric in Supportman now →

Five minutes to live, no IT ticket required.

See pricing