Skip to content
Supportman
Support Operations

The Five Dimensions of Customer Support Quality

Imagine closing out the week at 92% CSAT with three escalations still open on your desk. The survey says the team is doing well. The escalations say something slipped — and the CSAT number cannot tell you what, because a customer’s rating compresses everything into one tap: the product, the policy, the wait, the agent, the entire relationship.

A QA rubric decompresses it. Instead of asking how the customer felt, it asks whether the answer was technically correct, whether the issue actually got resolved, whether the explanation was easy to follow, whether the agent sounded like your company, and whether anyone stayed responsible through the handoffs. Those five questions are the five dimensions of support quality, and this article maps each one: what it measures, what observable evidence looks like, and where scoring gets hard.

Supportman evaluates every conversation across the quality dimensions you define, so you can see exactly where your team is strong and where it slips.

This is the conceptual map. For the step-by-step build process — collecting examples, testing with reviewers, calibrating AI against humans, critical-failure rules — see How to Build a Customer Support QA Rubric.

Outcome signals versus quality drivers

Customer ratings capture the customer’s judgment, which is exactly why they resist diagnosis: the customer may be rating the product, the policy, the wait, the outcome, the agent, or the whole relationship, and the score does not say which.

A QA rubric examines the behaviours inside the interaction instead: what the agent knew, what the agent did, what the customer needed, whether the need was resolved, and how much effort the customer had to spend.

Research on Actionable Conversational Quality Indicators formalises this distinction for dialogue systems: rather than grading a conversation as broadly good or bad, it tags specific, improvable behaviours in the dialogue — indicators chosen precisely because a team can act on them. The five dimensions below apply the same principle to support conversations, human or AI.

1. Technical knowledge

Technical knowledge asks whether the information was correct, relevant, and safe. It is the easiest dimension to score because the evidence sits in the transcript and can be checked against documentation.

Useful criteria include:

  • The answer matches current product behaviour and policy.
  • Troubleshooting steps are logically ordered.
  • The agent does not invent unsupported facts.
  • Links and instructions lead to the intended resource.
  • Important limitations and risks are disclosed.

Technical accuracy should not reward unnecessary detail. An answer can be correct and still burden the customer with information they do not need, so measure accuracy alongside relevance.

2. Problem resolution

Resolution asks whether the interaction moved the customer to the outcome they needed.

Look for:

  • The agent identified the real goal, not only the opening question.
  • The answer addressed every material part of the request.
  • The agent verified assumptions where necessary.
  • A credible next step, owner, and timeframe were supplied when immediate resolution was impossible.
  • The conversation was not closed prematurely.

Resolution is not always “the feature now works.” A product limitation may prevent the ideal result, and the agent can still deliver strong resolution quality by explaining the constraint, offering the best available path, and keeping ownership.

Open conversations are the tricky case. Do not defer scoring until closure — score the trajectory instead: has the real goal been established, is there a named owner and a stated timeframe, and is it obvious whose move is next, the company’s or the customer’s? A conversation can be unresolved and still on track, and a closed conversation can score poorly precisely because it was closed too early.

3. Communication

Communication measures how easily the customer can understand and act on the answer. It is also the dimension most often written as vague two-word criteria — “clear structure,” “appropriate brevity” — that no two reviewers apply the same way. Write it instead as checks a reviewer can verify by pointing at the transcript:

  • The direct answer to the customer’s question appears in the first two sentences of the reply, before background or caveats.
  • Instructions with more than two steps are numbered, in the order the customer will perform them.
  • The reply uses the customer’s vocabulary, not internal feature names or team jargon the customer never used.
  • Every link says what the customer will find there, and no step depends on a resource that was not linked.
  • After one read, the customer knows what to do next and who acts first.
  • The reply does not restate information the customer already provided beyond a one-line confirmation.

Whether “brief” means 80 words or 300 depends on the channel — decide per channel and write the threshold down, so reviewers are checking a number rather than a feeling. Evaluation research such as LLM-Rubric takes the same approach, scoring qualities like naturalness and concision as separate questions rather than one overall impression; the rubric build guide covers how to apply that structure in practice.

4. Brand voice and tone

Tone asks whether the interaction feels respectful, human, and aligned with the company.

Useful criteria include:

  • Acknowledges impact when appropriate
  • Matches the customer’s level of formality
  • Remains calm during frustration
  • Avoids defensiveness and blame
  • Sounds natural rather than scripted
  • Does not use warmth to avoid giving a direct answer

Empathy should be situational. Requiring an apology or emotional phrase in every interaction can make support sound mechanical.

The standard is not “use more empathy words.” It is “respond appropriately to the customer in front of you.”

5. Ownership and handoff quality

Ownership measures whether responsibility stays clear from first reply to resolution — and it is the hardest of the five to score, because the evidence spans multiple messages, multiple days, and sometimes multiple reviewers. A single reply can look flawless while the thread around it shows a customer being passed between three teams.

Look for:

  • The customer is not asked to repeat information already provided.
  • Internal handoffs preserve context.
  • The agent explains what will happen next.
  • Escalations identify an owner.
  • Follow-up commitments are kept.
  • The customer is not bounced between teams without guidance.

Two scoring rules make this dimension workable. First, score ownership at the conversation or ticket level, never per message — the failure mode is the gap between messages, not any one of them. Second, check commitments against what actually happened later: “I’ll follow up tomorrow” only counts as evidence when tomorrow’s message exists. That usually means the reviewer needs the full thread and any linked escalation, which is why ownership reviews take longer and why it is worth budgeting for them separately.

Ownership belongs in the rubric because a conversation can contain polite, correct messages and still feel abandoned.

Scoring one ticket across all five dimensions

Here is an illustrative ticket. A customer writes: “I can’t find the invoice export button — I need all of last quarter’s invoices for our accountant by Friday.” The export feature exists, but it is gated to the Business plan and this customer is on Starter. The agent replies warmly, correctly explains the plan gating, offers to email a one-off CSV “as a workaround, just this once,” and closes the conversation.

Scored on a 1–5 scale:

DimensionScoreWhy
Technical knowledge5Correctly identified the plan gating instead of “troubleshooting” a button that was never going to appear, and named the plan that includes the feature.
Problem resolution4The real goal (invoices to the accountant by Friday) was understood and met. Loses a point for not addressing the recurring need — this customer will hit the same wall next quarter.
Communication4Answer first, one clear next step. The plan explanation ran three paragraphs where one would do.
Brand voice and tone5Matched the customer’s urgency, no scripted apology, no defensiveness about the paywall.
Ownership2Closed the conversation with the commitment still open: nobody is named to send the CSV, no deadline is stated, and if the agent is out Thursday the promise silently dies.

This customer almost certainly rated the interaction 5 — they were treated kindly and promised what they needed. The QA score surfaces two things CSAT never will: an unowned commitment, and an off-policy workaround (“just this once”) that was never logged anywhere. That gap between the rating and the evaluation is the entire argument for scoring both.

Weighting the dimensions by team type

The five dimensions should not automatically receive equal weight — the right split depends on what a failure costs your operation. The mechanics of weighting, critical-failure rules, and separating agent performance from product or policy failures are covered in the rubric build guide; the short version is to weight what the agent controls and route everything else to product feedback rather than agent scores.

What this article adds is how the weights shift by team type. Three illustrative profiles:

DimensionRegulated fintechHigh-volume consumerPremium B2B concierge
Technical knowledge35%20%20%
Problem resolution25%35%25%
Communication15%25%15%
Brand voice and tone10%10%15%
Ownership15%10%25%

The fintech profile weights accuracy heavily because a wrong answer there is a compliance incident, not just a bad experience — and it is the kind of team that also needs a critical-failure rule capping the total score when guidance is unsafe. The consumer profile weights resolution and communication because volume punishes repeat contacts and long replies more than anything else. The concierge profile weights ownership because its customers are paying specifically to never be bounced between teams.

If you cannot explain why your top-weighted dimension deserves its weight in one sentence about customer harm, you have inherited someone else’s priorities.

Use CSAT and QA together

The most useful analysis compares outcome and process, because each quadrant demands a different response from the manager:

Customer ratingQA evaluationLikely interpretationWhat to do next
PositiveStrongExperience and execution alignedPull excerpts into the rubric’s example library — this is where your “exceeds expectations” anchors come from.
NegativeStrongProduct, policy, wait, or expectation is the likely causeTag as product feedback and exclude from coaching stats; route the recurring themes to the product team with the transcripts attached.
PositiveWeakCustomer accepted an interaction that still contains riskSample these in calibration sessions — they hide the failures CSAT will never surface, like off-policy workarounds and unowned commitments.
NegativeWeakClear recovery and coaching candidateRecover the customer first, then coach from the specific criterion that failed rather than the overall score.

The invoice ticket above lands in the Positive / Weak row: the customer got their files by Friday and likely rated the experience a 5, while the conversation carries an unlogged workaround and a promise nobody owns. Teams that only review low-CSAT conversations never see that row at all, which is where quality erodes quietly.

Measure what your team can improve

Supportman evaluates every eligible conversation against a custom set of weighted attributes and criteria — the five dimensions here, tuned to your weights — so the gap between “customers seem happy” and “the work was actually strong” shows up in your data instead of in next month’s escalations.

Score your team on these five dimensions →

Frequently asked questions

How many QA attributes should a support rubric contain?

Four to six is a useful starting range. Too few hide the cause of a score; too many create overlap and review fatigue.

Should CSAT be part of the QA score?

No — keep them side by side. CSAT is one subjective rating of the whole experience; folding it into the QA score lets it override five behavioural evaluations and hides the positive-rating, weak-QA conversations where risk accumulates.

Can AI score all five dimensions?

AI can evaluate conversation evidence at scale, but it should be calibrated against human-reviewed examples and monitored by attribute, channel, language, and issue type.

Five minutes to live, no IT ticket required.

See pricing