Skip to content
Log in
Team & Leadership

Customer Service Quality Assurance for Intercom Teams (2026)

customer service quality assurance

Customer service quality assurance (QA) is a repeatable process for scoring conversations against defined criteria, calibrating reviewers, and turning scores into coaching — not a quarterly spreadsheet nobody opens. For Intercom teams, QA works when the rubric is specific, coverage exceeds the ~10–20% of threads that get a CSAT rating, and agents see feedback within days of the conversation, not weeks.

This guide covers what to measure, how to run calibration, and where Supportman fits if you want AI scoring on every closed conversation instead of a manual sample.

CSAT response rates sit below 20% for most teams, which means manual QA on a sample misses most conversations. Supportman scores every closed Intercom thread against your rubric and posts weekly IQS trends to Slack so coverage is not limited to customers who happened to tap a rating.

What customer service QA actually means

QA answers one question: did this conversation meet the standard we promised customers? That standard should be written down as weighted attributes with observable criteria — not "be helpful" or "show empathy."

A mature QA programme has four parts:

  • Rubric: four to six attributes (resolution, product knowledge, tone, compliance) with a defined scale and critical-failure rules.
  • Scoring: human reviewers, AI scoring, or both on closed conversations.
  • Calibration: reviewers score the same thread independently until scores align within an agreed band.
  • Coaching: one behaviour change per agent per cycle, tied to a specific attribute — not a generic "do better."

See the five dimensions of customer support quality for a starting attribute set and how to build a rubric agents can actually use for the full worked example.

Why CSAT alone is not enough

CSAT tells you how a customer felt after one conversation. It does not tell you why the conversation went well or badly, and it only covers the minority of customers who submit a rating.

Signal What it tells you Blind spot
CSAT / DSAT Customer sentiment on rated threads Selection bias; policy issues scored against agents
Manual QA sample Detailed rubric scores on chosen threads 5–10% coverage; reviewer drift without calibration
AI QA (IQS) Rubric score on every closed conversation Needs a good rubric and periodic human calibration
Reopen rate Whether the issue actually stayed fixed Does not explain tone or knowledge gaps

Route negative CSAT to a #dsat Slack channel for fast recovery. Use QA scores for the coaching conversation the following week.

Build the rubric first

Do not buy a QA tool before you can score ten real conversations consistently. Start with:

  1. One-sentence purpose: this rubric evaluates whether we delivered a complete first-contact resolution for business customers.
  2. Four to six weighted attributes aligned to that sentence.
  3. Observable criteria per attribute (what a 3 vs 5 looks like on a real thread).
  4. A critical-failure rule (e.g. sharing another customer's data = automatic fail).

Run a calibration session: three reviewers score the same five Intercom conversations. Debate gaps until scores converge. That session is worth more than any software purchase.

When you add AI scoring, point it at the same rubric. Calibrate AI customer support QA explains how to tune prompts until machine scores match human judgment on a gold set.

Sampling vs full coverage

Approach Best for Trade-off
5–10% manual sample Teams under 8 agents with a dedicated lead Misses patterns in unrated, middle-quality threads
100% AI scoring + human review of exceptions Teams above 10 agents or with Fin in the queue Requires rubric maintenance and calibration cadence
CSAT-only Early-stage teams with no QA bandwidth Coaching on anecdotes, not representative data

If Fin handles part of your queue, score Fin and human conversations separately. AI-to-AI benchmarks differ from human CSAT targets — see Intercom CSAT benchmarks.

Calibration and coaching cadence

Monthly calibration: same five conversations, all reviewers, compare scores. Update rubric language when reviewers disagree by more than one point on the same attribute.

Weekly coaching: each agent gets one conversation review — lowest IQS or lowest DSAT from the prior week. Focus on one attribute. Link to a macro, help article, or policy clarification when the gap is systemic.

Quarterly rubric review: retire criteria nobody uses; add attributes for new product areas. Version the rubric so trend lines stay comparable.

Use the support QA coaching plan for a week-by-week rollout template.

QA workflow in Intercom and Slack

A practical stack for Intercom teams:

  1. Score closed conversations in Intercom (manual tags), a QA spreadsheet, or Supportman IQS against your rubric.
  2. Route DSAT ratings to Slack in real time so recovery happens while context is fresh.
  3. Report weekly team and per-agent IQS trends in Slack alongside volume and FRT.
  4. Coach in 1:1s using the same thread the score came from — paste the Intercom link, not a paraphrase.

Intercom's native Slack app handles conversation replies; it does not score quality or post rubric breakdowns. Pair it with a reporting layer — see connect Intercom and Slack.

Customer service QA FAQ

What is customer service quality assurance?

QA is a structured process for evaluating support conversations against defined criteria, calibrating reviewers, and using scores to coach agents. It complements CSAT by explaining why conversations succeeded or failed, not just whether the customer clicked a happy emoji.

How often should you QA customer service calls and chats?

Review at least 5–10% of conversations manually if that is your only method. With AI scoring on every closed thread, run human calibration monthly and agent coaching weekly on one conversation each.

What should a customer service QA rubric include?

Four to six weighted attributes with observable criteria, a defined rating scale, N/A rules, and a critical-failure definition. Start from resolution and product knowledge before adding softer attributes like tone.

Can AI do customer service QA?

Yes, when calibrated against human scores on a representative sample. AI works best for full coverage on async chat and email; keep human review for edge cases, compliance failures, and rubric updates.

How does QA relate to CSAT?

CSAT is customer-reported satisfaction on a subset of conversations. QA is expert evaluation against your standard, ideally on every thread. Use CSAT for customer-facing trends and QA for agent development and process fixes.

Under two minutes to live, no IT ticket required.

See pricing
Prefer us on Google