ع
Learn Tracks Reference Guides Saved
Capability Track Quality & the team

Quality & the team: the QA rubric, the coaching system, and the new-hire ramp

The playbooks show you how to score replies and ramp a new hire. This is the module where you master the judgment those recipes can't give you — and prove it on real artifacts.

13 min read · Updated 2026-06-30
Quality & the team: the QA rubric, the coaching system, and the new-hire ramp

Your team already has the qa-and-coaching and agent-ramp playbooks — the worked recipes for scoring a sample of replies against your standards and building a new hire’s first week from the docs you already own. This module is the layer above the recipe. It’s where you master why those two systems live or die based on the judgment behind them — and prove, against a real rubric, that you can run both.

It’s Module 3 of the certifiable Customer Support track, and it builds directly on the voice and policy docs you authored in M1. The QA rubric is derived from support-voice.md and support-policy.md; the onboarding system is built from the macros and help articles that sit behind those docs. Neither artifact is possible without M1’s foundation.

The voice file defines the standard; the QA rubric enforces it; the onboarding makes it the new hire’s starting point.

Quality at scale — the discipline the recipe can’t teach

Every support team drifts. Not because agents are careless, but because a queue is constant pressure and a standard that lives only in a style guide is a standard nobody applies under load. The reply that goes out under pressure is the reply that felt right in the moment — which is often a bit more hedged, a bit less warm, a bit more “please let me end this ticket” than the voice doc intended.

QA exists to make the drift visible before it compounds. But the failure mode of QA isn’t too little scoring — it’s scoring that produces the wrong outputs. There are three places the system breaks:

  • The rubric measures the wrong thing. A rubric that comes from Claude’s general sense of good support scores against a generic standard, not your standard. The agent who nailed the Mizan voice — warm, direct, owns the problem in the first sentence — scores “needs-work” because the generic rubric wanted a more formal opening. The rubric must be derived from your docs: your voice file, your policy doc, your macros. What you score against is what you get.
  • Findings land on individuals, not on the process. A QA pass that produces a list of who got what score is surveillance, not quality. The output that actually improves a team is: “Fourteen of twenty replies contain a phrase that’s on the never-say list — that’s a training and macro problem, not fourteen people problems.” Systemic patterns have systemic fixes. Routing them to the macros and the help center means the team doesn’t drift back in sixty days.
  • Coaching notes are gotcha-shaped. A coaching note that leads with a score and a list of failures makes a person defensive, which means they don’t use it. A coaching note that leads with one specific thing they did well, then gives one or two concrete things to try next — in plain language, not in rubric language — is the note someone reads twice and acts on. The format matters as much as the content.

The recipe in the playbook walks you through the mechanics. This module is the judgment that decides whether the mechanics produce a system or a one-off exercise.

Building a rubric from your own documents

The rubric is not a list of good support principles. It is a direct translation of your support-voice.md and support-policy.md into scored criteria — so that the only way to pass the rubric is to sound like your team and follow your policy. That specificity is the whole point.

Four to five criteria is the right scope. More than five and a score becomes a document nobody can hold in their head; fewer than four and you miss the dimensions that actually predict quality. Each criterion needs:

  • A name. Short enough to appear in a score table.
  • A one-line definition. What does it actually measure?
  • A pass/needs-work description. Not a grade — a binary with a reason. “Pass: the first sentence names the customer’s problem. Needs-work: the first sentence is an apology or a greeting.” That specificity makes the score a piece of evidence you can verify, not a judgment call you have to trust.

The four natural criteria for a team that has both a voice doc and a policy doc are: voice warmth (did the reply sound like us?), policy adherence (did the reply stay inside what we’re authorized to do?), resolution quality (did the reply actually close the issue, or did it kick it forward?), and over-promise check (did the reply commit to anything a senior needs to approve, or invent a timeline the team hasn’t confirmed?).

The rubric quality test: once you have a rubric draft, score the same batch of five replies twice — once today, once two days from now. If the scores are consistent, the rubric is doing its job. If they’re different, the criterion definitions are ambiguous and need sharpening before you run the full batch. A rubric that produces different scores on the same reply isn’t a rubric — it’s an opinion that changes.

Reading the patterns — team-wide vs individual

The scored batch is not the output. It’s the input to the real output: the team-wide pattern analysis that tells you what to fix and where.

Sorting a batch by score and reading from the bottom up is the wrong move. It finds the low scorers. What you want to find is the low criterion — the column where even the high scorers are failing. That column is the systemic gap. It tells you something is wrong with the process, the macros, the policy doc, or the training — not with the people.

The logic is this: if one agent over-promises on timelines, that’s a coaching conversation with one person. If twelve of twenty replies over-promise on timelines, that’s a sentence missing from support-policy.md and a macro that needs a clear refund-window line. Coaching twelve people one at a time produces twelve slightly-improved replies. Fixing the macro and the policy doc produces every reply going forward.

The pattern analysis has three outputs, in this order:

  1. Team-wide patterns, ranked by frequency. “14/20 replies used a phrase on the never-say list. 8/20 replied to a VAT question without linking the help article. 5/20 committed to a refund timeline.” The count is what makes a pattern real versus a cluster of coincidences.
  2. Routing to the fix. Each systemic pattern gets routed to the place that actually fixes it — a macro update, a policy edit, a new help article, a training point for new hires. The test: if we fix this in the macro, does the QA score on this criterion improve next quarter without another coaching conversation? If yes, route it to the macro.
  3. Individual coaching notes. These come last, because the coaching note should frame individual gaps against the systemic ones — “this is something the whole team is working on” rather than “you failed.” The note leads with one specific thing that went well (with a quoted line), then gives one or two things to try next. No scores, no grades, no piling on.

The key discipline: team-wide findings are public, shared with the whole team as process findings. Individual coaching notes are private, one-to-one. Never publish individual scores.

Ramping a new hire on the real standard

The failure mode of new-hire onboarding in support is the osmosis model: the new agent watches, asks, and gets things wrong on real customers for a month before their replies start sounding like the rest of the team. That month of inconsistency is costly in two directions — the customers who got below-standard replies, and the agent who built habits from watching the team’s average rather than from the team’s ceiling.

The insight that changes this: a new hire who learns from your docs doesn’t inherit the team’s average. They inherit the standard. Your support-voice.md and support-policy.md are the ceiling, not a description of what actually goes out most Tuesdays. If the ramp is built from those docs — plus the macros and help articles that are the practical expression of those docs — the new hire starts closer to the ceiling than most agents who’ve been in the queue for six months.

Three things make a ramp from docs into a real ramp:

  • The cheat sheet maps to real articles. A cheat sheet that says “for VAT questions, see the help center” is not useful. A cheat sheet that says “for VAT export issues: see /help/vat-export (the article), use the vat-export-steps macro (in macros.md), escalate if the customer says the export has the wrong period” is infrastructure. The [no article yet] and [no macro yet] gaps in the cheat sheet are a bonus output — they’re the documentation debt the ramp makes visible.
  • Practice tickets teach judgment, not just voice. Reading the voice doc is passive. Drafting a reply to a practice ticket, then comparing it to a model answer written in the team’s voice — that’s active recall. And two or three of the practice tickets should be ones where the right answer is to escalate, not to solve. A new hire who’s never practiced an escalation before they face one live will fill the gap with something — and it usually isn’t the escalation.
  • The ramp has gates. Shadow first. Then supervised replies (a senior approves before sending). Then solo on easy tiers, with a senior reviewing after. Each gate has a concrete checkpoint — “can solve the top five issue types without checking the cheat sheet” is a checkpoint; “settling in” is not. The gates exist because a new hire given more responsibility than they’ve demonstrated is being set up to fail publicly, on real customers.

The ramp is built once per hire cycle, not once for all time. Re-run it with a fresh ticket export every time someone new joins, so the top-issue cheat sheet reflects the current queue rather than the queue from eighteen months ago.

The Arabic team standard

For MENA-based support teams — and Mizan’s team is UAE-based, with customers who write in Arabic — the quality standard must be bilingual in fact, not just in intention. A QA system that only scores English replies is a QA system for half the team.

Four adaptations make the system genuinely bilingual:

  • The QA rubric applies to Arabic replies too. “Voice warmth” in Arabic means the reply sounds like warm Gulf MSA — not stiff fus’ha, not a machine translation of an English template. The “owns the problem in the first sentence” criterion applies in Arabic: the opening should name the customer’s problem, not begin with a greeting block. Score Arabic replies on the same criteria, because the standard is the same. The rubric language in Arabic stays in Arabic for the scoring session — don’t force agents to evaluate Arabic replies against English criterion names.
  • Coaching notes for Arabic-first agents are written in Arabic. A coaching note written in English to an agent whose working language is Arabic is a note that lands at half strength. Write the note in the language the agent actually thinks in. The structure is the same — one thing done well, one or two things to try next — but the language matches the agent.
  • The new hire with Arabic tickets gets Arabic practice tickets. A practice set that’s all English tickets doesn’t prepare a new hire for the Arabic queue. Build at least half the practice set from anonymized Arabic tickets — real cases from the queue — so the new hire practices the voice in both languages before going live.
  • The cheat sheet has Arabic-language macros where they exist. If the macro library has Arabic variants, the cheat sheet should link to them explicitly. A new hire who doesn’t know an Arabic macro exists for the top billing question will re-draft it from scratch, and the voice drift starts on day one.

The Arabic team standard isn’t an add-on. It’s the difference between a QA system that improves the team and one that improves the team’s English replies while the Arabic queue drifts unchecked.

Your assignment

Build the three deliverables for Mizan’s support team — or for your own team if you have the assets. Open your support folder in Claude Desktop, approve the reads of support-voice.md, support-policy.md, macros.md, and your ticket export in the “Ask permissions” prompt, and work in the chat. No terminal needed.

Module 3 deliverables — quality & the team

1. qa-rubric.md
   A 4–5 criterion rubric derived directly from support-voice.md
   and support-policy.md — not Claude's general idea of good support.
   Each criterion:
     - A name (short)
     - A one-line definition of what it measures
     - A "pass" description with a concrete example
     - A "needs-work" description with a concrete example
   A rubric-quality note: describe how you would test whether the
   rubric is consistent (score the same batch twice; if scores change,
   which criterion is ambiguous and what would sharpen it?).

2. batch-findings.md
   A scored QA batch and team-wide findings report for a 20-reply
   sample. Includes:
     - The per-reply scores (pass / needs-work per criterion, with
       the quoted line that earned each needs-work)
     - A summary table: per criterion, how many of 20 passed
     - Team-wide patterns: at least 2, ranked by frequency, each
       with a count and 2 example lines
     - For each pattern: the routing (macro update / policy edit /
       help article / training point) and a draft of the actual fix
     - One coaching note per agent: leads with one specific thing
       done well (quoted), gives 1–2 things to try next — no grades,
       no piling on
   Explicitly separate systemic patterns from individual one-offs.

3. onboarding.md for a new support hire
   A reusable first-week ramp built from your real docs. Includes:
     - Top-5 issue cheat sheet: issue, help-center article (by title),
       macro (by name), one-line recognition cue, escalation line.
       Mark [no article yet] / [no macro yet] where gaps exist.
     - 5 practice tickets from anonymized real cases (not invented),
       with model answers in a separate section so they're not spoiled.
       At least 1 practice ticket where the right answer is escalation.
     - Day-1 through Day-5 ramp outline: shadow → supervised → solo,
       with a concrete checkpoint at each stage.

Bilingual teams: produce qa-rubric.md and onboarding.md in Arabic
too — authored, not translated. Arabic coaching notes are written in
Arabic. Arabic practice tickets come from real Arabic tickets.

M1’s support-voice.md and support-policy.md are the foundation this module builds on. If you haven’t completed M1, do it first.

How it’s graded — the rubric

Your three files are scored against five criteria. Each is meets / nearly / not yet — and “nearly” on any one is a revise, not a pass.

Module 3 rubric

1. The QA rubric comes from         Every criterion traces to a named
   YOUR docs                        section of support-voice.md or
                                    support-policy.md. No generic
                                    "professional tone" or "customer
                                    satisfaction" language that could
                                    have come from anywhere.

2. Findings are team-wide,          At least 2 named patterns with
   not individual gotchas           a count and a routing to the fix.
                                    Individual coaching notes are warm
                                    and do not read as score cards.
                                    Systemic and individual are
                                    explicitly separated.

3. Coaching notes are warm          Each note leads with one specific
   and actionable                   genuine positive (quoted line),
                                    then gives 1–2 concrete things to
                                    try. No grades, no defensiveness
                                    triggers. The note is the kind of
                                    feedback you'd actually want to
                                    receive.

4. Cheat sheet maps to real         Every issue in the top-5 cheat
   docs                             sheet links to a real article by
                                    title and a real macro by name —
                                    or is explicitly marked [no article
                                    yet] / [no macro yet]. No invented
                                    articles or macros.

5. Practice tickets are from        All 5 practice tickets are adapted
   anonymized real cases            from real tickets (anonymized to
                                    [customer]), not invented scenarios.
                                    At least 1 escalation case. Model
                                    answers are in the team's real
                                    voice, not a generic support voice.

The bar, shown — a worked model answer (Mizan)

You don’t have to guess what “meets” looks like. Here is a passing excerpt for Mizan — yours doesn’t need to look like this, it needs to clear the same bar.

qa-rubric.md — Layla Al-Nasser, Mizan CX Lead (excerpt)

Derived from: support-voice.md (§ Opening, § Ownership, § Phrasing to
avoid) and support-policy.md (§ Solo-tier authorizations, § Manager-
tier authorizations, § Never)

Criterion 1 — Voice warmth
  What it measures: does the reply open by naming the customer's
  problem and take ownership in the first sentence? Does it avoid
  corporate filler?
  Pass: the first sentence names the problem back ("Your VAT export
    hasn't generated correctly — I'm on it") with no apology block
    and no "per our policy" language.
  Needs-work: the reply opens with a greeting, a generic apology, or
    "as a one-time courtesy" before the customer's issue appears.
    Example of needs-work: "Dear valued customer, we apologize for
    any inconvenience this may have caused..."

Criterion 2 — Policy adherence
  What it measures: does the reply stay inside the solo-tier
  authorizations in support-policy.md? Does it avoid the Never list?
  Pass: refunds cited are ≤ AED 150 (solo tier), no roadmap promises,
    no invented order numbers or ship dates.
  Needs-work: the reply offers a refund that requires manager approval,
    promises a fix by a date engineering hasn't confirmed, or uses
    language from the Never list ("I can confirm this will be resolved
    by...").

Criterion 3 — Resolution quality
  What it measures: does the reply close the issue or create a clear
  next step, rather than forward it back to the customer or defer?
  Pass: the issue is resolved in this reply OR the next step is stated
    plainly ("I'm escalating this to our billing team — you'll hear
    back within 24 hours").
  Needs-work: the reply ends with "please let us know if you need
    anything else" without having resolved the original question, or
    asks the customer a clarifying question that should have been in
    the first reply.

Criterion 4 — No over-promise
  What it measures: does the reply avoid committing to timelines,
  refund amounts, or feature delivery that haven't been confirmed?
  Pass: timelines are stated as ranges from policy ("5–7 business
    days"), not as specific dates; no ETA is given for a bug fix
    unless engineering has confirmed one.
  Needs-work: "I'll make sure this is fixed by end of week" when
    no such commitment exists; "your refund will process in 24 hours"
    when the actual SLA is 5–7 days.

Rubric quality note: To test consistency, I'll score the same 5 replies
today and again in two days. If criterion 3 (Resolution quality) shifts —
which is the most judgment-dependent — I'll sharpen the "clear next
step" definition to include examples of what does and doesn't count as
a next step (e.g., "I'm escalating" = yes; "feel free to contact us
again" = no).
batch-findings.md — team-wide pattern excerpt (Mizan, 20-reply sample)

Summary table (pass out of 20):
  Voice warmth:       12 / 20
  Policy adherence:   16 / 20
  Resolution quality: 14 / 20
  No over-promise:    11 / 20

Team-wide patterns (systemic, not individual):

Pattern 1 — "as a one-time courtesy" and other never-say phrases
  Count: 14 of 20 replies contain at least one phrase from the
  never-say list in support-voice.md. This is the most widespread
  gap in the sample.
  Examples:
    "As a one-time courtesy, I'll waive the fee this time."
    "We apologize for any inconvenience this may have caused."
  Routing: MACRO UPDATE. Add the never-say list to the top of
  macros.md as a red-block header so it's visible every time an
  agent opens the file. Draft fix:
    ⛔ NEVER USE: "as a one-time courtesy" / "any inconvenience" /
    "per our policy" / "please be advised" — see support-voice.md §
    Phrasing to avoid for the full list.

Pattern 2 — No help-article link on resolvable issues
  Count: 17 of 20 replies to a question that has a help-center
  article do not link the article.
  Examples:
    A VAT export question answered in full with no link to
    /help/vat-export.
    A bank sync question with no link to /help/bank-sync-troubleshoot.
  Routing: RUBRIC UPDATE + MACRO UPDATE. Add "links the help article
  (when one exists)" to criterion 3 (Resolution quality) so it's
  scored every QA pass. Add a line to the top 5 macros: "Link:
  [article title and URL]" as a fill-in placeholder so agents see
  the prompt before sending.
  This is a process gap, not a people gap — no one put "link the
  article" in the rubric or the macro, so no one did it.

Note on individual one-offs: Two replies in the sample committed to
specific refund dates that aren't in policy (criterion 4). These are
individual conversations, not a team pattern — coaching notes address
them one-to-one below, not as a systemic finding.

Coaching note — (agent initials redacted in this excerpt):
  "Your reply to [customer] about the bank sync failure was a clean
  example of the Mizan voice — you named the problem in the first
  sentence and owned it without hedging. One thing to try: on the
  next VAT question, drop the help-center link into the reply before
  you send. This is something the whole team is adding to the habit
  — it gives the customer a place to go if the issue recurs without
  needing to open another ticket. Grab the URL from the macro."
onboarding.md — Hani Al-Rashidi, Support Agent, Mizan (excerpt)

Welcome, Hani. This document is your first week. Everything here
comes from the docs the Mizan support team already uses — so you're
learning the real standard from day one, not a generic one.

--- TOP 5 ISSUES CHEAT SHEET ---

Issue                 | Article (title + link)          | Macro               | Recognize it when...             | Escalate when...
----------------------|---------------------------------|---------------------|----------------------------------|------------------
VAT export / filing   | VAT Export Guide (/help/vat-    | vat-export-steps    | "My VAT export is wrong / blank  | Customer says wrong
issues                | vat-export)                     |                     | / won't generate"                | *period* or filed
                      |                                 |                     |                                  | already — manager
Login / password      | Reset Your Password (/help/     | password-reset      | "I can't log in / forgot         | Account locked after
reset                 | password-reset)                 |                     | password"                        | 5 attempts — manager
Billing / invoice     | Reading Your Invoice (/help/    | billing-query       | "What is this charge?" / "I      | Refund > AED 150 —
questions             | invoice-guide)                  |                     | don't recognize this amount"     | manager
Reconciliation help   | Reconcile Your Accounts (/help/ | [no macro yet]      | "My accounts don't balance" /    | Customer has
                      | reconciliation)                 |                     | "transactions are missing"       | accountant involved
Plan upgrade /        | [no article yet]                | plan-change         | "I want to upgrade / downgrade   | Annual plan change —
downgrade             |                                 |                     | my plan"                         | always manager

[no article yet] gaps — flagged for Layla: reconciliation macros
and plan-change help articles are the two documentation gaps this
ramp exposed.

--- DAY 1–5 RAMP ---

Day 1 — Read only
  Read support-voice.md (full). Read support-policy.md (full).
  Read macros.md (skim the names; read the top 5 in full).
  Read this cheat sheet. No tickets today.
  End of day: tell Layla the one thing in policy you're not sure
  about. That question is the deliverable.

Day 2 — Shadow
  Shadow Layla on 5–8 live tickets. For each: after the reply goes
  out, write down what you would have said and compare. What's
  different? Ask why.
  Practice tickets 1–3 (from the practice set below): draft your
  reply before reading the model answer.

Day 3 — Shadow + supervised
  Shadow one more session (morning). Afternoon: draft 3 supervised
  replies — Layla reads and approves before they send. Not her
  edits; her approval or a note on what to change.
  Practice tickets 4–5.
  Gate checkpoint: can you resolve the top 3 issue types
  (VAT, login, billing) using the cheat sheet without asking?
  If yes, Day 4. If not, one more shadow day.

Day 4 — Supervised solo (easy tier)
  Handle easy-tier tickets solo. Layla reviews each reply *after*
  it sends (not before). Flag anything you were uncertain about.
  No annual-plan billing, no double-charge tickets, no refund
  requests above AED 150 — those are manager tier, not yet.

Day 5 — Review and debrief
  Layla reviews Day 4 replies using the QA rubric.
  Debrief: three things Hani got right, two things to try next week.
  One thing that goes back to the macro or the policy doc as a fix
  (not a coaching note — a process fix).
  If the Day 4 batch clears 80% pass on all four criteria: Hani is
  ready for the full easy-tier queue, supervised spot-checks weekly.

What you’ve proven — and what’s next

Clear the rubric and you’ve built three things your team will use beyond the module: a QA system that compounds (each round, the patterns route back to macros and policy, so the next round drifts less), a coaching cadence that makes feedback something agents look forward to rather than dread, and a ramp that turns your standard into the new hire’s starting point rather than the thing they drift toward after six months.

The credential you’re earning is “Certified Customer Support with Claude” — and Module 3 is where the track turns from individual-reply quality into team-wide systems.

From here the track moves into the two hardest support disciplines:

  • Module 4 — CSAT & voice of the customer: reading the score drop correctly (self-selection bias in a response rate of 18% is not a small problem), diagnosing the behavioral drivers behind it (over-hedging on timelines is a different fix than slow first-touch resolution), and routing VoC themes to product and process without over-claiming on qualitative data.
  • Module 5 — Incident communications and the bug escalation, then the capstone — one complete incident response and the full “Support-in-a-Box” artifact package, graded into the certificate.

If you’re building this across a growing team, the operating guide is the data, privacy, and approval layer that goes underneath all of it — what Claude is allowed to see, how customer PII stays inside your workspace, and who approves which categories of reply before they ship.

supportcustomer-supportqacoachingonboardingnew-hirequality-assurancevoicepolicycertificationassessmentarabicbilingualdesktopteams

Questions people ask

How is this different from the free qa-and-coaching and agent-ramp playbooks?
The playbooks are the recipe — the steps and prompts to score a batch of replies and build a first-week ramp. This module is mastery plus proof: the judgment the recipe can't give you (why you must derive the rubric from YOUR docs and not Claude's idea of good support, why team-wide patterns and individual coaching notes are completely different outputs that require different judgment calls, why a cheat sheet that maps to real articles is infrastructure and one that maps to invented articles is a liability), a real assignment you complete for a real team, and a rubric you're graded against. The playbook gets you through one QA pass; the module gives you the system that runs every quarter and compounds — and a credential that says you can build it.
Do I need Module 1 before this module?
Yes — M1 is load-bearing. The QA rubric is derived from support-voice.md and support-policy.md, which are M1's core deliverables. Without those docs, you're scoring against a generic idea of good support, which is not this module. If you haven't done M1, start there.
Do I need a real team and real tickets, or can I use Mizan?
Both work. Bring your own team's reply export and ticket data and the assignment doubles as real infrastructure your team inherits — the onboarding.md you build is the one your next hire actually uses. If you'd rather learn on neutral ground first, use Mizan's team and the ticket data given here, then redo it for your own context. Either way, the judgment is transferable.
How is it graded, and who grades it?
Against the explicit rubric in this module — the QA rubric comes from your docs, findings are team-wide patterns not individual gotchas, coaching notes are warm and actionable, the cheat sheet maps issues to real articles and macros, and practice tickets come from anonymized real cases. In a cohort a reviewer scores your three deliverables against that rubric; the worked Mizan model answer here shows you the bar before you submit.
What if I'm a solo support agent with no team to QA?
The QA system still applies — solo or team, your replies should be consistent across a batch, and the rubric tells you whether they are. Score your own batch of 10–15 recent replies against the rubric. The coaching notes become self-coaching notes. The onboarding.md becomes the ramp doc you hand to the first person you hire. The discipline is the same; the scale is smaller.