Your team already has the qa-and-coaching and agent-ramp playbooks — the worked recipes for scoring a sample of replies against your standards and building a new hire’s first week from the docs you already own. This module is the layer above the recipe. It’s where you master why those two systems live or die based on the judgment behind them — and prove, against a real rubric, that you can run both.
It’s Module 3 of the certifiable Customer Support track, and it builds directly on the voice and policy docs you authored in M1. The QA rubric is derived from support-voice.md and support-policy.md; the onboarding system is built from the macros and help articles that sit behind those docs. Neither artifact is possible without M1’s foundation.
The voice file defines the standard; the QA rubric enforces it; the onboarding makes it the new hire’s starting point.
Quality at scale — the discipline the recipe can’t teach
Every support team drifts. Not because agents are careless, but because a queue is constant pressure and a standard that lives only in a style guide is a standard nobody applies under load. The reply that goes out under pressure is the reply that felt right in the moment — which is often a bit more hedged, a bit less warm, a bit more “please let me end this ticket” than the voice doc intended.
QA exists to make the drift visible before it compounds. But the failure mode of QA isn’t too little scoring — it’s scoring that produces the wrong outputs. There are three places the system breaks:
- The rubric measures the wrong thing. A rubric that comes from Claude’s general sense of good support scores against a generic standard, not your standard. The agent who nailed the Mizan voice — warm, direct, owns the problem in the first sentence — scores “needs-work” because the generic rubric wanted a more formal opening. The rubric must be derived from your docs: your voice file, your policy doc, your macros. What you score against is what you get.
- Findings land on individuals, not on the process. A QA pass that produces a list of who got what score is surveillance, not quality. The output that actually improves a team is: “Fourteen of twenty replies contain a phrase that’s on the never-say list — that’s a training and macro problem, not fourteen people problems.” Systemic patterns have systemic fixes. Routing them to the macros and the help center means the team doesn’t drift back in sixty days.
- Coaching notes are gotcha-shaped. A coaching note that leads with a score and a list of failures makes a person defensive, which means they don’t use it. A coaching note that leads with one specific thing they did well, then gives one or two concrete things to try next — in plain language, not in rubric language — is the note someone reads twice and acts on. The format matters as much as the content.
The recipe in the playbook walks you through the mechanics. This module is the judgment that decides whether the mechanics produce a system or a one-off exercise.
Building a rubric from your own documents
The rubric is not a list of good support principles. It is a direct translation of your support-voice.md and support-policy.md into scored criteria — so that the only way to pass the rubric is to sound like your team and follow your policy. That specificity is the whole point.
Four to five criteria is the right scope. More than five and a score becomes a document nobody can hold in their head; fewer than four and you miss the dimensions that actually predict quality. Each criterion needs:
- A name. Short enough to appear in a score table.
- A one-line definition. What does it actually measure?
- A pass/needs-work description. Not a grade — a binary with a reason. “Pass: the first sentence names the customer’s problem. Needs-work: the first sentence is an apology or a greeting.” That specificity makes the score a piece of evidence you can verify, not a judgment call you have to trust.
The four natural criteria for a team that has both a voice doc and a policy doc are: voice warmth (did the reply sound like us?), policy adherence (did the reply stay inside what we’re authorized to do?), resolution quality (did the reply actually close the issue, or did it kick it forward?), and over-promise check (did the reply commit to anything a senior needs to approve, or invent a timeline the team hasn’t confirmed?).
The rubric quality test: once you have a rubric draft, score the same batch of five replies twice — once today, once two days from now. If the scores are consistent, the rubric is doing its job. If they’re different, the criterion definitions are ambiguous and need sharpening before you run the full batch. A rubric that produces different scores on the same reply isn’t a rubric — it’s an opinion that changes.
Reading the patterns — team-wide vs individual
The scored batch is not the output. It’s the input to the real output: the team-wide pattern analysis that tells you what to fix and where.
Sorting a batch by score and reading from the bottom up is the wrong move. It finds the low scorers. What you want to find is the low criterion — the column where even the high scorers are failing. That column is the systemic gap. It tells you something is wrong with the process, the macros, the policy doc, or the training — not with the people.
The logic is this: if one agent over-promises on timelines, that’s a coaching conversation with one person. If twelve of twenty replies over-promise on timelines, that’s a sentence missing from support-policy.md and a macro that needs a clear refund-window line. Coaching twelve people one at a time produces twelve slightly-improved replies. Fixing the macro and the policy doc produces every reply going forward.
The pattern analysis has three outputs, in this order:
- Team-wide patterns, ranked by frequency. “14/20 replies used a phrase on the never-say list. 8/20 replied to a VAT question without linking the help article. 5/20 committed to a refund timeline.” The count is what makes a pattern real versus a cluster of coincidences.
- Routing to the fix. Each systemic pattern gets routed to the place that actually fixes it — a macro update, a policy edit, a new help article, a training point for new hires. The test: if we fix this in the macro, does the QA score on this criterion improve next quarter without another coaching conversation? If yes, route it to the macro.
- Individual coaching notes. These come last, because the coaching note should frame individual gaps against the systemic ones — “this is something the whole team is working on” rather than “you failed.” The note leads with one specific thing that went well (with a quoted line), then gives one or two things to try next. No scores, no grades, no piling on.
The key discipline: team-wide findings are public, shared with the whole team as process findings. Individual coaching notes are private, one-to-one. Never publish individual scores.
Ramping a new hire on the real standard
The failure mode of new-hire onboarding in support is the osmosis model: the new agent watches, asks, and gets things wrong on real customers for a month before their replies start sounding like the rest of the team. That month of inconsistency is costly in two directions — the customers who got below-standard replies, and the agent who built habits from watching the team’s average rather than from the team’s ceiling.
The insight that changes this: a new hire who learns from your docs doesn’t inherit the team’s average. They inherit the standard. Your support-voice.md and support-policy.md are the ceiling, not a description of what actually goes out most Tuesdays. If the ramp is built from those docs — plus the macros and help articles that are the practical expression of those docs — the new hire starts closer to the ceiling than most agents who’ve been in the queue for six months.
Three things make a ramp from docs into a real ramp:
- The cheat sheet maps to real articles. A cheat sheet that says “for VAT questions, see the help center” is not useful. A cheat sheet that says “for VAT export issues: see /help/vat-export (the article), use the vat-export-steps macro (in macros.md), escalate if the customer says the export has the wrong period” is infrastructure. The
[no article yet]and[no macro yet]gaps in the cheat sheet are a bonus output — they’re the documentation debt the ramp makes visible. - Practice tickets teach judgment, not just voice. Reading the voice doc is passive. Drafting a reply to a practice ticket, then comparing it to a model answer written in the team’s voice — that’s active recall. And two or three of the practice tickets should be ones where the right answer is to escalate, not to solve. A new hire who’s never practiced an escalation before they face one live will fill the gap with something — and it usually isn’t the escalation.
- The ramp has gates. Shadow first. Then supervised replies (a senior approves before sending). Then solo on easy tiers, with a senior reviewing after. Each gate has a concrete checkpoint — “can solve the top five issue types without checking the cheat sheet” is a checkpoint; “settling in” is not. The gates exist because a new hire given more responsibility than they’ve demonstrated is being set up to fail publicly, on real customers.
The ramp is built once per hire cycle, not once for all time. Re-run it with a fresh ticket export every time someone new joins, so the top-issue cheat sheet reflects the current queue rather than the queue from eighteen months ago.
The Arabic team standard
For MENA-based support teams — and Mizan’s team is UAE-based, with customers who write in Arabic — the quality standard must be bilingual in fact, not just in intention. A QA system that only scores English replies is a QA system for half the team.
Four adaptations make the system genuinely bilingual:
- The QA rubric applies to Arabic replies too. “Voice warmth” in Arabic means the reply sounds like warm Gulf MSA — not stiff fus’ha, not a machine translation of an English template. The “owns the problem in the first sentence” criterion applies in Arabic: the opening should name the customer’s problem, not begin with a greeting block. Score Arabic replies on the same criteria, because the standard is the same. The rubric language in Arabic stays in Arabic for the scoring session — don’t force agents to evaluate Arabic replies against English criterion names.
- Coaching notes for Arabic-first agents are written in Arabic. A coaching note written in English to an agent whose working language is Arabic is a note that lands at half strength. Write the note in the language the agent actually thinks in. The structure is the same — one thing done well, one or two things to try next — but the language matches the agent.
- The new hire with Arabic tickets gets Arabic practice tickets. A practice set that’s all English tickets doesn’t prepare a new hire for the Arabic queue. Build at least half the practice set from anonymized Arabic tickets — real cases from the queue — so the new hire practices the voice in both languages before going live.
- The cheat sheet has Arabic-language macros where they exist. If the macro library has Arabic variants, the cheat sheet should link to them explicitly. A new hire who doesn’t know an Arabic macro exists for the top billing question will re-draft it from scratch, and the voice drift starts on day one.
The Arabic team standard isn’t an add-on. It’s the difference between a QA system that improves the team and one that improves the team’s English replies while the Arabic queue drifts unchecked.
Your assignment
Build the three deliverables for Mizan’s support team — or for your own team if you have the assets. Open your support folder in Claude Desktop, approve the reads of support-voice.md, support-policy.md, macros.md, and your ticket export in the “Ask permissions” prompt, and work in the chat. No terminal needed.
Module 3 deliverables — quality & the team
1. qa-rubric.md
A 4–5 criterion rubric derived directly from support-voice.md
and support-policy.md — not Claude's general idea of good support.
Each criterion:
- A name (short)
- A one-line definition of what it measures
- A "pass" description with a concrete example
- A "needs-work" description with a concrete example
A rubric-quality note: describe how you would test whether the
rubric is consistent (score the same batch twice; if scores change,
which criterion is ambiguous and what would sharpen it?).
2. batch-findings.md
A scored QA batch and team-wide findings report for a 20-reply
sample. Includes:
- The per-reply scores (pass / needs-work per criterion, with
the quoted line that earned each needs-work)
- A summary table: per criterion, how many of 20 passed
- Team-wide patterns: at least 2, ranked by frequency, each
with a count and 2 example lines
- For each pattern: the routing (macro update / policy edit /
help article / training point) and a draft of the actual fix
- One coaching note per agent: leads with one specific thing
done well (quoted), gives 1–2 things to try next — no grades,
no piling on
Explicitly separate systemic patterns from individual one-offs.
3. onboarding.md for a new support hire
A reusable first-week ramp built from your real docs. Includes:
- Top-5 issue cheat sheet: issue, help-center article (by title),
macro (by name), one-line recognition cue, escalation line.
Mark [no article yet] / [no macro yet] where gaps exist.
- 5 practice tickets from anonymized real cases (not invented),
with model answers in a separate section so they're not spoiled.
At least 1 practice ticket where the right answer is escalation.
- Day-1 through Day-5 ramp outline: shadow → supervised → solo,
with a concrete checkpoint at each stage.
Bilingual teams: produce qa-rubric.md and onboarding.md in Arabic
too — authored, not translated. Arabic coaching notes are written in
Arabic. Arabic practice tickets come from real Arabic tickets.
M1’s support-voice.md and support-policy.md are the foundation this module builds on. If you haven’t completed M1, do it first.
How it’s graded — the rubric
Your three files are scored against five criteria. Each is meets / nearly / not yet — and “nearly” on any one is a revise, not a pass.
Module 3 rubric
1. The QA rubric comes from Every criterion traces to a named
YOUR docs section of support-voice.md or
support-policy.md. No generic
"professional tone" or "customer
satisfaction" language that could
have come from anywhere.
2. Findings are team-wide, At least 2 named patterns with
not individual gotchas a count and a routing to the fix.
Individual coaching notes are warm
and do not read as score cards.
Systemic and individual are
explicitly separated.
3. Coaching notes are warm Each note leads with one specific
and actionable genuine positive (quoted line),
then gives 1–2 concrete things to
try. No grades, no defensiveness
triggers. The note is the kind of
feedback you'd actually want to
receive.
4. Cheat sheet maps to real Every issue in the top-5 cheat
docs sheet links to a real article by
title and a real macro by name —
or is explicitly marked [no article
yet] / [no macro yet]. No invented
articles or macros.
5. Practice tickets are from All 5 practice tickets are adapted
anonymized real cases from real tickets (anonymized to
[customer]), not invented scenarios.
At least 1 escalation case. Model
answers are in the team's real
voice, not a generic support voice.
The bar, shown — a worked model answer (Mizan)
You don’t have to guess what “meets” looks like. Here is a passing excerpt for Mizan — yours doesn’t need to look like this, it needs to clear the same bar.
qa-rubric.md — Layla Al-Nasser, Mizan CX Lead (excerpt)
Derived from: support-voice.md (§ Opening, § Ownership, § Phrasing to
avoid) and support-policy.md (§ Solo-tier authorizations, § Manager-
tier authorizations, § Never)
Criterion 1 — Voice warmth
What it measures: does the reply open by naming the customer's
problem and take ownership in the first sentence? Does it avoid
corporate filler?
Pass: the first sentence names the problem back ("Your VAT export
hasn't generated correctly — I'm on it") with no apology block
and no "per our policy" language.
Needs-work: the reply opens with a greeting, a generic apology, or
"as a one-time courtesy" before the customer's issue appears.
Example of needs-work: "Dear valued customer, we apologize for
any inconvenience this may have caused..."
Criterion 2 — Policy adherence
What it measures: does the reply stay inside the solo-tier
authorizations in support-policy.md? Does it avoid the Never list?
Pass: refunds cited are ≤ AED 150 (solo tier), no roadmap promises,
no invented order numbers or ship dates.
Needs-work: the reply offers a refund that requires manager approval,
promises a fix by a date engineering hasn't confirmed, or uses
language from the Never list ("I can confirm this will be resolved
by...").
Criterion 3 — Resolution quality
What it measures: does the reply close the issue or create a clear
next step, rather than forward it back to the customer or defer?
Pass: the issue is resolved in this reply OR the next step is stated
plainly ("I'm escalating this to our billing team — you'll hear
back within 24 hours").
Needs-work: the reply ends with "please let us know if you need
anything else" without having resolved the original question, or
asks the customer a clarifying question that should have been in
the first reply.
Criterion 4 — No over-promise
What it measures: does the reply avoid committing to timelines,
refund amounts, or feature delivery that haven't been confirmed?
Pass: timelines are stated as ranges from policy ("5–7 business
days"), not as specific dates; no ETA is given for a bug fix
unless engineering has confirmed one.
Needs-work: "I'll make sure this is fixed by end of week" when
no such commitment exists; "your refund will process in 24 hours"
when the actual SLA is 5–7 days.
Rubric quality note: To test consistency, I'll score the same 5 replies
today and again in two days. If criterion 3 (Resolution quality) shifts —
which is the most judgment-dependent — I'll sharpen the "clear next
step" definition to include examples of what does and doesn't count as
a next step (e.g., "I'm escalating" = yes; "feel free to contact us
again" = no).
batch-findings.md — team-wide pattern excerpt (Mizan, 20-reply sample)
Summary table (pass out of 20):
Voice warmth: 12 / 20
Policy adherence: 16 / 20
Resolution quality: 14 / 20
No over-promise: 11 / 20
Team-wide patterns (systemic, not individual):
Pattern 1 — "as a one-time courtesy" and other never-say phrases
Count: 14 of 20 replies contain at least one phrase from the
never-say list in support-voice.md. This is the most widespread
gap in the sample.
Examples:
"As a one-time courtesy, I'll waive the fee this time."
"We apologize for any inconvenience this may have caused."
Routing: MACRO UPDATE. Add the never-say list to the top of
macros.md as a red-block header so it's visible every time an
agent opens the file. Draft fix:
⛔ NEVER USE: "as a one-time courtesy" / "any inconvenience" /
"per our policy" / "please be advised" — see support-voice.md §
Phrasing to avoid for the full list.
Pattern 2 — No help-article link on resolvable issues
Count: 17 of 20 replies to a question that has a help-center
article do not link the article.
Examples:
A VAT export question answered in full with no link to
/help/vat-export.
A bank sync question with no link to /help/bank-sync-troubleshoot.
Routing: RUBRIC UPDATE + MACRO UPDATE. Add "links the help article
(when one exists)" to criterion 3 (Resolution quality) so it's
scored every QA pass. Add a line to the top 5 macros: "Link:
[article title and URL]" as a fill-in placeholder so agents see
the prompt before sending.
This is a process gap, not a people gap — no one put "link the
article" in the rubric or the macro, so no one did it.
Note on individual one-offs: Two replies in the sample committed to
specific refund dates that aren't in policy (criterion 4). These are
individual conversations, not a team pattern — coaching notes address
them one-to-one below, not as a systemic finding.
Coaching note — (agent initials redacted in this excerpt):
"Your reply to [customer] about the bank sync failure was a clean
example of the Mizan voice — you named the problem in the first
sentence and owned it without hedging. One thing to try: on the
next VAT question, drop the help-center link into the reply before
you send. This is something the whole team is adding to the habit
— it gives the customer a place to go if the issue recurs without
needing to open another ticket. Grab the URL from the macro."
onboarding.md — Hani Al-Rashidi, Support Agent, Mizan (excerpt)
Welcome, Hani. This document is your first week. Everything here
comes from the docs the Mizan support team already uses — so you're
learning the real standard from day one, not a generic one.
--- TOP 5 ISSUES CHEAT SHEET ---
Issue | Article (title + link) | Macro | Recognize it when... | Escalate when...
----------------------|---------------------------------|---------------------|----------------------------------|------------------
VAT export / filing | VAT Export Guide (/help/vat- | vat-export-steps | "My VAT export is wrong / blank | Customer says wrong
issues | vat-export) | | / won't generate" | *period* or filed
| | | | already — manager
Login / password | Reset Your Password (/help/ | password-reset | "I can't log in / forgot | Account locked after
reset | password-reset) | | password" | 5 attempts — manager
Billing / invoice | Reading Your Invoice (/help/ | billing-query | "What is this charge?" / "I | Refund > AED 150 —
questions | invoice-guide) | | don't recognize this amount" | manager
Reconciliation help | Reconcile Your Accounts (/help/ | [no macro yet] | "My accounts don't balance" / | Customer has
| reconciliation) | | "transactions are missing" | accountant involved
Plan upgrade / | [no article yet] | plan-change | "I want to upgrade / downgrade | Annual plan change —
downgrade | | | my plan" | always manager
[no article yet] gaps — flagged for Layla: reconciliation macros
and plan-change help articles are the two documentation gaps this
ramp exposed.
--- DAY 1–5 RAMP ---
Day 1 — Read only
Read support-voice.md (full). Read support-policy.md (full).
Read macros.md (skim the names; read the top 5 in full).
Read this cheat sheet. No tickets today.
End of day: tell Layla the one thing in policy you're not sure
about. That question is the deliverable.
Day 2 — Shadow
Shadow Layla on 5–8 live tickets. For each: after the reply goes
out, write down what you would have said and compare. What's
different? Ask why.
Practice tickets 1–3 (from the practice set below): draft your
reply before reading the model answer.
Day 3 — Shadow + supervised
Shadow one more session (morning). Afternoon: draft 3 supervised
replies — Layla reads and approves before they send. Not her
edits; her approval or a note on what to change.
Practice tickets 4–5.
Gate checkpoint: can you resolve the top 3 issue types
(VAT, login, billing) using the cheat sheet without asking?
If yes, Day 4. If not, one more shadow day.
Day 4 — Supervised solo (easy tier)
Handle easy-tier tickets solo. Layla reviews each reply *after*
it sends (not before). Flag anything you were uncertain about.
No annual-plan billing, no double-charge tickets, no refund
requests above AED 150 — those are manager tier, not yet.
Day 5 — Review and debrief
Layla reviews Day 4 replies using the QA rubric.
Debrief: three things Hani got right, two things to try next week.
One thing that goes back to the macro or the policy doc as a fix
(not a coaching note — a process fix).
If the Day 4 batch clears 80% pass on all four criteria: Hani is
ready for the full easy-tier queue, supervised spot-checks weekly.
What you’ve proven — and what’s next
Clear the rubric and you’ve built three things your team will use beyond the module: a QA system that compounds (each round, the patterns route back to macros and policy, so the next round drifts less), a coaching cadence that makes feedback something agents look forward to rather than dread, and a ramp that turns your standard into the new hire’s starting point rather than the thing they drift toward after six months.
The credential you’re earning is “Certified Customer Support with Claude” — and Module 3 is where the track turns from individual-reply quality into team-wide systems.
From here the track moves into the two hardest support disciplines:
- Module 4 — CSAT & voice of the customer: reading the score drop correctly (self-selection bias in a response rate of 18% is not a small problem), diagnosing the behavioral drivers behind it (over-hedging on timelines is a different fix than slow first-touch resolution), and routing VoC themes to product and process without over-claiming on qualitative data.
- Module 5 — Incident communications and the bug escalation, then the capstone — one complete incident response and the full “Support-in-a-Box” artifact package, graded into the certificate.
If you’re building this across a growing team, the operating guide is the data, privacy, and approval layer that goes underneath all of it — what Claude is allowed to see, how customer PII stays inside your workspace, and who approves which categories of reply before they ship.