ع
Learn Tracks Reference Guides Saved
playbook

QA the queue and coach the team to one voice

Score a sample of recently-sent replies against your voice and policy docs, surface where the whole team is drifting off-voice or over-promising, and turn it into warm, specific coaching notes — growth, not a gotcha audit.

medium ~45 min
when to reach for this

You can feel the team drifting — replies that hedge every refund, that never link the help article, that sound like four different companies depending on who answered. But you can't read every ticket, and the old way of catching it is a manager spot-checking a handful and leaving a few people feeling watched. This system samples a batch of recently-sent replies, scores them against your support-voice.md and support-policy.md, and finds the patterns that are TEAM-WIDE — the systemic gaps a process can fix, not the individual ones a person should feel bad about. The point is coaching and consistency, never surveillance: it hands each agent specific, kind, actionable notes the way you'd want feedback, and it routes the recurring gaps back into the macros, the help center, and new-hire training so QA fixes the cause instead of nagging the symptom.

gather this first
  • A representative sample of recently-sent replies — say 20–40 across the team, exported as replies.csv (or pasted in). Anonymize the customer to [customer] first, but keep the agent name or initials so coaching notes land with the right person. In Claude Desktop, drop the file into the chat (or open the folder it lives in) so Claude can read it.
  • Your support-voice.md and support-policy.md — the same two docs every reply is supposed to inherit. They're what you score against, so the rubric measures real standards, not Claude's idea of good support.
  • A one-line note on what you already suspect is drifting ("we over-promise on refund timelines"), so Claude can confirm or challenge it rather than start blind.
the workflow
  1. Build the QA rubric from your own voice and policy

    Don't score against a generic idea of good support — derive the rubric straight from support-voice.md and support-policy.md so it measures YOUR standards. Keep it to 4–5 criteria, each a simple pass / needs-work with a reason, so the scores stay readable and fair.

    you ask
    Read support-voice.md and support-policy.md. Draft a lightweight QA rubric of 4–5 criteria for scoring a support reply against these docs — covering tone and warmth, factual accuracy, policy adherence, and whether the reply actually resolved the issue. For each criterion give a one-line definition and what 'pass' vs 'needs-work' looks like. Don't score anything yet — I want to confirm the rubric first.

    what you get back A short, named rubric — e.g. "Warmth (pass: owns it plainly, no corporate filler), Accuracy (pass: no claim the agent couldn't verify), Policy (pass: matches support-policy.md), Resolution (pass: one clear next step, issue actually closed)." Tightly scoped to your docs, not a generic checklist.

    Confirm the rubric reads like your standards before you score against it — a wrong rubric makes every score below it wrong in the same direction.

  2. Score a representative sample against the rubric

    Now run the sample through the rubric. Claude drafts the scores; a lead spot-checks. Ask for the reason on every 'needs-work' so a score is a piece of evidence you can verify, not a number you have to trust blind.

    you ask
    Score each reply in replies.csv against the rubric, anonymizing the customer to [customer]. For each reply give pass / needs-work per criterion with a one-line reason, and quote the exact line that earned a 'needs-work' so I can check it. Keep the agent name attached. End with a short table: per criterion, how many of the sample passed.

    what you get back A scored sample — each reply rated per criterion with the offending line quoted — plus a summary table like "Warmth 31/36 pass, Policy 22/36, Resolution 28/36." The low column is your first clue where the team is drifting.

    These scores are a draft for a human to verify, not a verdict. Spot-check a handful — especially every 'needs-work' — against the real reply before any of it reaches a person.

  3. Find the team-wide patterns, not the individual slips

    This is the whole point. A one-off bad reply is a person having a rough day; the same gap across half the team is a process problem. Ask Claude to separate the systemic patterns from the individual noise, because the fix is completely different — you change the macro, not the person.

    you ask
    Looking across the whole scored sample, what are the TEAM-WIDE patterns — gaps that show up across many agents, not one or two? Rank them by how widespread they are, with the count and 2 example lines each. Explicitly separate the systemic patterns (most of the team does this) from individual one-offs. Examples of systemic: everyone hedges refund timelines, nobody links the relevant help article.

    what you get back A ranked list of process problems — "22 of 36 replies never link a help article; 18 hedge the refund window instead of stating it" — clearly split from the handful of individual one-offs, so you fix causes instead of correcting people one reply at a time.

  4. Draft warm, specific per-agent coaching notes

    Turn the scores into the kind of feedback you'd actually want to receive: specific, kind, and forward-looking. Lead with what worked, point to a real example, and give one or two concrete things to try next — not a list of everything wrong.

    you ask
    For each agent in the sample, draft a short coaching note in a warm, human voice — the way I'd want a manager to give me feedback. Structure each as: one genuine thing they did well (with a quoted line), then one or two specific things to try next, each tied to a real example from their replies. No scores or grades in the note, no piling on — growth, not a gotcha. Frame the systemic gaps as 'something the whole team is working on,' not their personal failing.

    what you get back One note per agent — "Your apology to [customer] about the late order was genuinely warm and owned it cleanly. One thing to try: when you mention a refund, state the timeline straight ('5–7 business days') instead of 'as soon as possible' — that's a team-wide thing we're tightening up." Specific, kind, and actionable.

    Read each note as if it were about you. If it would make you defensive instead of better, soften it before it goes out — and you, not Claude, decide whether a coaching conversation happens at all.

  5. Route the systemic gaps to the right fix

    Coaching changes behavior; fixing the cause removes the need. Take each team-wide pattern and send it to where it actually gets solved — a new macro, a help-center article, a tweak to support-policy.md, or a training point for new hires — so the gap stops recurring instead of getting re-coached every quarter.

    you ask
    For each systemic pattern we found, recommend the single best fix and route it: a new or revised macro for the canned-response library, a help-center article, a specific edit to support-policy.md, or a new-hire training point. For each, write a one-line rationale and a draft of the actual fix (the macro text, the article outline, or the exact policy line). Flag [confirm] anywhere the right answer depends on a real policy I haven't given you.

    what you get back A short routing table — "'nobody links help articles' → add a link-the-article line to the top 5 macros; 'refund hedging' → one clear sentence in support-policy.md, drafted here for you to approve" — so the QA pass ends by improving the system, not just the people.

    This step is what makes QA compound: each round, a few recurring gaps become permanent fixes in the macros, the help center, or the policy — so next quarter's sample drifts less by design.

make it your own
  • Tighten the source of truth: if the rubric keeps catching the same ambiguity, the gap is often in the docs themselves — feed it back into Set the support voice and policy every reply inherits so the standard you score against is sharper next round.
  • Convert gaps into reusable answers: route the 'nobody links the article' and 'everyone re-types this' patterns into Build a canned-response library you'll actually reuse and Build a help center from the questions you actually get, so the systemic gap becomes a macro or article instead of a recurring coaching note.
  • Onboard against the patterns: fold the recurring gaps into Ramp a new support hire in their first week so new agents learn the team's hard-won voice from day one instead of drifting into the same habits.
  • Pair it with the numbers: run it alongside Find what's really driving your CSAT — CSAT tells you a score dropped, QA explains the 'why' behind it, so you fix the behavior instead of guessing at it.
  • Automate the sampling (Power Track): if you QA on a fixed cadence, save the rubric-and-score prompts as a /qa custom command or a scheduled agent (see the Playbook's Features tab) that pulls a weekly sample and drafts the scores before your review slot. Custom commands and scheduled agents are the opt-in Power Track — on Desktop you run the same prompts by hand each week until you're ready to wire it together.
watch out for
  • Every score is a draft for a human to verify, never a verdict. Spot-check the sample — especially each 'needs-work' — against the real reply before any number reaches a person; a confidently wrong score that becomes a coaching note erodes trust faster than no QA at all.
  • This data is doubly sensitive: the replies contain customer PII (anonymize the customer to [customer] first) and they're about real teammates, so keep the file and the scores private, share notes one-to-one, and frame everything as growth — coaching toward one voice, not a surveillance log of who's failing.
  • Claude drafts the scores and the coaching language; a human owns the call. Whether a pattern is real, whether a note is fair, and whether to have a conversation at all are management decisions — Claude makes the read faster, it doesn't make the judgment.
  • Score the system, not just the people. If a gap shows up across most of the team, that's almost always a missing macro, an unclear policy, or a training hole — route it to the fix in the last step instead of coaching the same thing into thirty people one at a time.

you'll end up with A scored sample of real replies against your own voice and policy, the team-wide patterns surfaced and separated from individual noise, a warm and specific coaching note per agent, and the systemic gaps routed to the macros, help center, and policy that will keep them from recurring — a QA pass that grows the team toward one voice and fixes causes, instead of an audit that just catches people.

Questions people ask

Isn't this just surveilling my team?
No — and the design is deliberately the opposite. The output isn't a leaderboard of who's failing; it's per-agent coaching notes that lead with what worked and give one or two specific things to try, plus a routing step that sends systemic gaps to a process fix instead of to a person. The replies and scores stay private, notes go one-to-one, and team-wide patterns are framed as 'something we're all tightening,' not anyone's personal failing. It's QA as coaching toward one voice — growth, not a gotcha.
How do I make sure the scores are fair before I act on them?
Treat every score as a draft to verify, not a verdict. The rubric is built from your own `support-voice.md` and `support-policy.md` so it measures your standards, and each 'needs-work' quotes the exact line that earned it — so you can spot-check a handful (especially the low scores) against the real reply in minutes. A lead verifies before anything reaches a person.
What's the difference between a systemic gap and an individual one — and why does it matter?
A systemic gap shows up across many agents (everyone hedges refunds, nobody links the help article); an individual one is one or two people. It matters because the fix is completely different: a systemic gap is a missing macro, an unclear policy, or a training hole — you fix the process and it stops recurring. Coaching the same thing into thirty people one reply at a time just nags; the last step routes systemic gaps to the macros, help center, or policy instead.
How is this different from CSAT?
CSAT tells you a score dropped; QA explains why. CSAT is the customer's verdict on the outcome, while this samples the actual replies and scores them against your voice and policy — so when a CSAT dip shows up, QA shows you the behavior behind it (the hedged refund, the missing next step) that you can actually coach and fix. They pair: run them together for the number and the reason.