ع
Learn Tracks Reference Guides Saved
playbook

Find what's really driving your CSAT

Join the survey scores, the free-text comments, and the underlying tickets to find what actually moves your CSAT — split into 'fix the process' vs 'fix the reply' — and hand leadership the 2–3 levers worth pulling, framed as hypotheses to verify, not proof.

medium ~40 min
when to reach for this

CSAT moved, and now everyone's anxious about a number that tells them nothing they can act on. The score is a symptom — the cause is buried across three places at once: the ratings, the comments people left, and the tickets behind those ratings. Was it slow resolution? A specific issue type? A channel that's letting people down? A policy that's frustrating, or just a reply that landed wrong? This system joins those three sources, finds what actually drives the low (and the high) ratings, and separates the drivers into "fix the process" vs "fix the reply" so each goes to the right place. The point is to turn a scary number into a short, prioritized list of what to change — with every finding framed as a hypothesis to verify, because survey data is biased and correlation isn't cause.

gather this first
  • The survey export as csat.csv — at minimum a score column and the free-text comment; a response date, the ticket ID, and the channel make everything richer. In Claude Desktop, drop the file into the chat (or open the folder it lives in) so Claude can read it. Scrub names, emails, and account numbers to [customer] before you share it.
  • The matching tickets as tickets.csv if you can join them — issue type, time-to-resolve, channel, and first-contact-resolution are the attributes that turn a comment into a driver. The shared key is usually a ticket ID or email; if you only have the survey, you can still do the comment analysis.
  • What you already suspect is hurting the score — a slow queue, a recent policy change, a rough channel — written in one line, so Claude can confirm or challenge it rather than rediscover it from scratch.
the workflow
  1. Load it, then sanity-check the sample before reading anything into it

    A CSAT export is a biased sample — the people who felt strongly (usually the angry ones) self-select into answering. Before a single conclusion, make Claude show you the response rate, the score distribution, and the sample size, and name that bias out loud, so nobody treats a loud minority as the whole customer base.

    you ask
    Read csat.csv. First confirm the row count, the date range, and the score distribution (how many at each rating). Then estimate the response rate if I tell you total ticket volume, and call out the self-selection bias explicitly — who's most likely to have answered, and how that should temper everything we read into this. Don't analyze drivers yet; I just want to know how trustworthy this sample is.

    what you get back A confirmation ("840 responses, May 1–31, scores skew bimodal: 320 at 5★, 210 at 1★") and a plain caution — "this is roughly a 12% response rate and the distribution is U-shaped, which is classic self-selection: delighted and furious customers answer, the satisfied-but-quiet middle mostly didn't. Read every finding as 'true of respondents,' not 'true of all customers.'"

    A small sample or a low response rate can mislead badly — if it's a few dozen responses, treat the whole exercise as directional, and say so when you present it.

  2. Read the comments and cluster the reason behind each score

    "Negative" and "positive" tell you nothing to fix. The useful unit is the reason — the actual driver behind the rating. Have Claude cluster the verbatim comments by what people are reacting to, not by sentiment, so a flat pile of opinions becomes a handful of named drivers.

    you ask
    Read the free-text comments. Cluster them by the REASON behind the score — the actual driver, not just positive vs negative. For low scores, what specifically frustrated people (slow resolution, a confusing reply, a policy they hit, the wrong answer)? For high scores, what specifically delighted them? Give me the top drivers on each side with a count and one real comment per cluster, scrubbed to [customer]. Flag any cluster that's fuzzy or thin.

    what you get back Two sets of named drivers with counts — low: "slow resolution (88), unhelpful first reply (54), refund policy frustration (40)"; high: "fast fix (130), agent went the extra mile (70)" — each with an anonymized real quote and a confidence flag where the grouping is loose.

  3. Correlate the drivers with ticket attributes — as ranked hypotheses, never proof

    If you can join the tickets, the comments become testable. Does a low score track with long resolution time? A particular issue type? A channel? Have Claude rank what correlates with low (and high) scores — but pin it firmly as correlation, not cause, with a confidence note, because a pattern in a biased sample is a lead, not a verdict.

    you ask
    Join csat.csv to tickets.csv on the ticket ID. Then tell me which ticket attributes correlate with low scores and which with high — look at time-to-resolve, issue type, channel, and first-contact-resolution. Rank the correlations strongest-first, each with the size of the effect and a confidence note. Be explicit that these are correlations and possible drivers to VERIFY, not proven causes — and warn me where the sample for a slice is too small to trust.

    what you get back A ranked list framed as hypotheses — "strongest signal: tickets resolved in >48h average 2.1★ vs 4.3★ under 24h (likely a real driver, large sample); chat scores lower than email but the chat sample is only 30 responses — treat as a weak hint." Correlations, with confidence and a caution, not conclusions.

    Correlation isn't cause. "Slow tickets score low" might mean slowness frustrates people — or that the hard, doomed cases are both slow and low-scoring for the same upstream reason. Claude's job is to surface the candidate; yours is to test it.

  4. Split the drivers into process vs interaction so each goes to the right fix

    A driver is only useful if you know who can fix it. "Slow resolution" and "a painful refund policy" are process problems — staffing, tooling, policy. "A confusing reply" and "the wrong tone" are interaction problems — coaching and QA. Splitting them sends each to the right owner instead of a vague "do better."

    you ask
    Take the drivers we found and sort them into two buckets. PROCESS: things outside any one reply — slow resolution, a frustrating policy, a tooling gap, a channel that's under-resourced. INTERACTION: things in the reply itself — confusing wording, wrong tone, an unhelpful or incorrect answer. For each driver give the bucket, the count, the likely owner (ops/policy vs coaching/QA), and how confident you are in the placement.

    what you get back Two labeled lists — Process: "slow resolution (88) → ops/staffing; refund policy friction (40) → policy owner." Interaction: "unhelpful first reply (54) → QA/coaching; cold tone on chat (22) → coaching" — each with a count, an owner, and a confidence flag so the borderline ones are visible.

  5. Write the leadership summary — where it stands, the top drivers, the action per driver

    Package it so a busy leader gets the story in thirty seconds: where CSAT is, the 2–3 levers that would actually move it, the evidence behind each, and the recommended action — with the caveats kept intact, so nobody mistakes a hypothesis for a finding.

    you ask
    Write a leadership summary, paste-ready for our channel. Open with where CSAT stands (current score, trend, the sample caveat in one line). Then the top 2–3 drivers worth acting on, each with: the evidence (count + correlation), whether it's process or interaction, the recommended action, and the owner. End with a one-line caveat that these are hypotheses from a biased sample to verify, not proven causes. Keep it factual, cite counts everywhere, and keep it under a screen.

    what you get back A tight summary — current score and trend with the sample caveat, then 2–3 prioritized drivers each with evidence, bucket, recommended action, and owner, closing on the "verify before you act" line — that reads in thirty seconds and survives a skeptical leader who asks "how do you know?"

    The recommended actions are Claude's proposals from a biased sample — a human decides which to fund. Present them as "here's what the data suggests we test," not "here's what's wrong."

make it your own
  • Pull the qualitative and the volume in beside it: the comment clusters here are the same signal the Turn a month of tickets into a voice-of-customer report playbook builds at scale, and the issue-type drivers line up with Find what the queue is really about — run all three and a driver that shows up in scores, comments, and volume is one you can trust more than any single source.
  • Route the tone drivers to coaching: any interaction-bucket driver — cold replies, confusing answers, wrong tone — is the input to QA the queue and coach the team to one voice, which turns "replies score low" into specific coaching, not a vague memo.
  • Escalate a recurring product driver: if a low-score cluster keeps pointing at the same broken thing, it's not a CSAT problem, it's a bug — hand it to Turn a cluster of tickets into a bug engineering will act on with the count and the score evidence attached.
  • Make it the CSAT slot in the weekly cadence (Power Track): this analysis is one input to Run a weekly support operating system — drop it in as the satisfaction read each week. If you run the same export repeatedly, a /csat custom command or a scheduled agent (see the Playbook's Features tab) can draft the first cut. Custom commands and scheduled agents are the opt-in Power Track — on Desktop you run these same prompts by hand each cycle until you're ready to automate.
watch out for
  • Survey data self-selects — the people who answer skew to the delighted and the furious, with the quiet satisfied middle missing. Every finding is "true of respondents," not "true of all customers," and a low response rate makes that gap worse. Keep that caveat attached to the number the whole way up.
  • Correlation is not cause. "Slow tickets score low" is a hypothesis — slowness might frustrate people, or the hardest cases might just be both slow and unsatisfying for a shared upstream reason. Treat every driver Claude surfaces as a lead to verify, and say so when you present it.
  • A small sample can mislead with total confidence — a 2-star average on a channel with 30 responses is a hint, not a fact, and Claude will happily rank it next to a 500-response finding if you let it. Make Claude flag thin slices, and discount them yourself.
  • A CSAT export is dense with PII — comments name people, mention account details, and sometimes quote private complaints. Scrub names, emails, and account numbers to [customer] before uploading, and keep every quoted comment anonymized in the summary.
  • Claude proposes the drivers and the actions; a human decides what to do. The recommended levers are inferences from a biased sample — present them as "what the data suggests we test," not "what's wrong," and own the call to act.

you'll end up with A short, prioritized read on what's actually moving your CSAT — the top 2–3 drivers split into process vs interaction, each with its evidence, owner, and recommended action — framed honestly as hypotheses to verify, so a number that made everyone anxious becomes a plan a leader can act on.

Questions people ask

Can Claude tell me WHY my CSAT dropped?
Not as proof — it can tell you what most likely drove it, framed as hypotheses to verify. Survey data self-selects and correlation isn't cause, so "slow tickets correlate with low scores" is a strong lead, not a confirmed reason. Claude surfaces and ranks the candidate drivers with confidence notes; you verify the top one or two before you act on them.
Why split drivers into 'process' vs 'interaction'?
Because they go to different owners. A slow queue or a painful policy is a process problem — ops, staffing, policy. A confusing or cold reply is an interaction problem — QA and coaching. Sorting drivers into the two buckets turns a vague "improve CSAT" into a specific fix routed to the person who can actually make it.
My sample is small — is this still worth doing?
Yes, but as a directional read, not a verdict. Make Claude show you the response rate and flag any thin slice (a channel with 30 responses, say), and present the whole thing as "what this suggests we look at," not proof. A small sample sharpens what to investigate next; it doesn't settle anything on its own.
How do I handle the PII in the survey comments?
Scrub names, emails, and account numbers to `[customer]` before uploading — comments are often where the most personal detail lives, including private complaints. Keep every quote anonymized in the leadership summary too; the drivers don't need anyone's name attached to be convincing.