Review season is where fairness quietly breaks at scale. Recency bias rewards whoever shipped something in the last three weeks, the halo effect lets one strong trait carry a whole rating, and the person who writes the most polished self-assessment edges out the colleague who just did the work — and on top of all that, "exceeds" means something different from every manager. The fix isn't a better form; it's a discipline: anchor every review to one competency bar, weight evidence over eloquence, keep genuine disagreements visible instead of averaging them away, run a deliberate bias pass, and calibrate ratings across people so the same rating means the same thing. Claude synthesizes the scattered inputs into an evidence-anchored draft; the manager and the calibration room make every call.
- The competency framework and level expectations for this person's role — what "meets" and "exceeds" actually require at their level. This comes straight from the Build the competency framework every role inherits playbook; without one shared bar, every rating is a private opinion.
- The inputs, each clearly labeled: the self-review (
self-review.md), peer feedback (peer-feedback.md), the manager's own notes (manager-notes.md), and the real work artifacts (shipped projects, docs, metrics) — opened in a Claude Desktop workspace your company has approved for people data, with each file read approved in the 'Ask permissions' prompt. - What the rating scale actually means and what each rating is tied to — comp, promotion, or a PIP. A rating with real consequences attached needs a heavier touch than one that's purely developmental.
-
Anchor to the competency bar before reading a single input
Open the folder with the framework and the inputs in Claude Desktop and start in the chat — no terminal needed. Before any synthesis, ground Claude in this person's level expectations, so "strong" means "strong at what this level requires" — not a warm general impression. Approve the read of the framework file in the 'Ask permissions' prompt first.
you askRead the competency framework and the level expectations for this person's role and level. Don't open the review inputs yet. Read back to me in bullets: each competency, and what 'meets' versus 'exceeds' actually looks like at THIS level. That's the bar I want every later judgment measured against.what you get back A short, accurate read-back of the competencies and the meets/exceeds bar for this specific level — so the rest of the cycle is anchored to the framework, not to whoever wrote the most or most recently.
Anchoring to the bar first is what stops the next step from rewarding eloquence or recency instead of evidence against what the role actually requires.
-
Synthesize the inputs with evidence weighted equally
Now feed all four inputs and ask for an evidence-anchored read — explicitly telling Claude not to let the longest self-review or the most recent month dominate. Map the evidence to each competency so you can see where the case is strong, thin, or missing.
you askNow read self-review.md, peer-feedback.md, manager-notes.md, and the linked work artifacts. For each competency in the framework, summarize the evidence from all four sources and how it maps to the meets/exceeds bar. Weight the evidence equally — do NOT let the longest self-review or the most recent few weeks dominate, and don't let a confident self-assessment outweigh quieter peer evidence. Tie every point to a specific source.what you get back An evidence map — each competency with the corroborating evidence from across the four inputs, scored against the bar, every claim tied to its source — rather than a vibe-based summary that rewards whoever wrote most or shipped last.
-
Surface disagreement and recency bias — don't average them
An averaged read hides exactly what the manager and calibration room need to weigh. Pull the genuine splits into the open, and separately flag any judgment that's really driven by the last few weeks rather than the whole review period.
you askWhere do the inputs genuinely DISAGREE about this person — for example, the self-review claims something peers don't corroborate, or two peers read the same trait differently? Don't average it into a middle rating. Name each split, quote both sides, and say what evidence would resolve it. Then, separately, flag anything that looks driven by the last few weeks rather than the whole review period — recency bias.what you get back A short list of true disagreements with both sides quoted and a tiebreaker for each, plus a separate set of recency flags ('strong rating rests mostly on the Q4 launch; little evidence from the first three quarters') — kept visible, not blended away.
This is the heart of fairness. A real cycle makes disagreement and recency visible and resolvable, not invisible inside an averaged number.
-
Run the bias pass as its own step
Run a dedicated check for language that isn't about the work. Personality, "attitude", time away or parental leave, and anything that proxies a protected characteristic are where bias hides — naming them lets the manager set them aside on purpose. Flag only; never let it rewrite the verdict.
you askScan all the inputs for anything that isn't job-related evidence: comments on personality, 'attitude', likability, communication 'style' standing in for substance, time away / parental or medical leave, or anything that could proxy a protected characteristic (age, gender, background, accent, caregiving). For each, quote the phrase, say why it isn't job-related, and suggest the work-based question it should have been instead. Flag only — do NOT change any rating or rewrite the verdict.what you get back A flag list — each non-job-related phrase quoted, why it isn't evidence, and the work-based question it should have been ('not a team player' → what specifically did collaboration require here, and what did they do or not do?) — for the manager to consciously set aside.
This surfaces language to reconsider; it never overrules the manager. You and the calibration room still make every call.
-
Draft the review and a calibration brief — no auto-rating
Pull it together into the review narrative and a neutral calibration brief — and stop short of a final number. Claude produces the evidence-anchored narrative against each competency and a suggested-evidence summary, not a verdict. Review the accept/reject diff before anything is saved; the manager owns the rating and calibration owns consistency.
you askDraft the review: for each competency, an evidence-anchored narrative of where this person lands against the meets/exceeds bar, citing the sources. Then a SUGGESTED-EVIDENCE summary per competency (the strength of the case as 'strong / mixed / thin' — NOT a final rating or number). Finally a neutral one-page calibration brief I can read alongside other people's: the bar, the evidence, the open disagreements, the bias flags. Do NOT assign a final rating — that's mine and the calibration room's to set.what you get back An evidence-anchored review narrative per competency, a suggested-evidence summary that stops short of a number, and a neutral calibration brief built to sit side by side with the rest of the team — so the same rating ends up meaning the same thing across people.
The manager owns the rating; calibration owns consistency. Claude drafts the evidence and the brief, never the verdict — the same equal-weighting discipline as Turn five interviewers into one fair verdict.
- Calibrate the whole team: run the five steps identically for each person, then open all the calibration briefs side by side and ask Claude to line them up — where does the same rating rest on visibly different evidence, and where is one manager's 'exceeds' another's 'meets'? Same bar, same synthesis, same bias pass per person is what makes a rating mean one thing across the team instead of laundering manager-to-manager drift.
- Pair with the delivery message: once the rating is set, hand the narrative to the Draft a sensitive message you can stand behind playbook to write the actual review conversation and the growth message — honest, specific, and humane, so a fair assessment lands as a fair conversation.
- Spin a forward growth plan: ask Claude to turn the thinnest competencies and the open disagreements into a concrete development plan for next period — the gaps this cycle surfaced become the goals the next cycle measures, so reviews compound instead of resetting.
- Arabic / bilingual: where Arabic is the working language, synthesize and draft the review and the delivery message in Arabic — adapt the register and the sentiment to land naturally for the person, rather than machine-translating an English draft. Save the flow as a `/review` skill (Power Track): capture the five-step prompt sequence as a reusable skill or custom command (see the Playbook's Features tab) so every manager runs the same anchored, bias-checked cycle. Skills and custom commands are the opt-in Power Track — on Desktop the whole cycle runs fine as a shared set of saved prompts each manager works through in order.
- People data stays in an approved workspace, full stop. Reviews, self-assessments, peer feedback, and anything touching comp never go into a tool your company hasn't cleared for people data — confidentiality is the whole job here, not a footnote.
- Recency and halo bias are the default failure mode, not the edge case. Left alone, the rating drifts toward the last few weeks and toward one strong trait — anchor every judgment to the whole period and the competency bar, which is exactly what steps one and three are for.
- Calibrate, or it isn't fair. A cycle that skips calibration just launders manager-to-manager drift into official ratings — the same 'exceeds' meaning three different things. Running the briefs side by side is what makes a rating portable across the team.
- Anything a rating is tied to — comp, promotion, a PIP, or a termination — goes past HR and legal before it lands, and the number is decided by the manager and the calibration room, never by Claude. Claude synthesizes the evidence; people decide the consequence.
you'll end up with A fair, evidence-anchored, bias-checked, and calibrated set of reviews in about half a day per cycle — each rating measured against one competency bar rather than eloquence or recency, genuine disagreements kept visible, non-job-related language flagged and set aside, and the same rating meaning the same thing across the team — with the number, and every consequence attached to it, firmly a human decision.
Questions people ask
- Does Claude decide the rating?
- No. Claude reads the inputs against the competency bar, maps the evidence, surfaces disagreement and bias, and drafts the narrative and a calibration brief — but it explicitly stops at a 'strong / mixed / thin' read of the evidence and never assigns a final number. The manager sets the rating and the calibration room confirms consistency; Claude only does the synthesis underneath that decision.
- Is it safe to put performance data and self-reviews into Claude?
- Only inside a Claude Desktop workspace your company has approved for people data, with each file read approved in the 'Ask permissions' prompt. Reviews, peer feedback, and anything touching comp are among the most sensitive data you handle — never paste them into a general chat or a tool that hasn't been cleared. The confidentiality rule isn't a nice-to-have here; it's the first pitfall for a reason.
- How does this actually reduce bias?
- Three ways, deliberately separated. Anchoring to the competency bar before reading anything stops 'strong' from meaning a general good impression. Weighting evidence equally stops the longest self-review or the most recent month from dominating. And the dedicated bias pass quotes any non-job-related language — personality, 'attitude', time away, anything proxying a protected characteristic — and proposes the work-based question it should have been. It flags for the manager to set aside consciously; it never silently rewrites a rating.
- Can it calibrate ratings across different managers?
- That's the point of the calibration brief. Run the same five steps per person, then put the neutral briefs side by side and ask Claude to show where the same rating rests on visibly different evidence, or where one manager's 'exceeds' is another's 'meets'. Claude surfaces the inconsistencies; the calibration room resolves them. Same bar, same synthesis, same bias pass per person is what makes a rating portable across managers instead of drifting team to team.