Every manager carries a private idea of what "senior" means, so job descriptions invent requirements, reviews grade on vibes, and promotions come out inconsistent and legally shaky. The fix is one competency framework — named competencies, a level ladder, and the observable behaviors that show each level — built once and inherited by every job description, interview scorecard, performance review, and promotion case, so the whole org judges people against the same bar. This is the what good looks like; the JD playbook decides what a single role asks for, but it gets sharper once "the bar" finally points at one document. Build this early — most other People & HR playbooks tighten up the moment a real definition of each level exists to anchor to.
- The role family or function this is for — eng, sales, support, ops — and a rough note on what you'd expect at each level, typed straight into the chat in Claude Desktop or dropped in as
role-notes.md. - Any leveling docs or rubrics you already have — an old ladder, the last performance rubric, even an informal "here's how we think about senior" doc — so Claude matches your reality instead of inventing a generic one.
- A Claude Desktop workspace your company has approved for people-related work. This is role-level definition, not individual employee records — but keep it tidy and treat it as people-adjacent.
-
Name the competencies and the level ladder — before any behaviors
Open the folder with your notes in Claude Desktop and ask in the chat — no terminal needed. Force the foundation first: which 4–6 competencies actually matter for this role family, and the level ladder you'll grade against. Don't write a single behavior yet — get the skeleton right before you fill it in.
you askRead role-notes.md. For this role family, propose 4–6 core competencies that genuinely matter — what we actually hire, level, and promote on — and the level ladder we should use (e.g. IC1–IC5, or map to our existing levels). Challenge any competency that's really a personality trait or a duplicate. Don't write behaviors per level yet — just name the competencies and the ladder, and tell me which competencies you're unsure belong.what you get back A short list of 4–6 named competencies plus a clean level ladder — the grid you'll fill in next. If a competency can't be named without hand-waving, it probably isn't one.
Resist jumping to behaviors. A framework built on the wrong competencies just grades the wrong things consistently.
-
Define observable behaviors per level — not adjectives
Now fill the grid: for each competency at each level, what you'd actually see someone doing. This is the heart of the framework, and it has one rule — behavioral and evidence-able, never adjectives. "Strong" and "senior" describe nothing you can assess; a thing you'd watch them do does.
you askFor each competency, write the observable behavior at each level of the ladder. Hard rule: describe what you'd actually see the person doing — concrete, evidence-able actions — not adjectives like "strong", "senior", "owns ambiguity", or "high impact". If a cell reads like a personality grade instead of a behavior, rewrite it as the specific thing you'd witness. Flag any cell you can only describe in adjectives — that's a level we can't assess fairly yet.what you get back A filled competency × level grid where every cell is an observable behavior — and a flag on any cell that's still an adjective in disguise, so you can sharpen it before it reaches a review form.
A level you can't describe in observable behavior is a level you can't assess fairly — it'll get graded on vibes. Treat every flagged cell as unfinished, the same way the messaging framework treats a pillar with no proof.
-
Calibrate the boundaries between adjacent levels
The hardest and highest-value part: the line between adjacent levels. If two managers read "IC2 vs IC3" and land on different levels for the same person, the framework hasn't done its job. Make every transition crisp enough to survive two independent readers.
you askGo level-by-level on the boundaries. For each pair of adjacent levels (IC2→IC3, IC3→IC4, …), make the difference crisp enough that two managers reading it independently would level the same person the same way. For each boundary, name the single behavior that most clearly separates the lower level from the higher one. Then flag every transition that's still fuzzy or where the wording could be read two ways, and tighten it.what you get back Sharpened boundaries with a clear "this is what moves you from N to N+1" for each step on the ladder, and the still-fuzzy transitions flagged so you can resolve them before anyone is leveled against them.
This is the part that makes leveling defensible. A boundary two people read differently is a promotion dispute waiting to happen.
-
Pressure-test for bias and legality
A dedicated pass, because this is where a framework quietly goes wrong. Have Claude hunt for language that proxies for protected characteristics, tenure standing in for skill, "culture fit", or anything that isn't job-related — so the bar screens for ability, not background. Anything legally loaded gets routed to a human.
you askScan the whole framework for bias and legality. Flag: language that could proxy for a protected characteristic (age, gender, nationality, disability, parental status); anywhere tenure or years are standing in for actual skill; "culture fit" or "fits the team" phrasing; and any behavior that isn't genuinely job-related. For each flag, quote the wording, say why it could screen for background instead of ability, and suggest a job-related rewrite. Mark anything that touches protected characteristics or local labor law as "route to HR/legal" — don't try to resolve those yourself.what you get back A table of flags — coded language, tenure-as-proxy, culture-fit phrasing, non-job-related criteria — each with the reason and a job-related rewrite, plus a clearly marked set of items to route past HR or legal before the framework is used.
Read every flag yourself, and send the legally loaded ones to a real person. A bias scan is a prompt, not a guarantee, and labor law is not something to settle in a chat.
-
Assemble the one-page framework per role family
Pull it into one short canonical artifact the rest of the People & HR path inherits. Keep it tight — a framework nobody can hold in their head doesn't get used, and an unused framework is just every-manager's-own-bar with extra steps.
you askAssemble everything into one competency-framework.md for this role family: the named competencies, the level ladder, the observable behaviors per level, the calibrated boundaries (the one behavior that separates each adjacent pair), and a short note on how it's used (which JD, scorecard, review, and promotion case inherit it). Keep it to one page per role family — short enough that a manager can actually hold it in their head.what you get back A single, paste-ready
competency-framework.md— the source of truth the JD, the interview scorecard, the review, and the promotion case all inherit, short enough to be used rather than filed.Save this file. This is the spine the rest of the path hangs off — the JD pulls its requirements from it, the interview scorecard turns its competencies into rows, the review grades against it, and the promotion case argues against it.
- Feeds the JD and the hiring kit: hand
competency-framework.mdto the Write a job description that widens the pool playbook so requirements come from the named competencies, not a wishlist — then into the Run a hire from open to offer interview kit, where each competency becomes a scorecard row. The bar you write here is the bar candidates are scored against. - One format, many functions: keep a separate framework per function (eng, sales, support) but share one format across all of them — same competency-per-level grid, same calibration discipline — so a promotion case in support reads the same way as one in engineering and your review cycle stays consistent across the org.
- Arabic & bilingual: for Arabic-speaking teams, localize the level expectations and register — what "senior" looks like, the tone of the behaviors — rather than machine-translating the English. Ask Claude to adapt it for your context, then have a fluent people leader own the final wording, so the bar reads naturally in Arabic instead of like a translated American rubric.
- Make it a reusable skill (Power Track): load
competency-framework.mdas a reusable skill or a/levelcustom command (see the Playbook's Features tab) so every manager levels against the same framework instead of their own head. Custom commands and skills are the opt-in Power Track — on Desktop the same thing runs fine as a shared saved prompt every manager pastes in.
- Behaviors, not adjectives. A level described in adjectives — "strong", "senior", "owns ambiguity" — can't be assessed, so it gets graded on vibes. If you can't say what you'd see the person doing, the cell isn't finished.
- Calibrate the boundary or it's useless. If the line between two adjacent levels is fuzzy, two managers will level the same person differently and your promotions become indefensible. The boundary between levels is where the real work — and the real value — is.
- Bias and legality go past a real person. Proxies for protected characteristics and tenure-as-proxy-for-skill must be caught, and anything touching protected characteristics or local labor law goes past HR or legal before the framework is used. The bar should screen for ability, not background.
- One owner, versioned — or it drifts. A framework everyone can edit drifts back into every-manager's-own-bar within a quarter. Give it a single owner, treat changes as a reviewed version bump, and Claude drafts while the people leaders own the bar and the final wording.
you'll end up with A one-page competency framework — named competencies, a level ladder, observable behaviors per level, calibrated boundaries, and a bias-and-legality pass — inherited by every job description, interview scorecard, performance review, and promotion case, so the whole org finally judges people against the same bar instead of each manager's private one.
Questions people ask
- How is this different from a job description?
- A job description defines one specific role — its title, its day-to-day, its must-haves. The competency framework defines the bar *across* roles: what each level looks like, in observable behavior, for a whole role family. The JD inherits from it — your *Write a job description that widens the pool* requirements should come from the framework's competencies, not be invented per posting — so every role in the family is hiring and leveling against the same definition.
- Where do the \"observable behaviors\" come from if we've never written them down?
- From what your best managers already judge on implicitly. Claude drafts the behaviors from your rough notes and any old leveling docs, then you and the people leaders sharpen them against real examples — the engineer who clearly operates at the next level, the one who's close but not there. You're not inventing a standard from scratch; you're making the one already in managers' heads explicit, observable, and shared.
- Who should own the competency framework?
- One owner — usually a senior people leader or the function head, not a committee. The whole value is that it's a single source of truth; if everyone can edit it, it drifts back into every-manager's-own-bar within a quarter. Treat changes as a reviewed version bump, and have Claude draft while the people leaders own the bar and the final wording.
- Is it safe to use Claude for this — does it involve employee data?
- This is lower-risk than most People & HR work because it's role-level definition, not individual records — you're describing what "IC3" looks like, not grading a named person. Still, keep the work in a workspace your company has approved for people-related work, and never paste real performance data or named-employee examples into the chat while building the template. Build the bar from the framework; apply it to real people only inside your approved system.