By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
An assessment centre is effective when each exercise is deliberately mapped to a small number of pre-defined competencies, every competency is observed in more than one exercise, and multiple trained assessors record behaviour before they rate it. Its recurring failure mode is that scores cluster by exercise rather than by competency — candidates look consistent within a role-play and inconsistent across exercises — which means the centre is measuring the task rather than the underlying skill. Calibration, assessor rotation and a structured integration discussion are what separate a centre that produces comparable evidence from one that produces a long day and a set of impressions.
What an assessment centre is actually for
An assessment centre (assessment center in US spelling) is not a place. It is a method: a set of job-relevant exercises, observed by multiple trained assessors, scored against competencies that were defined before the exercises existed. The US Office of Personnel Management describes it as employing "multiple assessment methods and exercises to evaluate a wide range of competencies" — the plural is the point. One exercise scored by one person is an interview with extra steps.
The reason organisations accept the cost is that a centre can show behaviour rather than description. A candidate can tell you how they handle a difficult stakeholder; an in-basket exercise and a role-play show what they actually do when the stakeholder is impatient and the information is incomplete. That evidence is comparable across candidates in a way that a conversation is not, provided the design work was done.
Everything below is design work, and none of it is about the venue or the length of the day. A centre that runs six hours and produces four assessor opinions is not stronger than one that runs two and produces a scored, traceable record against a competency map.
The classic problem: measuring the exercise, not the competency
The best-documented weakness of assessment centres is that their scores tend to cluster by exercise rather than by competency. A candidate scores consistently high across every dimension in the group discussion and consistently mid across every dimension in the in-basket — which is the opposite of what the design intends. If "influencing" is a real, stable competency, an influencing score from the role-play and an influencing score from the presentation should agree with each other more than they agree with the other dimensions of their own exercise.
Charles Lance's 2008 review of this evidence is blunt about the implication: assessment centres often do not work the way they are supposed to, because what the assessors are picking up is performance on a specific task, situationally shaped, rather than a trait that travels. That does not make the centre worthless. It changes what you are allowed to conclude from it, and it changes how you report the result.
The practical consequence is that a single overall competency score, averaged across exercises, can hide the thing you most need to see. If a candidate handles pressure well in a written prioritisation task and poorly in a live customer role-play, the average tells you nothing and the two observations tell you a great deal — particularly if the job is mostly live conversation. Keep the exercise-level evidence visible, and treat the roll-up as a summary of it rather than a replacement for it.
Exercise-to-competency mapping is the real design task
Before writing any exercise, write the matrix: competencies down one axis, exercises across the other, and a mark in each cell where that exercise is genuinely capable of eliciting that behaviour. The discipline is in what you leave blank. A common failure is a matrix where every exercise claims to measure every competency, which usually means no exercise was designed to measure anything in particular.
Two rules keep the matrix honest. First, each exercise should carry a small number of competencies — three or four is a working ceiling, because an assessor watching a twenty-minute role-play cannot reliably track eight distinct behaviours at once. Second, each competency should appear in at least two exercises, in different formats, so you can see whether the signal survives a change of task. That second rule is the direct answer to the exercise-effect problem above: it is what makes the inconsistency visible instead of averaging it away.
The matrix also does the compliance work. The Uniform Guidelines on Employee Selection Procedures and the AERA/APA/NCME Standards both approach content-based evidence the same way: you show that what is measured is a representative part of the job, and you can document the chain from the job analysis to the task to the scoring standard. A matrix nobody wrote down is a chain you cannot show later. Obligations vary by jurisdiction and none of this is legal advice, but the documentation habit is the same everywhere.
- Competencies come from a role blueprint, not from the exercises you already own.
- Cap each exercise at three or four competencies; blank cells are a design decision, not a gap.
- Every competency appears in at least two differently shaped exercises.
- Each cell names the observable behaviour, not the adjective — "states the trade-off before choosing" rather than "shows judgement".
- Anything the exercise cannot elicit is explicitly out of scope for that exercise, and assessors are told so.
Multiple trained assessors — and what "trained" means
Multiple assessors is not the same as multiple opinions. The purpose of a second assessor is not a second vote; it is an independent observation of the same behaviour, recorded separately, so that disagreement becomes visible rather than absorbed. If two assessors discuss the candidate before either has written anything down, you have one observation with a witness.
Assessor training has a specific content, and it is behavioural rather than motivational. Assessors are trained to observe, record, classify and only then evaluate: write what the candidate did and said, assign it to a competency, and rate it last. Most assessor error happens when those steps collapse into one — the rating is formed in the first two minutes and the notes are then written to support it. Training also covers what not to score: the typo in a task about judgement, the accent in a task about clarity of reasoning, the confidence in a task about accuracy.
Rotation matters too. No candidate should be seen by the same assessor in every exercise, because a strong or weak first impression then follows them through the whole day. Rotating assessors across exercises is the cheapest available control on that, and it costs nothing but a scheduling constraint.
Calibration, and the integration discussion
Calibration happens before the centre runs, not after it goes wrong. The session is short: assessors independently score two or three real sample responses — one clearly strong, one clearly weak, one genuinely borderline — then reveal all scores at once and discuss only the criteria where they diverged by more than a level. The output is not agreement in the room; it is sharper wording in the rubric and saved anchor examples attached to each level.
The integration discussion at the end of the centre is where a lot of evidence quietly evaporates. Done well, it walks competency by competency, each assessor reads their recorded behaviour before any rating is stated, and disagreements are resolved by returning to the evidence and the rubric wording. Done badly, it starts with an overall impression from the most senior person in the room and works backwards. The order of speaking is not a nicety; it decides what the meeting measures.
Record the disagreements rather than smoothing them. A competency where two trained assessors, working from the same rubric, land two levels apart is telling you something about the rubric, the exercise, or the candidate — and you cannot tell which if the only surviving artefact is the agreed number.
What you can observe, and what you should not claim
There are things a well-run centre lets you measure about itself. Inter-rater agreement on a competency — an intraclass correlation across assessors who scored the same candidates on the same criterion — tells you whether the rubric is being read the same way, and it is computable whenever at least two assessors overlap on enough candidates. Internal consistency across the items of an objectively scored exercise tells you whether that exercise is coherent. Coverage — which competencies were observed, in how many exercises, by how many assessors — is bookkeeping, and it is the first thing an auditor asks for.
There are also things you should not claim. Running a structured centre does not entitle anyone to say the result is validated, normed or benchmarked, and it does not license a claim that the score forecasts later job performance. What you can defend is content-based evidence: this competency came from the role, this exercise elicits it, this rubric level describes it, these assessors applied it and here is how much they agreed. Where the data is thin — one assessor, five candidates — the honest report says the coefficient is not estimable rather than printing a number that looks like precision.
The same restraint applies to fairness. Comparing outcomes across groups is worth doing, and small-group differences are a prompt to investigate the exercise and the rubric, not a verdict. Legal obligations around adverse impact differ by jurisdiction and this is not legal advice; the useful internal posture is that a flag opens a review of the design, and the design documentation is what makes that review possible.
Evidence, or a long day
The separation between the two kinds of centre is visible before anyone arrives. An evidence-producing centre has a written competency matrix with deliberate blanks, exercises capped at three or four competencies, each competency observed twice in different formats, assessors calibrated on real samples with anchor examples, assessors rotated across exercises, behaviour recorded before it is rated, and an integration discussion that reads evidence before it states scores.
A long day has a full schedule, experienced people, a debrief, and a decision that everyone in the room believes in and nobody can reconstruct six weeks later. The cost is identical. The difference in what you can defend — to a hiring manager, to an unsuccessful candidate, to an auditor — is total.
If you are building one now, start with the matrix and the calibration session. They are the cheapest artefacts in the whole exercise, and they decide whether the rest of the day produces anything you can use.
- assessment centres
- exercise design
- assessor training
- calibration
- competency mapping
Key takeaways
- Map exercises to competencies before writing the exercises; blank cells in the matrix are a design decision, not an omission.
- Observe every competency in at least two differently shaped exercises — that is what makes the exercise effect visible instead of averaging it away.
- Train assessors to observe, record and classify before they rate, and rotate them so no candidate is seen by the same person all day.
- Calibrate on real sample responses before the centre runs, and keep anchor examples attached to each rubric level.
- Report what you can compute — assessor agreement, coverage, exercise-level evidence — and do not claim validation, norming or forecasting of job performance.
Sources & further reading
- Lance (2008), Why Assessment Centers Do Not Work the Way They Are Supposed To — Industrial and Organizational Psychology
- U.S. Office of Personnel Management, Assessment & Selection: Assessment Centers
- Standards for Educational and Psychological Testing (AERA, APA, NCME) — Joint Committee
- Uniform Guidelines on Employee Selection Procedures, 29 CFR Part 1607 — eCFR
