By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
A job assessment is worth a candidate's time only when it measures something the role actually requires, takes no longer than the decision it informs, and produces evidence a named person uses. Many assessments fail all three: they test recall rather than the work, add days to a process without shortening any other stage, and end in a number no interviewer opens. The exceptions share a checkable structure — a task drawn from the job itself, a rubric agreed before anyone is scored, and an output that visibly changes who advances.
The objection, stated as well as its critics state it
Start with the version of the complaint that is hardest to answer, because it is the one candidates are actually making. They were sent a forty-minute assessment that asked them to recall syntax, formula names or policy definitions — things that, on the job, they would look up in eight seconds and never think about again. The assessment did not ask them to do the work. It asked them to prove they had memorised its vocabulary, and a strong candidate knows the difference immediately.
Then there is the calendar. An assessment inserted between the application and the first conversation adds a stage, and stages have queues. The candidate waits for the invitation, waits for an uninterrupted hour, then waits again for someone to read the result. Nothing downstream got shorter to pay for it — the same four interviews still happen afterwards. A process that was three weeks is now four, and the people most likely to leave during the extra week are the ones holding another offer.
The third part of the objection is the quietest and the most damning: nobody used the result. An hour of work produced a number, and the panel that made the decision never opened it — or opened it, saw 74, and had no idea whether 74 was good. When a hiring team cannot say what a score changed, the honest description of that assessment is a filter nobody trusted, applied to people who did not get to skip it.
Four ways an assessment wastes everyone's time
Those complaints have distinct causes, and separating them matters, because they have different fixes and only one of them is about the questions.
That last distinction is the useful one. Only the first failure below is a question-writing problem; the other three are process design. It is why swapping test vendors so rarely rescues a process people already dislike — the new questions arrive into the same empty slot in the workflow.
- It measures something the role does not require: recall instead of judgment, speed instead of accuracy, or a tool the team stopped using two years ago.
- Its length is out of proportion to the decision it informs. Ninety minutes to decide whether someone gets a thirty-minute screening call is a bad trade, however good the ninety minutes are.
- It has no rubric, so the score is one reviewer's impression written down. Two reviewers reading the same answer disagree, and neither can say why.
- Its output has no consumer. If no named person reads the result before a decision, the assessment is not part of the process — it is a stage happening next to it.
Position in the funnel decides most of the cost
An assessment's cost to a candidate is not fixed. It depends on what they have already invested and what they are being asked to invest next. Thirty minutes after a promising conversation with the hiring manager reads as a reasonable next step. The same thirty minutes before any human contact reads as an unpaid audition for a company that has not yet shown any interest in return.
That asymmetry decides who drops out. Candidates running several processes triage by effort per unit of signal, and an early, long, unexplained assessment is the easiest thing on the list to abandon. So a badly placed assessment does not merely annoy strong candidates — it removes them, and it removes them non-randomly. The pool that survives is the pool with fewer alternatives, which is the opposite of what the assessment was bought to achieve.
The fix is usually not a shorter test. It is moving the assessment to the point where the candidate already knows the role is real, and paying for it by removing a stage rather than adding one. If an assessment does not let you delete a later interview, it is worth asking out loud what it is for.
What the exceptions have in common
The assessments that earn their place are not distinguished by being harder, longer or more scientific-sounding. They are distinguished by four fairly boring properties, and every one of them can be checked before a single candidate sees the thing.
The first property carries most of the weight. The US Office of Personnel Management's guidance on work samples describes the mechanism plainly: tasks that mirror the work carry a high degree of content validity, and candidates tend to perceive them as fair — the part hiring teams forget is a benefit rather than a nicety. A support candidate answering a genuinely difficult customer message is doing the job for twenty minutes. There is nothing to memorise, and very little to resent. The same guidance is equally plain about the cost: work samples take real effort to build and to administer, which is precisely why they should be reserved for decisions that deserve the effort.
- The task is drawn from the job. Someone who has done the role can look at it and recognise the work, not a proxy for the work.
- The rubric exists before anyone is scored. What counts as a strong answer is written down first, so the standard is not reverse-engineered from the candidates who happened to apply.
- The output is specific enough to argue with. Not a single number but criterion-level evidence a hiring manager can disagree with in a sentence that names something concrete.
- Someone is named as its reader, and that person's decision visibly moves when the evidence changes.
The evidence is more modest than the sales deck
It is worth being honest about how strong the underlying research actually is, because the assessment industry has quoted numbers more confidently than the literature supports. In 2022 Sackett, Zhang, Berry and Lievens re-examined the meta-analytic validity estimates that had underpinned selection practice since the late 1990s and found that widely repeated statistical corrections had inflated them. Their revised figures are lower across the board. Structured, job-related methods still come out ahead of unstructured ones, but the honest summary is 'better than the alternative', not 'solved'.
The legal framing points the same direction. The EEOC's guidance on employment tests states that a selection procedure with a disparate impact on a protected group must be shown to be job-related and consistent with business necessity, and the Uniform Guidelines set out how that showing is made. Obligations differ by jurisdiction, none of this is legal advice, and no claim of compliance is being made here. But the underlying instruction is a sound design rule wherever you hire: be able to say which part of the job this measures, and why that part matters.
This is also why SkillCort describes content validity — the traceable chain from a role blueprint to a task to a rubric criterion — and reports reliability where there is enough data to compute it, rather than asserting that an assessment forecasts later job performance. That is a smaller claim than the category usually makes. It is also one that can be shown to you, item by item, which is the only kind worth making.
A short test for your own test
If you want to know whether your assessment sits in the majority or among the exceptions, four questions settle it faster than any vendor comparison.
An assessment that survives those four is doing work, and candidates generally tolerate work that is visibly connected to the job. One that fails them is not redeemed by better questions or a nicer interface, and the candidates complaining about it are describing your process accurately.
So the honest answer to the question in the title is: often yes. That is not an argument against measuring skill. It is an argument against measuring it out of habit, at the wrong moment, with an output nobody is waiting for.
- Would you be comfortable showing a candidate the rubric before they start? If not, you are measuring something you cannot defend.
- Can you name the stage this assessment removed? If it only added one, the process got longer with nothing paying for it.
- Point at the last decision it changed — not a decision it agreed with, one it changed.
- Six months from now, could you reconstruct why one candidate was rejected and another advanced from the record rather than from memory?
- candidate assessments
- assessment design
- work samples
- candidate experience
- skill-based hiring
Key takeaways
- Most of the complaint is accurate: many assessments test recall instead of the work, extend the process without shortening it, and end in a score nobody consumes.
- Only one of the four common failures is about the questions. The rest are process design, which is why changing vendors rarely fixes a process people already dislike.
- Position in the funnel drives most of the cost. An early, long assessment removes candidates who have other options, and it removes them non-randomly.
- The exceptions share four checkable properties: a task drawn from the job, a rubric agreed in advance, an output specific enough to argue with, and a named reader.
- If an assessment does not let you delete a later stage, ask what it is for.
Sources & further reading
- U.S. Equal Employment Opportunity Commission, Employment Tests and Selection Procedures
- U.S. Office of Personnel Management, Assessment & Selection: Work Samples and Simulations
- Sackett, Zhang, Berry & Lievens (2022), Revisiting Meta-Analytic Estimates of Validity in Personnel Selection — Journal of Applied Psychology
