By Daniel Whitmore, MSc in Industrial and Organizational Psychology
In short
The early warning signs of a bad hire usually appear in the first six to ten weeks: work that is technically correct but poorly prioritised, a ramp that stalls after the training period, repeated surprise about what the job involves, and panel doubts that resurface once the person starts. Each signal is more useful as a diagnosis of the selection process than of the person, because each one points to something the process failed to observe — decision-making under realistic constraints, actual working method rather than described experience, or a role definition the panel never agreed on. Read that way, a mis-hire is a defect report on the hiring design, and the fix belongs upstream in the next assessment rather than downstream in the next performance conversation.
Read the signal as a defect report, not a verdict
The obvious way to handle an early warning sign is to treat it as information about the individual and manage accordingly. That is sometimes right and always incomplete. A single hire is one data point about a person and one data point about a process that will select dozens more people the same way. Only the second is repeatable.
So the framing here is diagnostic. For each signal, the question is: what would the selection process have had to observe in order to see this coming, and did it have any way to observe it? Most of the time the answer is that the evidence simply was not collected — the process measured articulation, credentials or rapport, and the thing now going wrong was never put in front of anyone in a form they could score.
This is also the honest limit of the exercise. A signal at week six does not prove the assessment was wrong; onboarding, team context, manager clarity and workload all shape early performance, and some people ramp slowly and then outperform. What a signal does reliably tell you is where to look. If the same signal keeps appearing across hires from the same process, you are no longer looking at individuals. For the economics of that pattern and the full remediation sequence, our hiring guide on how to reduce mis-hires — linked at the end of this post — covers the cost model and the structural fixes; this post stays on the signals themselves.
Signal: the work is correct, the judgement is not
The person can do the task. Given a well-specified ticket, the output is fine. What is missing is the decision layer: which of five open items matters most this morning, when to escalate rather than keep trying, when the request as written is not the request that should be answered. Managers describe this as needing to "stay close" to work they expected to hand over.
This traces almost directly to what the process measured. Knowledge questions, credential screens and experience-led interviews all reward the ability to produce a correct answer to a stated problem. None of them presents an unstated problem, an incomplete brief or a conflict between two reasonable priorities — which is exactly the material the job turns out to be made of. If judgement was never elicited, its absence is not a surprise; it is an unmeasured variable that happened to land badly.
The design fix is a task that contains a trade-off rather than a right answer, and a rubric that scores the reasoning as well as the outcome. A prioritisation exercise with more work than time, a customer request that cannot be granted as asked, a data set with a discrepancy the brief does not mention: each produces a scoreable decision. What you look for is whether the candidate names the constraint before choosing — the observable behaviour a knowledge test cannot reach.
Signal: the ramp stalls after the training period
Structured onboarding hides a lot. While there is a curriculum, a buddy and a queue of supervised work, most hires look similar. The signal appears at the handover point: the person who was tracking well in week four is still asking week-four questions in week nine, and the questions are about the same underlying skill rather than about new territory.
The process failure here is usually that the assessment measured a description of the work rather than the work. Candidates who interview well are, among other things, good at narrating past performance, and that skill has only a loose relationship to executing the task in front of them. The selection literature is consistent on the direction of the point: structured, job-relevant methods — structured interviews and work samples, in the modern meta-analytic re-estimates summarised by Sackett and colleagues — carry more signal than unstructured conversation, though how much depends on the job and on how carefully the method is built.
The practical test of your own process is simple. Take the thing the stalled hire cannot yet do, and ask whether any stage of the hiring process would have shown it. If the honest answer is "they would have described it well," the gap is a work sample that nobody ran.
Signal: they are doing a different job than the one you hired for
This one is easy to misread as attitude. The person is working hard and producing things nobody asked for; they optimise the part of the role they find interesting; their sense of what "good" looks like does not match the team's. Six weeks in, someone says the phrase that gives it away: "I thought this role was more about…"
Trace it back and you rarely find a candidate who misunderstood. You find a panel that never agreed. Three interviewers each described the job in their own words, each weighted the competencies differently in their own head, and the offer went to the candidate who scored highest against whichever version happened to be loudest at the debrief. Nothing was written down that could have caught the disagreement, because the role definition existed only as a job advert and four private mental models.
The upstream fix is a role blueprint agreed before sourcing: the competencies that matter, their relative weight, and what each looks like at the level being hired. It takes an hour. Its value shows at the debrief, where a disagreement about a candidate becomes a disagreement about a weighted competency rather than a contest of impressions — and again at week six, when the person knows what the job is because the process was clear about it.
Signal: the panel's doubt comes back
Someone in the process had a concern. It was raised, it was not resolved, and the hire happened anyway — often because the concern was expressed as a feeling and feelings lose to enthusiasm in a debrief. Then the person starts, the same concern reappears as a real problem, and the interviewer who raised it remembers.
Two different process defects produce this. In the first, the concern was never testable: "I'm not sure about their attention to detail" is not something a debrief can adjudicate, because there is no evidence to point at. In the second, the concern was testable and got overridden — a scored criterion where one evaluator was two levels below the rest, and the group averaged the gap away instead of examining it. Averaging is how a process quietly discards its most informative signal.
Both fixes are procedural rather than cultural. Calibrate evaluators on real sample responses before live scoring, so a level means the same thing to everyone. Score criterion by criterion with written rationale, so a doubt has to attach to evidence. And treat a two-level disagreement as a trigger for a short conversation about that evidence — it is telling you something about the rubric, the task or the candidate, and you cannot tell which once it has been averaged.
- Require written rationale per criterion, not an overall impression.
- Reveal all evaluator scores at once; discussion before independent scoring measures conformity.
- Examine any criterion where evaluators land two or more levels apart.
- Record the dissent alongside the decision so it is retrievable later.
Signal: the hire is surprised by the job
The quietest signal, and often the earliest: the person is not underperforming so much as recalibrating. The volume is higher than they pictured, the tooling is different, the autonomy is greater or smaller. Some of this is unavoidable. Much of it means the process never showed them the work, only talked about it.
An assessment that resembles the job does two jobs at once. It gives you evidence, and it gives the candidate an accurate preview — which is why a well-built work sample sometimes produces a self-withdrawal, and why that outcome is a success rather than a loss. A candidate who reads a realistic support queue exercise and concludes the pace is not for them has saved everyone a quarter.
This is also where candidate experience stops being a courtesy and becomes measurement quality. A task that is unclear, disproportionately long, or hostile to accommodations produces a poor sample of the person's ability and a poor preview of the role. Proportional design serves accuracy, not just goodwill.
Turning signals into a process change
One mis-hire is an anecdote and should be treated as one. Making the signals countable takes two unremarkable habits: record the decision and its evidence at the time it is made, and revisit it after a fixed interval with an outcome noted against it. Without the first, the review six months later is a memory contest. Without the second, nothing closes the loop.
Then watch for repetition rather than severity. Three hires from the same role who all stalled at the handover point is a statement about the assessment; three hires who each struggled differently is probably three ordinary situations. Where a pattern appears, change one thing in the process and watch the next cohort — a work sample added, a criterion reweighted, a calibration session reinstated — rather than redesigning everything at once and learning nothing about which change worked.
Two cautions. Small numbers are unstable: outcome data on a handful of hires supports a hypothesis, not a conclusion. And selection procedures carry legal obligations that differ by jurisdiction — the EEOC's guidance on employment tests and selection procedures orients you to the US framework, this is not legal advice, and no claim of compliance is being made here. The posture that survives scrutiny is the one that improves hiring anyway: document what was measured, why it was job-relevant, and what the evidence was.
- bad hires
- early warning signs
- selection process design
- work samples
- evaluator calibration
Key takeaways
- Treat an early warning sign as a defect report on the selection process; the individual case is one data point, the process runs again next month.
- Work that is correct but poorly prioritised means judgement was never elicited — the fix is a task with a trade-off, scored on reasoning.
- A hire doing a different job than you intended usually traces to a panel that never agreed the competency weights before scoring started.
- A doubt that resurfaces after the start date was either untestable or averaged away; both are fixed by calibrated, criterion-level scoring with written rationale.
- Count repetitions, not severity: change one thing in the process and watch the next cohort rather than redesigning everything at once.
Sources & further reading
- U.S. Equal Employment Opportunity Commission, Employment Tests and Selection Procedures
- Sackett, Zhang, Berry & Lievens (2022), Revisiting Meta-Analytic Estimates of Validity in Personnel Selection — Journal of Applied Psychology
- U.S. Office of Personnel Management, Assessment & Selection: Structured Interviews
