01Part 1
The current hiring system is failing most of the people in it.
The majority of hiring decisions, made with resumes, conversations and gut feel, produce outcomes that disappoint both the organisation and the person hired. The reasons new hires fail break down like this:
| Reason for failure | Share of cases |
|---|---|
| Coachability | 26% |
| Emotional intelligence | 23% |
| Motivation | 17% |
| Temperament | 15% |
| Technical skill | 11% |
Technical skill, what the resume shows and most interviews test, accounts for the smallest share of hiring failures. Leadership IQ
82%of hiring managers saw warning signs and hired anywayLeadership IQ, 2020 The reasons: time pressure, confidence in their ability to coach the candidate, or simply wanting the search to be over.
What this costs
31%US employee engagement in 2024, a ten-year lowGallup, 2024 60%of organizations reported longer time-to-hireGoodTime, 2024
02Part 2
Why standard hiring methods fail.
57%of the time a typical interview picks the right hirePsychological Bulletin Unstructured interviews, the most common format, have a predictive validity of about 0.20. A validity score is the correlation between an assessment method and actual job performance. Zero means no relationship. One means perfect prediction. At 0.20, an unstructured interview is marginally better than guessing. Sackett et al., 2022
Three documented mechanisms:
- First-impression lock-in. 60%of interviewers decide within the first 10 minutesCheckster, 2022 What happens in the remaining time is mostly confirmation: looking for evidence to support the initial judgment, discounting anything that contradicts it.
- Unconscious bias. 50%of hiring decisions are influenced by unconscious biasCheckster, 2022 Affinity bias, the halo effect, similarity bias. A 2024 peer-reviewed study confirmed that interviewers who pay attention to irrelevant applicant characteristics produce biased judgments that lead to hiring discrimination. Wingate, 2024 Audit studies found a 50%callback gap based on the name on the resumeBertrand & Mullainathan, 2004 and that racial hiring discrimination has not declined since 1989hiring discrimination unchanged sinceQuillian et al., 2017.
- Resume and experience filters don't work. 31 secondsmedian time a recruiter spends on a resumeLerner, 2024 The filters applied in that time, years of experience, GPA, school prestige, are consistently among the weakest predictors of performance. They reliably screen out certain candidates, just not the wrong ones. 73%of applicants are not qualified for the roleSparkHire, 2024
The resume problem
85 years of research show that resumes are essentially ineffective at predicting job performance.
"The school you went to and the places you've worked are poor predictors of performance."
Google's internal research confirmed this, and led to the removal of degree requirements across most roles.
03Part 3
What the research says actually works.
The landmark reference is Schmidt and Hunter (1998), which synthesised research from 1913 to 1998 across 85 years of published personnel selection data. It remains the most cited paper in industrial-organisational psychology. It was updated and partially revised by Sackett, Zhang, Berry and Lievens (2022).
| Method | Validity coefficient |
|---|---|
| Work sample + cognitive ability | 0.63 |
| Structured interview + cognitive ability | 0.63 |
| Cognitive ability alone | 0.51 |
| Structured interview alone | 0.44 |
| Integrity / conscientiousness test | 0.41 |
| Situational judgment test | 0.35 |
| Unstructured interview | 0.20 |
| Years of experience | 0.18 |
| GPA / school prestige | 0.10 |
Why situational judgment tests are the right foundation
A situational judgment test presents the candidate with realistic scenarios, situations they would actually encounter in the role, and evaluates how they respond. Research from multiple meta-analyses shows:
- They add validity over cognitive ability and personality tests alone. They capture information other methods miss.
- They are less prone to faking than self-report. Ask someone "are you coachable?" and most say yes. Put them inside a scenario where a manager gives critical feedback and you get a different kind of data.
- They measure judgment quality, not just knowledge. Knowledge-based ("what should you do") and behavioural-tendency ("what would you do") items are more predictive together than either alone.
- Delivery that immerses the candidate outperforms plain text. This is why Job Strategy AI is built for active scenarios rather than multiple-choice trivia.
- They are particularly effective for early-career and customer-facing roles where formal experience may be limited but judgment matters immediately.
Key limitation, acknowledged: situational judgment tests capture intended behaviour, not guaranteed actual behaviour. This is why Job Strategy AI uses them as part of a broader picture of signal, never as the only measure. Convergent signals across several scenarios and dimensions are more reliable than any single test.
Why conscientiousness matters more than almost any other trait
Of the Big Five, conscientiousness is the trait with meaningful, consistent predictive validity across roles and industries. The reason is structural: it predicts follow-through. Whether someone will execute when it's inconvenient, meet commitments, and hold standards under pressure. Coachability, the number one reason new hires fail, is a behavioural expression of it. It shows up in how someone handles feedback they disagree with, not how they describe their approach to feedback. PMC, 2019
04Part 4
Money motivation is not a red flag.
A multi-study paper in the Journal of Vocational Behavior (three independent studies across industries) found intrinsic motivation was associated with positive work outcomes (performance, commitment, less burnout, less turnover intention), while extrinsic motivation was negatively related or unrelated. Crucially, the two were only moderately negatively correlated: they trade off, but both can coexist. Gagné et al., JVB
The applied finding: once compensation exceeds a subsistence threshold, intrinsic factors become the stronger motivators. Pay still matters. A mismatch creates immediate dissatisfaction. Above a reasonable threshold, it stops being what drives performance.
What this means for assessment: asking "are you here for the money?" is a low-quality signal. What our assessments surface instead is whether the candidate engages with the actual content of the work, responds with genuine curiosity, and shows the specific care for the role that intrinsic drivers produce. These patterns show up in how someone explores a problem, not in what they say their motivation is.
05Part 5
What Job Strategy AI measures, and why.
Our assessments read candidates across five behavioural dimensions. Each maps to a documented predictor of job performance.
What do they do when they're told they're wrong?
Coachability
- Research basis
- The number one cause of new hire failure: 26% of failures come down to not being able to accept or apply feedback. It is the behavioural expression of conscientiousness.
- What we measure
- How a candidate responds when a scenario presents a correction, a challenge to their first approach, or feedback that they were wrong. Not whether they say they're open to feedback. What they actually do with it.
- The signal we look for
- Adjusting the approach, no defensive language, and taking in new information even when it means changing course.
How do they decide when there is no obvious answer?
Decision quality under ambiguity
- Research basis
- Cognitive ability is the single strongest individual predictor of job performance (0.51). Its power comes from processing new information and solving problems when there is no obvious answer.
- What we measure
- How candidates move through scenarios with missing information, competing priorities, or unclear success criteria. Most real decisions look like this.
- The signal we look for
- Structured thinking, assumptions stated out loud, priorities named, and honesty about what they don't know.
What happens when a person, not a problem, is the hard part?
Emotional signals and interpersonal judgment
- Research basis
- 23% of new hire failures are emotional intelligence issues. As a standalone measure it predicts at about 0.30 in customer-facing, collaborative and leadership roles, and rises sharply combined with other signals.
- What we measure
- How candidates respond to interpersonal tension inside scenarios: a colleague in conflict, a frustrated customer, a manager under pressure, a team that disagrees. Response patterns, not self-description.
- The signal we look for
- De-escalation over escalation, taking the other person's view, directness without harshness, and a response that fits the size of the situation.
Do they think the way this job actually requires?
Role-specific judgment
- Research basis
- Work sample tests are the most powerful predictor in the literature when combined with cognitive ability (0.63). Watching someone do the work beats hearing them talk about the work.
- What we measure
- Scenarios built from the actual thinking the target role demands. A product role tests prioritisation under constraint. A sales role tests objection handling. A support role tests escalation judgment.
- The signal we look for
- Pattern recognition specific to the role, not generic professional behaviour. Whether their natural orientation matches the real cognitive demands of the job.
Will they follow through when it's inconvenient?
Conscientiousness
- Research basis
- The second strongest individual predictor after cognitive ability. Organisation, dependability and persistence. About 0.41 as a standalone predictor when measured with integrity tests.
- What we measure
- Follow-through inside scenarios, handling of deadline pressure, response to competing demands, and consistency across several scenario touchpoints.
- The signal we look for
- Does the approach in scenario three match what they signalled in scenario one? Inconsistency under pressure is itself a signal.
06Part 6
How to read Job Strategy AI results.
Results represent patterns of judgment across several scenarios. Not a personality type, not a prediction of success, and not a ranking against other candidates. No assessment can guarantee a hire. What a well-designed situational assessment can do is raise the baseline probability of a good decision by giving you more relevant signal before the conversation.
The results surface
- Which dimensions showed a clear, consistent signal
- Which showed variance or tension across scenarios
- Specific scenario moments that produced notable responses
- A short narrative per dimension, not a score on its own
The results don't tell you
- Whether to hire this person. That's your judgment.
- How this person ranks against others on a number line
- Anything about their motivation to work for you specifically
For hiring teams
The results are most useful as a pre-conversation brief, not a post-interview validation. Review them before your first live conversation. Use them to know where to go deeper, where you already have confidence, and where a scenario created an ambiguous signal worth exploring. Most of the hiring managers who saw warning signs and ignored them did so because the signs came mid-interview, after a positive impression had formed. A pre-conversation brief separates the data from the relationship. Leadership IQ, 2020
Questions to bring into the interview when a read is mixed or thin:
- Coachability: Tell me about a time you received feedback you disagreed with. What did you do?
- Decision quality: Walk me through how you approached a decision where you did not have all the information.
- Emotional signals: How do you handle a situation where a colleague's approach conflicts with how you would do it?
- Role judgment: Take me through the last time you had to choose what not to do in this kind of role.
- Conscientiousness: Tell me about a commitment that became inconvenient to keep. What happened?
For candidates
Read your results honestly. The dimensions where you showed clear, strong signals are your strongest positioning in a role of this type. The dimensions where signals were mixed or thin are worth reflecting on, not as weaknesses, but as places where your natural approach may need adjusting for this kind of environment. A result that shows tension isn't a rejection. It's a conversation starter. Most good interviews go to the places where the answer isn't obvious.
07Part 7
Why this matters beyond hiring.
Most of this document is framed around organisational cost. But there's a person on the other side. Someone takes a job they're wrong for. They spend eighteen months struggling, failing quietly, being managed out, or leaving under pressure. They lose confidence and carry it into the next search.
The charismatic candidate who lands a role they weren't right for doesn't win long-term either. They just fail later, more visibly. And the talented candidate who got filtered out because they didn't interview well, and went to a competitor: that's the loss nobody measures.
Job Strategy AI exists because both outcomes are preventable when the evaluation method is built on what the research says about human performance, not on what the last fifty years of hiring tradition assumed.
Appendix
Sources.
Every number on the site links back to one of these rows.
| Source | What it establishes |
|---|---|
| Leadership IQ Leadership IQ, "Hiring for Attitude" study. 5,247 hiring managers, 312 organizations, 20,000+ new hires tracked over three years. Findings reported in Fortune and Forbes. | 46% of new hires fail within 18 months. 89% of those failures are about attitude, not skill. |
| Leadership IQ, 2020 Leadership IQ, 2020 HR Executive Survey. 1,400 respondents. | 82% of hiring managers saw early warning signs in a candidate and hired them anyway. |
| Schmidt & Hunter, 1998 Schmidt, F.L. & Hunter, J.E. (1998). "The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings." Psychological Bulletin, 124(2), 262-274. | The 85-year meta-analysis of selection method validity. The foundation of the predictive validity hierarchy. |
| Sackett et al., 2022 Sackett, P.R., Zhang, C., Berry, C.M. & Lievens, F. (2022). "Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range." Journal of Applied Psychology. | Updated validity estimates. Separates structured from unstructured interviews and situational judgment test subtypes. |
| Checkster, 2022 Checkster (2022). Reference Checking Report. Data on hiring bias and how quickly interviewers form opinions. | 60% of interviewers decide within the first 10 minutes. Nearly 50% of hiring decisions are influenced by unconscious bias. |
| CareerBuilder, 2024 CareerBuilder Annual Hiring Survey (2024). US employers on bad hire frequency and cost. | 75% of employers made a bad hire in the past year. Average loss $17,000, up to $240,000 for executive roles. |
| SHRM SHRM Human Capital Benchmarking Reports. | Cost per hire ($4,700), replacement cost (50% to 200% of annual salary), manager time spent on underperformers (17%). |
| US Dept. of Labor U.S. Department of Labor estimate of the cost of a bad hire. | A bad hire costs about 30% of the employee's first-year wages. |
| Gallup, 2024 Gallup (2024). State of the Global Workplace. | A disengaged employee costs $3,400 for every $10,000 in salary. US engagement sat at 31%, a ten-year low. |
| TestGorilla, 2024 TestGorilla, State of Skills-Based Hiring Report (2024). 2,000+ employers surveyed. | Retention, mis-hire, diversity and cost-to-hire improvements after moving to skills-based assessment. |
| GoodTime, 2024 GoodTime Annual Hiring Insights Report (2024). | Talent acquisition teams met their hiring goals 48% of the time. 60% reported longer time-to-hire. |
Lerner, 2024 Lerner, J. (2024). Median resume review time. | Recruiters spend a median of 31 seconds reviewing a resume. |
| SparkHire, 2024 SparkHire (2024). Applicant qualification data. | 73% of applicants are not qualified for the roles they apply to. |
| Gagné et al., JVB Gagné, M. et al. Journal of Vocational Behavior. Three independent cross-sectional and cross-lagged studies across industries on intrinsic and extrinsic motivation. | Intrinsic motivation predicts positive work outcomes. Extrinsic motivation is neutral to negative on its own. |
| McKinsey McKinsey & Company research on intrinsic motivation, cited via the Workstars 2020 research compilation. | Intrinsically motivated employees show 46% higher job satisfaction and 32% greater organizational commitment. |
| Christian et al. Christian, M.S., Edwards, B.D. & Bradley, J.C. Meta-analysis of situational judgment tests across delivery modes. | Situational judgment tests add validity beyond cognitive and personality tests. Video delivery outperforms text. |
| PMC, 2019 PMC / NCBI (2019). Situational judgment tests as measures of personality; conscientiousness research summary. | Conscientiousness is a valid predictor of performance in educational, work and life outcomes. |
| Psychological Bulletin Psychological Bulletin meta-analysis on employment interview accuracy. | A typical employment interview leads to the correct hire 57% of the time. |
| eSkill (meta-analyses) eSkill, citing meta-analyses on emotional intelligence as a selection measure. | Emotional intelligence has a standalone predictive validity of about 0.30 in customer-facing, collaborative and leadership roles. |
| Cogn-IQ.org Cogn-IQ.org, citing resume audit studies on school prestige. | Resumes with elite university names get about twice the callbacks with no corresponding performance advantage. |
| Laszlo Bock, Google Laszlo Bock, former SVP of People Operations at Google, on Google's internal hiring research. | "The school you went to and the places you've worked are poor predictors of performance." |
| Bertrand & Mullainathan, 2004 Bertrand, M. & Mullainathan, S. (2004). Resume audit study on name-based callback bias. | A 50% callback gap based on the name on the resume. |
| Quillian et al., 2017 Quillian, L. et al. (2017). Meta-analysis of field experiments on hiring discrimination. | Racial hiring discrimination has not declined since 1989. |
| Wingate, 2024 Wingate (2024). International Journal of Selection and Assessment. Bias in the employment interview: riskiest characteristics and information sources. | Interviewers who attend to irrelevant applicant characteristics produce biased judgments that lead to discrimination. |
Research brief version 1.0, September 2026. For platform users and partner distribution.