LPL Logo

Build the Interview First

Twelve steps for the ideal case, and an eight-hour version for one person with a diary that is already full


The most confident people in the room

Almost everyone who conducts job interviews believes they are good at it. That is not a cheap observation, it is a measured one. Chapman and Zweig (2005), surveying 812 interviewees and 592 interviewers from over 502 organisations, found that fewer than 34% of interviewers in one sample and 28% in the other had received any formal interview training, and that interviewers were confident they could identify the best candidates regardless of how much structure was involved. Confidence was not related to method. It was simply present.

I find that finding uncomfortable rather than funny, because I recognise the instinct in myself. Sitting across from someone for forty minutes produces a very strong feeling of having learned something. The question this article is about is whether that feeling corresponds to anything, what can be done to make it correspond to something, and how much of that work is realistic for a person who has a role to fill next month and no research department.

This article is the practical half of that question. It gives you the ideal build first, in twelve steps, and then rebuilds the same thing for one person with eight hours. It is written to be read on its own: nothing below assumes any prior reading, and the terms it depends on are explained as they arrive.

Two things it deliberately does not do. It does not argue the evidence for each instruction, which is a long argument and has its own article. And it does not settle where a psychometric instrument belongs in a hiring process, which has another. Both are named at the end.

What you need to know before starting

Structuring an interview works, and it is not a small effect. A loosely run conversation and a carefully designed interview are not variations on one method. They behave like different methods with different accuracy.

The numbers everyone quotes changed in 2022, and they went down. For twenty-odd years the field ran on figures published by Schmidt and Hunter in 1998, which put both general mental ability tests and structured interviews at around .51. A re-analysis by Sackett, Zhang, Berry and Lievens (2022) argued that a statistical correction applied throughout that older work had been systematically overdone, and republished the estimates lower. Structured interviews came out at .42, which now places them at the top of the list, and cognitive ability tests came out at .31. If you see .51 in a vendor deck, you are looking at a figure the field has moved away from.

Adding a questionnaire to an interview helps only when the two measure different things. The reason a combined process beats either half is not that both halves are good. It is that they are weakly related to one another, so each one adds information the other missed. Where they overlap, you have paid twice for one predictor.

Five terms carry the rest.

A validity coefficient. When researchers say a selection method has a validity of .42, they mean a correlation. It runs from 0, meaning the method ranks candidates in an order unrelated to how they later perform, to 1, meaning it ranks them perfectly. In the revised estimates the strongest selection methods sit between about .3 and .42, and the weakest sit near zero. Nothing in employment gets close to 1.

A correlation is an abstraction, so here is what .42 means in practice. Suppose half the people who apply to you would turn out to be good hires if you took them, and you can afford to take one applicant in ten. Selecting at random, you would expect half your hires to work out. Selecting on a method with a validity of .42, roughly 79% of them would. That is the Taylor and Russell (1939) model, which converts a validity coefficient into a hit rate given how selective you can be and how good the applicant pool is. It also makes the limit visible: at the same base rate, hiring eight applicants out of ten, a validity of .42 lifts your hit rate from 50% to about 56%. If you are hiring most of the people who apply, interview design is not where your money should go.

Reliability, and interrater reliability in particular. Reliability is consistency. If two trained interviewers watch the same candidate and score independently, how much do their scores agree? That figure caps validity. A method whose raters disagree with each other cannot predict anything well, because most of what it is measuring is the rater.

Structure. In this literature, structure is not a synonym for “we use the same questions”. Campion, Palmer and Campion (1997) identified fifteen distinct components of structure and split them into two families. Content structure governs what information gets elicited: basing questions on an analysis of the job, asking the same questions of everyone, limiting improvised follow-ups, using better question types, using more questions, controlling what the interviewer sees beforehand, and holding the candidate’s own questions until the end. Evaluation structure governs how that information gets judged: rating each answer rather than forming an overall impression, using anchored rating scales, taking detailed notes, using multiple interviewers, using the same interviewers throughout, not discussing candidates between interviews, training interviewers, and combining the scores by formula rather than by discussion.

That list is the single most important idea in the field, because it explains why “we do structured interviews” tells you almost nothing. Levashina, Hartwell, Morgeson and Campion (2014), content-analysing 104 interviews reported in the research literature between 1997 and 2010, found that published studies implement about six of the fifteen on average (M = 5.74, SD = 2.83). That is a figure about what researchers have studied rather than what employers do, and the spread around it is wide. Everything below is a decision about which six, or which ten.

Incremental validity. The variance in performance a new measure explains over and above what you already had. A questionnaire can correlate respectably with performance and still add nothing, if everything it captures was already captured by the interview. This is the question worth asking of any instrument you are considering adding, and it is why Step 2 exists.

Adverse impact. A selection process has adverse impact when it selects people sharing a protected characteristic at a materially lower rate than others. In the United Kingdom this is the territory of indirect discrimination under section 19 of the Equality Act 2010, which turns on whether a criterion or practice that disadvantages a group can be shown to be a proportionate means of achieving a legitimate aim. It is not a US-only concern and it is not confined to cognitive tests.

Trait EI and ability EI. Two different constructs with a shared name. Ability emotional intelligence treats emotional understanding as a form of intelligence and is measured by tests with correct answers. Trait emotional intelligence, the construct underlying the TEIQue family, is a constellation of emotional self-perceptions located at the lower levels of personality hierarchies (Petrides, 2009). It is a personality construct, measured by self-report, and it is not claiming to be an intelligence. Figures from one do not transfer to the other.

Part One: The Ideal Build

This is the ideal build. Part Two of this article makes it affordable. Read this first anyway, because you cannot sensibly cut something you have not seen whole.

Twelve steps in four stages. The order matters more than it looks: several of these steps constrain the ones after them, and doing them out of sequence is how processes end up internally inconsistent.

Stage A: Foundations

Step 1. Establish what the job requires

Produce a description of the work before producing questions about it. Two routes are standard.

The critical incident technique, from Flanagan (1954), asks people who know the job to describe specific occasions when someone did the work notably well or notably badly, with enough detail that the consequences are clear. Flanagan’s own definition of a critical incident is “any observable human activity that is sufficiently complete in itself to permit inferences and predictions to be made about the person performing the act”, qualifying as critical where “the purpose or intent of the act seems fairly clear to the observer and where its consequences are sufficiently definite to leave little doubt concerning its effects.”

Competency modelling works the other way, starting from the capabilities the role demands and working down to observable behaviours.

Three practical rules regardless of route.

Collect ratings from each expert separately and combine them arithmetically. Do not run the group to a consensus in the room. A workshop that ends in agreement may have produced consensus, or it may have produced deference to the most senior person present, and from outside those look identical.

Write task-level statements rather than broad activity statements, and have them rated on frequency and importance.

Use trained analysts as well as job holders where you can afford to. They tend to agree with each other about which capabilities matter and to disagree about how much, which means the choice of source matters more for setting a standard than for setting a priority.

Output: four to six competencies, each with two or three concrete behavioural indicators drawn from real incidents.

Step 2. Allocate each competency to exactly one method

Write a two-column table. Left column, the competencies from Step 1. Right column, the single method that will assess each one.

This is the step nobody writes down and it is the one that decides whether the rest of the process adds up. If two methods both assess the same competency, one of them is redundant, and you should know which before you build it.

As a rule of thumb, technical knowledge and craft belong in a work sample or a knowledge test, reasoning belongs in a cognitive measure if the role genuinely demands it, and applied behaviour under real conditions belongs in the interview. A personality-style questionnaire belongs in the right-hand column only if you can say what it is measuring that nothing else in the process measures.

Also write down, explicitly, what you are not assessing. Every process has gaps. Undocumented gaps get filled by impressions.

Output: a competency-to-method map, and a short list of what the process does not cover.

Step 3. Choose the architecture and check the arithmetic

Two architectures exist. Compensatory means every candidate completes everything and a combined score decides, so strength in one area can offset weakness in another. Multiple hurdle means each stage screens, and only survivors continue.

Default to compensatory. Use hurdles only where volume genuinely forces it, and when you do, write down the proportion of candidates passing each stage, because the overall selectivity of your process is those proportions multiplied together. A tight first stage leaves the later, better assessment with very little influence over who is actually hired.

Then check the economics before building anything. Two numbers you do not control determine what any of this is worth: how selective you can be, and how good the applicant pool is. If you are hiring most of the people who apply, no amount of interview design will earn its build cost, and the money belongs in attraction instead.

Output: a written process map with the pass rate expected at each stage.

Stage B: The instrument

Step 4. Write the questions

Choose the format by role complexity. Situational questions pose a hypothetical dilemma drawn from the job analysis and ask what the candidate would do. Past-behaviour questions ask what the candidate did do in a comparable real situation. For senior and complex roles, use past-behaviour questions. That is a validity finding: situational interview validity falls for high-complexity jobs while past-behaviour validity holds across complexity levels (Huffcutt, Conway, Roth and Klehe, 2004). For lower-complexity roles either format works, and in my experience situational questions are easier to key.

Write two questions per competency, both derived from the incidents collected in Step 1.

Write your follow-up probes in advance, two per question, and use them identically for every candidate and for every question type. Improvised probing is the most common way a structured interview quietly stops being one.

Do not use brainteasers. A brainteaser is a puzzle or estimation question with no connection to the work, asked to see how someone thinks under pressure. “How many windows are there in New York?” is the standard example. It is not built from a job analysis and it cannot be scored against one, which is reason enough not to use it.

Do not let candidates ask their questions until the assessed portion is over, and do not read the CV immediately before scoring. Both are content components of structure for the same reason: they introduce information that contaminates the comparison between candidates.

Output: eight to twelve questions, each tagged to one competency, each with pre-written probes.

Step 5. Build the anchored rating scale

For every question, write what a weak answer, an adequate answer and a strong answer actually contain, in the language of the job. Not “shows good judgement”, but the specific things a good answer would include.

The classical method for this is retranslation: one group of experts generates behavioural examples, a second group independently sorts each example into the competency it best fits, and only examples that a clear majority place in the same category survive. It is expensive and it is the reason well-built scales work.

Use five points, and write anchors for at least points 1, 3 and 5.

Rate each answer separately. Do not collect one overall impression per interviewer. Step 6 explains why this instruction is conditional on the one that follows it.

Output: a scoring guide with an anchored scale for every question.

Step 6. Fix the scoring rule before the first interview

This is the step with the largest effect on accuracy and the one most often left informal.

Every interviewer scores every answer independently, on paper or in a form, before any discussion takes place.

Scores are combined arithmetically, using equal weights for each competency unless you have a local validation study that justifies otherwise. Equal weighting is not a compromise, it is a defensible choice that holds up well against the alternatives.

The combined score is produced before the debrief meeting, and the debrief is used to check whether the process was followed, not to revise the number.

If you want interviewers to record an overall gut impression, and there are good reasons to let them, treat it as one more input into the formula rather than as a veto over the formula’s output.

If anyone overrides the ranking, require the override to be written down with its reason, and review those overrides later against what actually happened.

Output: a score sheet with the arithmetic already built in, and a written decision rule.

Step 7. Decide who interviews, and train them

Use the same interviewers across all candidates for a role. This is worth more than adding people.

Panels are good for accountability and documentation, and they are not a validity intervention. Panel members rate one single performance, so they agree with each other far more than independent interviewers do, .77 against .53 in Conway, Jako and Goodman (1995), and that agreement is not evidence of accuracy. Meta-analytic validity for panel interviews is no higher than for individual interviews. If you use a panel, have every member score independently before anyone speaks.

Train interviewers on the scale, not on interviewing. Frame-of-reference training, which shows raters examples of performance at each anchor level and calibrates them against a target score, has the best meta-analytic evidence behind it (Roch, Woehr, Mishra and Kieszczynska, 2012). That evidence concerns the accuracy of ratings rather than the validity of interviews.

If you require notes, require behavioural notes, meaning a record of what the candidate said and did. Give people a template that makes anything else awkward. Unstructured general note-taking is not a neutral addition.

Output: a named interviewer set, a two-hour calibration session, and a note template.

Stage C: Data from instruments

Step 8. Decide the role each psychometric instrument plays

Any instrument you add can occupy one of three roles, and each carries a different evidence requirement. Choose one per instrument, in writing.

Role one, a source of interview content. The instrument is administered, a qualified person interprets the report, and the report generates lines of enquiry that the interviewer explores through behavioural questions scored on the anchored scale. Nothing is decided on the score. The score is a hypothesis and the interview is the evidence. The evidence bar here is reliability and interpretability, because the instrument is not predicting anything, it is directing attention.

Role two, a scored input to the decision. The instrument’s score enters the formula from Step 6 and moves the ranking. The evidence bar here is high and specific: criterion evidence from a predictive design against a defined performance measure in a comparable population, evidence about how the instrument behaves when the person answering it wants the job, and monitoring of subgroup differences in your own data.

Role three, post-decision. Onboarding, development, coaching, team formation, self-awareness. The evidence bar is interpretive value or evidence of change.

For a trait emotional intelligence measure specifically, the defensible placement on current evidence is role one and role three. The evidence for that placement, and the finding that changed my own view of it, is set out in The Interview Drifts: The Instrument Does Not.

One boundary worth setting in writing. Whatever instrument you use and whichever role you assign it, the ranking should be driven by scores you can trace back to observed behaviour against a written standard. An instrument’s output belongs in role one or role three unless it has cleared the role two bar in your own setting.

And one thing that cuts across all three roles. Administer it identically to everyone and keep the results, whichever role you have assigned it. A standardised measure given consistently is stable over years. Global trait EI on the TEIQue correlates between .62 and .77 with itself at retest intervals running from one month to four years (Zadorozhny, Petrides, Jongerling, Cuppello and van der Linden, 2024). That stability makes it useful for two purposes that have nothing to do with any individual decision: it gives you an external reference against which to notice that your interviewing has drifted, and it accumulates into the dataset that makes a real validation possible in three or four years. Both are free if you are administering the instrument anyway, and neither can be reconstructed later. Treat the first of them as control-chart logic rather than as a tested finding, because no study has examined it, and note its two assumptions: the stability of scores is not the same as the stability of the underlying trait, and the comparison holds only if your applicant pool is not itself shifting.

Output: an instrument-to-role table, and a decision to keep the data.

Step 9. Handle response distortion structurally

If any self-report instrument is in the process, assume that candidates who want the job will present themselves more favourably than they would in a development context, and design so that this does not matter much.

Warn candidates that responses may be checked or discussed. The effect is modest but real, a mean difference of about a fifth of a standard deviation in the literature Dwight and Donovan (2003) reviewed, and it costs nothing.

Do not attempt to correct scores statistically for social desirability. Corrected scores do not approximate honest ones (Ellingson, Sackett and Hough, 1999), and correcting or removing suspected fakers moves the mean performance of those hired by roughly a tenth of a standard deviation (Schmitt and Oswald, 2006).

Have a written retest policy with a waiting period and a rule about which score counts, before you need one.

Where an instrument is used in role one, inflation works against the candidate rather than for them, because a claim made on a questionnaire becomes something an interviewer probes for evidence of.

Output: a candidate-facing statement, and a retest rule.

Stage D: Operating it

Step 10. Build the legal and access requirements in from the start

This is a summary of the terrain, not legal advice, and it is UK-first.

Reasonable adjustments before the first candidate, not on request at the last minute. A time limit is a practice that can put a disabled candidate at a substantial disadvantage. Extra time, alternative formats and assistive technology are adjustments to consider, and the employer cannot require the candidate to pay the cost of the adjustment.

No health questions before an offer, except in the narrow circumstances the Equality Act permits, one of which is establishing what adjustments a candidate needs in order to take an assessment. That channel exists for adjustments. It is not a route to health information that then informs the decision.

Proportionality is the test that actually bites. A defensible aim does not rescue a disproportionate method, particularly where a workable alternative was available and had been asked for.

If any stage sifts candidates without meaningful human involvement, build in the safeguards. Section 80 of the Data (Use and Access) Act 2025 replaced Article 22 of the UK GDPR with new Articles 22A to 22D, and those provisions took effect on 5 February 2026 (SI 2026/82, reg. 2(j)), with decisions taken before that date governed by the previous position. For ordinary personal data the position is now permission plus mandatory safeguards rather than a general prohibition. The safeguards, in Article 22C, are information about the decisions taken and routes to make representations, to obtain human intervention and to contest the decision. Telling candidates in advance that a stage is automated is a separate duty, arising from the transparency provisions in Articles 13 and 14. Where special category data such as health information is involved, restrictions still apply, and a solely automated significant decision is also barred where the processing rests on the recognised legitimate interests basis.

If you have a human review some candidates’ scores, have them review everyone’s. Inconsistent human review is worse than none, and it is simultaneously a measurement problem, a fairness problem and a data protection problem.

Output: an adjustments process, a candidate privacy notice, and a note of which stages are automated.

Step 11. Choose and document your validation route

Four routes exist and you should be able to name which one you are on.

Criterion-related validation means correlating your scores against later performance in your own organisation. It is the strongest route and it usually requires more hires than a single team has, which is a matter of statistical power rather than diligence: detecting a realistic observed correlation of .20 at 80% power takes roughly 194 cases, and 50 cases buys you about 28%. The Thirty-Point Gap sets out that arithmetic.

Content-oriented evidence means demonstrating that the interview samples the work directly. This is a legitimate route for a structured interview built from a job analysis. It is not a legitimate route for a personality measure, because a trait is not a sample of the work.

Validity generalisation means relying on meta-analytic evidence for the method, conditional on the job being comparable to the jobs in that evidence.

Transportability means relying on a validation study conducted elsewhere, conditional on documented job similarity.

If you have done none of these, document that too. A process with a documented job analysis, documented question development, anchored scales and a written scoring rule is defensible as a well-built process. It is the unearned claim of criterion validity that fails, not the process.

Two practical notes. If you use manager ratings as your performance measure, collect them for the research rather than pulling them out of the appraisal system, because appraisal ratings are made under different incentives and are measurably less consistent between raters: interrater reliability for overall performance is .45 for ratings collected administratively and .61 for ratings collected for research (Salgado and Moscoso, 2019). And do not lead a business case with a monetary utility estimate. In two experiments with experienced managers, presenting a utility analysis reduced support for adopting a valid selection procedure rather than increasing it, and the effect held when the analysis was delivered by a recognised authority on the technique (Latham and Whyte, 1994; Whyte and Latham, 1997).

Output: a one-page validation statement.

Step 12. Audit the artefacts

Structure is not self-sustaining. Interviewers who want discretion over their questions, and who want to build interviews efficiently, use less of it, while those who have been through interviewing training use more (Lievens and De Paepe, 2004). Nothing in the process announces when that has happened.

You cannot survey your way to knowing whether your process is intact. In the study that tested this directly, managers’ stated intentions to interview in a structured way were unrelated to how they actually interviewed, while their intentions to interview unstructured tracked their practice (van der Zee, Bakker and Bakker, 2002). Look at the artefacts instead.

Once a year, and after any hiring surge, check: whether the questions actually used matched the approved set; whether probes were the written ones; whether scores were recorded before any discussion; whether interviewers saw CVs or test scores before scoring; whether notes are behavioural; every recorded override and what happened to that hire; subgroup pass rates at each stage; and every adjustment request and how it was handled.

And build in the concession that protects the rest. Ring-fence a period of unstructured, unscored conversation after the assessed portion. It gives interviewers the rapport they want, it gives candidates a better experience, and it keeps it out of the scoring.

Output: an annual audit against the artefacts, and a process that people will actually keep using.

Part Two: Eight Hours, One Person

Everything above describes a process built by specialists. Very few people reading this have a panel of subject matter experts, a validation sample or a research budget. Most have a role to fill, a diary that is already full, and a genuine wish to do this properly.

So this part answers a different question. If you have eight hours of preparation and you are the only person building it and probably the only person conducting it, what do you build?

The answer is not a diluted version of the ideal build. Some components are cheap and carry most of the benefit. Others are expensive and carry benefit you cannot capture alone anyway. The eight-hour build keeps the first group, drops the second, and is honest about the difference.

What survives, and why. The components that survive are the ones whose benefit does not depend on having many people: writing questions from real incidents, writing anchored scales, scoring each answer, and combining arithmetically. Those four are almost free once you know to do them, and between them they cover the content and evaluation families.

What is dropped, and why that is survivable. The subject matter expert panel, retranslation with a second independent group, panel scoring, interrater calibration and local validation are all dropped. Every one of them is dropped because it requires people you do not have, and dropping them has a specific, statable consequence: you will not know your interrater reliability and you cannot claim criterion validity. You can still say the process was built from the work, applied identically and scored by rule, which is the honest and defensible claim.

The eight-hour plan

Block 1 (45 minutes): write down what good looks like at twelve months

Not a job description. Write, in plain sentences, the five or six things this person must be able to do well by the end of their first year for the hire to have been worth making. Then, for each, write one real occasion you can remember where someone did that thing notably well, and one where someone did it badly. If you cannot recall a real occasion, the item is probably an aspiration rather than a requirement, and you should cut it.

That is your critical incident collection, compressed. It is far weaker than a proper job analysis. It is enormously better than starting from a job advert.

Use an occupational database if you want a prompt for things you have forgotten. Do not let it supply the answer.

Output: four to six competencies, each with two remembered incidents.

Block 2 (30 minutes): decide what the interview will and will not carry

Take your four to six competencies and mark each one as: the interview carries this, a work sample carries this, or nothing in this process carries this.

Be ruthless. Four competencies assessed properly beat eight assessed vaguely. If the role has a technical core, a small work sample carries it better than any interview question, and it is the highest-return hour in this whole plan.

Write the “nothing carries this” list down and keep it. Those are the gaps you are accepting knowingly, which is the only kind worth having.

Output: a one-page allocation, including what you are not assessing.

Block 3 (90 minutes): write eight questions

Two per competency, both derived from the incidents in Block 1.

For a senior or complex role, write past-behaviour questions: describe a specific occasion when you had to do this, what the situation was, what you did, and what happened. For a junior or lower-complexity role, situational questions are easier to write and easier to key.

Write two follow-up probes under each question, and commit to using only those. This is the cheapest structural improvement available to a single interviewer, because improvised probing is how a solo interview quietly turns into a conversation.

No brainteasers. No questions about hobbies as a proxy for character. No questions you would not be willing to ask every candidate.

Output: eight questions with sixteen pre-written probes, on one page.

Block 4 (90 minutes): write the anchors

This is the highest-value block in the plan and the one people skip.

For each question, write in one or two sentences what a 1, a 3 and a 5 answer actually contains. In the language of the job, not in the language of competency frameworks. “Names a specific trade-off they made and what it cost” is an anchor. “Demonstrates commercial awareness” is not.

Writing anchors does something that reading about anchors does not: it forces you to decide what you are looking for before you meet anyone who might change your mind. Half the value arrives during the writing.

Output: a scoring guide with three anchors per question.

Block 5 (30 minutes): build the score sheet and fix the rule

Build a single sheet: eight rows, one per question, a 1 to 5 box for each, a space for behavioural notes beside each, and a total at the bottom.

Then write the decision rule on the sheet itself, in one line, and hold yourself to it. For example: score every answer immediately after the candidate leaves and before doing anything else; sum the scores; do not revisit an earlier candidate’s sheet.

If you have a second person available for even one stage, have them score independently from the same sheet and average. Do not discuss before scoring. Two independent scores averaged is worth more than an hour of debate.

Output: one score sheet, with its own rule printed on it.

Block 6 (30 minutes): decide the psychometric element and write the probes

If you are using a trait EI profile or similar, this block decides what it is for. On the evidence set out in The Interview Drifts: The Instrument Does Not, there are two defensible answers for a single hiring manager, and they are not mutually exclusive.

The first is as interview content. The report is read by someone qualified to interpret it, the interpreter identifies two or three areas worth exploring, and those become probes attached to your existing questions rather than new questions of their own. A low self-reported score on emotion regulation does not become a mark against the candidate. It becomes “tell me about a time when a project went wrong late in the day”, scored on the same anchored scale as everything else.

Write the three probes now, before you see any candidate’s report, so that you are not constructing the enquiry around the person.

If nobody available is qualified to interpret the report, do not use it this way. An uninterpreted profile in the hands of an untrained reader is worse than no profile, because it produces confident conclusions from an instrument that does not support them.

The second answer is easier and is available to everyone, including anyone without an interpreter. Administer the instrument identically to every candidate, keep the results, and do not use them in the decision at all.

That sounds like doing nothing, and it is the highest-return thirty minutes in this whole plan on a three-year view. A standardised measure administered consistently holds its rank order over years: on the TEIQue, global trait EI correlates between .62 and .77 with itself at retest intervals from one month to four years. So the series gives you two things you cannot get any other way. It is an external reference against which to notice that your own interviewing has drifted, because if a stable measure holds steady while your interview scores move, the interviewing changed rather than the candidates, provided your applicant pool has not shifted underneath you. No study has tested that use, so treat it as a prompt to go and audit rather than as proof of anything. And it is the beginning of the only dataset that will ever let you check whether your process works: in three years you either have consistent scores paired with recorded outcomes, or you have nothing to correlate. You cannot go back and collect it.

You do not need to interpret any of it today. You need to collect it identically and store it lawfully, which is a decision, not a skill.

Output: either three profile-derived probes attached to existing questions, or a decision to administer and store consistently, or both.

Block 7 (45 minutes): the work sample, if the role has one

If Block 2 assigned anything to a work sample, build it now. Keep it to sixty to ninety minutes of candidate time, make it a genuine slice of the actual work, and write its scoring anchors the same way you wrote the interview’s.

A small, honest, well-scored work sample is probably the single most valuable component available to a small organisation, and work samples and interviews are the two methods candidates rate most favourably, ahead of every test-based method (Hausknecht, Day and Thomas, 2004; Anderson, Salgado and Hulsheger, 2010).

Output: one task, one scoring guide, one time limit.

Block 8 (30 minutes): access, law and the candidate-facing note

Write four things down.

The adjustments line, sent with every invitation: what the process involves, and an open invitation to tell you what adjustments are needed. Offer this to everyone rather than waiting to be asked, and make clear that you do not expect the candidate to pay for it.

The health question boundary: you ask what adjustments someone needs to take part. You do not ask about health, and you do not follow up on health information volunteered.

What happens to the data: what you collect, how long you keep it, who sees it, and, if any stage is automated, that it is, with a route to a human.

Your timings, if anything is timed, and whether they are flexible. A time limit that has no real job justification should simply be removed. That takes five minutes and removes the most common proportionality problem.

Output: a short candidate-facing note you can reuse for every role.

Block 9 (45 minutes): pilot on one person

Ask a colleague or a friendly contact to sit through the questions. You are not testing them. You are testing whether the questions produce answers you can score.

Two things always surface. At least one question will be ambiguous and will need rewriting. And at least one anchor will turn out to be unusable because real answers do not divide the way you imagined.

Fix those two things. Do not fix everything.

Output: a revised question set that has survived contact with a human.

Block 10 (15 minutes): write the brief for yourself

One page. The order of the interview, the time per question, the reminder to score before doing anything else, the reminder not to read the CV immediately beforehand, and the reminder that the friendly unstructured chat happens at the end and is not scored.

You will need this more than you think, because on the day you will be tired and the candidate will be charming.

Output: a one-page run sheet.

Total: 7 hours 30 minutes, with thirty minutes of slack, which you will use.

What to do on the day

Run every candidate through the same questions in the same order with the same probes. Score immediately afterwards, before your next meeting and before speaking to anyone about the candidate. Sum, do not weigh. Keep the sheets.

Save the conversation about salary, the team and the company for the unscored period at the end, and let it be genuinely relaxed. It is good for the candidate, it satisfies your own need for rapport, and it stays out of the score.

If two candidates are within a point or two of each other, treat them as tied rather than ranked. Your process is not precise enough to separate them, and pretending otherwise is where bias re-enters. Break genuine ties on something you can articulate and record, ideally a second work sample or a second independent scorer.

What you have, and what you do not

You have a process built from the actual work, applied identically to everyone, scored against written standards, combined by rule, with a documented approach to adjustments and data. On the evidence, that is most of the distance between a conversation and a selection procedure, and it took a working day.

You do not have a known interrater reliability, a criterion validation, or a defensible standard for what a given total score means in absolute terms. Say so if anyone asks. It is a considerably better position than the alternative, which is a confident claim resting on nothing.

And one honest caution about your own scores. A single interviewer scoring against self-written anchors has no external check. The mitigations, in order of value, are: a second independent scorer for at least the final shortlist; recording the interview with consent so a second person can score later; and, failing both, scoring immediately and never revising an earlier candidate after meeting a later one.

Where the rest of the argument is

This article is the instructions. Two things it leaves out have their own pieces.

Why each instruction takes the form it does, and where the evidence behind it runs out, is in The Thirty-Point Gap. Structured interviews top the validity table at .42, and they carry the widest credibility interval of any predictor near the top of it, .18 to .66 (Sackett et al., 2023). That article is about what separates the two ends: job analysis reliability, the combination rule, question format, rating scales, panels, the arithmetic of validation, and why structure is not self-sustaining.

Where a psychometric instrument belongs, and what happens to a self-report when the person answering wants the job, is in The Interview Drifts: The Instrument Does Not. It also covers what the law now requires of an automated or partly automated process, which changed on 5 February 2026.

Build the interview first. It is the component that carries the decision and nothing substitutes for it.


References

  1. Anderson, N., Salgado, J. F., & Hulsheger, U. R. (2010). Applicant reactions in selection: Comprehensive meta-analysis into reaction generalization versus situational specificity. International Journal of Selection and Assessment, 18(3), 291-304. https://doi.org/10.1111/j.1468-2389.2010.00512.x
  2. Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655-702. https://doi.org/10.1111/j.1744-6570.1997.tb00709.x
  3. Chapman, D. S., & Zweig, D. I. (2005). Developing a nomological network for interview structure: Antecedents and consequences of the structured selection interview. Personnel Psychology, 58(3), 673-702. https://doi.org/10.1111/j.1744-6570.2005.00516.x
  4. Conway, J. M., Jako, R. A., & Goodman, D. F. (1995). A meta-analysis of interrater and internal consistency reliability of selection interviews. Journal of Applied Psychology, 80(5), 565-579. https://doi.org/10.1037/0021-9010.80.5.565
  5. Data (Use and Access) Act 2025, c. 18, s. 80 and Sch. 6. https://www.legislation.gov.uk/ukpga/2025/18/section/80
  6. Data (Use and Access) Act 2025 (Commencement No. 6 and Transitional and Saving Provisions) Regulations 2026, SI 2026/82. https://www.legislation.gov.uk/uksi/2026/82/made
  7. Dwight, S. A., & Donovan, J. J. (2003). Do warnings not to fake reduce faking? Human Performance, 16(1), 1-23. https://doi.org/10.1207/S15327043HUP1601_1
  8. Ellingson, J. E., Sackett, P. R., & Hough, L. M. (1999). Social desirability corrections in personality measurement: Issues of applicant comparison and construct validity. Journal of Applied Psychology, 84(2), 155-166. https://doi.org/10.1037/0021-9010.84.2.155
  9. Equality Act 2010, c. 15, ss. 19, 20, 21, 39, 60. https://www.legislation.gov.uk/ukpga/2010/15
  10. Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51(4), 327-358. https://doi.org/10.1037/h0061470
  11. Hausknecht, J. P., Day, D. V., & Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57(3), 639-683. https://doi.org/10.1111/j.1744-6570.2004.00003.x
  12. Huffcutt, A. I., Conway, J. M., Roth, P. L., & Klehe, U.-C. (2004). The impact of job complexity and study design on situational and behavior description interview validity. International Journal of Selection and Assessment, 12(3), 262-273. https://doi.org/10.1111/j.0965-075X.2004.280_1.x
  13. Latham, G. P., & Whyte, G. (1994). The futility of utility analysis. Personnel Psychology, 47(1), 31-46. https://doi.org/10.1111/j.1744-6570.1994.tb02408.x
  14. Levashina, J., Hartwell, C. J., Morgeson, F. P., & Campion, M. A. (2014). The structured employment interview: Narrative and quantitative review of the research literature. Personnel Psychology, 67(1), 241-293. https://doi.org/10.1111/peps.12052
  15. Lievens, F., & De Paepe, A. (2004). An empirical investigation of interviewer-related factors that discourage the use of high structure interviews. Journal of Organizational Behavior, 25(1), 29-46. https://doi.org/10.1002/job.246
  16. Petrides, K. V. (2009). Technical manual for the Trait Emotional Intelligence Questionnaires (TEIQue). London Psychometric Laboratory.
  17. Roch, S. G., Woehr, D. J., Mishra, V., & Kieszczynska, U. (2012). Rater training revisited: An updated meta-analytic review of frame-of-reference training. Journal of Occupational and Organizational Psychology, 85(2), 370-395. https://doi.org/10.1111/j.2044-8325.2011.02045.x
  18. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. https://doi.org/10.1037/apl0000994
  19. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283-300. https://doi.org/10.1017/iop.2023.24
  20. Salgado, J. F., & Moscoso, S. (2019). Meta-analysis of interrater reliability of supervisory performance ratings: Effects of appraisal purpose, scale type, and range restriction. Frontiers in Psychology, 10, 2281. https://doi.org/10.3389/fpsyg.2019.02281
  21. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262-274. https://doi.org/10.1037/0033-2909.124.2.262
  22. Schmitt, N., & Oswald, F. L. (2006). The impact of corrections for faking on the validity of noncognitive measures in selection settings. Journal of Applied Psychology, 91(3), 613-621. https://doi.org/10.1037/0021-9010.91.3.613
  23. Taylor, H. C., & Russell, J. T. (1939). The relationship of validity coefficients to the practical effectiveness of tests in selection: Discussion and tables. Journal of Applied Psychology, 23(5), 565-578. https://doi.org/10.1037/h0057079
  24. van der Zee, K. I., Bakker, A. B., & Bakker, P. (2002). Why are structured interviews so rarely used in personnel selection? Journal of Applied Psychology, 87(1), 176-184. https://doi.org/10.1037/0021-9010.87.1.176
  25. Whyte, G., & Latham, G. P. (1997). The futility of utility analysis revisited: When even an expert fails. Personnel Psychology, 50(3), 601-610. https://doi.org/10.1111/j.1744-6570.1997.tb00705.x
  26. Zadorozhny, B. S., Petrides, K. V., Jongerling, J., Cuppello, S., & van der Linden, D. (2024). Assessing the temporal stability of a measure of trait emotional intelligence: Systematic review and empirical analysis. Personality and Individual Differences, 217, 112467. https://doi.org/10.1016/j.paid.2023.112467

Note: this article is a summary of the terrain rather than legal advice. Statutory positions are stated as at 31 August 2026.

Shopping Cart

Your cart is empty