← SAKIZLI AI
Article30 Sept 2026 · 26 min read21 / 21Members · Subscription

Candidate Selection: The Human in the Process

Who actually reaches human review?

FFurkan SakızlıAI researcher & tutor · independent
On pale paper a row of glass cards with circles and half circles joined by a line; a larger card carries a golden ring, and a branch leads two cards downwards to an empty end point
Who stays on the main route and who branches off
Image generated with AI

Imagine a small organisation receiving many applications for a single vacancy. A new system sorts the documents by estimated suitability. In the interface, the HR role initially sees only the profiles with the highest scores. The other applications are not deleted, but they disappear from the view in which the next selection round actually takes place.

At the end, a person reads the shortlist, conducts interviews and signs off on the decision. Can we therefore say that a human made the selection? The answer depends on what “involved” means. Did the person see every application? Could they review and change the sorting? Did they know why a profile was not visible? Had the ranking itself already become so decisive that nobody would have considered an application outside the top group?

These questions lead to the decisive point in the process: the threshold at which a person or organisation can review only those applications that a system has made visible beforehand. The selection decision does not begin only when an offer is made. It also begins with deciding which documents are processed, which characteristics count, how they are ordered and who drops out of the shortlist.

This does not mean that every digital tool used in recruitment is impermissible. A good system could relieve teams, apply requirements more consistently and make initial screening more manageable. It means that “a human is still involved” is not, by itself, evidence of fairness, lawfulness or independent judgement. The real question is: which decision does the system influence, who can recognise and change that effect, and who is responsible for the consequences?

Diagram: on the left two profile cards, then a tall card with a large dot; then the route splits – a light card passes a blue vertical threshold to a review card and on to an end point, while paler cards drop below the threshold; a dashed golden arc leads from them back to the review card
The threshold decides who reaches human review – and a return route brings hidden profiles back
Image generated with AI

Selection is a chain, not a signature

An application rarely passes through just one selection moment. Even the job advertisement can be influenced by a delivery system: who sees the advert, to whom is it recommended, and which groups are particularly targeted? This is followed by receipt and formatting of documents, automatic checks of certain minimum requirements, sorting, ranking, summarisation, human review, interview, selection and feedback.

At each of these points, it can be decided whether a profile remains in the process. A system need not explicitly send anyone a rejection to shape the selection. It is enough for a low score to exclude an application from the next human review. A linguistically persuasive summary can also influence later review if the HR role no longer reads the original documents. A tool can therefore prepare a decision without formally completing it itself.

Anyone who asks only about the last click or signature overlooks the process that came before it. Conversely, asking “Is AI being used?” is not enough either. A tool may sort a file, check a required qualification, summarise text, assess candidates or generate a ranking. These functions have different consequences. What matters is the specific purpose, the point at which the tool intervenes and its actual effect in the process.

A selection chain can be described with six questions:

1. What task is to be performed? Distributing adverts, organising documents, checking minimum conditions, comparing performance or drawing up a shortlist?

2. What information is used for this? Required qualifications, work samples, CV data, public profiles or inferred characteristics?

3. What output does the system produce? A summary, a flag, an exclusion, a threshold, a rank or a recommendation?

4. What happens to lower-scoring profiles? Are they displayed, checked at random or on professional grounds, set aside, or never considered by a person?

5. What can the human role actually do? Does it have the time, information, expertise and authority to reject a recommendation or reopen the entire list?

6. How can an error be raised and corrected? Is there a clear route for accountability and complaints, and can an applicant add to the way their own qualifications are represented?

This does not make the process fair automatically. The questions do, however, prevent a single human step at the end from making the selection effects that came before it invisible.

The small company in the example

A wholly fictional small company is looking for someone to coordinate technical service orders. The work includes liaising with customers, keeping records and prioritising cases. Management is considering a digital tool that sorts profiles by a suitability score. The HR role is initially to see only the top group.

According to the internal job description, the role requires communication skills, careful record-keeping and the ability to pass technical information on to the right people. A particular university degree is not specified as mandatory. Even so, the ranking might give substantial weight to familiar job titles or formal qualifications. A synthetic profile has relevant experience from a non-linear career path and demonstrates the required skills through concrete tasks. If the system mainly expects a matching title or degree field, this profile could end up near the bottom.

The case does not claim that a real system works exactly this way or that every unusual ranking is discriminatory. It asks a question about the process: if the HR role sees only the top profiles, how can it tell whether the score reflects a genuinely necessary criterion or a proxy only loosely connected to the actual work?

Management has an understandable reason for considering assistance. A small team cannot spend unlimited time on initial screening. Consistently applied criteria could reduce the workload; a structured overview could make incomplete applications quicker to spot. The potential benefit belongs in the assessment. It does not prove, however, that the criteria are suitable, that all relevant skills remain visible, or that the HR role can reliably correct the system output.

Small organisations in particular should not imitate an elaborate audit they cannot carry out legally or professionally. Before purchasing, however, they can prepare a clear job analysis, ask about the system's function and supporting evidence, identify a manual alternative, and decide which role is allowed to open or stop a ranking. An unclear answer from the provider is itself important information: it shows which assumption has not yet been tested.

What does “fair” mean in selection?

Fairness is not a single metric. At least four questions sit alongside one another:

Relevance: Are skills considered that are genuinely needed for the role?

Equal treatment: Are comparable applications treated according to comparable rules?

Access: Can people with different educational and career paths show their relevant skills in the expected form at all?

Accountability: Can the organisation explain why one profile was reviewed further and another was not?

These demands can be in tension. A strictly standardised form can make comparison easier, but represent relevant experience from a different career path poorly. Individual human review can recognise an unusual qualification, but it takes time and can be inconsistent because expectations vary. An automatically generated ranking can process all documents in the same way, but consistent processing is not the same as a good criterion.

The philosophical problem is therefore not simply whether a machine is objective or a person subjective. A person can assess an application on the basis of an unclear gut feeling. An algorithm can reproducibly apply a rule defined in advance. The more important question is whether the reasons used are professionally relevant, disclosed and reviewable, and whether an affected person can correct an inaccurate representation.

Turning an application into a suitability score compresses a multidimensional person into a selection representation. The score can be useful for a limited comparison. But it is neither the person nor a complete description of what they could do in a particular job. A selection process does not owe applicants a job offer. It does owe them treatment that is understandable, work-related and non-discriminatory within the applicable rules. That obligation does not disappear when software sorts documents.

This also applies to the target on which a system was trained or tuned. If an organisation uses “successful former employees” as its model, it must ask who was included in that group and under what conditions success was measured. If the target is a later assessment by managers, that assessment may be shaped by the culture, support and promotion opportunities of the former workplace. A prediction may be statistically associated with a past measure without that measure being the only legitimate standard for future applicants.

A serious review therefore asks about both the system and the organisation: What work task is being described? Which skill is genuinely required? What forms may that skill take? Which applications become invisible? And who is prepared to revise a selection when its rationale does not hold up?

Human involvement is no guarantee

Two controlled studies illuminate different aspects of human involvement. A German online study simulated algorithmic pre-selection and compared a biased recommendation with an unbiased one. In the biased condition, approximately 60 per cent of participants failed to recognise the bias when asked directly. At the same time, participants overall edited algorithmic assessments more often in that condition. This is not simple proof that people always follow a recommendation: it shows instead that changing, recognising and correcting are different processes. Someone who changes a score has not necessarily understood the underlying bias; someone who notices a bias does not automatically correct it at every stage. The study used a simulated task and participants without professional recruitment roles. It does not measure real hiring decisions.[1]

An experiment by Wilson and co-authors studied 528 people working with AI recommendations for 16 occupational groups in a simulated CV screening. In certain conditions, participants selected the candidate groups favoured by the simulated AI up to 90 per cent of the time. A brief intervention using an implicit association test also influenced selection in the reported conditions. This figure describes a particular experimental set-up, not the proportion of real HR decision-makers in all companies who adopt AI recommendations. Its value here lies in the process question: when a recommendation is shown early, it can steer selection even if the decision officially remains with a human.[2]

An audit study by Gaebler and colleagues worked with 1,373 applications for teaching positions at a public school district in Texas; the full analysis included 801 applications with CVs and transcribed video interviews. The researchers had eleven language models produce standardised assessments and changed signals associated with gender and racial attribution in synthetic counterfactuals. The study reports assessment differences depending on the model and threshold. It also shows that simple group ratios alone can be inconclusive: an observed difference does not by itself explain why it arises. To the researchers' knowledge, the school district itself did not use AI for this pre-selection. The study measured model assessments, not hiring decisions. The data and legal context come from the United States and the education sector; direct transfer to German companies would not be justified.[3]

Taken together, the studies say neither “people do not help” nor “algorithms are always unfair”. They show that the effects of recommendations, interfaces, criteria and review conditions need to be investigated empirically. Human oversight is an organisational task: a role must see the original documents or sufficiently informative evidence, have time to review them, be allowed to reject the assessment and know how to recognise an inaccurate or systematic effect. Formal approval by a person does not replace these conditions.

The legal framework is specific, but not a blanket verdict

The legal position in this article is stated as at 30 September 2026. The relevant jurisdictions, roles and actual technical functions must be determined separately for any real selection process.

The AI Act lists certain systems for the recruitment and selection of natural persons as a high-risk area in Annex III, point 4(a). These include, in particular, systems for targeted job adverts, analysing or filtering applications, and assessing candidates. This does not amount to a blanket ban on AI in recruitment. Classification depends on the intended purpose and use. Article 6(3) provides limited exceptions for certain systems that perform only a narrow procedural or preparatory task, or improve the result of a human activity that has already been completed, where the statutory conditions are met and the system does not materially influence the decision. This exception is not a blanket exemption for deployers: the applicable conditions and the provider's documented classification must be reviewed by a suitably qualified person. The provision treats profiling of natural persons separately; such systems remain high-risk under this exception. A system that creates a ranked group and removes every profile below it from human review cannot therefore be dismissed without more as a mere sorting tool. A qualified review must consider the specific purpose and actual influence.[4]

Amending Regulation (EU) 2026/1744 sets the application date of Chapter III, Sections 1 to 3—with the exception of Article 6(5)—for high-risk systems under Article 6(2) and Annex III as 2 December 2027. It also concerns other stages of application, for example for certain products under Annex I. This does not mean that all AI use is unrestricted until then or that data protection and equal-treatment law are suspended. Nor is it a reason to postpone necessary procurement and process questions: an organisation must know which function it is using, which data are processed and who makes which decision.[5]

The General Data Protection Regulation must be considered for the processing of personal data in applications regardless of whether Article 22 applies. Article 22(1) concerns decisions based solely on automated processing that produce legal effects or similarly significantly affect a person. Paragraphs 2 and 3 set out exceptions and safeguards. The presence of a human in the process does not automatically rule out a solely automated decision if that contribution is merely formal and the assessment is not actually reviewed or changed. Conversely, not every display or sorting tool itself decides about a person. The actual process is decisive. Other GDPR duties may also apply, including those concerning purpose, legal basis, transparency, data minimisation and, where appropriate, a data protection impact assessment.[6]

In Case C‑634/21, the Court of Justice of the European Union ruled on a credit-scoring system and its possible determining role in a decision by a third party. The case does not concern recruitment. At most, it offers limited interpretive guidance that a score cannot be considered outside the decision-making process when other parties in practice treat it as decisive.[7]

Under Germany's General Equal Treatment Act (AGG), applicants fall within the personal scope of employment protection. The Act contains prohibitions on discrimination, requirements for job advertisements and rules on the burden of proof where facts give rise to a presumption of discrimination. Whether a specific selection decision falls within these provisions depends on the actual circumstances and legal assessment. A low score, a gap in a CV or an unclear system description does not, by itself, prove an AGG violation. But it would be equally wrong to dismiss possible discrimination by saying that the system “only sorted”.[8]

The useful practical conclusion is more modest and more precise: before real-world use, the organisation must assess the function, legal status, use of data, criteria, exceptions, transparency and responsibilities on a case-by-case basis. This article is not legal advice and does not classify a specific product.

F v5: turning a general question into a reviewable brief

The project-specific question and prompt optimisation model F v5 follows the sequence A → B → D → C. In this article, it does not answer a question about fairness. It changes the question so that a team can trigger the review it needs.

A – Clarify the initial question. Management initially asks: “Can an AI system make our selection more objective?” The question contains two undefined terms. What does “selection” mean here: initial screening, ranking or the final offer? And what does “more objective” mean: criteria applied identically, better prediction, less bias or reasons that can be understood? A brief GROW situational check helps establish the current position: the goal is a work-related and manageable selection; the reality is an interface that shows only one ranked group; the options include a structured human process and a combination of rule-based screening with support later in the process; the next step is to request details of the function and evidence before entering applicant data.

The refined guiding question is: “Which stage of selection does the proposed system influence for this role, which data and criteria shape the ranking, which profiles can the human role see, and can it independently bring an unseen application back into the review?” This makes it possible to examine where the decision actually arises.

B – Open up the context, hypotheses and counterarguments. The first plausible hypothesis is that a standardised tool can reduce repetitive screening work and apply requirements more consistently. A second is that if a ranking mainly rewards existing job titles, qualifications or similar past profiles, other evidence of the same required skill may be overlooked. A third concerns the interface: even a knowledgeable HR role might defer to a plausible-looking ranking if the reasons are not visible or there is no time for a counter-check. A fourth concerns the existing process: manual screening may also be inconsistent and perpetuate unfavourable routines.

The hypotheses are not equivalent to findings. The team needs information about the criteria, how the ranking is generated, the training or validation data, performance for the specific role, the treatment of non-linear career paths and the user interface. It must also identify a viable alternative, such as a pre-agreed matrix of criteria with human review instead of an opaque shortlist.

● Members only

Read the full article and download all files with a membership.

Unlock full article + downloads → Subscribe

0 comments

● Loading comments…

Sign in to comment · become a member →