← SAKIZLI AI
Article25 Sept 2026 · 28 min read7 / 7Members · Subscription

Judging Under Time Pressure

What Meaningful Human Oversight Needs

FFurkan SakızlıAI researcher & tutor · independent
Five frosted glass tiles with blue circle and dot symbols on white paper, linked by a blue line; between the second and the fourth tile the line splits into a direct path and a detour via a higher tile
The direct route and the detour through review
Image generated with AI

Someone clicks “Confirm.” On paper, the AI suggestion has been reviewed. But did that person see the original request? Was there time to understand the recommendation? Could they disagree, pause the case, or send it to someone with more context? What happens when the queue grows faster than the review can be completed?

These are questions about the substance of human oversight. An organisation can require a review and still design a workflow in which nobody can meaningfully review. The click is visible; the attention it requires is not. That creates an ethical tension: AI should make work easier, but a human role is not preserved merely by adding a person to the process when they lack time, information, or room to act.

Central claim: Human oversight is not a property of a checkbox. It is an organisational capability. It requires relevant information, adequate competence, real review time, authority to intervene, and a safe route for cases that do not fit the routine. Review depth should reflect possible harm, reversibility, and uncertainty. If the necessary capacity is unavailable, the process must be able to narrow, delay, reroute, or pause the task.

This article connects philosophical questions about practical judgment with controlled studies of time pressure and AI-assisted decisions. A wholly fictional SME case shows how the Question Expansion Model and the Consultation Model can turn the abstract demand for “human control” into concrete working rules. The proposed matrix is an editorial tool, not a scientifically validated risk scale.

For: Small and medium-sized enterprises, course facilitators, teams reviewing or integrating AI outputs, and learners who want to distinguish accountability from a click.

Learning outcome: You will be able to distinguish substantive review from its mere presence, make a capacity gap visible, and design review rules that account for harm, reversibility, uncertainty, competence, and available time.

1. The Review That Exists on Paper

A small, entirely fictional IT service company uses AI to sort incoming requests into categories such as “routine,” “request for clarification,” and “urgent.” Management requires a staff member to confirm every classification. It sounds like a clear safeguard: the AI does not decide alone.

Then the queue grows. In this fictional exercise, 54 requests arrive during a three-hour work window. A careful review takes an average of 75 seconds because someone must read the original text, retrieve context, and assess the category. That amounts to 67.5 minutes of review time. Yet after customer calls, questions, and interruptions, the team has at most 45 minutes available for this task. These figures are illustrative assumptions, not observations from a real company.

Under these conditions, “review everything” becomes an impossible instruction. Some entries are skimmed. Others receive a confirmation mark because clearing the queue matters more than inspecting the details. A request about access that is due to expire soon is labelled routine. The model need not have offered an especially persuasive explanation; the request simply disappeared among the others.

The important question is no longer, “Did a human confirm it?” It is: “Which cases need what kind of review, who can perform it with which information and in what time, and what happens when that time is not available?” The focus moves from individual failure to the conditions that make judgment possible.

The case is entirely fictional. It describes no real person, customers, organisation, or product environment. Its figures only illustrate an arithmetic mismatch: 54 requests multiplied by 75 seconds equals 4,050 seconds, or 67.5 minutes. If at most 45 minutes are scheduled, 22.5 minutes are missing. The organisation has not simply added a safeguard; it has demanded more review work than its own schedule permits.

There are several ways to address this gap. The company can add review time, limit volume, reserve deeper checks for selected cases, or activate a non-AI route. It may also find that the AI step offers no net benefit. What it cannot reasonably do is recast an impossible instruction as an individual reviewer’s weakness.

2. Judgment Is More Than Applying a Rule

In the Nicomachean Ethics, Aristotle distinguishes theoretical knowledge from practical wisdom, which concerns changing situations and concrete action. General rules can offer direction, but in a particular situation someone must recognise which details matter and which action serves the relevant good purpose. Practical wisdom is not simply a rule that executes itself in every case. [1]

This does not mean Aristotle anticipated a theory of AI oversight. The narrower connection is this: a classification may capture general patterns, while the meaning of an individual case also depends on what is at stake. “Access blocked” could mean a routine password reset or, just before a deadline, a serious disruption to the work process. A system that sees only an excerpt may not know the history or the consequences of delay.

This account of judgment is not an invitation to idealise people. Humans make mistakes, tire, overlook evidence, and respond to incentives. A person working through an overloaded queue is not automatically more reliable than a system. A purely human process can also be slow, inconsistent, or unfair. Any AI-supported solution must therefore be tested against the actual task and its consequences.

The ethical standard is not to insert a human as often as possible. It is to design decisions so that an organisation can explain why a given level of review is sufficient and who can act when something goes wrong. A human role is useful when it contributes relevant perception, expertise, or intervention. Without these, “human in the loop” is only a label.

Autonomy also depends on more than formal permission to decide. A person needs reasons, a real alternative, and the ability to disagree without penalty when a recommendation does not fit. If work pressure makes disagreement practically impossible, a right to intervene may exist in a process document but be hollow in practice.

3. What Research Shows About Time Pressure and AI Assistance

Experiments can control selected aspects of time pressure and interface design. They do not automatically predict everyday office work. The findings below challenge both “people always trust AI more under pressure” and “a human review is always better.”

3.1 Two Review Sequences in a Laboratory Task

Rieger and Manzey studied a luggage-screening task involving X-ray images in two laboratory experiments. In each experiment, 60 analysed participants searched for particular objects. Participants worked either without decision support or with a system set to 95% or 75% reliability. They had 4.5 or 9 seconds to inspect each image. Participants were recruited through the psychology department’s system at Technische Universität Berlin; the task was a controlled laboratory exercise. [2]

In Experiment 1, participants saw the system’s advice before inspecting the image themselves. Under the shorter time limit, performance fell with and without support. Reliance on the system—in this task, accepting the negative signal that no target object was present—did not simply increase under pressure; in this condition it fell slightly. In Experiment 2, participants inspected the image and gave an initial response before seeing the recommendation. The negative effect of time pressure largely disappeared in both assistance conditions, while it remained in the manual condition. [2]

The result is interesting but narrow. The altered timing and the additional chance to revise a decision may themselves have influenced performance. That was part of the experimental design. The study therefore does not show that “think first, then see AI” is universally superior. It shows that sequence and room to act can change the collaboration.

Another result complicates a familiar assumption: in the tested conditions, the joint performance of human and system was often below that of the highly reliable decision aid alone. This does not mean people should always be removed from oversight. It means that their participation must serve a purpose; adding a human does not guarantee better performance. An inappropriate correction can make a good system recommendation worse. [2]

3.2 Speed Can Rise While Reliance on AI Also Rises

Swaroop and colleagues ran two controlled online experiments using newly designed logic puzzles. Participants selected an option in a fictional “medicine for an alien” task. The study could vary task difficulty, AI quality, and time pressure. Experiment 1 included 159 of 207 starters in its analysis; Experiment 2 included 316 of 403. [3]

The researchers compared two sequences. With AI-before, the recommendation appeared before the person’s decision. With AI-after, the person made an initial decision and then saw the AI suggestion. Experiment 1 did not find a statistically significant direct increase in AI-before overreliance under time pressure (p = .13). A prior period of time pressure also affected later behaviour, so Experiment 2 assigned time-pressure conditions between participants. There, AI-before was faster than AI-after under time pressure (39 versus 56 seconds on average) and showed higher overreliance (0.56 versus 0.42), while accuracy was similar (0.72 versus 0.75). The study defined overreliance as following a wrong or suboptimal AI suggestion. [3]

This does not show that early AI suggestions are harmful in every situation. In the puzzle task, an early recommendation saved time and could offer an accuracy advantage compared with no assistance. The more limited point is that the trade-off between speed and independent review cannot be reduced to one design goal. A sequence that gives more room for an independent first judgment may reduce acceptance of wrong suggestions, but it may also take longer.

The authors also note a transfer limit: the tasks were novel, fully described on a screen, and not high stakes. That differs from work requiring domain knowledge, accumulated context, incomplete records, or decisions with social consequences. We use the study to motivate local testing, not as a ready-made workplace rule. [3]

3.3 A Thinking Pause Has Benefits and Costs

Buçinca, Malaya, and Gajos studied cognitive forcing interventions: designs that encourage people to assess a recommendation rather than accepting it immediately. Their online experiment with 199 participants used a simplified food-ingredient selection task. For some incorrect AI suggestions, certain interventions improved error correction compared with simple explanations. Across all task instances, however, the objective performance measures did not differ significantly between the two AI designs. The more effective interventions also received less favourable subjective ratings, and their benefit varied with people’s motivation to engage in effortful thinking. [4]

The study does not show that a mandatory pause will work in every company. It highlights a practical design question: every additional check consumes time and attention. A safeguard should be used where it offers a plausible benefit, and its side effects should be measured. If people must complete a multi-step justification for every routine case, the safeguard itself can become a bottleneck.

3.4 What These Studies Do Not Establish

The three studies use different tasks, groups, time windows, and forms of AI advice. They do not provide a single error rate for current language models, agents, or SME workflows. In particular, they do not prove that time pressure always increases trust or always lowers accuracy. Depending on the task and when advice appears, reliance may rise, fall, or barely change. AI support can speed some decisions or buffer performance loss; it can also encourage acceptance of wrong advice or unnecessary changes to correct advice.

The practical conclusion is therefore not a psychological judgment about individual employees. It is that a workflow must be tested under the conditions in which it will actually be used: case volume, difficulty, available time, prior knowledge, error costs, display sequence, and the ability to interrupt the process safely.

4. Why a Confirmation Field Does Not Prove Oversight

Review can provide substantive control only when several conditions come together. The reviewer must see the relevant case and the evidence needed to assess the recommendation. They need enough competence to recognise a problem, enough time for the required depth of review, and authority to change, stop, or reroute the case. There must also be a fallback if review or the system cannot proceed.

If one of these conditions is missing, the confirmation mark can hide the real bottleneck. A button does not say how long anyone read. A short comment does not prove that the person understood the case. Permission to disagree has little value if performance measures reward speed alone.

Downstream metrics can also mislead. A low override rate could mean suggestions are good. It could mean nobody has time to disagree. A high override rate could signal poor system performance, a new task, or appropriate professional corrections. None of these numbers explains what happened by itself. Review rates, sampled errors, reason categories, and workload need to be interpreted together.

Organisations should collect process data only as far as needed for a clear purpose. Aggregate information about queue age, unclear categories, corrections, and unavailable review time can help improve the rules. It should not become a covert individual monitoring system. If personal work data are unavoidable, purpose, legal basis, transparency, and access require a separate review. This article is not legal advice.

5. The Question Model Changes the Assignment

The fictional company’s first question is: “Did a person confirm every AI classification?” It sounds measurable, but leaves “confirmed” undefined. The Question Expansion Model F v5 helps expand the question through A → B → D → C. The model is not evidence and does not establish truth or effectiveness.

A · Clarify the Starting Question

What exactly was reviewed: the category, the original text, the rationale, the risk, or the route to the next team? Did the AI appear before or after the reviewer formed an initial view? Which cases and which version of the workflow are in scope? Who reviews, what competence do they have, and how much time is genuinely available?

The result is a more useful question: Can a qualified person review the information that matters to the next step in time, change the recommendation if necessary, and route the case safely? A click may be one signal within a larger process, but it does not define the process.

B · Allow Multiple Explanations

An urgent request may be misclassified for several reasons. The system may lack context. The AI suggestion may appear so early that it shapes the first impression. The reviewer may not understand the categories. The queue may exceed available capacity. Or the reviewer may be allowed to disagree in theory but lack authority to escalate.

A short 5 Whys trace can open the issue: Why was the urgent request not recognised in time? Because it remained in the routine queue. Why did it remain there? Because the classification was confirmed. Why was it not examined more closely? Because the schedule did not allow time for every case. Why was that schedule set? Because the volume estimate counted a confirmation per item, but did not account for review depth or interruptions. Why was the estimate made that way? Because expected time savings were estimated without measuring errors and escalation work.

This chain is a hypothesis, not proof of causation. The team should check alternatives: Is the volume correct? Was review time measured realistically? Do errors happen even when time is sufficient? Is the right role involved? Are categories and escalation rules clear? The data used to answer these questions should be as non-personal as possible.

D · Keep Uncertainty and Revision Visible

Before reaching a conclusion, the team should separate what it observed from what it suspects. A long queue is observable. The claim that it caused a particular error is an additional causal assumption. A click log proves a click, not the quality of review. If counterexamples appear, the first interpretation must remain revisable.

The case note might say: “The current confirmation requirement does not demonstrate adequate review capacity. It remains unclear whether misclassification is mainly caused by missing context, display sequence, or volume.” This avoids unsupported blame while preserving a concrete investigation.

● Members only

Read the full article and download all files with a membership.

Unlock full article + downloads → Subscribe

0 comments

● Loading comments…

Sign in to comment · become a member →