The bounded pilot
How a small team trials AI, compares it with a non-AI baseline, and actually stops after a poor result

An AI suggestion can become habitual very quickly in everyday work. At first, people say, “We are only checking once whether it helps.” A few weeks later, the old way of working is no longer practised, the tool is built into templates, and no one knows exactly who would have ended the trial. That is not a small introduction with an open outcome; it is an introduction without a clear decision.
A bounded pilot starts earlier. It is a time-limited comparison between specific ways of working. Before it begins, the team decides what it will try, what it will not try, which work remains possible without AI, what counts as benefit, which errors or burdens matter, and what finding triggers a pause. The outcome is not a success story but a reasoned decision: end it, redesign it more narrowly, clarify a professional question, or continue examining it under tightly specified conditions.
This article develops a wholly fictional case. “Nordwerk” is not a real company. The case contains no personal, customer, employee, or contractual data. It concerns only artificial request cards that are internally assigned to one of four processing paths. There is no external contact, dispatch, or access to inventory or customer systems. This limitation is precisely what makes the case useful for a course or a small team: the group can work on the design of a decision without using a productive process as a practice field.
The guiding question is: How can a small team trial an AI use in a sufficiently time-limited way that benefit, additional work, errors, and harms are compared with a non-AI baseline, and a negative result really leads to a stop?
The answer is deliberately modest. An eight-week pilot provides indications for the tested cards, roles, and conditions. It proves neither general effectiveness nor safety, ethical appropriateness, or legal compliance. A good average figure can conceal relevant individual errors. The pilot therefore needs, alongside an overall balance, error classes, separate analysis of difficult cases, a usable fallback, and a hard stop threshold agreed in advance.
Not “Is the AI good?”, but “Which work decision is defensible here?”
In the example, Nordwerk initially processes spare-parts enquiries as an internal short note. Every artificial request card first needs to be assigned. There are four paths:
1. Standard case: The card contains an unambiguous reference; the case handler can start the normal internal checking routine.
2. Missing reference: Essential information is missing; the card is set aside for clarification.
3. Professional review: The information is present, but the assignment requires specialist knowledge or involves a contradictory reference.
4. Stop/exception: The card contains a contradiction, an unprovided-for category, or a signal that may not be processed in the test.
The team is considering an AI assistant that only suggests one of these four paths. The AI may not trigger anything. A person accepts or rejects the suggestion in a test table. The cards, reference solution, and criteria are artificially created. This keeps the decisive question clear: in this limited part of the work, does the suggestion reduce total work without producing errors, rework, or unequally distributed checking burden?
This formulation distinguishes this article from a single oversight test. An oversight test can show whether a person actually stops, overrides, or escalates when faced with a known erroneous output. That is important, but it is not yet a pilot. In addition, the pilot asks about comparison, duration, case mix, effort over several weeks, and a concluding decision. The ability to exercise oversight is assumed here as a precondition of the shadow arm; it is not carried out again as a separate test.
Nor is a use-case card alone sufficient. It can record purpose, data zone, roles, and ways of stopping. This article uses this preparatory work, but adds a testable trial logic: a manual starting position, a structured non-AI rule card, an AI shadow arm, an independently specified reference, and rules for evaluating both paths.
This is not distrust of every technical aid. It is a question of accountability. Anyone claiming the benefit of a new tool must also see the work it newly creates: preparing inputs, checking suggestions, looking up references, correcting errors, handling exceptional cases, and returning the process to the non-AI baseline after a stop. If that work disappears because it is not counted, an apparent time saving results.
Why a quick average figure is not enough
Philosophically, this concerns more than measurement methodology. A pilot distributes time, attention, and risk. Management may welcome faster completion, while a reviewing person carries additional checking work. A frequent suggestion that is easy to correct can lower the average time while sending rare cases down the wrong path. If only the average is reported, different consequences are treated as though they were interchangeable.
A defensible decision does not have to compress these differences into one moral number. It should, however, disclose which goods are in tension: working time, reliability, ability to object, professional care, and who bears the burden when in doubt. Reversibility is a practical value here. It means that the team does not destroy the manual way of working before it has sufficient reasons to change it. A stop is therefore not a personal failure of the people who prepared the pilot. It is one of the intended decisions.
The strongest counterargument is nevertheless that a small team cannot afford an elaborate comparison. Documentation takes time. With few cards, the case mix, learning effects, and chance can shape the result. A shadow operation demands duplicated work. If a team measures for eight weeks while others are already working faster, caution itself can become harmful.
This objection is strong. It argues against a ritual of spreadsheets, not against every bounded test. The appropriate response is not an ever larger control system. It is this: measure only as much as is needed for the specific decision; explain the measurement in advance; retain a real non-AI option; and keep the conclusion narrow. If the effort of the pilot already exceeds its possible benefit, the decision before it starts can be “no pilot.” A short, tightly bounded comparison can also show that a better rule card, a clearer reference, or additional learning time helps more than AI.
What sources suggest for a pilot — and what they do not prescribe
The NIST AI Risk Management Framework is a voluntary framework, not a legal norm. It does not require an eight-week duration, a minimum number of cases, or a general stop button. It does, however, recommend documenting the deployment context and risk tolerances, recording metrics and test conditions, making limits on transferability visible, and considering usable non-AI alternatives when treating risks.1 For Nordwerk, this yields a working rule, not a NIST requirement: the comparison needs a real manual alternative, not merely the question of whether the AI output sounds plausible.
Why a single overall figure is insufficient is illustrated by an original study in nursing education and clinical simulation. In it, 450 nursing students and twelve registered nurses assessed historical cases with and without AI recommendations. AI support could improve or worsen joint human-AI performance, depending on whether the respective algorithm was particularly correct or particularly misleading. The authors warn that a representatively weighted or averaged evaluation can trade frequent small benefits against rarer, larger harms and thereby conceal the serious consequences.2 This is not a measurement in a spare-parts operation. It does, however, support the cautious design idea of analysing difficult and error-prone cards separately rather than allowing them to disappear into the mean.
A second original study observed 38 intensive care physicians in six simulated sepsis scenarios per person. After making their own dosing decision, they saw a safe or deliberately unsafe AI recommendation and decided again. Unsafe recommendations were stopped more often than safe ones, but not without exception; including requested second opinions, the proportion of unsafe recommendations stopped or escalated was higher; this difference was not statistically significant. In addition, the physicians changed their decision after AI suggestions in 105 of 228 runs.3 These figures are not a target value for Nordwerk. The transferable idea is narrower: in a pilot, the final outcome is not the only point of interest. Changes after a suggestion, overrides, escalations, and missing interventions are also observable process findings.
A third study, preregistered online experiments involving a total of 3,800 participants estimating house prices, examined clearly presented versus less transparent linear models. The easily understandable model helped participants mentally follow its predictions. It did not, however, appropriately improve following recommendations; for large errors on atypical data points, participants detected and corrected the errors less well.4 This too is not a test of operational request assignment. It does prevent an easy conclusion: an explanation, dashboard, or approval click is not a quality measurement. The pilot must examine whether problematic suggestions are detected and handled correctly against an independent reference.
The EU AI Act belongs here only in a narrower sense. For high-risk AI, the Regulation requires providers to maintain continuous, iterative risk management; it also contains requirements for post-market monitoring. Under the conditions stated, deployers must inform and suspend use when there is a risk. For high-risk AI under Annex III, the corresponding obligations apply from 2 December 2027 under the time frame used here, and for certain systems under Annex I from 2 August 2028.5 This does not create a general pilot duty for every SME or a statutory eight-week rule. Whether these rules apply in a real case depends on the system, role, intended purpose, risk classification, and date. Nordwerk therefore makes no compliance claim.

The eight-week Nordwerk pilot
The case uses three comparison arms. They are not an attempt to have several AI products compete against one another. They make different ways of working visible.
| Arm | Way of working | Purpose of the comparison |
|---|---|---|
| M: manual baseline | A person reads the card and accessible work information, and assigns it without decision support. | Shows the existing baseline, including normal checking. |
| R: non-AI rule card | A structured card leads through fixed questions: Reference present? Contradiction? Specialist feature? Exception? | Tests whether process clarity already achieves the benefit without AI. |
| K: AI shadow arm | The AI suggests a path. A person checks it against the same work information and records acceptance, correction, or stop. | Tests the benefits and risks of the human-AI configuration. |
All three arms use artificial cards from the same category of task. The AI remains in the shadows: its output creates no subsequent process outside the test table. This means deliberately difficult cards can also be examined without a real case being routed incorrectly.
Before the start, one designated person creates 240 cards. Each card contains work information available to all arms to the same extent: for example, a product detail or an explicitly missing reference. Separately from this, the reference role specifies the correct one of the four paths as a hidden answer key. The processing role sees the key only after completing its assignment; otherwise, the trial would measure reading off the solution. The reference role may not reword cards during processing. There are 144 standard cases, 48 with missing work references, 24 requiring professional review, and 24 stop/exception cards. This case mix is a fictional course decision, not a statement about real enquiry frequencies. Each week, every arm receives 10 new cards: six standard cases, two cards with a missing reference, one professional review, and one exception. Before the start, the cards within each class are also stratified by predefined difficulty and randomly allocated to M, R, and K. The classes and difficulty levels remain equally represented each week; the same card versions do not move into two arms. Across eight weeks, at most 80 cards per arm and at most 240 processing instances in total are planned. The order within a week is also randomized. The allocation reduces obvious differences, but does not eliminate all bias in this small, artificial set.
The briefing and finalization of the hidden answer key take place before week one. In weeks one to seven, the agreed comparison runs with ten cards per arm; week eight contains the final ten cards per arm, feedback, and the final assessment. If K is paused earlier, the comparison for all three arms ends. The weeks completed in parallel up to that point form the shared evaluation window; the remaining artificial cards are not processed. In everyday work, new work would return to the previously secured manual path. Neither the prompt nor the rule card, roles, or answer key may be silently changed during the ongoing run. If a genuine need for change arises, K is paused and the change is noted as a new question. A later, clearly separate run can test a new version.
The roles are kept small:
Pilot owner: plans dates, protects time for processing, and convenes the final decision.
Reference owner: creates the hidden answer key in advance and releases only reasoned corrections during evaluation.
Review role: processes the cards and may reject any AI suggestion or pause the K arm.
Log owner: counts times and error classes without collecting raw texts or personal details.
Fallback role: switches to the manual M path when paused and records that the K arm is no longer used.
One person can hold several roles in a small team. The hidden answer key and ongoing processing should not, however, be in the same hands or visible at the same time. The work information usable by all arms is separate from it. The aim is not formal independence as in a clinical study. It is to prevent the person who has seen the AI suggestion from inventing or reading off the correct assignment to fit it.
Decide in advance what is counted
The central metric is total working time per completed card. It begins when the card is opened and ends only when the assignment, including checking, correction, log entry, and any necessary rework, is complete. If a reference was missing and the “missing reference” path was correctly chosen, the time for that decision is included. It would be misleading to count only the seconds until the AI suggestion and remove another person’s checking from the denominator.
For each arm, the simple accounting is: Total time = initial processing + checking + correction + log + rework. The relative time gain of K compared with R is (mean total time R − mean total time K) / mean total time R. A negative number would mean additional time instead of savings.
The word “gain” only describes the measured value. It is not a decision rule. A fast K arm can still fail.
For that reason, the cards are additionally counted by error class. The categories are fixed before the first card:
| Error class | Definition in the Nordwerk case | Consequence in the pilot |
|---|---|---|
| F1: deviation corrected before completion | In M, R, or K, an incorrect path first arises; the processing role corrects it before completion. | Correction time counts in every arm; not yet a final quality error. |
| F2: incorrect path completed in the log | The test card is documented as complete contrary to the reference. | Quality error; the reference owner reviews the cause. |
| F3: impermissible further processing of an exception | A stop/exception card receives a normal or professional assignment instead of a stop. | Hard safety breach; immediate pause. |
| F4: incorrect handling of a missing reference | A reasoned assignment is claimed despite a missing reference. | Hard quality breach; immediate pause. |
| F5: checking burden | In K, separately logged checking time exceeds five minutes per card in at least three of K’s ten cards in one week, or the pre-reserved checking time is missing for two K cards in the same week. | Pause at weekly review; no expansion without protected checking time. |
A log may also record whether a K suggestion was accepted, corrected, ignored, or escalated. This is a K-specific diagnostic, not an error class comparable between M, R, and K. A high number of acceptances proves neither quality nor trust; it can also show that suggestions are barely checked. A high number of corrections can provide a useful signal about system limits, or simply indicate additional work. It gains meaning only through the subsequent comparison with the hidden answer key, time, and error consequence. F3 and F4 are serious subsets of an incorrectly completed path and are not added to F2 as a separate count in a total error figure.
● Members only
Read the full article and download all files with a membership.
Unlock full article + downloads → Subscribe0 comments
● Loading comments…