SAKIZLI AI
Article28 Jul 2026 · 17 min read21 / 23Members · Subscription

A pilot deployment is not yet controlled real-world testing

Once real people, decisions or consequences are involved, the word "pilot" is no longer a safety architecture.

AI ActGovernanceRiskEvaluation
FFurkan SakızlıAI researcher & tutor · independent
A bright test path moves a modular AI system from simulation through a transparent control gate into bounded real-world testing with measurement points, reversal control and exit; an uncontrolled shortcut is blocked in amber
Only effective boundaries turn „let us try it“ into a responsible experiment

Once real people, decisions or consequences are involved, the word „pilot" is no longer a safety architecture.

Organisations like pilots because the term sounds small, reversible and educational. A limited audience, short duration, dashboard and feedback sessions appear to create realistic testing. Yet a pilot may already be operational: people change behaviour, staff rely on recommendations, data flows into other systems and errors cause actual disadvantage.

Controlled real-world testing begins with a verifiable test architecture. It defines purpose, population, duration, roles, safeguards, stop rules, reversal, data use and evidence. Only effective boundaries turn „let us try it" into a responsible experiment.

Separate four operating modes

Different experiments are often called pilots:

1 · Simulation: synthetic or historical cases with no influence on current people or processes.

2 · Shadow mode: real input is processed beside the existing workflow, but the output drives no decision.

3 · Controlled real-world test: output may influence action inside a bounded setting under special supervision, consent and reversal.

4 · Regular operation: the system enters its intended workflow and full operating and monitoring regime.

These are not merely maturity levels. Shadow mode can process sensitive data or indirectly influence staff. A test becomes operation when results are used, retained or transferred beyond its plan.

Article 60 creates a narrow route

Article 60 of the EU AI Act concerns real-world testing outside regulatory sandboxes for certain Annex III high-risk systems before market placement or putting into service. Providers or prospective providers may test alone or with deployers when all conditions are met. Article 5 prohibitions and other law remain applicable.

Conditions include a submitted plan, involvement or approval of market surveillance authorities, registration, Union establishment or representative, safeguards for third-country transfers, time limits, protection of vulnerable groups, role agreements, informed consent, qualified oversight and effective reversal of outputs.

Each legal condition needs a technical or organisational control point.

The plan is an executable contract with reality

A good plan identifies system version, sources, populations, locations, workflows and decision types. It defines hypotheses, metrics, thresholds and stop criteria, including normal, failure and extreme cases.

„Test prioritisation accuracy" is too vague. A stronger statement defines 500 voluntary cases during a period, an independent reference, mandatory confirmation before action, and automatic pause when error or group-difference thresholds are exceeded.

The plan also states non-goals. Data cannot silently enter staff assessment, model training or marketing analysis. Negative boundaries prevent scope creep.

Consent is a process, not a checkbox

Article 61 requires voluntary, specific, informed and unambiguous consent in the relevant cases. Subjects need understandable information about the test, objectives, benefits, risks, duration, conditions, rights and withdrawal. They can generally withdraw without detriment or explanation and request immediate permanent deletion of personal data.

A long privacy notice before a click does not create informed choice. Information needs accessible language and format without improper pressure. Dependency relationships, including employment, require particular scrutiny of whether consent is genuinely voluntary.

Consent is versioned. Records show which information was accepted, when, how withdrawal stops flows and which completed activities remain unaffected. Exit must work as reliably as entry.

Reversal must be technically demonstrated

Predictions, recommendations and decisions must be effectively reversible and disregarded. A reviewer on an organisation chart is insufficient.

The test demonstrates that a qualified person sees output in time, understands it, has counter-evidence, can deviate without penalty and actually stops downstream action. Tool-using systems need transaction boundaries, approval gates, checkpoints or compensating actions.

Disturbance tests include wrong identity, stale data, conflicting evidence, control failure, unauthorised tool use and consent withdrawal during processing. A button that stops only the interface while background actions continue is not a stop.

Oversight requires capacity and authority

Reviewers need expertise, time, training and decision rights. A person responsible for hundreds of simultaneous cases, or unable to enforce disagreement, provides decorative oversight.

The plan defines coverage, caseload, escalation, substitution, conflicts and authority. It measures not only override rates but whether meaningful errors are detected and acted upon.

Automation bias can be tested with plausible but deliberately incorrect recommendations. This reveals whether humans review or merely confirm.

Vulnerability needs active safeguards

Age and disability can create particular vulnerability; context, dependency, language and accessibility can create others. The first question is whether inclusion is necessary and proportionate.

Where included, subjects may need adapted information, accessible interaction, tighter thresholds, extra human support and rapid complaints. Averages must not hide systematic harm to smaller groups.

Time limits prevent permanent pilots

Testing lasts no longer than necessary and generally no more than six months. One additional period of up to six months requires prior notification and justification. This is a maximum, not a target.

Every test has an end date and exit state: deactivate, move into another authorised phase, or enter operation after required procedures. A pilot that simply continues is uncontrolled deployment.

Incidents trigger mitigation, suspension or termination

Serious incidents are reported to the market surveillance authority. Providers adopt immediate mitigation; otherwise testing is suspended or terminated. Termination requires prompt recall.

Incident readiness is rehearsed before launch: detect, preserve logs and versions, stop, inform subjects and authorities, correct downstream actions. An unattended mailbox is not an incident process.

Liability does not vanish during testing. Providers remain liable under applicable Union and national law for damage caused. Pilot economics must include safeguards, reversal and remediation.

Evidence should explain differences

Average accuracy is insufficient. Evidence connects case, version, input state, output, human decision, action, outcome and deviation. Relevant groups and scenarios are analysed without collecting unnecessary personal data.

Predefined metrics prevent selective reporting. Negative results and near misses remain visible. Model or threshold changes create new comparison groups so different systems are not blended into one comforting number.

Market transition is a separate gate

A successful test is not automatic production approval. Conformity requirements, technical documentation, risk and quality systems, registration, instructions, oversight and sector duties still need completion.

Monitoring then transitions into post-market monitoring. Population, volume and failures change in operation. The plan covers drift, complaints, incidents, corrective action and re-evaluation. Real-world testing starts evidence-based learning; it does not end it.

A technically possible test may still be organisationally unacceptable

A readiness gate evaluates the full protection capacity, not only the model. Are enough qualified reviewers available? Can withdrawal and deletion be executed within the promised time? Are budget and on-call capacity sufficient for incidents, reversal and possible remediation? Is the independent reference trustworthy?

If any capability is missing, the test is reduced, returned to simulation or postponed. Technical functionality is not permission to expose people to a poorly controlled learning process. The decision and its rationale belong in the evidence record.

Test economics should include supervision time, accessibility, participant support, secure logging, independent evaluation, incident response and exit work. A cheap pilot that externalises these costs is not efficient; it merely transfers risk to subjects and operators.

Worksheet: Design controlled real-world testing

1. Classify simulation, shadow mode, real test or operation.

2. Define population, version, location, duration and non-goals.

3. Set hypotheses, metrics, thresholds and exit criteria.

4. Design consent, withdrawal and deletion.

5. Demonstrate oversight, stopping, reversal and incident response.

6. Protect vulnerable groups and inspect group differences.

7. Plan registration, authority contact and final reporting.

8. Define the gate to conformity and post-market monitoring.

All materials to download — the topic overview and the worksheet:

Scope: This article presents a professional control pattern, not legal advice. Article 60 concerns certain Annex III high-risk systems outside a regulatory sandbox; product-specific, national, ethical, data protection and liability rules still require assessment. Not every operational pilot is real-world testing under that Article. Editorial review date: 17 July 2026.

Members only

Read the full article and download all files with a membership.

Unlock full article + downloads → Subscribe

0 comments

Loading comments…

Sign in to comment · become a member →