SAKIZLI AI
Article28 Jul 2026 · 17 min read20 / 40Members · Subscription

A pilot deployment is not yet controlled real-world testing

Once real people, decisions or consequences are involved, the word "pilot" is no longer a safety architecture.

AI ActGovernanceRiskEvaluation
FFurkan SakızlıAI researcher & tutor · independent
A bright test path moves a modular AI system from simulation through a transparent control gate into bounded real-world testing with measurement points, reversal control and exit; an uncontrolled shortcut is blocked in amber
Only effective boundaries turn „let us try it“ into a responsible experiment
Image generated with AI

Once real people, decisions or consequences are involved, the word „pilot" is no longer a safety architecture.

Organisations like pilots because the term sounds small, reversible and educational. A limited audience, short duration, dashboard and feedback sessions appear to create realistic testing. Yet a pilot may already be operational: people change behaviour, staff rely on recommendations, data flows into other systems and errors cause actual disadvantage.

Controlled real-world testing begins with a verifiable test architecture. It defines purpose, population, duration, roles, safeguards, stop rules, reversal, data use and evidence. Only effective boundaries turn „let us try it" into a responsible experiment.

Separate four operating modes

Different experiments are often called pilots:

1 · Simulation: synthetic or historical cases with no influence on current people or processes.

2 · Shadow mode: real input is processed beside the existing workflow, but the output drives no decision.

3 · Controlled real-world test: output may influence action inside a bounded setting under special supervision, consent and reversal.

4 · Regular operation: the system enters its intended workflow and full operating and monitoring regime.

These are not merely maturity levels. Shadow mode can process sensitive data or indirectly influence staff. A test becomes operation when results are used, retained or transferred beyond its plan.

Article 60 creates a narrow route

Article 60 of the EU AI Act concerns real-world testing outside regulatory sandboxes for certain Annex III high-risk systems before market placement or putting into service. Providers or prospective providers may test alone or with deployers when all conditions are met. Article 5 prohibitions and other law remain applicable.

Conditions include a submitted plan, involvement or approval of market surveillance authorities, registration, Union establishment or representative, safeguards for third-country transfers, time limits, protection of vulnerable groups, role agreements, informed consent, qualified oversight and effective reversal of outputs.

Each legal condition needs a technical or organisational control point.

The plan is an executable contract with reality

A good plan identifies system version, sources, populations, locations, workflows and decision types. It defines hypotheses, metrics, thresholds and stop criteria, including normal, failure and extreme cases.

„Test prioritisation accuracy" is too vague. A stronger statement defines 500 voluntary cases during a period, an independent reference, mandatory confirmation before action, and automatic pause when error or group-difference thresholds are exceeded.

The plan also states non-goals. Data cannot silently enter staff assessment, model training or marketing analysis. Negative boundaries prevent scope creep.

Consent is a process, not a checkbox

Article 61 requires voluntary, specific, informed and unambiguous consent in the relevant cases. Subjects need understandable information about the test, objectives, benefits, risks, duration, conditions, rights and withdrawal. They can generally withdraw without detriment or explanation and request immediate permanent deletion of personal data.

A long privacy notice before a click does not create informed choice. Information needs accessible language and format without improper pressure. Dependency relationships, including employment, require particular scrutiny of whether consent is genuinely voluntary.

Consent is versioned. Records show which information was accepted, when, how withdrawal stops flows and which completed activities remain unaffected. Exit must work as reliably as entry.

Reversal must be technically demonstrated

Predictions, recommendations and decisions must be effectively reversible and disregarded. A reviewer on an organisation chart is insufficient.

The test demonstrates that a qualified person sees output in time, understands it, has counter-evidence, can deviate without penalty and actually stops downstream action. Tool-using systems need transaction boundaries, approval gates, checkpoints or compensating actions.

Disturbance tests include wrong identity, stale data, conflicting evidence, control failure, unauthorised tool use and consent withdrawal during processing. A button that stops only the interface while background actions continue is not a stop.

Oversight requires capacity and authority

Reviewers need expertise, time, training and decision rights. A person responsible for hundreds of simultaneous cases, or unable to enforce disagreement, provides decorative oversight.

The plan defines coverage, caseload, escalation, substitution, conflicts and authority. It measures not only override rates but whether meaningful errors are detected and acted upon.

Automation bias can be tested with plausible but deliberately incorrect recommendations. This reveals whether humans review or merely confirm.

Vulnerability needs active safeguards

Age and disability can create particular vulnerability; context, dependency, language and accessibility can create others. The first question is whether inclusion is necessary and proportionate.

Members only

Read the full article and download all files with a membership.

Unlock full article + downloads → Subscribe

0 comments

Loading comments…

Sign in to comment · become a member →