Automate only what you can verify
The boundary of useful automation is not where a system can act, but where its outcomes can still be checked reliably.

Automation is often judged by feasibility: Can the system perform the process? That is not enough for responsible AI work. An agent can research, draft, modify files or prepare decisions while producing outcomes that look convincing. The decisive measure is whether we can determine, at reasonable cost, that an outcome is correct, complete and fit for purpose. Where that check no longer works, technical capability may continue—but responsible autonomy ends.
The real boundary appears after the output
In conventional automation, success and failure are often observable. A file was transferred or not, a field satisfies a rule or not. Generative systems create open-ended outcomes: a plan, evaluation, summary or recommendation. Such outputs may be formally polished and still be substantively incomplete.
Planning must therefore continue beyond the action. Every automated action needs a verification path describing how a good outcome is recognised, which foundation applies and what happens when the check cannot reach a clear decision. Without that path, automation is accelerated uncertainty.
Verifiability does not mean absolute truth. It means that a claimed outcome can be checked against observable criteria, sources or consequences. The less verifiable an output, the smaller its permitted effect should be.
Three classes of task require three boundaries
The first class is deterministically testable. Filenames follow a pattern, required fields exist, totals reconcile or a test passes. A system can act with considerable autonomy here when permissions are limited and errors are reversible.
The second class is assessable but not fully computable. A text should contain all approved claims, a draft should suit an audience or research should expose conflicting sources. This requires rubrics, reference examples, sampling and traceable approval. Automation may prepare much of the work, but it cannot close every quality question by itself.
The third class depends on situated judgement. It includes ethical, strategic, legal or interpersonal decisions whose quality relies on circumstances that cannot be fully formalised. AI can prepare options, consequences and open questions. The binding decision remains with an accountable person.
Verification fatigue is a hidden price of automation
A system can generate more outcomes in minutes than a person can inspect carefully. Production becomes cheaper while control becomes the bottleneck. If every output must be fully reworked, labour has not disappeared; it has shifted into a more exhausting form.
Fatigue changes decisions. After the tenth plausible result, attention falls. Reviewers skim, approve under time pressure or focus on visible defects while substantive gaps survive. A human approval step is therefore not automatically an effective safeguard.
Good automation limits verification load as well as error. It produces less, groups deviations, prioritises high-risk cases and demands human attention where it changes the outcome. The question is not whether a person can inspect everything, but whether the process leaves that person capable of judgement.
Acceptance criteria must exist before the run
When quality is defined only after the output appears, its presentation can set the standard. Elegant prose becomes its own evidence. A better approach starts with a short verification contract: desired outcome, authoritative sources, three to five acceptance criteria, prohibited deviations and behaviour under uncertainty.
Criteria must be observable. “The text is good” cannot be tested reliably. “Every number points to an approved source”, “missing information is marked open” and “no change occurs outside the working directory” can. Concrete criteria can be assigned to a test, a rule or a human decision.
Negative criteria matter too. They describe outcomes that remain unacceptable despite a polished surface: invented evidence, hidden uncertainty, irreversible changes or recommendations outside the brief. A system becomes safe through its behaviour at the boundary, not through the ideal case.
Verification needs several layers
No single method catches every error. Deterministic checks work for formats, completeness, value ranges and technical tests. An independent model review can look for contradictions, missing perspectives and unsupported leaps. Source checks reconnect claims to evidence. People finally evaluate context, consequences and acceptability.
● Members only
Read the full article and download all files with a membership.
Unlock full article + downloads → Subscribe0 comments
● Loading comments…