What Evidence Warrants a Decision?
Keeping empirical findings, values, organisational objectives, and risk tolerance distinct—and connecting them with reasons

Guiding question: When does the available evidence support an organisational decision—and which additional reasons must be made explicit?
The point at which “more evidence” is not a complete answer
A small, entirely fictional continuing-education provider is considering using a language model to pre-sort general course enquiries. The intention sounds reasonable: routine work should take less time so that the team has more time for advice. Three statements are made in a meeting: “The model seems remarkably accurate in our examples.” “We do not want to overlook prospective participants who phrase things differently.” “If we do not become faster, we will lose enquiries.” Each statement points to something relevant. None, by itself, answers whether the organisation should use the system.
“Accurate” is initially a claim about observed outputs. It requires a defined task, comparison standards, data, and a limited statement of the cases to which the finding applies. “Do not overlook anyone” is a normative commitment: it says which types of error and which affected groups should carry particular weight. “Become faster” is an organisational objective. It may be sensible, justified, and in conflict with other objectives. And whether a remaining risk of error is tolerable is a matter of risk tolerance that must fit the use, consequences, and opportunities for correction.
The difficulty begins when these registers slide into one another unnoticed. A positive test is then treated as moral justification; an economic objective is reinterpreted as an alleged measurement; a particularly cautious value standard is treated as “poor data quality”; or a good average result is supposed to excuse every individual consequence. All of this can happen without anyone deliberately misleading others. That is precisely why a sound decision needs a visible structure of justification.
The central claim of this article is: Empirical evidence can support factual claims and constrain courses of action. But it does not by itself determine which consequences count, which objective resources should serve, or what risk of consequences an organisation may bear in light of stated uncertainty. A defensible AI decision makes these different reasons visible, examines their relationships, and identifies the person or role accountable for the decision.
The four-register worksheet presented here is a pedagogical and editorial tool. It is neither a statistical scale nor a validated governance procedure, and it does not replace expert review. Its purpose is more modest: to prevent a single metric or persuasive formulation from appearing to carry the entire decision.
Four registers, four different tasks
A decision under uncertainty becomes clearer when four questions are answered separately. They may be brought together in the conclusion; the reasons should remain recognisable.
| Register | Question | Justifies | Does not justify |
|---|---|---|---|
| Empirical evidence | What do sources, observations, or tests show—and under what conditions? | The strength, scope, and limits of a factual claim | Which harms count morally or which risk is acceptable |
| Normative reasons | Which rights, interests, duties, distributions, and harms must be taken into account? | Why particular consequences matter for those affected and which constraints apply | That a factual claim is true |
| Organisational objective | What problem does the organisation want to solve, and how is that objective justified? | Why an option deserves attention and resources | An exception to rights, duties of protection, or applicable rules |
| Risk tolerance | What remaining risk of consequences may the organisation bear in this specific use and under the stated uncertainty? | A use-specific boundary for acting, experimenting, or stopping | The quality or truth of the underlying evidence |
1. Evidence needs a scope
“We tested it” is not yet a sufficient finding. A testable statement requires at least: What was tested? On which cases? Against which reference? With which version and procedure? Which cases were missing? Which error types became visible, and which possibly did not? A finding can be useful for a narrowly limited task without establishing a general characteristic “of the model.”
Evidence therefore does not simply take the form of “yes” or “no.” It can support a claim more strongly or more weakly, rule out particular uses, and open further questions. Uncertainty should not be given an artificial numerical value when the data and methods do not support such a number. “Unknown” is information about the state of review, not a placeholder that may automatically be filled with an optimistic assumption.
Two thresholds must be distinguished here. The epistemic threshold asks what statement the finding justifies: for example, whether a claim is well supported, only provisionally plausible, or still open. The action threshold asks which option is defensible in light of possible consequences, duties of protection, and opportunities for correction. A team may continue learning with limited evidence for a low-consequence, reversible exercise and regard the same evidence as insufficient for real automated sorting. Conversely, a clear finding may speak against use without requiring the organisation to resolve every uncertainty completely. The thresholds are related, but they answer different questions.
This separation prevents a common false precision: a test can provide a finding about a narrowly defined task. It does not automatically determine whether its scope is sufficient for the intended decision. Each conclusion should therefore make clear whether it is a description (“X occurred in these cases”), a prediction (“under these conditions we expect Y”), or a recommendation (“we should choose option Z”). Anyone moving from one type of statement to another must state the additional assumptions.
2. Normative reasons identify those affected and the consequences
An error is not merely a deviation in a table. It affects people, teams, processes, or public goods in different ways. Incorrect sorting may mean only a small rerouting; it may also mean that an urgent or linguistically unusual enquiry is seen later. These consequences need not all carry equal weight. Those who assess them should state which interests and duties of protection are decisive and whose perspective is missing so far.
Values are not a decorative list at the end of a project plan. They can determine which types of error need especially close attention before a decision and which side effects may not disappear as mere cost items. In its Recommendation on the Ethics of Artificial Intelligence, UNESCO describes human dignity and human rights, inclusion, proportionality, participation, responsibility, and human oversight, among other things, as normative reference points. The Recommendation is a global, standard-setting, legally non-binding framework. It does not prescribe an automatic decision in an individual case and does not establish either the effect or permissibility of a particular system.
3. An organisational objective is not evidence
An organisation may set objectives, such as accessibility, reliability, learning quality, or a sustainable workload. These objectives must be weighed against one another in the concrete situation. Efficiency may matter, but it does not yet show whether AI is the right solution in this case. The cause of a bottleneck could equally lie in unclear responsibilities, inconsistent templates, or insufficient time planning.
That is why the problem formulation comes before the tool: What should improve, and for whom? Which non-AI option could achieve the same improvement? An objective makes an investment understandable. It does not turn an untested benefit into an empirical fact.
4. Risk tolerance must fit the use
Risk tolerance does not mean the quality of the evidence, nor is it a freely selectable approval number. It is the justified boundary up to which an organisation may bear risks of consequences in a particular use case in pursuit of an objective. How large those risks are depends, among other things, on the stated uncertainty in the evidence. Relevant questions are: How serious would a possible harm be? How many people would be affected? Is the step reversible? Can a person intervene in time? Is there an effective route for complaint or correction? Which rules and rights limit the room for manoeuvre independently of business interests?
The NIST AI Risk Management Framework 1.0 (2023) treats risk tolerance as context- and use-dependent and does not prescribe a universally applicable threshold for organisations. AI RMF 1.0 is a voluntary, non-sector-specific management framework; it replaces neither applicable legal norms nor the justification of the relevant weighing of values and harms. NIST currently notes that AI RMF 1.0 is being revised. For this article, it is therefore the published 2023 version, not a purportedly completed 2026 revision.
One practical consequence follows: a low-consequence, fully reversible exploration may require different evidence than an automated external-facing use, a selection decision, or a process that affects rights and opportunities. More uncertainty is not automatically indefensible; it does, however, require a narrower action, clearer safeguards, or postponement. Sometimes the responsible decision is precisely to refrain from deciding for the time being.

A philosophical tension: values in science or values only in action?
The question of whether and where values enter scientific judgments is philosophically contested. Richard Rudner argues that the consequences of possible errors play a role when accepting or rejecting empirical hypotheses: the threshold of “enough evidence” sometimes presupposes how seriously different errors matter. Heather Douglas further develops the argument from inductive risk. On this account, non-epistemic values can indirectly justify how strictly methods must be chosen, data characterised, and findings interpreted when the consequences of possible errors are taken into account. They do not replace observation and may not substitute a desired finding for an established one.
A strong opposing position defends a sharper separation. Richard Jeffrey proposes that accepting or rejecting a hypothesis should not be treated as a separate scientific act: evidence changes degrees of belief; action decisions additionally consider consequences and preferences. This is a strong separationist position. An organisation can therefore say both: “We do not yet know whether the claim is sufficiently supported” and “Even if it is probably true, the intended use is not defensible because of its consequences.”
Daniel Steel does not take a simple Jeffrey position. He defends the distinction between epistemic and non-epistemic values precisely as a basis for distinguishing legitimate from illegitimate influences of non-epistemic values when dealing with inductive risk. The comparison makes the philosophical tension more precise: on the one hand, the question of whether consequences may help determine the evidence threshold; on the other, the attempt to separate epistemic assessment and practical decision more strongly.
This contrast cannot be resolved with a model form. Rather, it sharpens two distinct levels:
1. Evidence level: What statement follows from the data and methods—and what does not?
2. Action level: Which option is defensible in light of possible consequences, duties, and objectives?
The four-register worksheet does not decide which philosophical position is ultimately correct. It takes the opposing position seriously by documenting evidence quality and the normative action threshold separately. The question of consequences may make transparent why an organisation requires more or less assurance. It may not retroactively change the content of a measurement result. This makes the value conflict visible without reducing every factual question to mere opinion.
Why an overall score is often the wrong shortcut
It is tempting to translate strength of evidence, benefit, risk, and effort into a single score. One value promises comparability and a quick ranking. But if criteria and weightings are not substantively justified, set in advance, and understandable to those affected, the score merely moves the conflict into a formula. A serious objection can then disappear arithmetically behind several small positive points.
An overview can place options side by side, but not every decision needs to result in a ranking. Especially with incomplete or non-comparable data, exclusion conditions are often more helpful: no use without a clarified purpose; no external-facing use without human review; no real dataset while data flows and responsibility remain unclear; stop when a defined error case occurs or oversight no longer functions. These conditions also require justification. They make visible which boundary should not be traded off against an efficiency gain.
A qualitative term such as “low risk” is also insufficient on its own. The organisation must explain whose risk is meant, who bears the consequences, and who can challenge the decision. Where there is disagreement, the dissent belongs in the record rather than disappearing into an average.
Worked example: a fictional continuing-education provider
Starting point
A small continuing-education team handles general enquiries about course content, dates, participation requirements, and organisational procedures. The existing sorting is inconsistent. A language model is discussed as a possible aid. The team has no representative local evaluation, no defined reference classification, and no evidence that use would shorten processing time or improve quality. In a brief classroom simulation, several artificially created examples appear plausible. This demonstration is not a product-performance test: it contains no representative real distribution, no systematic error analysis, and no robust baseline.
Completing the four registers
Empirical evidence. At present, the only established point is that a limited tabletop exercise with synthetic enquiries can be run. It permits questions about roles, sorting rules, and possible handoffs. It does not establish how a particular model would classify real enquiries, how frequently errors would occur, or whether time would be saved. The claim “AI improves our process” remains open.
Normative reasons. No one should systematically receive help later because of unusual phrasing, a lack of technical terms, or a language pattern. Enquiries affecting participation or accessibility may not silently disappear as “routine.” Staff need an understandable way to intervene. Those affected and the possible consequences of errors need closer examination before moving beyond a simulation.
Organisational objective. The team wants to handle enquiries more reliably and reduce repetitive sorting work. The objective is not “use AI.” A consistent manual template can be a plausible first step, because inconsistent rules are already visible as a process problem.
Risk tolerance. A purely process-based exercise with synthetic cases and without calling a real system is reversible. Sorting real enquiries, or automated replies, would be qualitatively different steps. Until purpose, data flow, assessment standard, responsibility, appeal route, and stopping criterion are clarified, no production use or external-facing use will be approved.
Options instead of “AI or standstill”
| Option | What it enables now | What it does not establish or solve | Status in the simulation |
|---|---|---|---|
| A. Standardise manual sorting rules | Address the visible process problem directly; clarify responsibilities and templates | Whether an AI system would additionally provide benefit | Immediately sensible and reversible |
| B. Conduct a synthetic tabletop exercise | Make handoffs, edge cases, missing criteria, and questions for a later review visible | Real model performance, real error rates, or productivity gains | Approved as a learning and governance exercise |
| C. Real-data pilot or automated reply | Could later examine a more realistic use | May not start without a clarified purpose, data/contractual situation, measurement design, and oversight | Deferred; no approval |
The reasoned decision is therefore neither “AI is proven useful” nor “AI may never be used.” The team first standardises manual sorting and conducts a synthetic tabletop exercise. It documents not a supposed accuracy rate, but open working questions: Which enquiry needs human advice? Which category is too broad? Who corrects an incorrect assignment? Which cases must be prioritised regardless of any automation? Such a tabletop exercise is neither a test with real system outputs nor evidence of later accuracy. It can, however, make the question for a future investigation narrower and more testable.
● Members only
Read the full article and download all files with a membership.
Unlock full article + downloads → Subscribe0 comments
● Loading comments…