← SAKIZLI AI
Article25 Sept 2026 · 25 min read1 / 7Free · Public

AI ethics begins before the first prompt

This article develops the editorial thesis that ethical decisions about AI can be made before anyone opens the model. Which problem counts as worth solving, and who may bear the consequences of a proposed solution?

FFurkan SakızlıAI researcher & tutor · independent
A hand places translucent cards in a row: a globe, three blue dots, a tall pane with a dark disc, then blank circles and a sheet with a blue dot
What goes onto the line is decided before the first prompt
Image generated with AI

A small spare-parts business, entirely fictional here, wants to handle delivery and complaint enquiries more quickly. Management proposes that a language model should read incoming messages and send replies straight away. A perfectly worded instruction could control the tone, length and format of those replies. It would still not settle whether a delayed delivery, a complaint and a change of address are the same kind of work at all.

With a change of address, a wrong assignment can lead to the wrong delivery. With a complaint, an overly confident reply can intensify a dispute. A question about delivery status may not need a language model at all, but a more reliable status from the enterprise-resource-planning system. The ethical question is therefore initially not, “How do we prompt the AI?”, but, “Which decision is to be prepared or made here, by whom, and with what information?”

This view is not an invitation to stand still. It creates the basis for a small, bounded trial. The NIST AI Risk Management Framework, Version 1.0 includes, among other things, documented purposes, conditions of use, potential impacts and assumptions in determining context; in governance it describes clear roles. NIST expressly describes the framework as voluntary guidance, not as an approval stamp for an individual business (pp. 2, 22 and 25 of the report). The decision card proposed here is our educational implementation of these ideas, not an official NIST template.

Four questions that can easily blur into one another

An employee considers it wrong for a customer to receive a persuasive but unreviewed rejection. That is a moral judgement: from her perspective, something is at stake. Ethics begins here as reasoned examination: which interests conflict, which argument supports fast service, which supports slower review, and what would an opposing position change? Law asks a different question: which specific duties and rights apply to the proposed process in the relevant jurisdiction? Governance is the organisational answer to who may decide, review, stop and improve.

This distinction is a working proposal, not a complete philosophical taxonomy. An act that is lawful can be unwise or unfair for good reasons. A well-intentioned act can at the same time fail at a specific legal boundary. And an operating instruction that says “a human checks everything” creates no effective oversight if that person has neither time nor information nor authority to intervene.

Choosing the problem distributes benefits and burdens

One might object that a problem definition merely describes what already exists. In a business, it is also a selection. Anyone who makes “response times are too long” the only problem will probably measure minutes. Anyone who adds “incorrect information” must take an interest in a different kind of review. Anyone who includes “inaccessible complaint channels” asks about the people absent from the standard process. None of these descriptions is value-free. Each suggests which cases count as success and which as unwanted deviation.

This is not a claim that there are no objective observations. A waiting time can be observed. What is disputed is which observation should matter to the decision. An average could improve even while rare, consequential misdirections increase. The averaging would be arithmetically correct and still inadequate as the only decision aid. A team therefore needs, alongside a desired benefit, an explicitly named harm threshold. Only then does a goal description become testable without hiding moral questions in a metric.

Two ethical examinations come into tension here. A consequence-oriented perspective asks whether the solution overall reduces waiting time and wrong routes and reaches more people with scarce resources. A rights- and duties-oriented perspective asks which replies must not go out unreviewed even with a good overall result, and who can challenge an error. A fair-distribution perspective adds: are the people who already receive poor service in exceptional cases the very people who bear the new burden? These lenses do not automatically give the same answer. This article uses them as questions for the decision, not as a mathematical vote between philosophical schools.

For the business, this means it is not enough to put a declaration, “we act responsibly,” at the start of a project document. The team must formulate the purpose so that someone can dispute the choice of goal. Why should the fastest first reply of all things be the success criterion? Would a slower but correct path of handling better match the stated concern? Which effects do not count in the metrics but must appear in review? These objections alone can change the assignment before a single line of prompt is written.

This framing has a scholarly parallel: Selbst and co-authors (2019), in their theoretical work on fairness in sociotechnical systems, criticise the error of assessing a social standard only at the model and its inputs/outputs while leaving out human and institutional parts. Applying this to our spare-parts business is an analogy: correctly classifying an invented message would still not prove that the customer process is fair or helpful. The paper supplies neither a measured effect size nor a ready-made method for this small and medium-sized enterprise (SME).

The case contains empirical questions about the process, normative questions about the distribution of error consequences, and institutional questions about responsibility. The NIST framework describes AI systems as sociotechnical: risks and benefits arise from technology and deployment context, operation and societal conditions (NIST, p. 1). The UNESCO Recommendation on the Ethics of Artificial Intelligence requires responsibility for AI-assisted decisions to be assigned to human or organisational actors (paras. 35 and 42). The Recommendation is an international normative text. By itself, it establishes neither the duties that apply in a particular case nor proof that our method works.

“Faster” is not yet a sufficient goal

In the fictional business, the opening question is: “How can we have AI answer all customer enquiries immediately?” It already contains a prior decision: the reply is to go out automatically. “Immediately” also sounds unambiguous while concealing different measures. Does it mean the first acknowledgement, the time to the right department, substantive resolution, or the feeling of being taken seriously?

Here we use a small, expressly limited application of the provided question-expansion model [Fragenerweiterung v5, original PDF pp. 1–3]. Its order is A → B → D → C: first clarify, then develop competing interpretations, then address uncertainty and errors, and finally formulate a revised question. A later article in this series follows the complete model sequence. Here, a single traceable change is enough.

A — Clarify. The team separates three goals: faster routing, fewer incorrect pieces of information, and good accessibility for customers with non-standard concerns. It asks what work is actually taking time at present. Repeatedly asking “why?” suggests, as a testable hypothesis, that missing order numbers and unclear responsibilities are slowing the process. The questioning technique does not prove the cause; the business would need to observe and analyse its processes for that.

B — Alternatives. Three explanations remain open: perhaps good access to status information is mainly missing; perhaps the forms are poor; perhaps preliminary sorting of non-critical messages really helps. A fourth option is a better input form without AI. This comparison rule is our addition to the case, not an explicitly named F step. The decision now no longer rests on an impressive text example, but on competing, testable explanations.

D — Errors and follow-up question. There is still no robust observation of how often messages are assigned wrongly. Data flows and whether a complaint must receive separate treatment are also unresolved. The team can name the gap instead of issuing an apparent “confidence score” for the best solution. Without calibration and a data basis, such a value from a prompt would not be a probability.

C — Revised question. “Can a bounded trial using invented messages show whether AI sorts only non-critical enquiries better for internal review than an improved form, without sending messages itself?” This question is narrower. It defines an alternative, a responsibility, an observable result and a boundary. In closing, the model connects it to a compact GROW action plan: Goal = more reliable internal routing; Reality = wrong routes and the data situation are still unknown; Options = form or bounded AI trial; Way forward = the specialist team sets comparison cases, management names stop authority and documents the evaluation. It does not yet decide that a trial with real customer data is permissible or sensible.

The NIST framework expressly states that responsible actors should assess whether AI is appropriate or necessary for a purpose at all (NIST, p. 13). Our non-AI comparison follows editorially from this; it is not proof that a form is already better in this business.

What actually changes through the new question

The opening question calls for a system for all messages and an action outward. The revised question initially permits only internal pre-sorting of artificially created cases. This difference is operational: a sending project becomes a learning project; “all enquiries” become predefined, non-critical case classes; “fast” becomes a combination of correct routing, time to handling and documented exceptions. The benchmark is a revised input form. The specialist team must not decide only after the test which standard it wants to use for success.

The revision also rules out an obvious fallacy. If a language model produces convincing replies for ten invented examples, it has shown neither the distribution of real concerns nor the consequences of wrong routing. Conversely, a failed short test would not prove that every conceivable AI approach is unsuitable. The small trial is meant to illuminate a narrowly described question. Other data, roles and intervention rights would need separate examination later. The question is better because its falsifiability and limits are clearer, not because model F already guarantees a correct solution.

Those using the question-expansion model in a course can conduct an additional counter-test: half a group receives the role “management, goal: waiting time,” and the other half “customer service, goal: correct treatment of exceptions.” Both work on the same opening question in A, B and D. Only at C do they compare their new questions. The educational evaluation does not seek a uniform formulation; it records which assumption each group saw first, which it overlooked and what additional material it would need for a decision. This makes question expansion tangible as a method of reasoned disagreement, rather than a technique for elegant wording.

Who is missing when management decides alone?

The change of perspective begins with asking who experiences a mistaken decision. A customer may wait longer because her message is classified wrongly. An employee may have to correct automated drafts under time pressure. A person with an unusual concern may not find themselves in the prescribed categories. And management bears costs when a fast process creates new complaints.

The consultation model (K) provides another working framework for this [original PDF pp. 7–8; detailed case scenarios pp. 20–33]. Its basic methodological cycle has four steps: exploration, reflection and analysis, decision-making/action recommendations, and feedback/evaluation. In later case sequences, a separate implementation/pilot phase and a final evaluation follow the development of recommendations. This article uses only the first movement. In the original, exploration is chiefly inquiry and data gathering. Our educational affected-persons card adds: who is affected, from which position is the problem described, and whose objection would change the opening question? The full governance and pilot work that follows K belongs to later articles in this series.

Perspectives cannot be replaced by a single questionnaire. A business should also ask who can raise an objection safely. Management may prioritise speed; customer service sees the extra work caused by wrongly assigned cases; customers need a way to report errors. A bounded pilot can make these differences visible but does not guarantee agreement. The UNESCO Recommendation addresses its statements on ethical impact assessments mainly to states and public authorities and calls for pluralistic and inclusive perspectives (paras. 50–53). It does not prescribe a particular consultation process for this fictional business; the route outlined here is our proposal.

There is also an uncomfortable counter-question: can an organisation shift responsibility onto the person at the screen although that person could hardly influence the technical and organisational design? In her original article “Moral Crumple Zones”, researcher Madeleine Clare Elish describes this risk in complex automated systems. Her argument supplies no measured error rate for our SME example. It supports asking whether a mere signature below an AI output would already be a fair allocation of responsibility in this case.

Participation intervenes in the plan

In consultation, it would be easy to list four roles and then continue the same pilot unchanged. That would fulfil the form and miss the point. Suppose customer service points out that complaints often look linguistically like simple status questions but can require a binding commitment. Then the test boundary must change: a suspected complaint goes into a separate manual class; the pre-sorting may not treat it as a routine case for an automatic reply. If, instead, an affected person reported that the new input form was hardly usable with their information, the non-AI alternative itself would have to be redesigned. Both objections can overturn the first recommendation. Without a documented change, “participation” is only an attendance list.

For a small business, consultation can be brief: one person from processing, one person with specialist responsibility and one deliberately sought exception case are enough as a start to exploration, not as representative consultation of everyone affected. The record must show which objection led to which change and which remained open. If there is no safe way to raise an objection or revise a decision, participation is recorded as a gap. Management may reject a dissenting recommendation with reasons; it must not erase dissent from the record.

A possible objection must also have a reachable addressee. For decisions with possible effects on safety, rights or freedoms, the UNESCO Recommendation links transparency with the possibility of effectively challenging, reviewing and, where appropriate, correcting them (paras. 37–38). The OECD Recommendation on AI points to information for adversely affected people so they can challenge an output and to context-appropriate mechanisms to override, repair or decommission (§§ 1.3–1.4). For this article, this yields an independent design question: whom could a person contact after a misdirection, and who could actually remedy the cause? Neither recommendation establishes an automatically applicable right of complaint for the fictional business; that would need separate examination for the specific jurisdiction.

This is where the boundary between the three models becomes visible. F changes the question. The K-informed affected-persons card examines perspectives and can thereby change the trial boundary; organisational governance follows on from it. The separate advisory model R would structure the specialist review path when a concrete productive data flow is involved. None of the three models alone implies approval. The handover is precisely the lesson: anyone who jumps directly from F to a test with real messages skips the perspectives and specialist questions that the new question has made visible.

Four connected cards on a light background: a ring with a centre point, three people, a network next to sliders, and a person next to a pause switch
Question, people affected, system and stop authority — four cards that carry a decision only when connected
Image generated with AI

A serious counterargument: too much clarification also has a cost

A small team cannot treat every internal activity like a public impact assessment. If every test must go through a lengthy committee, useful improvements remain undone. Especially in a business with limited resources, faster service can itself be an ethically relevant good: less waiting, less repetitive work, more time for difficult cases.

The counterargument holds. It justifies a proportionate process: first a narrow question, a short comparison with the existing process, a clearly bounded trial with no external effects, and a stop signal for unexpected harm. It does not justify assuming that an attractive demo output has already understood the business, the people affected or the responsibilities. The NIST framework connects risk management to context and risk tolerance (including MAP 1 and GOVERN 1). We infer from this that the size of a company alone does not replace this contextual assessment.

Proportionality needs a boundary, not only speed

An elaborate investigation is not always the best first step. In the example, a short initial clarification meeting that records the affected case classes and stop authority is enough at first. The team can then create invented messages and compare two processes. This proportion holds only because no real customer is addressed and no real message is processed. Moving to production data or external communication is a new decision. This is so even if technical pre-sorting succeeded in the exercise case.

A fair examination records effort on both sides: how much time does it take to formulate the question and document exceptions? How much rework does a wrong routing create? Management could rightly say that an elaborate decision apparatus makes a small internal test uneconomic. An employee could with equal justification object that savings on the first reply are bought with additional correction work. Both are possibilities to examine, not facts collected here. Even a simple record sheet can make both kinds of effort visible.

In practice, three decision thresholds can be distinguished. Threshold 1: Is the problem definition clear enough to learn from artificial cases? Threshold 2: For a real pilot, have purposes, roles, data flows, paths for affected people and a comparison standard been gathered, and have missing facts and necessary specialist reviews been documented as an assignment for R Phase 1? Threshold 3: Do actual observations justify a limited external effect or broader operation? The first yes does not replace the two later yeses. Keeping these thresholds separate in the project record makes it easier to see where a demo success is silently reinterpreted as production approval.

A decision card instead of a wall of values

For the fictional business, the first meeting can be documented on one page:

FieldCompleted working version for the invented case
DecisionMay messages be pre-sorted internally? Automatic sending remains outside the test.
Purpose and measureCorrect routing, waiting time until handling by the responsible function, and traceable wrong routes; compare with the existing process in advance. No improvement is claimed without measurement.
GROW handoverG: more reliable internal routing; R: baseline and error routes remain open; O: form or bounded AI trial; W: specialist team prepares invented comparison cases, management names stop authority and documents the result.
People affectedCustomers with standard and exception cases, customer service, specialist team and management; a reachable channel for feedback.
AlternativeA better input form and clearer responsibility rules without a language model.
Material and boundaryAt first, only newly invented test messages; no real customer data and no external message.
Responsible peopleManagement decides on the bounded test; the specialist team assesses routing; customer service documents exception cases; one named person may stop it.
Stop signalA message leaves the test without authorisation, a real message enters it, or a consequential exception case is presented as routine: pause the test and examine the incident.
Open mattersBaseline, contractual and data-protection questions before any possible test with real data, participation of affected people, approval for a later pilot.

This is not a pilot that has been carried out. The card separates a first learning exercise from a production decision. Once real data, a service provider or actual customer communication is seriously contemplated, the separate five-phase advisory model (R) begins with its first phase: preparation/anamnesis of the facts [original PDF pp. 1–3, 20]. At that stage, the team first gathers roles, data flows, contracts, jurisdiction and timing; what is missing is documented as a review assignment. Only then can analysis, solution development, implementation/documentation and evaluation/follow-up proceed. The model supplies no blanket legal advice here; suitable specialists must be involved for the concrete expert assessment.

From the card to a testable mini-trial

A fillable card helps only when it protects decisions against later glossing over. For a teaching exercise, the team can design twelve wholly invented messages in three classes before the test: four unambiguous status questions, four enquiries with missing information, and four cases with a plausible complaint or exception signal. The numbers describe only the exercise design, not a real case distribution. Beforehand, participants specify for each message which specialist class and treatment a responsible human would expect. Disagreement about the intended treatment is recorded as an open dispute case, not resolved afterwards in favour of a system.

Two teams then receive the same set of cases. Team A works with the improved input form and clear responsibility rules. Team B receives internal AI pre-sorting only; its suggestions are initially not sent and are reviewed before every use. For each message, the proposed class, human correction, time spent and reason for escalation are documented. The order can alternate between teams so that an obvious learning effect from using the same cases is visible. A joint evaluation discussion looks at correct routing and missed exception cases; a simple success rate would be too coarse for the stop signal.

Before the first run, the team decides which errors interrupt the trial. Treating a complaint as a simple status question although its text requires a decision or commitment is a serious boundary case in the exercise design. An accidental sending or a real data set immediately triggers a pause. A single exercise error, by contrast, would not automatically prove against every form of AI; it first prompts an inquiry into cause: was the class unclear, the instruction misleading, the case unsuitable, or the system actually unreliable? Each explanation would require a different revision.

The trial arrangement proves little about everyday operations. Twelve invented messages do not represent a real customer population, and the people handling them know the exercise. This very limitation protects against a false approval. The learning gain is to formulate the later review decision precisely: which additional cases would be needed, how would exceptions be recorded, who could introduce real data at all, and which specialist review is still open before that step? If there is no clear advantage over the form, the non-AI solution remains seriously in contention.

A teaching block with visible revision

For a course, a 90-minute session can run as follows. First, learners spend five minutes writing their spontaneous prompt idea and marking the deployment decision already made within it. Pairs then work through F-A/B/D with a competing hypothesis and an open possibility of refutation. A second group takes an overlooked perspective for the K-informed affected-persons card and tries to change the test boundary. Only then does the first group formulate F-C/GROW as the revised question and next step. In plenary, the prize is not for the most elegant wording but for the demonstrably changed decision. Finally, groups complete the card and write a sentence they could explain factually to a sceptical SME team.

A deliberately uncomfortable case serves as the checkpoint: pre-sorting speeds up the clear status questions but wrongly assigns an unusually worded complaint case. One group must apply the pre-set stop rule; the other must formulate the strongest fair argument for another, narrower trial. Both must state which claim from the artificial cases is permissible and which is not. A good submission can conclude “no AI for now.” It can likewise propose a new, clearly bounded learning exercise, provided data, responsibility, comparison and the possibility of revision are described.

What was actually decided in the end

Three kinds of claim, three kinds of examination

The small trial separates claims that are easily mixed in AI projects. Empirically, the question is whether, in a specified set of cases, one variant routes more often correctly and what rework results. This requires defined cases, observations and comparison rules. Normatively, the question is whether a gain in clear cases can justify a possible disadvantage for people with rare concerns. This requires reasons, perspectives of affected people and a named boundary; measurement alone does not decide the value conflict. Institutionally, the question is who approves the trial, who can object, who acts when there is an error, and whether the responsible person actually has access, time and authority. This requires roles and verifiable paths of intervention.

This separation prevents two mirror-image shortcuts. Anyone who measures a small time gain has not yet proved an ethical justification. Anyone who convincingly defends a value such as fairness does not thereby know how often an error occurs. And even agreement on goals does not guarantee that someone in the business receives the rights needed for correction. A careful team keeps the three questions alongside one another and joins them only in a reasoned decision.

This approach also matters for later agentic systems. An agent could suggest the case class, complete the card or request missing information. It should not invent a missing baseline, mark an open value conflict as “resolved,” or infer real approval from a role field. For a later workflow, a permitted output would therefore be: “Facts incomplete; this observation and this specialist approval are missing; next human review step.” This restraint is a useful result when an apparently finished output would blur the limits of the test.

A decision against AI also needs documentation. If the new form and clear responsibilities are sufficient in the learning exercise, management can suspend the language-model path with reasons. If both variants show weaknesses, the next action may be another question: why do exceptions become visible only so late? The original request for a prompt has then at least exposed an organisational problem. The result of the ethics process is not a mandatory “AI yes/no” field, but a more robust description of what may responsibly be examined next.

After this first pass, the fictional team has not introduced AI. It has created a narrower subject for examination: internal pre-sorting using artificially created material, compared with a non-AI alternative; external communication and real customer data remain excluded for now. This makes clearer which observation could support or refute a later decision.

The thesis developed in this article is: ethics begins before the prompt because the problem definition already determines what counts as success, whose effort becomes visible and which error seems acceptable. A good prompt can support a carefully chosen task. The task itself needs reasons, counterarguments, perspectives of affected people and a way back when an assumption proves false.

Continuing the work

The completed and blank decision card develops the case further for SMEs and teaching. A later article in this series examines actual responsibility in AI-supported preparatory work; another presents the question-expansion model in full; another addresses the consultation model including governance; another designs the bounded pilot. These extensions are planned as further articles.

Sources and limits of application

NIST (2023), Artificial Intelligence Risk Management Framework (AI RMF 1.0), especially pp. 1–2, 13 and 21–25. Voluntary framework; the decision card is an independent educational derivation.

UNESCO (2021), Recommendation on the Ethics of Artificial Intelligence, especially paras. 35, 42 and 50–53. International recommendation, not a case-specific legal assessment.

Madeleine Clare Elish (2019), “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction”, Engaging Science, Technology, and Society 5, DOI: 10.17351/ests2019.260. Conceptual and case-analytic critique; no evidence of effectiveness for this teaching example.

Andrew D. Selbst et al. (2019), “Fairness and Abstraction in Sociotechnical Systems”, FAT ’19, DOI: 10.1145/3287560.3287598, especially section 2.1. Conceptual critique of overly narrow system boundaries; its case world is not our SME.

UNESCO (2021), paras. 37–38, and OECD (2019, amended 2024), Recommendation of the Council on Artificial Intelligence, §§ 1.3–1.4. Normative guidance on challenge and correction; no blanket business-law advice.

Three German method PDFs provided by the author: Fragenerweiterung v5 (A/B/D/C, pp. 1–3); Konsultationsmodell (four-step methodological cycle, pp. 7–8; scenarios, pp. 20–33); 5-Phasen-Beratungsmodell (pp. 1–3). They establish the form of the models, not their empirical effectiveness or the legal position.

HTMLTopic overview: AI ethics begins before the first prompt1 page↓DOCXWorksheet: A decision before the first prompt1 page↓

0 comments

● Loading comments…

Sign in to comment · become a member →