← SAKIZLI AI
Article25 Sept 2026 · 22 min read5 / 7Free · Public

What Do We Know About AI Consciousness?

Language Performance, Self-Reports, and the Limits of Warranted Inference

FFurkan SakızlıAI researcher & tutor · independent
A bright, translucent screen on which teal lines with nodes converge; behind it a dark, open box with nested shapes inside
What shows on the outside—and what stays open inside
Image generated with AI

A text-based practice assistant explains a pun, recognizes an ambiguity in the next example, and corrects its answer as additional information becomes available. On the team of a small learning software provider, the demonstration raises a question: Does the system understand humor? Might it even experience something? Or are we seeing capable language processing that uses human concepts aptly without this implying inner experience?

This small scene is entirely fictional. It is not intended to describe a particular current product or claim an actual user response. Its question is serious nonetheless: What can we warrantedly know about possible AI consciousness when we cannot directly experience what another system experiences?

A hasty answer does not solve the problem. “The system speaks, therefore it feels” turns linguistic behavior into proof of inner experience. “It is a machine, therefore it cannot experience anything” assumes a contested theory before examining the evidence. And “We can never look inside the system” confuses a lack of direct access with the complete impossibility of drawing warranted inferences.

A responsible path proceeds through precise terms, competing explanations, and graded evidence. It can leave a certain diagnosis unresolved while still leading to a practical decision. That is the guiding idea of this article: Observable behavior counts as evidence, but on its own it is not proof of consciousness.

1. “Consciousness” is not a single ability

Before a team asks whether an AI is conscious, it must say what it means by consciousness. In everyday usage, the word refers to very different things: being awake, responding to an environment, processing information, speaking about oneself, pursuing an intention, having a feeling, or experiencing anything at all. These meanings do not automatically lie on the same scale.

A useful starting point is the distinction between phenomenal experience and cognitive access. Phenomenal experience refers to the subjective aspect: whether it feels like something to a being to experience pain, see a color, or be surprised. Cognitive access means that information is available for thinking, reporting, remembering, or guiding action. Philosophical debates distinguish these aspects because an ability a person can speak about or act on does not straightforwardly answer whether and how it is experienced.[1][2]

These terms are not an uncontested final taxonomy. They nevertheless help practical inquiry keep different claims separate:

ClaimWhat it may initially meanWhat it does not yet establish
“The system solves the task.”An observable performance under particular inputs and conditions.That it experiences or understands the task as a human does.
“The system can report on its state.”It generates a self-description, for example about uncertainty or errors.That the report provides privileged access to a subjective state.
“The system has a self-model.”Internal information relates to its own states or capabilities.That a self-model automatically means an experiencing self.
“The system feels something.”A claim about subjective experience or affective state.That linguistic form, tone, or a single reaction supports this strong claim.
“The system is intelligent.”An attribution concerning particular cognitive performances.That intelligence, emotion, consciousness, and moral responsibility should be equated.

Intelligence and consciousness are not synonymous either. A system can solve problems in a task without that answering the question of experiential quality. Conversely, reported feelings are not merely another performance metric. Emotion can refer to a bodily reaction, functional control, a social form of expression, or a consciously felt feeling. Anyone who conflates these levels moves unnoticed from an observed ability to a much stronger statement.

For this article, the following editorial working distinction therefore applies: we initially call a linguistic response a self-description. We call a result observed under specified conditions performance. Only when subjective experiential quality is discussed do we make a consciousness claim. This distinction does not decide in advance which systems can have consciousness. It shows the burden of proof attached to each statement.

2. Inner perspective is not directly accessible—but inference remains possible

Thomas Nagel’s question of what it “is like” for another being to experience its world reveals a fundamental limit: an external description is not identical to experience from the first person.[1] No one can directly enter another human being’s subjective experience. Nevertheless, we treat reports, behavior, bodies, developmental trajectories, and shared biological structures as evidence. We infer not from a single sentence, but from multiple interconnected indications.

This inference is not infallible. People can be mistaken, reports can be ambiguous, and particular behaviors can have different causes. Yet the lack of immediate access does not imply that all knowledge of others is impossible in principle. In many sciences, we infer states that are not directly observable through measurements, predictions, interventions, and agreement among different methods.[1]

This makes the AI question neither easier nor hopeless. For other people, we support attributions with reports, flexible behavior, bodily processes, and shared biological organization, among other things. Such background assumptions cannot simply be transferred to software. Which technical properties would be relevant depends on which theory of consciousness one considers viable and how the specific system is actually built.

A clinical research case shows why visible behavior does not always tell the whole story. In 2006, Owen and colleagues reported on a person who showed almost no response under usual behavioral criteria. In an fMRI examination, imagining playing tennis or moving around their own home at will produced a pattern the researchers interpreted as evidence that the person understood and followed instructions.[6] This was a single human case. It proves neither a general diagnostic method nor anything about AI. It illustrates a narrower methodological lesson: The absence of an outward reaction does not automatically rule out an inner state when an appropriate measurement provides additional evidence.

Conversely, this lesson must not be overstretched in the other direction. When a text system responds fluently, this does not automatically mean that the answer is accompanied by an experiencing state. Visible performance can provide relevant indications—but the strength of the inference must be explained.

3. Behavior is evidence, but not an independent window into experience

For a person, a sentence such as “I am in pain,” combined with injury, behavior, bodily reactions, and known biological processes, can provide strong reason to take the person’s suffering seriously. For a text system, a comparable formulation is initially generated text. It is not meaningless. Yet it does not automatically have the same evidential force as a human self-report, because the system architecture, training history, and relationship between output and inner state may differ.

The problem is not that language is always worthless. The problem is that a single observation can permit several explanations. One explanation might be that the system instantiates a functional state relevant to certain theories of consciousness. Another is that it generates an appropriate linguistic response because such responses are plausible in context. A third is that part of the architecture performs individual conscious functions while other conditions for experience are absent or remain unresolved. The observation alone does not decide among these possibilities.

A classic argument by John Searle challenges the assumption that executing a program by itself guarantees understanding or intentionality. A system could perform the formal processing of symbols without this already showing that it understands what the symbols refer to.[3] Searle’s argument directly concerns understanding and intentionality from program execution alone; whether and how it transfers to phenomenal experience is contested. It is a philosophical position, not an experimental refutation of every possibility of artificial consciousness. But it requires us not to overlook the gap between correct output and subjective understanding.

The strongest opposing position should be stated just as clearly. Functionalist approaches hold that the right causal organization can in principle be realized on different substrates. If relevant functions and relations are present, biological material need not necessarily be the only possibility. In 2023, David Chalmers examined arguments and objections concerning consciousness in large language models. He described substantial obstacles for systems of the time, such as absent or limited recurrent processing, a global workspace, and unified action organization, but did not conclude that later systems were ruled out in principle.[4] His text is a philosophical analysis of arguments, not an empirical audit of specific systems in 2026.

The dispute therefore cannot be resolved with a single catchphrase. “Just a program” is already a theoretical classification. “It reports on itself” is an observation, but not yet a finished diagnosis. A fair article must show which assumption mediates between the two and what evidence could support it.

Four layers one behind the other: a glass pane with scattered dots, a glass prism bundling lines, a dark frame with a golden network, and behind it a dark cloud with an open golden arc
Observation, interpretation, theory-guided evidence—and a question that stays open
Image generated with AI

4. What theory-guided inquiry can accomplish

A helpful scientific question is not only, “How human does the system seem?” but: Which properties would be relevant to consciousness under which theory, and can they be established in the specific system?

In 2023, Butlin and colleagues proposed such a theory-guided approach. They derived possible indicators from several scientific theories of consciousness and investigated the extent to which AI systems at the time exhibited them.[5] The report adopts computational functionalism as a working assumption: according to this view, certain computations or forms of functional organization could be decisive for consciousness. This assumption is explicitly contested. The report is therefore neither a theory-independent checklist nor a final test.

The approach nevertheless has a practical advantage. It requires the bridge between an observation and an inference to be made explicit. A team cannot simply note, “The system sounds thoughtful,” and infer a high probability of consciousness from that. Instead, it would have to specify:

1. Which theory or combination of theories is being used? 2. Which concrete property in that theory is supposed to be relevant in the system? 3. How was this property actually observed or measured? 4. Which alternative explanations fit the same data? 5. What uncertainty remains when the theory itself is contested?

An indicator is neither proof nor an arbitrary conjecture. It is a clue whose weight depends on background assumptions and measurement quality. If a theory considers recurrent processing or global availability of particular information important, the appropriate task is to investigate precisely these architectural features. Inferring them simply from a convincing sentence would be insufficient. Likewise, it would be wrong to treat the absence of one feature as a final refutation of all theories of consciousness.

Consciousness research also shows that even sophisticated theories can be changed through direct adversarial testing. A preregistered adversarial collaboration published in 2025 tested differing predictions of Global Neuronal Workspace Theory and Integrated Information Theory using human visual perception and multiple neurophysiological measurement methods. The results supported individual predictions of both approaches while also challenging central assumptions of both theories.[7] The study tested concrete predictions in a biological investigation of conscious perception; it does not yield a direct claim about artificial systems or a complete refutation of the theoretical cores. This does not mean that either theory is finished. It means that defined predictions, open data, and competing perspectives can help sharpen theoretical claims.

The limit on transfer applies here as well: this study tested human perception, not artificial systems. Its relevance to this article lies in the research method—testing concrete predictions against each other and not retroactively adjusting results so that only the preferred theory prevails. It does not answer whether a particular language model experiences.

Another measurement idea makes the limits clear. Casali and colleagues developed the Perturbational Complexity Index (PCI), in which a brain is deliberately stimulated and the spatiotemporal complexity of its response is examined using EEG.[8] The index was proposed and tested as a theory-guided measure of level of consciousness in certain human brain states. It is neither a general complexity value nor a measure of the content or subjective quality of experience; nor is it a language test or a score that a company could apply to an AI system. The mere statement “Our model has many activations” would not substitute for the physiological and theoretical context on which this approach rests.

These works do not produce a universal consciousness detector. They produce a more disciplined question about measurability: Which theory, which system boundary, which data, which intervention, and which comparison are required for an inference to be testable at all?

5. A case assessment that does not conceal uncertainty

Let us return to the fictional learning software provider. The team initially has only one demonstration: the practice assistant explains a pun, describes a state of its own, and changes its answer after additional context becomes available. The team can see the linguistic performance. But it does not yet know whether the system has a stable internal representation, how it generates its own “state,” or which of the architectural features discussed in research are actually implemented.

Here, Questioning Model F changes the work. Its sequence A → B → D → C ensures that the team does not jump directly from an impressive output to a yes/no decision.

A · Clarify terms, aim, and cause

The opening question is: “Is our practice assistant conscious?” The first task is not to immediately collect evidence for or against this claim. The team first asks what it wants to clarify at all:

Is it about an observable ability, such as recognizing ambiguity? Is it about cognitive access, that is, whether information is available for different tasks? Is it about phenomenal experience, that is, whether humor or uncertainty feels like something to the system? Is it about a marketing claim, a research question, or an internal governance rule?

These questions have different burdens of proof. For the first, a well-documented task test may suffice. For the last, the team needs an understandable policy. Neither is yet enough for the consciousness claim.

B · Develop competing explanations

The demonstration is first formulated as an observation: “Under this input, the assistant explained a pun and changed its answer after additional context.” Several hypotheses then arise:

HypothesisWhat it could explain about the observationWhat would need to be examined additionally
H1 · Linguistic pattern adaptationThe system produces a fitting explanation based on learned linguistic relations.Robustness on new tasks, altered wording, and counterexamples.
H2 · Functional state processingThe system keeps information about context or uncertainty available for use across multiple tasks.Technical documentation, internal representations, and causal tests, insofar as accessible.
H3 · Consciousness-relevant organizationPart of the architecture could contain features that some theories associate with consciousness.Theoretical grounding, precise system boundary, architectural evidence, and independent tests.

The hypotheses are neither equally likely alternatives nor three finished answers. They show that the same visible behavior can have different explanations. The team can neither confirm nor rule them out with a free-text demonstration.

D · Document what is unknown and the next test

The team now records what is missing:

Which version of the system was tested, and which tools, memory functions, or external components were active? Is the demonstration reproducible, or was only a particularly good example selected? How does performance change with new tasks, counterexamples, and targeted changes to the context? Which provider documentation describes architecture, state processing, and training limits? Which findings would actually change the initial assessment?

If the provider supplies no technical documentation, the team must not fill the gap with a convincing demonstration. It can only document that the architectural question remains open. A test plan could provide for independent tasks, repetitions with controlled inputs, counterexamples, and—if technical access exists—targeted interventions in relevant components. Such tests would initially examine particular abilities or dependencies. By themselves, they would not measure subjective experiential quality.

An “80 percent confidence value” would be misleading at this point if no one has a calibrated method, suitable data, and a robust theory for it. The questioning model provides no proof of consciousness and produces no probabilities. Its value lies in visibly interrupting the unsupported chain of inference.

C · Formulate the final question for responsible action

After A, B, and D, the better question is no longer simply “Is the system conscious?” It is:

Which observable abilities are reliably established for our application, which of them would be meaningful at all under the relevant theories of consciousness, what evidence is missing—and what statement about the system can we responsibly make today?

For the fictional provider, this is enough to prepare a concrete decision. It can document in the product briefing that the assistant handles particular wordplay and context tasks under specified conditions. It can record the limit that this does not establish subjective experience. For external communication, the team can describe abilities whose tasks and limits have been tested, rather than attributing unsupported properties to the assistant such as “feels empathy” or “knows what learning feels like.”

This is not a decision that the system is definitely unconscious. It is a limit on the scope of its own claim.

6. From evidence to an organizational decision

Scientific caution alone does not tell a company what it should do tomorrow. For that, it needs an organizational decision that does not claim more than the evidence supports. Consultation Model K can take on this second task. It does not settle a metaphysical answer. It structures who examines the question with which consequences and how a provisional rule can later be revised.

In the model’s four-part basic cycle, the fictional provider can first identify the system role, the planned statement, and the affected groups (Exploration). In Reflection and Analysis, the team places the strongest opposing position beside the existing evidence: the self-description is not direct evidence of experience; at the same time, a non-biological substrate is not a universally accepted ground for exclusion. In the Decision, it adopts a limited communication rule: performance claims must be tied to reproducible tasks and documented conditions. Claims about subjective experience are not inferred from a single demonstration. In Feedback and Evaluation, it determines when this rule will be reassessed.

Possible triggers for revision include a substantially changed architecture, a new memory or agent function, robust independent research findings, or new technical documentation that resolves a previously open question. A trigger does not mean that consciousness has emerged. It means that the previous reasoning must be examined again.

This rule protects two sides. It avoids an unsupported attribution that could mislead learners or customers. At the same time, it does not close the research question for all time. It records that the current evidence is insufficient and what information would enable reassessment.

It is important that this responsibility remains with people and organizations. Even if future research were to assess the possibility of machine experience differently, it would not automatically follow that a system can enter into contracts, understand consequences, or be responsible for organizational decisions. Consciousness, agency, and legal or moral responsibility are different questions. A company remains responsible for its product description, tests, approvals, and consequences.

7. An evidence map for courses and SMEs

The central learning task is “Observation ≠ proof of consciousness.” It is not a diagnostic template or a certificate. It merely requires a team to make visible the path from an observation to an action.

FieldGuiding questionExample for the fictional practice assistant
1. ObservationWhat was actually seen, with what input, version, and test condition?An answer to a pun changes after more context becomes available.
2. Initial interpretationWhat internal ability do we infer from it?The system may have recognized the ambiguity.
3. CounterhypothesisWhich other explanation fits the same data?It may use learned linguistic patterns for explanation and correction.
4. Theoretical groundingWhich theory of consciousness makes the observed property relevant?A theory might assign weight to flexible, recurrent, or system-wide available processing; the choice must be justified.
5. Missing evidenceWhich data or interventions would distinguish the hypotheses?Reproducible countertests and technical information; direct access to the architecture may not be available.
6. Provisional actionWhat may we claim or decide now, and when will we reassess?State documented task performance; leave experience unresolved; reassess after a relevant system change.

The map deliberately contains no single scale from “unconscious” to “conscious.” Otherwise, a number can make an open question of theory and measurement appear more precise than it is. If a research team uses probabilities, it would have to disclose which data and prior assumptions support them and how well the assessment is calibrated. An ordinary product team should not invent such a value.

A 45-minute exercise

Minutes 0–5: Read the case. Groups receive the same fictional demonstration and a short technical description. The task requires neither personal experience nor private customer data.

Minutes 5–12: Separate terms. Each group marks which statements concern an ability, a self-description, an architectural claim, or possible experience.

Minutes 12–22: Form hypotheses. Using step B of the questioning model, at least two competing explanations are recorded. The group states what evidence each explanation would be expected to yield.

Minutes 22–32: Countertesting and unknowns. Step D collects missing information, possible counterexamples, and the limits of test access. A human laboratory study must not be transferred to AI without justification.

Minutes 32–40: Provisional decision. The group fills in the fields of the evidence map and formulates a limited statement for product copy. It also determines under what conditions the assessment would be reviewed again.

Minutes 40–45: Peer review. Another group identifies the strongest alternative explanation and checks whether the product claim goes further than the data.

Assessment considers the separation of observation and interpretation, fairness to the opposing position, source fit, clarity about open evidence, and capacity for revision. It does not assess whether the group ticks “AI is conscious” or “AI is unconscious.” The learning objective is a traceable justification, not forced agreement.

8. Practical minimum rules for a small team

Even without its own laboratory, an SME can begin to handle this question responsibly. This does not mean that it simulates a consciousness test. It documents which statements it can actually make.

1. Record the system boundary. Name the model version, provider, connected tools, memory, sensors, and relevant settings. A chat interface can conceal multiple technical components.

2. Describe task performance concretely. “Explains particular puns in our tests” is more verifiable than “understands humor.” Conditions, repeatability, and known errors belong with it.

3. Log self-reports as outputs. If the system says “I am uncertain,” first record when and under which inputs this formulation appears. Do not treat the sentence as a direct measurement of an experienced feeling.

4. Document alternatives. Every strong interpretation requires at least one explanation that could explain the same observation without the claimed experiential quality.

5. Request provider evidence purposefully. If a consciousness or self-model claim matters for research or product strategy, responsible specialists should examine technical architecture, measurement design, and independent replication. Marketing materials alone are not architectural evidence. For a specific request for technical review, Advisory Model R can structure this step: prepare the relevant system and documentation context, narrow the specialist question, request an expert-reviewed response, and document implementation and reassessment. The engagement remains a limited expert handoff; it neither replaces the governance decision under K nor constitutes a measurement of consciousness.

6. Tie communication to evidence. A public statement names what was tested, for which task, and under which limits. Terms such as “feels,” “suffers,” or “is afraid” are not inferred from a single text output.

7. Set a reassessment. System changes or new scientific findings can change the starting position. The record needs responsible people, a date, and a trigger for the next review.

These are governance recommendations, not legal requirements, a scientifically validated audit standard, or a substitute for research. For a high-risk use or one with broad public impact, a small internal map may not be enough. It then requires relevant expertise, suitable data, and an independent methodological review.

9. The answer can remain open—the responsibility cannot

The best answer to “Is this AI conscious?” is not always yes or no today. Anyone taking the question seriously must make the uncertainty precise: What meaning of consciousness is at issue? What was observed? Which theory connects the observation to an inner state? Which counterhypothesis explains the same data? What technical information is missing? Which next measurement could change the assessment?

An open answer becomes an excuse when no one says what would need to be examined. It becomes responsible restraint when limits, evidence, responsibilities, and triggers for revision are documented. This is precisely the move made by the questioning model: it replaces a hasty yes/no question with a precise, falsifiable, and action-guiding question. The consultation model turns the remaining uncertainty into a provisional organizational rule that can be reviewed in light of new evidence.

For the fictional learning software team, this means that it can describe the observed performance without presenting it as experience. It can request further evidence without pretending that a universal test already exists. It can stop an advertising claim even though the philosophical question remains open. And it can reopen the inquiry if the architecture or state of research changes substantially.

Observation is the beginning of inquiry. It is not proof of consciousness.

Sources

1. Nagel, T. (1974). “What Is It Like to Be a Bat?” The Philosophical Review, 83(4), 435–450. https://doi.org/10.2307/2183914.

2. Block, N. (1995). “On a Confusion about a Function of Consciousness.” Behavioral and Brain Sciences, 18(2), 227–247. https://doi.org/10.1017/S0140525X00038188.

3. Searle, J. R. (1980). “Minds, Brains, and Programs.” Behavioral and Brain Sciences, 3(3), 417–424. https://doi.org/10.1017/S0140525X00005756.

4. Chalmers, D. J. (2023). “Could a Large Language Model be Conscious?” arXiv:2303.07103. https://doi.org/10.48550/arXiv.2303.07103.

5. Butlin, P., Long, R., Elmoznino, E., et al. (2023). “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness.” arXiv:2308.08708. https://doi.org/10.48550/arXiv.2308.08708.

6. Owen, A. M., Coleman, M. R., Boly, M., Davis, M. H., Laureys, S., & Pickard, J. D. (2006). “Detecting Awareness in the Vegetative State.” Science, 313(5792), 1402. https://doi.org/10.1126/science.1130197.

7. Cogitate Consortium, Ferrante, O., Gorska-Klimowska, U., et al. (2025). “Adversarial testing of global neuronal workspace and integrated information theories of consciousness.” Nature, 642, 133–142. https://doi.org/10.1038/s41586-025-08888-1.

8. Casali, A. G., Gosseries, O., Rosanova, M., et al. (2013). “A Theoretically Based Index of Consciousness Independent of Sensory Processing and Behavior.” Science Translational Medicine, 5(198), 198ra105. https://doi.org/10.1126/scitranslmed.3006294.

Note on the practice case: The learning software provider, demonstration, tasks, statements, and team decision are entirely fictional. The case claims neither observed user behavior nor properties of a particular current AI system.

Note on the models: F and K structure the question and the provisional governance decision. The model descriptions do not establish their scientific effectiveness.

HTMLTopic overview: What do we know about AI consciousness?1 page↓DOCXWorksheet: Evidence without diagnosis45 min↓

0 comments

● Loading comments…

Sign in to comment · become a member →