← SAKIZLI AI
Article25 Sept 2026 · 26 min read9 / 9Members · Subscription

When Two Models Disagree

A comparison is a test design, not a vote on truth

FFurkan SakızlıAI researcher & tutor · independent
Two stacked paper cards whose lines converge in a shared point in the middle; from there a branch leads to a smaller card and on to an empty circle
Two answers, one checkpoint—and a path that may stay open
Image generated with AI

Two AI responses sit side by side. Both are fluently written, both propose a next step, both seem reasonable. Yet their recommendations exclude each other. The first promises that a second model reliably confirms content; the second requires differences to be checked first against an original source. Anyone who now asks which model is “right” is asking an understandable, but still overly broad, question.

A contradiction between models is first of all a reason to investigate. It may indicate a factual error. It may arise from an unclear task, different unstated assumptions, or a genuine trade-off in values. Nor does consensus reliably settle the matter: two systems may overlook the same gap or repeat a shared, false assumption. The number of matching outputs does not turn a sentence into knowledge.

This article therefore presents a bounded comparison design. It holds the task, provided information, output schema, and assessment standard constant. It separates checking facts from assessing style and evaluations. The system identity remains hidden during the initial blind assessment; presentation order is documented and deliberately swapped when a model judge is used. It also records when the design itself must be changed. The aim is neither a ranking nor a general statement about providers. A real run could provide only a traceable finding about a narrowly described task.

The ongoing practical case is entirely fictional. It describes neither a real company nor an actual model output, person, or teaching situation.

1. Why the dispute is philosophical

At first glance, a model comparison seems technical: two inputs, two responses, one table. Its core is epistemological. We must decide what kind of reason a response can provide at all.

A response can be an indication. It can propose an explanation, draw attention to a source, or make an overlooked option visible. That does not make it testimony in the strong sense. Testimony rests on an attributable practice of perceiving, checking, and correcting; an organization can rely on such practices because others can retrace and critique its path. A language model generates text without taking responsibility for its truth. Responsibility remains with the role that turns the result into a publication, recommendation, or workflow.

This does not mean model comparison is useless. Differences can distribute our attention. If response X names a source and response Y adds a condition, a review task emerges: Which assertion is actually supported? Which condition was included in the assignment? Which perspective is missing? The comparison can thus improve the quality of a human investigation. It does not replace that investigation.

The strongest philosophical mistake would be to turn agreement into epistemic authority: if two systems say the same thing, it must be more likely to be true. That can apply when several judgments genuinely combine independent, competent access to a matter and their sources of error are sufficiently different. That independence must not be assumed for language models. Training data, model families, instructions, evaluation schemes, and common text patterns can overlap. Even different providers do not guarantee that their errors are uncorrelated. Likewise, a dissenting model may ask a better question without thereby being “the better model” overall.

A more modest rule follows for practice: Disagreement creates an occasion for diagnosis; agreement does not end review. Which diagnosis is appropriate depends on the task and on the kind of disputed assertion.

2. What the research provides for this—and what it does not

Several original studies help sharpen the test design. None prescribes a ready-made comparison method for a small organization. Their findings relate to particular models, datasets, observation periods, and metrics. That is precisely why they are useful: they show which generalizations should be avoided.

In “Correlated Errors in Large Language Models,” Elliot Kim, Avi Garg, Kenny Peng, and Nikhil Garg examine more than 350 language models across two widely used leaderboards and a résumé-evaluation task. In one of the leaderboard datasets examined, models agreed in 60 percent of cases when both were wrong. The paper identifies shared architecture and provider, among other factors; it also reports high error correlation among larger, more accurate models with different architectures and providers. This is not proof that every possible pair makes the same errors in everyday use. But it refutes the convenient assumption that different names or providers already constitute an independent check. Kim et al., 2025

With HELM, Percy Liang and a large author team develop a framework that does not reduce comparison to a single metric. The study maps possible application scenarios and criteria, measures seven metrics in its evaluation at the time—including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency—across 16 core scenarios where possible, and adds targeted tests. The important idea for an SME trial is not to reproduce this extensive evaluation. It is this: before a comparison, one must decide which property matters for the specific use and which trade-offs the chosen standard hides. A win in readability does not answer a question about factual support or permissible data handling. HELM remains a study of its scenario and metric space, not a ranking for every future task. Liang et al., 2023

Using MT-Bench and Chatbot Arena, Lianmin Zheng and coauthors examine whether strong language models can assess open-ended chatbot responses. Their work compares model judgments with controlled and crowdsourced human preferences and reports high agreement between strong model judges of the time and those preferences. It also identifies limits such as position, verbosity, and self-preference bias, as well as limited reasoning. This does not mean a model judge is worthless. It can be a scalable indicator for clearly defined, open quality questions. Its assessment, however, is not an independent tribunal for facts; under the conditions examined, it approximates human preference. Zheng et al., 2023

In the preprint “Judging the Judges,” Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi examine position bias among model judges. They compare judges across MTBench and DevBench, multiple tasks and response models, and consider repeat stability, position consistency, and preference fairness. Their findings show that position bias in the judge models examined was not merely random and varied by judge, task, and quality gap. This does not yield a fixed correction formula for a small comparison. The practical implication is: if two responses are evaluated by a model, the order must be deliberately swapped and the result documented as potentially biased. Shi et al., 2024

The four studies therefore do not provide a simple rule such as “always use three models” or “always assess blindly.” They justify something narrower: comparison requires a described object, several separate criteria, a way of handling possible dependence, and a visible limit on the judge’s role.

3. Repair the question first: F in the order A → B → D → C

The F question model is not used here as decoration. It changes the trial. Its order is A → B → D → C. Only after clarification, competing hypotheses, and a revision loop does the bounded comparison question emerge.

A: What is actually in conflict?

“Which model is better?” combines at least four questions: Does it follow the instruction? Does it get facts right? Does it make uncertainty visible in a traceable way? Does it write clearly for the intended audience? Without separating these, a well-written but unsupported response can win because it sounds more convincing.

The fictional team’s initial question is therefore: “Which model should draft our course text?” After A, it is narrowed to: “Which of the two outputs, under the same case card, meets the specified limits for a public, source-bound course information text?” This is not yet an approval question. It is a question about two text outputs and a reference formulated in advance.

B: Competing hypotheses before viewing the responses

Before looking at the responses, the team records several possible explanations for a later difference.

HypothesisWhat the comparison might distinguishObservation that would weaken the hypothesis
H1: different weighting of objectivesOne response prioritizes a promotional benefit, the other bounded, source-close information.Both responses address the same objectives and limits; the difference lies only in wording.
H2: the task was understood differentlyAn ambiguous term in the assignment is interpreted differently.Both responses use the same permissible meaning and meet the same instructions.
H3: a secondary condition was overlookedOne response omits, for example, the study’s scope or the prohibition on promising learning effects.Both responses include every secondary condition; the deviation remains even with an explicit checklist.
H4: one output adds unsupported contentOne output claims learning effects, general model reliability, or other facts outside the source package.Every substantive assertion can be directly supported within the agreed source package.
H5: shared false assumptionBoth responses adopt the same unsupported conclusion or the same unclear reference.An independent original anchor supports the disputed claim within the stated scope, or the outputs differ in their error pattern.

These hypotheses are not probabilities. They prevent the first explanation found from immediately becoming the judgment. They also state which observation would weaken a hypothesis.

D: Change the question when the test cannot support it

The initial test question can fail. Perhaps the models did not receive the same input. Perhaps the reference itself is too sweeping. Perhaps a person can see only a difference in style although the test claims to measure facts. Then the result is not “unclear, but model A is narrowly ahead.” The question must be revised.

In this article’s case, the initial question “Which text is truer?” is rejected. It is too broad and contains neither an operationalized assessment criterion nor a reference anchor. The next version is: “Which output violates fewer of the requirements set in advance for the two subclaims in the fictional case card?” That changes the evaluation. Instead of asserting a diffuse property, the team can describe marked violations.

C: A bounded final question with an action step

The revised question is ultimately:

“Under the frozen case card V-01, does output X or output Y more fully meet the five requirements set in advance for a public course information text—and which open questions prevent its approval?”

The GROW step is small: the team first formulates a separate question-expansion round and then tests three additional, pre-specified course claims using the same rubric. It does not use the single run as a quality ranking. This step is a work plan, not an empirical claim about F’s effectiveness.

4. The minimal test contract

A fair comparison can never control every difference. It can, however, make visible what was held constant and what remains open. For a small, bounded trial, a test contract is sufficient.

FieldTo be determined before the run
Purpose and scopeWhich narrowly defined property is being examined? To which case does the finding apply?
Input cardCase description, permitted materials, prohibited assumptions, language, output format, and maximum length.
Model stateProvider name, model/version information, access method, date, and relevant settings, insofar as visible.
Reference anchorExpertly justified case note, original source, or explicit rule against which a subclaim is checked.
CriteriaSeparate categories for factual accuracy, instruction following, completeness, uncertainty marking, and style.
Blind codeOutputs initially only as X and Y; system identity remains hidden from the reviewer, and presentation order is recorded.
Order controlIf a model judge is used: check both presentation orders and note any deviation.
Revision logTrigger, change, time, and effect on the final wording.
Form of conclusionFinding, scope, open question, responsible role, and stop signal.

The reference is the most delicate part. It must not be formulated after reading the responses so that a preferred response wins. For a verifiable factual question, it requires a relevant original source and the relevant passage. For a local process question, it can be a bounded case rule decided in advance. In neither case is it neutral simply because it is written down. The team must explain why this rule or source fits the question and which perspective it does not include.

Blind coding provides limited protection against expectations: in three text tasks, Zhu and coauthors found that origin labels could markedly change preferences even though raters could not distinguish the text types in the blind condition. This supports blind codes for preference judgments, not for checking facts. Someone who does not know which output comes from which provider is less easily guided by assumptions about brands. But blind coding creates neither representative data nor a complete reference. If the reviewer recognizes the style or already knows the task very well, influence remains possible. Blind review therefore belongs in the protocol, not in a claim to now be objective. Zhu et al., 2025

Diagram without text: two equally sized output cards lead via blind codes to a shared reference, from there to five separate criteria points with a revision loop, and finally to three exits: finding, revision or pause
From disagreement to a bounded finding: blind codes, shared reference, separate criteria, revision
Image generated with AI

5. The fictional SME case

Course Workshop by the River GmbH is entirely fictional. The small continuing-education company is preparing a public course page on critically comparing AI responses. Two text systems receive the same assignment. They are to write a short information text from a narrow source package. The company, course page, roles, and all subsequent responses are fictional for this article.

The two subclaims deliberately differ. The first is a research claim: HELM developed a broad evaluation with multiple scenarios and metrics. The second is an editorial case rule: in the course, agreement between two responses does not replace checking a disputed subclaim against the original anchor. Kim et al. report correlated errors as an empirical finding under the conditions of their investigation. The separation between study and case rule is part of the comparison design.

For the exercise, the editorial team prepares this case card in advance. It serves a didactic blind-assessment simulation, not an actual model test.

FieldSpecified content
IdentifierV-01, fictional course-editorial case
Test mode and run metadataSimulation with invented texts; no real model calls, hence no model versions, access paths, settings, or run data exist. A real comparison would have to record this information in the test contract.
ObjectiveDraft a public information text of no more than 120 words for a fictional course page.
Permitted source notesQ1: HELM describes an evaluation across multiple scenarios and metrics; in the original study: 16 core scenarios and seven metrics, where measurable. Q2: Kim et al. observe correlated errors in their investigation of more than 350 language models.
Separate case ruleThe course exercise does not treat agreement between two responses as confirmation; disputed facts are checked against the relevant original anchor. This is the editorial rule of the case card, not a finding by Kim et al.
Identical task“Using Q1–Q2, formulate an accessible course information text. State both subclaims, mark their scope, promise no learning effect, and do not claim that two models confirm truth. End with a practical action that checks a subclaim against the original anchor.”
Output schemaHeadline and one paragraph; no more than 120 words; no further facts, sources, or promises of effect.
Reference set in advanceThe text must present HELM as a multi-criterion, scenario-bound evaluation framework; it must retain the separate case rule; it must not promise the course’s effectiveness or a general model ranking.
Error categoriesE1: contradicts Q1 or Q2; E2: adds an unsupported fact; E3: overextends the scope; E4: replaces the required source check with a style preference or omits it; E5: confuses preference, truth, and learning effect.

The editorial team saves the unchanged input card as Round 1. A question expansion may occur only after evaluation. If, for example, it adds that the text should explain the perspective of learners with little prior knowledge, this is Round 2 with a new version identifier. The changed question must not retrospectively be presented as an identical input and therefore as a fair tie.

6. Two invented outputs, one blindly reviewed table

The following responses are entirely invented example texts for this article. They come from no model call and are neither a transcript nor evidence of performance. The table simulates what a blind assessment could look like: in a real run, the system names would be hidden from the reviewer; here there are only fictional codes.

Output X

Two models make knowledge certain

HELM proves which AI system gives the best answer. That is why we compare two models in the course: if they agree, an assertion is reliably confirmed. Our method demonstrably makes learners better fact-checkers. They learn to select the friendliest and clearest AI response.

Output Y

Compare, check, limit

HELM evaluates language models across several scenarios and criteria; a result applies only within the framework examined. In the course, we do not compare two responses as a vote on truth. Differences and agreements become occasions to check subclaims against an appropriate original source and its limits. Learners record which question remains open before adopting a text.

The following table is a completed assessment matrix within the simulation. It does not show observed behavior of real systems.

CheckpointOutput XOutput YBlind finding against the reference fixed in advance
HELM subclaim“proves which system gives the best answer”several scenarios and criteria, bounded frameworkX: E1/E3. HELM is not proof of a general best-of ranking. Y: within the reference framework.
Treatment of agreement“reliably confirmed”occasion for source checkingX: E3/E5. The conclusion turns agreement into truth. Y: within the reference framework.
Learning effect“demonstrably better fact-checkers”describes a learning actionX: E2/E5. No source in the package supports this effect. Y: within the reference framework.
Practical learning actionselecting the “friendliest and clearest” responsecheck subclaims and the original source; note the open questionX: E4, because the required action replaces the source check required in advance with a style preference. Y: within the reference framework.
Scope and uncertaintyno limit“only within the framework examined”X: E3. Y: within the reference framework.
Stylecatchy, but promotional and absoluteaccessible and boundedStyle is recorded separately; it does not determine E1–E5.

After the simulated blind assessment, the fictional assignment is disclosed: in this teaching example, X stands for “Model A” and Y for “Model B.” There were no real model runs. In a genuine trial, such an assignment would be recorded but shown to the reviewer only after their factual judgment.

In the simulation, X violates several specified requirements: the text turns a broad evaluation into proof, agreement into confirmation, and claims a learning effect without a source. Y meets more of the requirements within this constructed assessment matrix. Neither actual model performance nor greater comprehensibility of a course page can be derived from this.

● Members only

Read the full article and download all files with a membership.

Unlock full article + downloads → Subscribe

0 comments

● Loading comments…

Sign in to comment · become a member →