When Two Models Disagree
A comparison is a test design, not a vote on truth

Two AI responses sit side by side. Both are fluently written, both propose a next step, both seem reasonable. Yet their recommendations exclude each other. The first promises that a second model reliably confirms content; the second requires differences to be checked first against an original source. Anyone who now asks which model is “right” is asking an understandable, but still overly broad, question.
A contradiction between models is first of all a reason to investigate. It may indicate a factual error. It may arise from an unclear task, different unstated assumptions, or a genuine trade-off in values. Nor does consensus reliably settle the matter: two systems may overlook the same gap or repeat a shared, false assumption. The number of matching outputs does not turn a sentence into knowledge.
This article therefore presents a bounded comparison design. It holds the task, provided information, output schema, and assessment standard constant. It separates checking facts from assessing style and evaluations. The system identity remains hidden during the initial blind assessment; presentation order is documented and deliberately swapped when a model judge is used. It also records when the design itself must be changed. The aim is neither a ranking nor a general statement about providers. A real run could provide only a traceable finding about a narrowly described task.
The ongoing practical case is entirely fictional. It describes neither a real company nor an actual model output, person, or teaching situation.
1. Why the dispute is philosophical
At first glance, a model comparison seems technical: two inputs, two responses, one table. Its core is epistemological. We must decide what kind of reason a response can provide at all.
A response can be an indication. It can propose an explanation, draw attention to a source, or make an overlooked option visible. That does not make it testimony in the strong sense. Testimony rests on an attributable practice of perceiving, checking, and correcting; an organization can rely on such practices because others can retrace and critique its path. A language model generates text without taking responsibility for its truth. Responsibility remains with the role that turns the result into a publication, recommendation, or workflow.
This does not mean model comparison is useless. Differences can distribute our attention. If response X names a source and response Y adds a condition, a review task emerges: Which assertion is actually supported? Which condition was included in the assignment? Which perspective is missing? The comparison can thus improve the quality of a human investigation. It does not replace that investigation.
The strongest philosophical mistake would be to turn agreement into epistemic authority: if two systems say the same thing, it must be more likely to be true. That can apply when several judgments genuinely combine independent, competent access to a matter and their sources of error are sufficiently different. That independence must not be assumed for language models. Training data, model families, instructions, evaluation schemes, and common text patterns can overlap. Even different providers do not guarantee that their errors are uncorrelated. Likewise, a dissenting model may ask a better question without thereby being “the better model” overall.
A more modest rule follows for practice: Disagreement creates an occasion for diagnosis; agreement does not end review. Which diagnosis is appropriate depends on the task and on the kind of disputed assertion.
2. What the research provides for this—and what it does not
Several original studies help sharpen the test design. None prescribes a ready-made comparison method for a small organization. Their findings relate to particular models, datasets, observation periods, and metrics. That is precisely why they are useful: they show which generalizations should be avoided.
In “Correlated Errors in Large Language Models,” Elliot Kim, Avi Garg, Kenny Peng, and Nikhil Garg examine more than 350 language models across two widely used leaderboards and a résumé-evaluation task. In one of the leaderboard datasets examined, models agreed in 60 percent of cases when both were wrong. The paper identifies shared architecture and provider, among other factors; it also reports high error correlation among larger, more accurate models with different architectures and providers. This is not proof that every possible pair makes the same errors in everyday use. But it refutes the convenient assumption that different names or providers already constitute an independent check. Kim et al., 2025
With HELM, Percy Liang and a large author team develop a framework that does not reduce comparison to a single metric. The study maps possible application scenarios and criteria, measures seven metrics in its evaluation at the time—including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency—across 16 core scenarios where possible, and adds targeted tests. The important idea for an SME trial is not to reproduce this extensive evaluation. It is this: before a comparison, one must decide which property matters for the specific use and which trade-offs the chosen standard hides. A win in readability does not answer a question about factual support or permissible data handling. HELM remains a study of its scenario and metric space, not a ranking for every future task. Liang et al., 2023
Using MT-Bench and Chatbot Arena, Lianmin Zheng and coauthors examine whether strong language models can assess open-ended chatbot responses. Their work compares model judgments with controlled and crowdsourced human preferences and reports high agreement between strong model judges of the time and those preferences. It also identifies limits such as position, verbosity, and self-preference bias, as well as limited reasoning. This does not mean a model judge is worthless. It can be a scalable indicator for clearly defined, open quality questions. Its assessment, however, is not an independent tribunal for facts; under the conditions examined, it approximates human preference. Zheng et al., 2023
In the preprint “Judging the Judges,” Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi examine position bias among model judges. They compare judges across MTBench and DevBench, multiple tasks and response models, and consider repeat stability, position consistency, and preference fairness. Their findings show that position bias in the judge models examined was not merely random and varied by judge, task, and quality gap. This does not yield a fixed correction formula for a small comparison. The practical implication is: if two responses are evaluated by a model, the order must be deliberately swapped and the result documented as potentially biased. Shi et al., 2024
The four studies therefore do not provide a simple rule such as “always use three models” or “always assess blindly.” They justify something narrower: comparison requires a described object, several separate criteria, a way of handling possible dependence, and a visible limit on the judge’s role.
3. Repair the question first: F in the order A → B → D → C
The F question model is not used here as decoration. It changes the trial. Its order is A → B → D → C. Only after clarification, competing hypotheses, and a revision loop does the bounded comparison question emerge.
A: What is actually in conflict?
“Which model is better?” combines at least four questions: Does it follow the instruction? Does it get facts right? Does it make uncertainty visible in a traceable way? Does it write clearly for the intended audience? Without separating these, a well-written but unsupported response can win because it sounds more convincing.
The fictional team’s initial question is therefore: “Which model should draft our course text?” After A, it is narrowed to: “Which of the two outputs, under the same case card, meets the specified limits for a public, source-bound course information text?” This is not yet an approval question. It is a question about two text outputs and a reference formulated in advance.
B: Competing hypotheses before viewing the responses
Before looking at the responses, the team records several possible explanations for a later difference.
| Hypothesis | What the comparison might distinguish | Observation that would weaken the hypothesis |
|---|---|---|
| H1: different weighting of objectives | One response prioritizes a promotional benefit, the other bounded, source-close information. | Both responses address the same objectives and limits; the difference lies only in wording. |
| H2: the task was understood differently | An ambiguous term in the assignment is interpreted differently. | Both responses use the same permissible meaning and meet the same instructions. |
| H3: a secondary condition was overlooked | One response omits, for example, the study’s scope or the prohibition on promising learning effects. | Both responses include every secondary condition; the deviation remains even with an explicit checklist. |
| H4: one output adds unsupported content | One output claims learning effects, general model reliability, or other facts outside the source package. | Every substantive assertion can be directly supported within the agreed source package. |
| H5: shared false assumption | Both responses adopt the same unsupported conclusion or the same unclear reference. | An independent original anchor supports the disputed claim within the stated scope, or the outputs differ in their error pattern. |
These hypotheses are not probabilities. They prevent the first explanation found from immediately becoming the judgment. They also state which observation would weaken a hypothesis.
D: Change the question when the test cannot support it
The initial test question can fail. Perhaps the models did not receive the same input. Perhaps the reference itself is too sweeping. Perhaps a person can see only a difference in style although the test claims to measure facts. Then the result is not “unclear, but model A is narrowly ahead.” The question must be revised.
In this article’s case, the initial question “Which text is truer?” is rejected. It is too broad and contains neither an operationalized assessment criterion nor a reference anchor. The next version is: “Which output violates fewer of the requirements set in advance for the two subclaims in the fictional case card?” That changes the evaluation. Instead of asserting a diffuse property, the team can describe marked violations.
C: A bounded final question with an action step
The revised question is ultimately:
“Under the frozen case card V-01, does output X or output Y more fully meet the five requirements set in advance for a public course information text—and which open questions prevent its approval?”
The GROW step is small: the team first formulates a separate question-expansion round and then tests three additional, pre-specified course claims using the same rubric. It does not use the single run as a quality ranking. This step is a work plan, not an empirical claim about F’s effectiveness.
4. The minimal test contract
A fair comparison can never control every difference. It can, however, make visible what was held constant and what remains open. For a small, bounded trial, a test contract is sufficient.
| Field | To be determined before the run |
|---|---|
| Purpose and scope | Which narrowly defined property is being examined? To which case does the finding apply? |
| Input card | Case description, permitted materials, prohibited assumptions, language, output format, and maximum length. |
| Model state | Provider name, model/version information, access method, date, and relevant settings, insofar as visible. |
| Reference anchor | Expertly justified case note, original source, or explicit rule against which a subclaim is checked. |
| Criteria | Separate categories for factual accuracy, instruction following, completeness, uncertainty marking, and style. |
| Blind code | Outputs initially only as X and Y; system identity remains hidden from the reviewer, and presentation order is recorded. |
| Order control | If a model judge is used: check both presentation orders and note any deviation. |
| Revision log | Trigger, change, time, and effect on the final wording. |
| Form of conclusion | Finding, scope, open question, responsible role, and stop signal. |
The reference is the most delicate part. It must not be formulated after reading the responses so that a preferred response wins. For a verifiable factual question, it requires a relevant original source and the relevant passage. For a local process question, it can be a bounded case rule decided in advance. In neither case is it neutral simply because it is written down. The team must explain why this rule or source fits the question and which perspective it does not include.
Blind coding provides limited protection against expectations: in three text tasks, Zhu and coauthors found that origin labels could markedly change preferences even though raters could not distinguish the text types in the blind condition. This supports blind codes for preference judgments, not for checking facts. Someone who does not know which output comes from which provider is less easily guided by assumptions about brands. But blind coding creates neither representative data nor a complete reference. If the reviewer recognizes the style or already knows the task very well, influence remains possible. Blind review therefore belongs in the protocol, not in a claim to now be objective. Zhu et al., 2025

5. The fictional SME case
Course Workshop by the River GmbH is entirely fictional. The small continuing-education company is preparing a public course page on critically comparing AI responses. Two text systems receive the same assignment. They are to write a short information text from a narrow source package. The company, course page, roles, and all subsequent responses are fictional for this article.
The two subclaims deliberately differ. The first is a research claim: HELM developed a broad evaluation with multiple scenarios and metrics. The second is an editorial case rule: in the course, agreement between two responses does not replace checking a disputed subclaim against the original anchor. Kim et al. report correlated errors as an empirical finding under the conditions of their investigation. The separation between study and case rule is part of the comparison design.
For the exercise, the editorial team prepares this case card in advance. It serves a didactic blind-assessment simulation, not an actual model test.
| Field | Specified content |
|---|---|
| Identifier | V-01, fictional course-editorial case |
| Test mode and run metadata | Simulation with invented texts; no real model calls, hence no model versions, access paths, settings, or run data exist. A real comparison would have to record this information in the test contract. |
| Objective | Draft a public information text of no more than 120 words for a fictional course page. |
| Permitted source notes | Q1: HELM describes an evaluation across multiple scenarios and metrics; in the original study: 16 core scenarios and seven metrics, where measurable. Q2: Kim et al. observe correlated errors in their investigation of more than 350 language models. |
| Separate case rule | The course exercise does not treat agreement between two responses as confirmation; disputed facts are checked against the relevant original anchor. This is the editorial rule of the case card, not a finding by Kim et al. |
| Identical task | “Using Q1–Q2, formulate an accessible course information text. State both subclaims, mark their scope, promise no learning effect, and do not claim that two models confirm truth. End with a practical action that checks a subclaim against the original anchor.” |
| Output schema | Headline and one paragraph; no more than 120 words; no further facts, sources, or promises of effect. |
| Reference set in advance | The text must present HELM as a multi-criterion, scenario-bound evaluation framework; it must retain the separate case rule; it must not promise the course’s effectiveness or a general model ranking. |
| Error categories | E1: contradicts Q1 or Q2; E2: adds an unsupported fact; E3: overextends the scope; E4: replaces the required source check with a style preference or omits it; E5: confuses preference, truth, and learning effect. |
The editorial team saves the unchanged input card as Round 1. A question expansion may occur only after evaluation. If, for example, it adds that the text should explain the perspective of learners with little prior knowledge, this is Round 2 with a new version identifier. The changed question must not retrospectively be presented as an identical input and therefore as a fair tie.
6. Two invented outputs, one blindly reviewed table
The following responses are entirely invented example texts for this article. They come from no model call and are neither a transcript nor evidence of performance. The table simulates what a blind assessment could look like: in a real run, the system names would be hidden from the reviewer; here there are only fictional codes.
Output X
Two models make knowledge certain
HELM proves which AI system gives the best answer. That is why we compare two models in the course: if they agree, an assertion is reliably confirmed. Our method demonstrably makes learners better fact-checkers. They learn to select the friendliest and clearest AI response.
Output Y
Compare, check, limit
HELM evaluates language models across several scenarios and criteria; a result applies only within the framework examined. In the course, we do not compare two responses as a vote on truth. Differences and agreements become occasions to check subclaims against an appropriate original source and its limits. Learners record which question remains open before adopting a text.
The following table is a completed assessment matrix within the simulation. It does not show observed behavior of real systems.
| Checkpoint | Output X | Output Y | Blind finding against the reference fixed in advance |
|---|---|---|---|
| HELM subclaim | “proves which system gives the best answer” | several scenarios and criteria, bounded framework | X: E1/E3. HELM is not proof of a general best-of ranking. Y: within the reference framework. |
| Treatment of agreement | “reliably confirmed” | occasion for source checking | X: E3/E5. The conclusion turns agreement into truth. Y: within the reference framework. |
| Learning effect | “demonstrably better fact-checkers” | describes a learning action | X: E2/E5. No source in the package supports this effect. Y: within the reference framework. |
| Practical learning action | selecting the “friendliest and clearest” response | check subclaims and the original source; note the open question | X: E4, because the required action replaces the source check required in advance with a style preference. Y: within the reference framework. |
| Scope and uncertainty | no limit | “only within the framework examined” | X: E3. Y: within the reference framework. |
| Style | catchy, but promotional and absolute | accessible and bounded | Style is recorded separately; it does not determine E1–E5. |
After the simulated blind assessment, the fictional assignment is disclosed: in this teaching example, X stands for “Model A” and Y for “Model B.” There were no real model runs. In a genuine trial, such an assignment would be recorded but shown to the reviewer only after their factual judgment.
In the simulation, X violates several specified requirements: the text turns a broad evaluation into proof, agreement into confirmation, and claims a learning effect without a source. Y meets more of the requirements within this constructed assessment matrix. Neither actual model performance nor greater comprehensibility of a course page can be derived from this.
● Members only
Read the full article and download all files with a membership.
Unlock full article + downloads → Subscribe0 comments
● Loading comments…