LLM 3 min read

Five Frontier LLMs Got the Same 1,000 Questions. They Disagreed on 67% of Them.

You ask ChatGPT something. It answers. You move on. That reflex just got harder to defend. Researchers ran 1,000 real fact-check claims past five frontier LLMs and found the models disagreed with each other on roughly two out of every three questions.

What 67% actually means

The setup was clean. Take a thousand claims that human fact-checkers had already adjudicated. Ask each of five top-tier models — the usual suspects from OpenAI, Anthropic, Google and friends — the same question. Compare answers.

The disagreement rate landed at 67%. Not “the models occasionally hedge differently.” Actually different verdicts. One model says true, another says false, a third refuses to commit. These aren’t sibling versions of the same model. These are the flagship products that millions of people now treat as a casual oracle.

“Which AI do you trust” is the new question

The deeper problem isn’t accuracy. It’s attribution.

For the last two years, “ChatGPT told me” or “Claude says” has quietly become a citation format. People paste model answers into Slack threads, news drafts, legal memos, even patient-facing notes. If five frontier models give five different answers to the same question, that citation is functionally a lottery ticket — the answer you got is an artifact of which model you happened to open, on which day, with which prompt phrasing.

That’s fine for asking what to cook for dinner. It’s a serious problem for journalism, law, and medicine, where “the AI verified it” is starting to do real epistemic work.

Why the models split

You’d think models trained on roughly the same internet would converge. They don’t, for at least three reasons.

Training data choices: each lab makes different calls about which sources to weight, which to filter out, what counts as a high-quality corpus. RLHF direction: the same factual question can produce different answers depending on whether a model was tuned to be cautious or assertive. Safety guardrails: some models punt on contested topics; others take a clear side. Stack those three differences and “the same question” stops being the same question by the time it reaches a verdict.

Where the disagreement clusters

The split isn’t uniform, and the pattern is the uncomfortable part.

On settled science, basic history, well-documented events — the models mostly agree. Where they fall apart is political claims, statistical interpretation, recent news, and anything with rhetorical nuance. In other words: exactly the categories where you’d actually want a second opinion. The areas with the most ambiguity are the areas where the models diverge most. The tool wobbles hardest precisely where you need it steady.

A more honest way to use these things

The study doesn’t argue for abandoning LLMs. It argues for using them with eyes open.

Three practical habits worth borrowing. First, for anything that matters, query two or three models and treat disagreement as a signal — if the frontier can’t agree, you probably need a human source. Second, chase the citation, not the summary: if a model gives you a fact, ask it for the source and verify the source exists and says what the model claims. Third, retire “the AI said so” from your vocabulary. It was always a shaky appeal to authority. Now it’s a quantified one.

The takeaway

What the study really exposes is a category error baked into how people talk about LLMs. We treat them as answer machines. They behave more like opinion machines — five of them in a room, disagreeing about most things, occasionally confident, occasionally wrong, never quite the same. For most of the last two years, we’ve been listening to whichever one we opened first. That habit is getting expensive.

LLM AI reliability fact-checking model comparison hallucination

Comments

    Loading comments...