Published August 7, 2026 | Version 1.1

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks

  • 1. Lenz Research
  • 2. ROR icon Bocconi University
  • 3. American College of Sofia

Description

Abstract

Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer.

On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.

In relation to disagreement, our results show that model confidence is not a reliable indicator of whether the panel will agree. The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44). The verdict that is given to an assertion will then depend heavily on which model one consults, which matters a great deal as more people turn to such models to verify information.

Key findings:

  • On 63% of claims (632 / 997; 95% CI 60–66%) at least one model dissents from the panel majority, or no majority forms at all.
  • On 23% of claims (232 / 997; 95% CI 21–26%), the two furthest-apart verdicts differ by at least two verdict categories — an actual dispute over the claim, not a difference in calibration.
  • The ordinal Krippendorff's α of 0.77, across 5 models on 997 claims, reflects structured but far from interchangeable judgement.
  • Disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.
  • The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44).

HTML rendering: https://lenz.io/research/llm-disagreement/v1.1
Harness, corpus, and raw results: https://github.com/lenzhq/lenz-research

Files

lenz-llm-disagreement-v1.1.pdf

Files (371.0 kB)

Name Size Download all
md5:249206400af657d77393c4fc00bd4f73
110.4 kB Preview Download
md5:19d1d1646da4c569a8ac79c5436e62d4
260.6 kB Preview Download

Additional details

Related works

Is identical to
Preprint: https://lenz.io/research/llm-disagreement/v1.1 (URL)

Software

Repository URL
https://github.com/lenzhq/lenz-research
Programming language
Python
Development Status
Active