Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
Authors/Creators
Description
Abstract
Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer.
On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.
In relation to disagreement, our results show that model confidence is not a reliable indicator of whether the panel will agree. The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44). The verdict that is given to an assertion will then depend heavily on which model one consults, which matters a great deal as more people turn to such models to verify information.
Key findings:
- On 63% of claims (632 / 997; 95% CI 60–66%) at least one model dissents from the panel majority, or no majority forms at all.
- On 23% of claims (232 / 997; 95% CI 21–26%), the two furthest-apart verdicts differ by at least two verdict categories — an actual dispute over the claim, not a difference in calibration.
- The ordinal Krippendorff's α of 0.77, across 5 models on 997 claims, reflects structured but far from interchangeable judgement.
- Disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.
- The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44).
HTML rendering: https://lenz.io/research/llm-disagreement/v1.1
Harness, corpus, and raw results: https://github.com/lenzhq/lenz-research
Files
lenz-llm-disagreement-v1.1.pdf
Files
(371.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:249206400af657d77393c4fc00bd4f73
|
110.4 kB | Preview Download |
|
md5:19d1d1646da4c569a8ac79c5436e62d4
|
260.6 kB | Preview Download |
Additional details
Related works
- Is identical to
- Preprint: https://lenz.io/research/llm-disagreement/v1.1 (URL)
Software
- Repository URL
- https://github.com/lenzhq/lenz-research
- Programming language
- Python
- Development Status
- Active