The AI industry is moving quickly, sometimes like a negotiator arriving too early in a still poorly lit room. However, it would be dangerous to transform these models into impeccable oracles, placed above reality. Current versions remain massive betas: powerful, useful, but still capable of confusing nuance, context and truth.

In brief
- The study compares five advanced AI models on 1,000 claims submitted by real users this year.
- AIs diverge sharply in 67% of fact checks performed during the full experiment.
- The Krippendorff score reaches only 0.639, well below modern scientific standards for algorithmic reliability.
- Unanimous consensus appears mainly on totally true or completely false statements now only.
When AI giants each negotiate their own reality
A study from Lenz Research shakes up the technology ecosystem. The researchers submitted 1,000 real-world claims to five advanced models: GPT-5.4, Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro with Search, and Sonar Pro. Each model had to choose between four verdicts: true, “mostly true”, “misleading” or false.
THE result is nothing like a simple countertop bug. In 672 cases out of 1,000, at least one AI diverges from the majority, or no strict majority appears. In other words, the models supposed to verify the facts do not sign the same contract with reality.
The report states:
These statements are not benchmark items with public responses; these are claims submitted by real users to a verification platform.
Source: Lenz Research report
This precision weighs heavily: AIs no longer play on marked ground, but in an open negotiation with rough facts.
Tech models crack as soon as nuance enters the deal
The problem is not limited to classic hallucinations, these involuntary lies served in a three-piece suit. Here, artificial intelligences sometimes read the same elements, then deliver incompatible judgments. In 34% of cases, the disagreement becomes substantial, with at least two categories of difference between models.
The Krippendorff score reaches only 0.639. In law as in science, this figure requires caution. It indicates real agreement, but too weak to treat these models as interchangeable judges. The threshold often used for solid reliability is around 0.8.
The report summarizes this divide:
The models converge towards definitive verdicts; the middle of the scale is where they fracture.
Source: Lenz Research report
Indeed, consensuses appear mainly at the extremes. Out of 328 unanimous agreements, only four relate to “misleading”. None concern “mostly true”.
When multiple machines check the same fact, the room becomes noisy
The examples cited show a concrete difficulty. A claim about the World Bank's active portfolio in Nigeria sharply divides models. GPT-5.4 chooses “mostly true”. Gemini 3 Pro responds “false”. Gemini 3 Pro with Search prefers “misleading”. The user therefore receives three different tickets at the same counter.
Another sensitive case: an assertion linked to Donald Trump, Iran and a request from Gulf allies. GPT-5.4 judges this to be false, Claude Opus 4.7 answers “mostly true”, Gemini 3 Pro answers false, while Gemini 3 Pro with Search answers true. For the reader, the promise of clarification becomes an algorithmic arbitration fair.
The study also reminds us that a majority of AI does not constitute legal truth. A dissident machine can be right against four others. This reservation concerns the media, teachers, tech companies and services that already automate their controls.
The figures that crack the AI showcase
- Five models tested on 1,000 recent real statements;
- Disagreement observed on 672 statements out of 1,000;
- Substantial disagreement noted in 34% of cases;
- Unanimous agreement obtained only on 328 statements analyzed;
- No “mostly true” consensus among the unanimous verdicts.
This study does not condemn AI; rather, it recalls its experimental status. Last September, Google artificial intelligence solved a mathematical problem deemed impossible. The paradox remains splendid: these systems can dominate scientific abstraction, then stumble in the face of ordinary human truths.
Maximize your Tremplin.io experience with our 'Read to Earn' program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
