Longitudinal study · US–Iran conflict

The Iran Arc

Three of five frontier models supply a statistic the text they were given never contained.

This page follows the EigenTrace convention: claims the instrument measured are marked and kept self-contained — each stands on its own evidence and needs no interpretation to hold. Claims that are argued — our reading of what the measurements mean — are fenced off and labeled, so a skeptical reader can reject every interpretation and find every measurement still standing. Where measurement ends, we say so.

Between May and September 2026, the EigenTrace broadcast gave four or five frontier models the same news story and the same instruction, and kept every summary. That archive answers a question a reader can check without trusting any model: when a summary states a number, was that number in the text the model was given?

On the Strait of Hormuz, three of the five often supplied a figure that was not. The strait is the kind of subject where a widely repeated background statistic — roughly a fifth of the world’s oil passes through it — sits close to hand, and the models differ sharply in whether they reach for it.

FINDING 05Three models add a Hormuz oil-share figure their input did not contain. ChatGPT and Gemini almost never do.

Measured

Restricting to summaries whose input mentioned the Strait of Hormuz, the rate at which a model supplied a figure for its share of world oil that was not in that input was: DeepSeek 65 of 199 (33%), Grok 61 of 200 (31%), Claude 52 of 163 (32%), Gemini 3 of 158 (1.9%), ChatGPT 0 of 192. ChatGPT had not stopped discussing oil: it mentions oil in 241 summaries in this corpus and adds the figure in none of them. On the 8 occasions it gives the share, the figure is in the text it was given.

Within the same stories, the average rate for DeepSeek and Grok minus the average rate for ChatGPT and Gemini is +0.128 (95% CI 0.100 to 0.156; 392 stories, May 2 to July 31). On the 59 stories from August 1 to September 10, held back until the comparison had been fixed in writing, it is +0.297 (CI 0.195 to 0.407; sign-flip permutation p = 0.0005, the floor for 2,000 permutations; Holm-adjusted 0.0025).

On the same story the difference runs one way only. With ChatGPT as the baseline, DeepSeek added the figure where ChatGPT did not in 52 stories and never the reverse (exact McNemar p = 4 × 10−16); the same count is 47 to 0 for Grok and 49 to 0 for Claude. On the held-out stories it is 20 to 0 for DeepSeek (p = 1.9 × 10−6) and 16 to 0 for Grok (p = 3.1 × 10−5).

The models do not agree on the number they add. Where two or more add one to the same story, the values differ in 34 of 47 stories in the exploration window: 20%, 21%, 20–30%, “about a fifth”, and 30% of “seaborne” oil.

The clearest instance is a July 11 live-blog story about US demands over the strait. The stored input runs 581 characters and contains no percentage. Four of the five call the strait a chokepoint; three attach a number to it.

ChatGPT
“The Strait of Hormuz is a critical chokepoint for global oil shipments.” — no figure.
Gemini
“The Strait of Hormuz is a critical chokepoint for global oil and gas shipments.” — no figure.
Claude
“The Strait of Hormuz handles roughly 20-30% of global oil trade.”
DeepSeek
“The Strait of Hormuz is a critical chokepoint for about 20% of global oil transit.”
Grok
“The Strait of Hormuz is the world’s most critical oil chokepoint; roughly 20-30% of global seaborne oil trade passes through it.”

An August 2 story — held out until the test was fixed — repeats the pattern with four models, Claude having returned nothing that day. Its stored input also contains no percentage.

DeepSeek
“Iran controls the strait, through which about 20% of global oil passes.”
Grok
“The Strait of Hormuz — through which roughly 20% of global oil trade passes — must be fully opened and kept open; …”
ChatGPT
No oil-share figure.
Gemini
No oil-share figure.
Segment files 20260711_004957_55a2e14a50a5 and 20260802_035233_a29da30e781c. Both were re-checked against the stored responses and the stored input text.
Argued · interpretation

Our reading is that the prompt asks for “concrete implications”, and three models answer by recalling a widely repeated background figure while two describe why the strait matters without putting a number on it. That is a plausible account of why. It is not the measurement. The measurement is only that the figure appears in the summary and not in the text the model was given, and that on the same stories the gap between the models runs one way.

Where measurement ends

This is not a claim that the figure is wrong. Figures of roughly a fifth of world oil, or up to 30% of seaborne oil, are commonly cited. The instrument does not judge accuracy. It records that a number the article did not supply was added and presented alongside the article’s facts, and that the models give different numbers for it.

What “not in the text given” means. A figure counts as unsourced only if its value appears nowhere in the stored story text — as digits, a number word, or a fraction word, so “a fifth” matches 20. The stored text is the title plus the feed summary plus the scraped body, which is at least as much as any model received, so the counts are lower bounds. The article beyond 2,000 characters was never stored, and the models never saw it either; whether the publisher’s full article carried a figure is unknown.

Claude’s share rests on the exploration window. Claude answered only 8 of the 59 held-out stories, so the held-out test compares DeepSeek and Grok against ChatGPT and Gemini; Claude’s rate is from May to July.

Held-out effect sizes are not comparable with the exploration sizes. Grok’s rate moves with its own dated changes (27 of 123 before May 21, 20 of 268 from May 21 to July 31, 16 of 59 held out), and DeepSeek’s summaries doubled in length on July 31 with no change in the pipeline. The held-out window confirms the direction and that the effect exists; it does not confirm the size. DeepSeek alone carries the contrast in every period.

“Almost never”, not “never”. Under this page’s corpus rule ChatGPT’s count is 0 of 440. In Iran stories whose titles fall outside that rule it adds the figure 4 times in 541, each time with no percentage or fraction in its input.

Scope. This is a Hormuz-bound behaviour and a trait of these models, visible in the Iran coverage. It says nothing about sides in the war.


What we tested against

The objections worth raising are the ones that say a per-model difference is really an artifact of something else — of when the stories ran, of who was in the panel, of how long the summaries are, or of the instructions each model was given. Where an objection could be turned into a test, we ran it.

Ruled out · Grok’s dated changes

Objection: Grok’s output changed twice inside the window — a silent change on May 15, then a Grok-only system prompt on May 21 — so the contrast may be an artifact of one vendor’s edits. Test: start the window after each change. Over the full window the difference is +0.135 from May 21 (CI 0.104 to 0.170) and +0.156 from June 1 (0.117 to 0.196); within the exploration window alone, where the headline +0.128 is measured, it is +0.099 (0.070 to 0.129) and +0.116 (0.077 to 0.155). ChatGPT added the figure 0 times in both.

Ruled out · shifting panel size

Objection: the panel was four models on some days and five on others, and a per-model rate can be distorted by who else was present. Test: restrict to the 282 stories where all five models answered. The difference is +0.099 (0.069 to 0.130), with ChatGPT at 0, Gemini 2, Claude 26, DeepSeek 33 and Grok 25.

Ruled out · summary length

Objection: a model that writes more has more room to add a number. Test: restrict to the 311 stories where ChatGPT’s summary is at least as long as DeepSeek’s. DeepSeek added the figure in 41 of them and ChatGPT in 0 (p = 9 × 10−13). The difference is not one model writing more.

Ruled out · the system prompts

Objection: the models were not given identical instructions, so the split may follow the prompt. Test: it does not. DeepSeek adds the figure and ChatGPT does not, under the same system prompt. Claude adds it and Gemini does not, and neither model had one.

Ruled out · kind of story

Objection: the contrast may live in one kind of story rather than in the models. Test: the difference holds in every stratum — NYT +0.156 (n = 144) and non-NYT +0.147 (n = 307); short-body +0.163 (n = 178) and long-body +0.141 (n = 273); live blogs +0.233 (n = 103) and single stories +0.125 (n = 348). ChatGPT is at 0 in all six.

Tested · the behaviour is tied to Hormuz

A placebo: in 325 war stories not about Iran the same difference is +0.008 (p = 0.26). Without Hormuz in the input it is +0.026.

Tested · the detector, and the bookkeeping

The matcher was checked against a deliberately wrong answer key: scored over every numeric specific in these summaries against a random other Iran article from within ten days, 91–93% of every model’s specifics come out unsourced, while scored against the model’s own input 3–32% do. The corpus was also rebuilt three ways — collapsing re-broadcasts by URL, by the first eight title words, and not at all — giving +0.129 to +0.144 in exploration and +0.29 to +0.30 held out. A hand check of 25 random detections from DeepSeek, Grok and Claude found 23 were the world-oil-through-Hormuz share and 2 were other unsourced energy shares.

Acknowledged limits we have not resolved

Two honest ones. We cannot see the whole article. Only the first 2,000 characters were ever stored, and that is also all the models received, so “absent from the source” is a statement about the text given, not about the publisher’s full page. One window, one topic. The held-out replication covers six weeks and one subject; whether the same models add unsourced specifics elsewhere is measurable with the same instrument and is not measured here.

Open invitation · help us harden it or break it

The code, prompts, model responses and raw measurements are public, and the two exhibit segments are named above so any reader can pull the same files. The attacks we have already run are listed in this section. The ones we would most like to see: run the detector on a subject other than Hormuz, where a background statistic is equally close to hand; hand-audit a larger random sample of detections; and test whether a model that adds the figure does so more when the input is thin. If any of these breaks the finding, it will be recorded at /withdrawals.


What this is, and what it is not

Measured · in sum

Across 455 stories of the 2026 US–Iran conflict, on inputs that named the Strait of Hormuz: three of five models supplied a figure anyway — DeepSeek 33%, Grok 31%, Claude 32% — while Gemini did so in 1.9% and ChatGPT in none of 192. The within-story difference is +0.128 in exploration and +0.297 on held-out stories, one-directional on the same story (52 to 0, 47 to 0, 49 to 0), and it survives later start dates, five-model panels, a length control, the differing system prompts, story type, three ways of counting duplicates, and a non-Iran placebo.

Argued · in sum

We read this as a difference in what a summarizer treats as part of the story: some models answer “what does this mean” by supplying the standard background number, and some answer it by naming the mechanism without a number. We find that reading natural, and it remains a reading. The measurement is the count and the direction.

What it is not

This is not a claim that any model hallucinated, that the figure is false, or that adding it is wrong — a reader may be better served by the number than without it. It is not a claim about the war, about sides in it, or about any lab’s intentions. It is a structural observation: when the same five models are deployed as the reading layer over the same story, some of them hand the reader a number the story did not contain, they do not agree on which number, and the reader cannot tell from the summary which figures came from the article.