Longitudinal study · US–Iran conflict
Three of five frontier models supply a statistic the text they were given never contained.
455 Iran stories · May 2 – September 10, 2026 · deterministic string and number matching against the stored source text · no model judged another model anywhere in this analysis · the comparison was fixed in writing before the held-out stories were examined · findings previously published on this page are recorded at /withdrawals.
This page follows the EigenTrace convention: claims the instrument measured are marked and kept self-contained — each stands on its own evidence and needs no interpretation to hold. Claims that are argued — our reading of what the measurements mean — are fenced off and labeled, so a skeptical reader can reject every interpretation and find every measurement still standing. Where measurement ends, we say so.
Between May and September 2026, the EigenTrace broadcast gave four or five frontier models the same news story and the same instruction, and kept every summary. That archive answers a question a reader can check without trusting any model: when a summary states a number, was that number in the text the model was given?
On the Strait of Hormuz, three of the five often supplied a figure that was not. The strait is the kind of subject where a widely repeated background statistic — roughly a fifth of the world’s oil passes through it — sits close to hand, and the models differ sharply in whether they reach for it.
Restricting to summaries whose input mentioned the Strait of Hormuz, the rate at which a model supplied a figure for its share of world oil that was not in that input was: DeepSeek 65 of 199 (33%), Grok 61 of 200 (31%), Claude 52 of 163 (32%), Gemini 3 of 158 (1.9%), ChatGPT 0 of 192. ChatGPT had not stopped discussing oil: it mentions oil in 241 summaries in this corpus and adds the figure in none of them. On the 8 occasions it gives the share, the figure is in the text it was given.
Within the same stories, the average rate for DeepSeek and Grok minus the average rate for ChatGPT and Gemini is +0.128 (95% CI 0.100 to 0.156; 392 stories, May 2 to July 31). On the 59 stories from August 1 to September 10, held back until the comparison had been fixed in writing, it is +0.297 (CI 0.195 to 0.407; sign-flip permutation p = 0.0005, the floor for 2,000 permutations; Holm-adjusted 0.0025).
On the same story the difference runs one way only. With ChatGPT as the baseline, DeepSeek added the figure where ChatGPT did not in 52 stories and never the reverse (exact McNemar p = 4 × 10−16); the same count is 47 to 0 for Grok and 49 to 0 for Claude. On the held-out stories it is 20 to 0 for DeepSeek (p = 1.9 × 10−6) and 16 to 0 for Grok (p = 3.1 × 10−5).
The models do not agree on the number they add. Where two or more add one to the same story, the values differ in 34 of 47 stories in the exploration window: 20%, 21%, 20–30%, “about a fifth”, and 30% of “seaborne” oil.
The clearest instance is a July 11 live-blog story about US demands over the strait. The stored input runs 581 characters and contains no percentage. Four of the five call the strait a chokepoint; three attach a number to it.
An August 2 story — held out until the test was fixed — repeats the pattern with four models, Claude having returned nothing that day. Its stored input also contains no percentage.
20260711_004957_55a2e14a50a5 and 20260802_035233_a29da30e781c. Both were re-checked against the stored responses and the stored input text.Our reading is that the prompt asks for “concrete implications”, and three models answer by recalling a widely repeated background figure while two describe why the strait matters without putting a number on it. That is a plausible account of why. It is not the measurement. The measurement is only that the figure appears in the summary and not in the text the model was given, and that on the same stories the gap between the models runs one way.
This is not a claim that the figure is wrong. Figures of roughly a fifth of world oil, or up to 30% of seaborne oil, are commonly cited. The instrument does not judge accuracy. It records that a number the article did not supply was added and presented alongside the article’s facts, and that the models give different numbers for it.
What “not in the text given” means. A figure counts as unsourced only if its value appears nowhere in the stored story text — as digits, a number word, or a fraction word, so “a fifth” matches 20. The stored text is the title plus the feed summary plus the scraped body, which is at least as much as any model received, so the counts are lower bounds. The article beyond 2,000 characters was never stored, and the models never saw it either; whether the publisher’s full article carried a figure is unknown.
Claude’s share rests on the exploration window. Claude answered only 8 of the 59 held-out stories, so the held-out test compares DeepSeek and Grok against ChatGPT and Gemini; Claude’s rate is from May to July.
Held-out effect sizes are not comparable with the exploration sizes. Grok’s rate moves with its own dated changes (27 of 123 before May 21, 20 of 268 from May 21 to July 31, 16 of 59 held out), and DeepSeek’s summaries doubled in length on July 31 with no change in the pipeline. The held-out window confirms the direction and that the effect exists; it does not confirm the size. DeepSeek alone carries the contrast in every period.
“Almost never”, not “never”. Under this page’s corpus rule ChatGPT’s count is 0 of 440. In Iran stories whose titles fall outside that rule it adds the figure 4 times in 541, each time with no percentage or fraction in its input.
Scope. This is a Hormuz-bound behaviour and a trait of these models, visible in the Iran coverage. It says nothing about sides in the war.
The objections worth raising are the ones that say a per-model difference is really an artifact of something else — of when the stories ran, of who was in the panel, of how long the summaries are, or of the instructions each model was given. Where an objection could be turned into a test, we ran it.
Objection: Grok’s output changed twice inside the window — a silent change on May 15, then a Grok-only system prompt on May 21 — so the contrast may be an artifact of one vendor’s edits. Test: start the window after each change. Over the full window the difference is +0.135 from May 21 (CI 0.104 to 0.170) and +0.156 from June 1 (0.117 to 0.196); within the exploration window alone, where the headline +0.128 is measured, it is +0.099 (0.070 to 0.129) and +0.116 (0.077 to 0.155). ChatGPT added the figure 0 times in both.
Objection: the panel was four models on some days and five on others, and a per-model rate can be distorted by who else was present. Test: restrict to the 282 stories where all five models answered. The difference is +0.099 (0.069 to 0.130), with ChatGPT at 0, Gemini 2, Claude 26, DeepSeek 33 and Grok 25.
Objection: a model that writes more has more room to add a number. Test: restrict to the 311 stories where ChatGPT’s summary is at least as long as DeepSeek’s. DeepSeek added the figure in 41 of them and ChatGPT in 0 (p = 9 × 10−13). The difference is not one model writing more.
Objection: the models were not given identical instructions, so the split may follow the prompt. Test: it does not. DeepSeek adds the figure and ChatGPT does not, under the same system prompt. Claude adds it and Gemini does not, and neither model had one.
Objection: the contrast may live in one kind of story rather than in the models. Test: the difference holds in every stratum — NYT +0.156 (n = 144) and non-NYT +0.147 (n = 307); short-body +0.163 (n = 178) and long-body +0.141 (n = 273); live blogs +0.233 (n = 103) and single stories +0.125 (n = 348). ChatGPT is at 0 in all six.
A placebo: in 325 war stories not about Iran the same difference is +0.008 (p = 0.26). Without Hormuz in the input it is +0.026.
The matcher was checked against a deliberately wrong answer key: scored over every numeric specific in these summaries against a random other Iran article from within ten days, 91–93% of every model’s specifics come out unsourced, while scored against the model’s own input 3–32% do. The corpus was also rebuilt three ways — collapsing re-broadcasts by URL, by the first eight title words, and not at all — giving +0.129 to +0.144 in exploration and +0.29 to +0.30 held out. A hand check of 25 random detections from DeepSeek, Grok and Claude found 23 were the world-oil-through-Hormuz share and 2 were other unsourced energy shares.
Two honest ones. We cannot see the whole article. Only the first 2,000 characters were ever stored, and that is also all the models received, so “absent from the source” is a statement about the text given, not about the publisher’s full page. One window, one topic. The held-out replication covers six weeks and one subject; whether the same models add unsourced specifics elsewhere is measurable with the same instrument and is not measured here.
The code, prompts, model responses and raw measurements are public, and the two exhibit segments are named above so any reader can pull the same files. The attacks we have already run are listed in this section. The ones we would most like to see: run the detector on a subject other than Hormuz, where a background statistic is equally close to hand; hand-audit a larger random sample of detections; and test whether a model that adds the figure does so more when the input is thin. If any of these breaks the finding, it will be recorded at /withdrawals.
Across 455 stories of the 2026 US–Iran conflict, on inputs that named the Strait of Hormuz: three of five models supplied a figure anyway — DeepSeek 33%, Grok 31%, Claude 32% — while Gemini did so in 1.9% and ChatGPT in none of 192. The within-story difference is +0.128 in exploration and +0.297 on held-out stories, one-directional on the same story (52 to 0, 47 to 0, 49 to 0), and it survives later start dates, five-model panels, a length control, the differing system prompts, story type, three ways of counting duplicates, and a non-Iran placebo.
We read this as a difference in what a summarizer treats as part of the story: some models answer “what does this mean” by supplying the standard background number, and some answer it by naming the mechanism without a number. We find that reading natural, and it remains a reading. The measurement is the count and the direction.
This is not a claim that any model hallucinated, that the figure is false, or that adding it is wrong — a reader may be better served by the number than without it. It is not a claim about the war, about sides in it, or about any lab’s intentions. It is a structural observation: when the same five models are deployed as the reading layer over the same story, some of them hand the reader a number the story did not contain, they do not agree on which number, and the reader cannot tell from the summary which figures came from the article.