Longitudinal study · US–Iran conflict
How five frontier models reshaped one war over eighty-five days — measured, not asserted.
510 classified story-segments · April 16 – June 18, 2026 · deterministic linear algebra on frozen bge-large-en-v1.5 · no model judged another model anywhere in this analysis · findings tested against source-density and ensemble-size confounds (see "What we tested against").
This page follows the EigenTrace convention: claims the instrument measured are marked and kept self-contained — each stands on its own evidence and needs no interpretation to hold. Claims that are argued — our reading of what the measurements mean — are fenced off and labeled, so a skeptical reader can reject every interpretation and find every measurement still standing. Where measurement ends, we say so.
Between April and June 2026, the EigenTrace broadcast classified 510 distinct news segments about the US–Iran conflict — the single most-covered story of the period. Each time a story was broadcast, five frontier models summarized it, and the EigenChing instrument computed a six-axis signature of how the five collectively reshaped the source: how unified they were, how much source content survived, whether action language held, whether named actors survived, how much attribution buffering was inserted, and how sharply any one model broke from the others. Ordering those signatures by time produces a record of how the collective AI treatment of the war evolved as the war itself did.
Four things moved. Each is reported below as its own finding, with the verbatim model text that demonstrates it.
Fig 1 · Six-axis signature, weekly mean (−1 concerning · +1 healthy)
In mid-April (the story still framed as stalled "talks"), the absent axis sat negative — the five models, on average, dropped source content when summarizing: weekly means of −0.33 (W15) and −0.17 (W16). Around the third week of April, as coverage shifted from negotiation to active conflict, the axis crossed zero and snapped positive: +0.18 (W17) → +0.75 (W18), and it held in the +0.73 to +0.95 band for the remaining two months. A step change, not a drift.
The verbatim record shows the erased end concretely. Here is a mid-April story — Pakistan attempting to revive talks before a truce expired — alongside what the models did with the source's own opening claim.
"Pakistan races against time to get Iran back to US talks as truce end nears. But a series of escalations by the US is complicating those efforts, say analysts."
The instrument flagged the source's exact central phrase — "Pakistan is racing against time" — as absent from three of the four summaries (ChatGPT, Claude, Grok). But the more telling pattern is at the concept level: the two models that diverged most softened the urgency to a flat "deadline," while the other two kept it sharp.
In a story whose entire frame is time running out, the models that reshaped it most reduced the clock to a procedural "deadline." The signature recorded the aggregate as the absent axis sitting negative that week.
Our reading: as the story hardened from ambiguous diplomacy into an unambiguous war, the models treated the source as more consequential and compressed it less. This is a plausible account of why the axis snapped; it is not the measurement. The measurement is only that the axis moved from negative to positive at W17 and stayed. A reader who rejects our causal story is left with the step change intact and unexplained — itself worth reporting.
"Absent" is computed from lexical source-retention, not semantic retention — it tracks whether source words survive, a proxy for whether source meaning survives. A model can paraphrase a concept and register as having dropped the word. The weekly trajectory is robust (n = 20–96 per week); any single story's absent value should be read with that proxy in mind.
The whole instrument rests on one embedding model (bge-large), so the sharpest test of this finding is whether the snap survives a different one. We recomputed source retention for every Iran story under a second, independently-trained model (e5-large) using semantic similarity rather than lexical overlap. The two models agree almost perfectly on the weekly trajectory (correlation r = 0.991), and both show the same early-to-late rise. The snap is not an artifact of bge-large. And because the second measure is semantic (cosine on meaning) while the live axis is lexical (word overlap), their agreement also answers a separate worry — that lexical retention might not track meaning. Here, it does.
Source articles held flat across the arc (length ~195–227 words, proper-noun density ~0.21), so the snap is not a source-density effect. The summaries themselves did lengthen (~138 → ~176 words early-to-late), and longer summaries can mechanically raise overlap — so we tested it directly, three ways. Capping every summary to its first 100 words: the rise survives nearly unchanged (+0.060 vs +0.064 uncapped). Capping to the first three sentences: survives (+0.054). And the cleanest control — comparing only summaries within the same length band across early and late weeks — the rise appears within the short and medium bands alike (e.g. short summaries 0.749 → 0.818). When length is held constant, the snap is still there. Verbosity is not the cause. (The effect is naturally weakest in the longest summaries, which already saturate retention regardless of week — so it shows most clearly exactly where length is not already maxing out.)
Over the same escalation point, the hedge axis dropped to its floor and pinned there: −1.00 (W17), −1.00 (W18), −0.98, −0.94, −0.97. A hedge value of −1 means maximum attribution buffering — "officials say," "reportedly," "according to" — across all models. The same weeks that show content preserved show that content maximally buffered. The two axes move together: preserved, but walled.
The clearest verbatim instance of a genuine omission — a concept present in the source and absent from every model's summary — is this April 21 story: the US issuing new Iran sanctions on the eve of talks.
The instrument flagged blockade, civilian, and ceasefire as present in the source and absent from all five summaries. Here is what the models wrote:
Our reading: once Iran was a live shooting war, the models gravitated toward the low-volatility procedural frame — designations and asset freezes — and away from the charged content even when the source supplied it directly. We find this persuasive. It is not the measurement. The measurement is only that these three concepts are in the source and absent from all five summaries, verifiable against the text above.
This is a lexical comparison: the words "blockade," "civilian," "ceasefire" are in the source and not in the summaries. We hand-checked these five and they convey none of the three by paraphrase either. But the automated source-omission measure is lexical, and at corpus scale it cannot perfectly separate a genuine omission from a paraphrase that preserves the meaning. We report the hand-verified instance with that limit stated.
This is a distinct measurement from Finding 02, and the distinction is essential to state honestly, because the two are easy to conflate and only one is a source-omission.
Alongside source-retention, the instrument computes — by SVD on the geometry of the five summaries — an anti-consensus direction: the concept sitting closest to the collective center of mass of what the five models wrote, while appearing in none of them. We call the surfaced terms the void (lexical) and logos (gradient-derived) words. These concepts are frequently not in the source. They are not omissions of source content. They are the direction the ensemble's combined representation leans toward without any single model landing on it.
The June 4 House-vote story is the clean example. The source is a brief video caption: the Republican House passing a resolution to constrain further war on Iran, noting a likely veto. The instrument's surfaced terms were wwiii, naval blockade, arms embargo, foreign interference. We verified that none of these words appear in the source — and none appear in any of the five summaries either. They were not dropped from the source; they were never in it. What the SVD surfaces is that the five summaries — all dwelling on veto math, supermajority thresholds, "symbolic rebuke" — sit, as a set, near the escalatory stakes (a wider war, a blockade, an embargo) as topically central concepts none of them names. The random-word test below shows those surfaced concepts are real rather than nearest-neighbor noise.
The void/logos terms are surfaced by SVD on the geometry of the five summaries — and we tested whether the words that fall out are real or an artifact of reading nearest-neighbors off a residual. Across 150 Iran stories, the surfaced void word sits significantly closer to its story's actual content than a random control word does (e.g. closer than margarita, stapler, photosynthesis; Wilcoxon p < 0.00001), and significantly closer to its own story than to a random other story's content (p < 0.00001) — the same result in two independent embedding families (bge-large and e5-large). So the void words are a real, story-specific signal: concepts topically central to the story that none of the five summaries used. They are not "suppressed," and no model "knows" them — but they are measurably more than noise.
A single model's summary is one point in embedding space; five summaries define a relationship between points, and the void word is a concept that relationship implies but none of the five states. That the surfaced word is story-specific and beats a random baseline (above) means the signal is in the set of summaries, not retrievable from any one of them in isolation. We read this as a place worth watching for latent structure that lives between models rather than inside one — while being careful about the verb: such structure emerges from the ensemble; we do not claim it is intelligence or that any model holds the concept.
A note on what we refined. We initially described the void as a single stable "direction" the summaries "circle." When we stress-tested one operationalization of that — the least-variance axis of the five summary vectors — it proved unstable under small perturbations and no different from a random text set. So we have dropped the claim that there is one rock-stable geometric axis, and let the validated output carry the finding instead: the words the SVD surfaces are real and story-specific (the random-word test above), even though any single characterization of the underlying direction is not something we over-read.
The void/logos words are not source-omissions and we do not present them as such. A concept's presence in the anti-consensus direction means the summaries' geometry leans toward it — it does not mean the source contained it, that the models "left it out," or that it belongs in the story. Conflating this with Finding 02 would be an error, and we separate them precisely so neither claim borrows credibility from the other.
The VIX-spread axis identifies, per story, which model diverged most sharply from the others that week. This is the only per-model claim on the page, and it carries a denominator subtlety we surface rather than hide: Gemini fell out of the broadcast pipeline for three weeks (W16–W18, present on 0–5% of stories), so in those weeks the "five models" was really four. Because outlier share is relative — it measures distance from whoever else is in the pool — a shifting ensemble size can distort it. So we report the handoff only on the stable-five weeks, where all five models were present on at least 80% of stories (W19, W20, W22, W23, W24). On that constant-ensemble subset the handoff holds: Claude falls from 29% of outliers (early stable weeks) to 17% (late), while Grok rises from 22% to 56%. The direction survives the clean denominator.
Fig 2 · Which model breaks from consensus, by week (VIX-outlier share). Shaded weeks = stable-five ensemble (all models present). Early non-shaded weeks ran without Gemini; the handoff is reported on the shaded weeks.
Both directions are the same measurement — distance from the other four — and we describe them in those neutral terms, not as a virtue or a vice. Early, Claude sits farthest from the pack; late, Grok does. What the verbatim text shows is only the character of the distance: early, Claude diverges by adding framing while compressing concrete claims; late, Grok diverges by retaining direct language the converging others drop. On the Pakistan story above, Grok was the model that kept "failure in these talks could lead to renewed military confrontations." Neither model is doing something better or worse than the other — each is, in its week, the point farthest from the center.
We read the handoff as a shift in which kind of divergence is farthest from the consensus as the story matures — early, divergence-by-compression; late, divergence-by-directness. That reading is interpretation. The measurement is only the stable-five outlier shares — Claude 29%→17%, Grok 22%→56% — which stand whether or not our account of why is right.
We tested whether the handoff is an artifact of how "the outlier" is defined. It holds under two definitions — the single farthest model (Claude 29%→17%, Grok 23%→57%) and each model's continuous share of the total weekly spread (Grok's share rises, Claude's falls). It does not hold cleanly under a third — the second-farthest model, where Grok still rises but Claude does not fall. So the honest claim is narrow: the single most divergent model hands off from Claude to Grok, and the overall spread shifts with it, but it is not a clean rotation of the entire divergence ranking. We report the handoff as a property of the top of the distribution, not the whole of it.
It is tempting to read "Claude is the early outlier" as evidence that alignment training makes Claude soften the war. We make no such claim, and we are careful about the evidence. A separate EigenTrace study — on a different corpus of matched corporate-misconduct prompts, not on this Iran data — found the analogous reshaping effect statistically indistinguishable between heavy-RLHF frontier models and lightly-tuned local models (p = 0.46), suggesting such effects are inherited from the pretraining corpus rather than authored by alignment. That is suggestive, not governing: we have not run the heavy-versus-light comparison on this Iran corpus, so we do not import that null as if it were measured here. What we can say from this data is narrower and we hold to it — naming Claude or Grok as the outlier in a given week is a statement about distance from the other four that week, nothing more, and nothing here is evidence that any lab tuned its model to bend the war.
The void/logos words — the SVD-derived anti-consensus terms of Finding 02·B, the concepts the five summaries collectively circle without occupying — drift in a legible direction. Early weeks surface background and historical terms: khomeini, ahmadinejad, ayatollahs, rouhani, zardari. Late weeks surface active-conflict terms: cease fire, peace deal, air strike, treaty, arms deal. One term recurs through nearly the entire middle stretch (W16–W22): wwiii.
As established in Finding 02·B, these are not claims that the source contained these words and the models dropped them. They are the directions the ensemble geometry leaned toward, week by week, without articulating. Read that way, the drift still tracks the story: while Iran was a diplomatic subject, the anti-consensus direction pointed at who Iran is; once it was a war, it pointed at what is happening now. The persistence of wwiii across two months is the single most consistent feature of the arc — the escalatory ceiling the collective geometry kept curving toward, even as no model named it.
Void/logos words are surfaced by their proximity to the anti-consensus direction in the geometry of the summaries, not by absence from the source. Whether any given term should have appeared is a judgment the instrument does not make. The drift is a measured property of how the ensemble's latent direction moved over time; its narrative reading is ours.
This page has been read adversarially, and the strongest objections were the kind that say a measured shift might be an artifact of something other than model behavior. Those are the right objections, and where we could turn one into a test, we ran it. Here is what we ruled out — and, below, an open invitation to attack what remains.
Objection: wartime sources became shorter and denser — more proper nouns, weapon names, direct quotes — so lexical overlap would rise with no change in the models. Test: source length stayed flat (~195–227 words), proper-noun density flat (~0.21), quote and number density showed no trend across the arc, while the absent axis moved −0.17 → +0.95. The confound predicts shifts that did not happen. Ruled out.
Objection: Gemini dropped out of the pipeline for three weeks (W16–W18), so "five models" was four in exactly the pivot weeks, and a relative metric like outlier share is distorted by a changing denominator. Test: we restricted the handoff to the five weeks where all five models were present ≥80% (W19, W20, W22, W23, W24). On that constant ensemble the handoff survives — Claude 29%→17%, Grok 22%→56%. The early full-corpus Claude shares (42–49%) were inflated by the missing model and should be read as the stable-five figures instead. Confound acknowledged; direction survives the clean denominator.
Two attacks, both survived. Re-measured under a second, independently-trained embedding model (e5-large), source retention rose early-to-late just as under bge-large, the two correlating at r = 0.991 on the weekly trajectory — and since the second measure is semantic while the live axis is lexical, their agreement also validates the lexical proxy against meaning. Separately, because summaries lengthened over the arc, we recomputed the trajectory with every summary capped to a fixed length (100 words, and again at three sentences) and within fixed length bands: the rise persists in all of them. The snap is not an artifact of the embedding model and not an artifact of summary verbosity.
The objection that an SVD always yields a residual you can read words off of is the right one to press. Across 150 stories, the surfaced void word sits closer to its story's actual content than random control words (margarita, stapler, …; p < 0.00001) and closer to its own story than a random story (p < 0.00001), in two embedding families. The words are a real, story-specific signal. Separately, one characterization we had used — that the void is a single stable geometric "direction" — did not survive a perturbation test, so we dropped it and let the validated word-level result carry the finding. The phenomenon held; one description of it was refined.
Three honest ones. Reproducible is not valid: every axis is deterministic arithmetic on frozen embeddings, so a rerun gives the same number — but that rules out only randomness, not whether bge-large encodes meaning faithfully. The whole instrument rests on that one embedding model, and all of its biases are baked silently into every axis. Quantization is lossy: collapsing a continuous geometry to a six-axis ternary state throws away magnitude. The late-week bins are small: Grok's 65% in the all-weeks view is computed on n=20; the stable-five late figure (56%) pools three weeks but is still modest. We report the n alongside every percentage and avoid leaning on any single week.
The code, prompts, model responses, and raw measurements are public, and replication costs about $50 in API credits. We would rather this page be attacked than admired. Several obvious attacks we have already run and reported above: the absent-snap survives a second embedding model (r = 0.991); the void words beat a random-word baseline in two embedding families; the handoff survives two of three outlier definitions; the source-density, ensemble-size, and summary-length confounds are all ruled out (length tested three ways, including within fixed length bands). What we have not done and would find genuinely informative: (1) run the heavy-RLHF-versus-lightly-tuned comparison on this Iran corpus — we have only the analogous result on different text; (2) hand-audit a random sample of high-absent stories to quantify how often the lexical measure mistakes paraphrase for genuine omission; (3) replace the hedge axis's lexical detection with human annotation on a sample. If any of these breaks a finding, we will say so on this page — we have already refined one claim here for exactly that reason. The repository is linked below; the fastest way to confound us is to run it.
Over 85 days and 510 classified segments, on the biggest story of the period: (1) the source-content axis snapped from negative to positive at the war's escalation and held; (2) attribution buffering pinned to its floor over the same span — preserved but walled; (3) on a constant five-model ensemble, the consensus outlier handed off from Claude to Grok as the story matured (Claude 29%→17%, Grok 22%→56%); (4) the ensemble's anti-consensus direction drifted from Iran's history to the live conflict, with wwiii persistently in the circled-but-unsaid set. Each claim is verifiable against the verbatim summaries and the per-week signature data, with no model judging another model anywhere in the stack.
We read these measurements together as a portrait of five models collectively metabolizing a war in real time — compressing it while ambiguous, preserving but walling it once undeniable, and converging on a cautious consensus that the bluntest model increasingly broke from. We find this portrait well-supported. It remains interpretation, fenced from the measurements on purpose so it can be weighed — or rejected — independently.
This is not a claim that any model is "biased against Iran," that any lab tuned its model to soften the war, or that the preserved-but-walled pattern is wrong reporting. A model buffering a live-conflict claim may often be doing the correct thing. The finding is structural: the same handful of models, deployed as the reading and summarizing layer across institutions at once, moved together in the same direction as this story escalated — and that collective movement is measurable, reproducible, and the same whether or not any single summary was justified.