AI & TechArtificial IntelligenceNewswireTechnology

AI Citation Test: Source Order Matters Less Than It Looks

▼ Summary

– Researchers Sriram Selvam and Anneswa Ghosh conducted a study on how source positioning affects AI search citations using the GPT-5.4 agent.
– The test involved replaying conversations with paired sources in different orders to isolate the impact of position versus content quality.
– Initial results showed a significant citation gap between top and fifth-positioned sources, though this conflated position with inherent relevance.
– When only the order of matched sources was swapped, the effect on citation rates was negligible and statistically insignificant.
– The findings suggest that while raw position gaps appear large, actual causal influence of ordering alone is minimal when source quality is controlled.

AI citation stability remains a complex variable for content creators, but new research suggests that source order plays a less definitive role than previously assumed. A recent preprint authored by Sriram Selvam and Anneswa Ghosh investigated how altering the position of sources within a search result affects which page an AI model chooses to cite. While initial data showed a significant disparity in citation rates between top and lower-ranked results, subsequent tests involving swapped source orders revealed minimal impact.

The study, posted to arXiv on September 14, examined a specific GPT-5.4 search agent powered by Exa. The experimental design utilized offline replayed conversations to ensure consistency, avoiding any live webpage edits that could introduce external variables. This controlled environment allowed the researchers to isolate specific formatting and positional factors without the noise of real-time web changes.

Experimental Methodology

To conduct the test, the researchers prompted the GPT-5.4 agent to answer 130 common questions using independent web searches. They recorded every message and search result from the 129 questions that received responses. From these transcripts, they identified pairs of pages that appeared in the same search outcomes and were both deemed to support the same factual claim. This screening process aimed to find situations where either page could be fairly cited, ensuring that credit differences stemmed solely from how the model apportioned recognition between two equally valid sources.

After filtering, 113 pairs remained. A blinded human verification confirmed 103 of these as genuine matches. The researchers then replayed each saved conversation four distinct ways. In these replays, one page was placed above or below the other, and its text was presented either as plain paragraphs or rewritten with headings, lists, or tables. Only the final answer was generated again during these re-runs.

Both versions of the text were created via AI rewrites of the original pages. Grok 4.3 handled nearly all rewrites, with GPT-5.4 serving as a fallback for one pair. A separate Grok review ensured factual accuracy across all iterations. The authors noted that because wording varied between the two versions, the test compared two rewrites rather than isolating formatting differences alone.

Position vs. Quality

In the initial analysis of saved transcripts, pages appearing first in a search call were cited 85.1% of the time, compared to just 42.8% for pages in the fifth position. This resulted in a raw difference of 42.3 percentage points. It is crucial to understand that ‘position’ here refers strictly to the order of the five Exa results returned in a single search, not a page’s organic ranking on Google or its general visibility on the live web.

Search providers typically prioritize more relevant pages at the top, meaning this raw difference reflects both position and inherent page quality. When the researchers moved the same page higher within its specific pair, the likelihood of it being cited increased by only 7.9 percentage points. However, this finding did not reach statistical significance after accounting for multiple testing variables.

A second testing set comprising 56 pairs, where only the order was switched, showed an estimated effect of 0.0 points, with a 95% confidence interval ranging from -5.4 to +5.4. The study clarifies that while position influenced citations in some instances, averages derived from raw position data are unreliable indicators of causality.

The Impact of Structured Formatting

Pages rewritten with headings and lists received an average of 0.50 more citation markers per answer compared to identical pages written as plain paragraphs. The 95% confidence interval for this metric ranged from 0.20 to 0.84. Answers in the test were heavily cited, with a median of 29 markers across six documents.

Crucially, the total number of citations per answer did not increase, nor did the count on the non-target page change significantly. The authors interpreted this as credit being focused more heavily on the structured page rather than a general boost in citation volume.

The primary metric planned before the experiment began was whether the page got cited at all. Using structured text increased this likelihood by 4.5 percentage points, though the 95% confidence interval (-1.4 to +10.4) indicated the result was not conclusive. The paper notes that the study could reliably detect effects of about 8.5 points or larger.

A stricter comparison, which kept every word identical but adjusted the layout to one sentence per list row, boosted citation rates across all 113 pairs. However, when this test was repeated with a subset, the effect reversed.

In the discussion section of the paper, the authors shared these insights:

“This is an attribution-sensitivity warning, not an optimization tactic.”

Model Instability and Randomness

The study also highlighted the inherent volatility of AI citation behavior. Researchers tested 120 responses again using the same inputs and found that the decision to cite or ignore the target page changed in 15% of cases, roughly one in seven. The average count effect remained consistent across these reruns, but the binary decision to cite was far less stable.

They estimate that about 45% of the variation in a single run’s effect is due to model randomness. Consequently, the authors recommend rerunning citation tests multiple times and sharing how consistent the results are across those runs. This instability aligns with earlier findings from SparkToro in January, which reported that ChatGPT and Google’s AI Overviews produced the same brand list less than 1% of the time when given the same prompt repeatedly.

Implications for Content Strategy

The raw position gap observed in this test was much larger than the average effect seen when researchers swapped source order. This contrasts with an Ahrefs report from May, which showed that pages cited by AI were about three times more likely to include JSON-LD schema. However, adding schema did not clearly increase citations in a controlled manner.

These findings raise questions about whether correlations seen in vendor reports or internal tracking have ever been tested by actively changing the variable. Relying on a single answer is a weak basis for labeling a citation as won or lost. Furthermore, the study cannot confirm if reformatting a live page boosts citations, since the rewrites only applied to text already retrieved, excluding the crawling, retrieval, and ranking processes that occur before the model sees the content.

Looking ahead, the researchers re-ran the saved searches on Grok 4.3 and discovered that structured rewrites leaned in the same direction. However, less than half of Grok’s first replies followed the correct citation format. The authors recommend further research to test each scenario multiple times, explore different search providers and models, and pay close attention to both the frequency of citations and whether a page is cited at all.

(Source: Search Engine Journal)

Topics

ai search bias 98% llm citation behavior 95% research methodology 85% search provider dynamics 80% statistical significance 75%
Show More