Overview
This research investigates the impact of repeated identical queries on large language models' (LLMs) brand recommendation generation and their accumulation of cited sources. The study contrasts LLM engines operating with and without web search capabilities to assess repertoire exhaustion and source accumulation dynamics. The core finding differentiates between the behavior of retrieval-enabled LLMs, which exhibit brand recommendation saturation relatively quickly, and non-retrieval LLMs, which continue to introduce novel brand recommendations over extended query sequences. Cited-domain accumulation, however, shows continuous growth across all tested horizons for both types of LLM engines.
Research Context
The study addresses the question of whether repeated identical buying questions lead to the exhaustion of a language model's brand recommendations, specifically depending on its retrieval mechanisms. This inquiry explores the generative and retrieval-augmented capacities of LLMs when subjected to consistent input, examining the diversity and novelty of their output over time. The investigation is framed around understanding how the presence or absence of web search in LLM engines influences the breadth of brand recommendations and the ongoing addition of cited domains.
Approach
The research design involved a substantial experimental setup comprising 300 question-engine cells. This configuration resulted from the permutation of 50 distinct questions with six different LLM engines. Each cell was subjected to 15 runs. The process involved open extraction over 1,470 adjudicated organizations. The six LLM engines were categorized based on their retrieval capabilities: five engines operated without web search, while one engine was retrieval-enabled. For deeper analysis, four deep cells were subjected to runs up to 24 iterations to monitor cited-domain accumulation.
To quantify the observed phenomena, the researchers employed specific statistical estimators: exact rarefaction and Chao2 richness. A parallel fixed-roster extraction methodology was utilized to compare with the open extraction approach. This parallel method aimed to reproduce flat curves based on identical responses, contrasting with the open extraction technique designed to track evolving repertoires and prevent artificial plateaus in the data.
Findings
- **Brand Recommendation Dynamics:**
- The five LLM engines answering without web search continued to add never-seen brands in 86-92% of cells even at run 15. These engines exhibited median repertoires of 15-31 organizations.
- In contrast, the single retrieval-enabled engine demonstrated brand list closure in a higher proportion of cases, with a median of 8 organizations and only 64% of cells still adding new brands by run 15.
- This pattern for the retrieval-enabled engine aligned with observations from four earlier deep cells, where web-search runs reached saturation by run 10.
- A single query run typically yielded 62-77% of the total brand set observed over five runs.
- Across all engines, the median question prompted 38 organizations, with a median of 15 organizations appearing in exactly one engine.
- **Cited-Domain Accumulation:**
- Cited-domain accumulation consistently increased across every tested horizon.
- Four deep cells continued to add domains even at run 24. For these cells, 59-84% of the Chao2 lower-bound estimate was observed, indicating substantial unobserved diversity.
- 44% of the retrieval engine's breadth cells were still adding domains at run 15.
- **Methodological Impact:**
- The use of fixed-roster extraction methodologies, when reproducing flat curves on identical responses, manufactured plateaus. These artificial plateaus were removed when open extraction was applied.
Why This Matters
The findings provide empirical data on the characteristics of LLM outputs under repeated identical query conditions, differentiating between models with and without web retrieval capabilities. The observed brand recommendation exhaustion in retrieval-enabled models and sustained source accumulation across all models offers insights into their internal mechanisms and how they interact with external information. This understanding can inform the design and application of LLMs, particularly in scenarios requiring consistent or novel recommendations and comprehensive source discovery.