Research
AI search research: primary source reviews, retrieval, and evidence
In brief. This is the primary-source research record for C-SEO Bench, SAGEO Arena, What Gets Cited, and seven other AI search studies. Their results support a pipeline-aware, query-specific view of retrieval, citation fitness, evidence, completeness, and direct lead answers.
Maintained by Jeff Patterson and Agency Enterprise · Updated August 30, 2026
Executive conclusion and overview
Generative engine optimization is real, but there is no universal recipe that guarantees citations. A defensible approach is to make a page technically accessible, aligned with the query, complete enough to answer it, and explicit about its evidence and source identity. Then measure repeatedly.
According to C-SEO Bench, changing source position in model context mattered more than any fixed rewrite it tested [1]. Later research narrowed the interpretation: being crawled, retrieved, reranked, and cited are different events. A tactic that helps one stage is neutral or harmful at another when the measured outcome changes.
This is why aiseo-audit 2.0 reports four pipeline stages and labels every factor by evidence strength. The audit is a transparent, repeatable heuristic. It is not a fitted estimate of citation probability and it does not claim to observe a proprietary engine’s index or ranking system.
The pipeline works by separating access, retrieval, citation fitness, and provenance. This means evidence from one stage does not silently justify a claim about another. The reason for the evidence tiers is to preserve each paper's scope. As a consequence, unsupported observations remain visible without entering the score.
The evidence is a record of measured outcomes, limits, and tool changes from each primary source review.
What makes content citable in AI search?
Citable content is defined as source material that directly answers the target query, supplies verifiable evidence, and remains direct when lifted into an engine-generated response. What Gets Citedtested 252,000 controlled trials and reported that topic match and product completeness behaved as gatekeepers, while other signals acted as smaller differentiators [2].
Position bias in LLM citations
Position bias refers to a language model favoring a source because of where it appears in the supplied context. According to C-SEO Bench, this effect was stronger than the tested content rewrites. The finding supports a firm limit: page edits cannot guarantee retrieval order or placement inside a private model context.
Retrieval-augmented generation sources
Retrieval-augmented generation refers to selecting sources before a model writes its answer. According to SAGEO Arena, aligned structural fields improved retrieval while body-only changes affected the later generation stage differently [3]. This is why the audit reports retrieval alignment separately from citation fitness.
Current findings
These conclusions recur across the reviewed literature or have a strong controlled design within a clearly stated scope.
Lead with the conclusion
A concise lead answer recurs across three reviewed studies.
Separate retrieval from citation fitness
Structural fields help a source enter context. Body evidence helps after retrieval.
Do not reward keyword repetition
Added keyword density hurt retrieval in tested benchmarks. The audit penalizes repetition but never rewards it.
Formatting counts are diagnostics
Lists, tables, and section counts support a review. Controlled tests do not justify points for their mere presence.
Engine behavior remains unstable
The same query produced different sources and citations under stable-looking settings.
Product pages need their own profile
Price, specifications, and comparisons affect product pages differently from informational pages.
Evidence tier refers to the stated strength and scope behind a factor. Supported means that direct outcome evidence exists in a relevant setting. Conditional means that the effect depends on the tested regime. Heuristic is defined as a useful proxy without isolated causal validation. Diagnostic is a type of visible observation that stays outside the score. These labels prevent an observation from silently becoming a promise.
Ten peer-reviewed paper reviews
Each review was completed against the primary paper. It records the experimental regime, metric, results, contradictions, scope limits, and the specific changes the paper justified in the tool.
C-SEO Bench: Does Conversational SEO Work?
Most fixed rewriting tactics did not transfer. Source position in context mattered more than any rewrite the study tested.
What Generative Search Engines Like (AutoGEO)
Engine-specific, query-aware rules beat one universal checklist. The study's visibility metric limits transfer beyond its tested setting.
SAGEO Arena
Aligned structural fields improved retrieval. Body-only rewrites hurt retrieval in tested cases even when they helped generation.
What Gets Cited: Competitive GEO in AI Answer Engines
A controlled study ran 252,000 trials. Topic match and product price acted as gates, while other signals had smaller effects.
Think Before Writing (FeatGEO)
Planning features before rewriting beat generic fluency edits. Query-specific decisions mattered.
Mind Reader
Breaking a query into its demands helped more than applying one broad content template.
IF-GEO
Feedback-led optimization beat static recipes. A deterministic page audit cannot reproduce the method directly.
From Experience to Skill (MAGEO)
Experience across queries guided changes in the study. Results still depend on the engine and domain.
Characterizing Web Search in the Age of Generative AI
Deployed engines change sources and differ from classic rankings. Measure more than once.
From Relevance to Authority
A production engine used authority as a retrieval goal. The study did not validate simple page proxies such as a byline.
Key takeaways and how to use the evidence
- The evidence is strongest when the paper directly measures the outcome named by the factor.
- Start with eligibility. A page an engine cannot fetch or parse cannot benefit from downstream improvements.
- Supply the real queries your audience uses. Query-blind checklists cannot distinguish a strong general page from the right page for a specific request.
- Fix stage constraints before polishing low-confidence signals. Retrieval problems and citation-fitness problems need different interventions.
- Preserve uncertainty. Treat conditional and heuristic factors as hypotheses to test on your pages, not universal rules.
- Re-measure. Deployed engines and fetched web content both change, so a single observation is not a durable outcome.
Bottom line: the baseline is the repeatable page result, while observed engine citations remain a separate outcome.
Run a query-aware check locally with no external AI service.
References
- C-SEO Bench: Does Conversational SEO Work?: NeurIPS 2025 Datasets & Benchmarks.
- What Gets Cited: Competitive GEO in AI Answer Engines: SIGIR 2026.
- SAGEO Arena: KDD 2026.