Research · Evidence

AI search evidence map: audit scores, tiers, reports, and stages

In brief. This is the evidence map that connects C-SEO Bench, What Gets Cited, and other primary research to every aiseo-audit score, tier, report, factor, measured outcome, and pipeline stage. Each record states what the evidence supports.

Maintained by Jeff Patterson and Agency Enterprise · Updated August 30, 2026

Five evidence tiers: overview

The tier communicates confidence; the point value communicates an internal audit weight. Neither one is a predicted increase in citation rate. The reason for this separation is that a research result and a software weight answer different questions. This means readers inspect the evidence without treating the score as an outcome forecast.

TierMeaningHow the audit treats it
supportedSupported refers to direct outcome evidence in a strong, relevant experimental regime.The audit scores it within the tested scope.
conditionalConditional means that credible evidence depends on a domain, stage, engine, or metric.The audit scores it conservatively and states the condition.
heuristicHeuristic is defined as a useful proxy or engineering judgment without isolated causal validation.It guides a review, but it is not a measured gain.
diagnosticDiagnostic refers to an observable property that research does not justify rewarding.The report shows it as 0/0 and excludes it from the score.
experimentalExperimental means that a feature or standard lacks stable outcome evidence.The report labels it explicitly and excludes it from causal claims.

Evidence is mapped to pipeline stages

A retrieval study cannot justify a citation-fitness claim by itself. Each evidence row names the pipeline stage and the metric the paper measured. A pipeline stage is a type of boundary that separates page access, retrieval alignment, citation fitness, and provenance.

TE

Technical eligibility

Fetch, extraction, and crawler-access prerequisites.

RA

Retrieval alignment

Signals associated with entering or ranking within context.

CF

Citation fitness

Properties measured after content is available to the generator.

PV

Provenance

Authorship and source-identity hygiene, labeled honestly as proxies.

Null results remain visible

Version 2.0 moved eleven formatting and presence checks out of the score. The report still shows those observations. Controlled research did not justify awarding points for their mere presence.

Diagnostic does not mean useless.

It means the tool reports what it saw without quietly converting an unvalidated observation into points.

The release traceability rule

Every factor needs two matching records: one in the public evidence map and one in the runtime registry. Both use the same tier and stage. A mismatch blocks the change and makes undocumented scoring changes visible in review. This works because reviewers compare a public research record with the code that produces the report.

The complete factor-level table remains available in the project’s EVIDENCE.md ↗.

How to verify an evidence record

The check works by following one claim from the report to the runtime registry and primary paper. As a consequence, a missing link becomes visible before a scoring change ships.

  1. Open the factor in the public evidence map.
  2. Follow the primary-source link and confirm the measured outcome.
  3. Check the domain, engine, data, and limits recorded for the study.
  4. Compare the tier and pipeline stage with the runtime registry.
  5. Reject a scoring change when either record is missing or inconsistent.

Key takeaways

  • The tier is the label for evidence strength and scope.
  • The point is an internal audit weight, not a predicted gain.
  • The stage is the part of the retrieval and citation pipeline a paper measured.
  • The diagnostic is visible but excluded from the score.
  • The map is the public record that keeps each factor inspectable.

Bottom line: each scored factor needs a matching source, tier, stage, and runtime record.

Sources

According to C-SEO Bench, source position outweighed the fixed rewrites it tested [1]. According to What Gets Cited, topic match and product completeness acted as gates in 252,000 controlled trials[2].

According to the project evidence map, each runtime factor records a tier and stage [3]. According to the migration guide, version 2 moved unsupported presence checks outside the score [4]. According to the version 2.0.0 release notes, the evidence model ships with the analyzer[5].

  1. C-SEO Bench, NeurIPS 2025.
  2. What Gets Cited, SIGIR 2026.
  3. aiseo-audit evidence map, project research record.
  4. Version 2 migration guide, project documentation.
  5. aiseo-audit 2.0.0 release notes, project release record.