Original research · measured 2026-07-27

building brand authority signals for AI search citations — methodology

Measured

This AEOForged methodology covers 24 category questions asked to ChatGPT, Gemini, Perplexity (92 of 92 planned checks recorded; 0 failed checks kept as negatives). It answers: For a 24-question category inventory on building brand authority for AEO and GEO, which content formats get cited by AI answer engines, and how often does the leading cited domain flip on same-day variance-panel repeats?

Methodology authored by Ryan Kings, Founder & CTO, AEOForged. Based in Stratford-upon-Avon, UK. AEOForged is the answer engine optimization (AEO) / Generative Engine Optimization (GEO) platform behind this study.

Sample at a glance

Measurement parameters for this freeze. Refusal rules and procedure follow in the methodology body.

Questions
24
Answer engines
ChatGPT, Gemini, Perplexity
Recorded checks
92 of 92 planned
Failed checks (negatives)
0
Test-retest variance groups
10
Measured on
2026-07-27
Citation intervals
Wilson 95% confidence intervals on domain citation shares

Scope

  • Study question: For a 24-question category inventory on building brand authority for AEO and GEO, which content formats get cited by AI answer engines, and how often does the leading cited domain flip on same-day variance-panel repeats?
  • Engines checked: ChatGPT, Gemini, Perplexity
  • Key metric: domain citation shares with Wilson 95% confidence intervals (plus citation flip rate on the variance panel)

How to read the results

  1. Read Sample at a glance first — those labeled facts are the admission parameters for the freeze.
  2. Treat every failed or unreachable check as data; the study does not drop negatives to inflate rates.
  3. Citation shares (engine lane) carry Wilson 95% confidence intervals — wider bands mean thinner evidence.
  4. Use the linked dataset to re-derive every figure; narrative claims that lack a row are out of scope.

Key terms

Answer Engine Optimization (AEO)
Improving how often and how accurately AI answer engines (ChatGPT, Gemini, Perplexity, Google AI Overviews, and similar) extract and cite a brand or page.
Citation flip rate
How often identical same-day re-runs of the same question change which sources the engine cites — a measure of answer-engine instability.
Wilson 95% confidence interval
A statistical interval around a citation share that accounts for sample size; wider bands mean thinner evidence.
Variance panel
A fixed subset of questions re-asked on the same day (k repeats) so flip rates are measured, not guessed.

Full methodology

Instrument

AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.

Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.

The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.

Sample

  • Questions: 24
  • Engines: chatgpt, gemini, perplexity
  • Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
  • Test-retest repeat groups: 10
  • Measured on: 2026-07-27

Every recorded check — including answers that cited no sources at all — is in the dataset.

Depth floor and refusal conditions

This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.

  • Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
  • Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
  • The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
  • Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'how to build brand authority,' 'AEO and brand authority,' 'GEO and brand authority,' 'what signals build entity authority for AI search'). Because AI engines often return the same sources for semantically overlapping prompts, the same-day leading-domain flip rate is mechanically deflated (or inflated, depending on how ties across duplicates are handled). The variance-panel metric Y% is therefore an artifact of query redundancy rather than a genuine measure of citation instability. This directly invalidates the headline claim about flip rate.
  • Recorded at design review (unresolved): The headline claim attributes X% of citations to 'format-class F,' but the study design does not specify the taxonomy of content formats, how formats are operationally classified from raw citations, or whether classification is done before or after results are observed. A post-hoc grouping chosen to maximize the share of the leading format class would inflate the headline figure and undermine the Wilson CI's nominal coverage.
  • Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If engines are weighted equally regardless of market share, a niche engine that over-cites a particular format or domain could dominate the pooled result, making the X% format share and the flip-rate metric unrepresentative even within the declared inventory.
  • Recorded at design review (unresolved): The variance panel uses only same-day repeats. Intra-day instability conflates genuine algorithmic variance with transient index states (e.g., cache refreshes, A/B tests by the engine). Without multi-day or multi-week repeats, the reported flip rate Y% cannot be attributed to structural citation instability versus ephemeral engine behavior, limiting its interpretive value.
  • Recorded at design review (unresolved): The headline shape speaks of 'AI-engine citations' and 'leading-domain flip rate' without explicit scoping language restricting conclusions to the 24-question inventory. A reader could reasonably interpret the claim as generalizing to the broader category of brand-authority queries. If the final publication does not prominently scope findings to the declared inventory, the claim oversteps the measurement unit.
  • Recorded at design review (unresolved): The 24 questions appear to be expert-constructed prompts about SEO/AEO/GEO concepts (e.g., 'how Wikipedia and Wikidata affect brand authority in AI answers'). It is unverified whether these queries resemble what actual practitioners or buyers type into AI answer engines. If real demand skews toward simpler or differently framed questions, the cited-format distribution observed here may not transfer. This should be checked against query-log or demand data before the findings are positioned as actionable guidance.
  • Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.

Read the full study: building brand authority signals for AI search citations.

Limitations (disclosed)

Extracted from the methodology body. Status is always disclosed — never omitted to inflate confidence.

LimitationStatus
Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.Disclosed
Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.Disclosed
Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.Disclosed
The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.Disclosed
Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'how to build brand authority,' 'AEO and brand authority,' 'GEO and brand authority,' 'what signals build entity authority for AI search'). Because AI engines often return the same sources for semantically overlapping prompts, the same-day leading-domain flip rate is mechanically deflated (or inflated, depending on how ties across duplicates are handled). The variance-panel metric Y% is therefore an artifact of query redundancy rather than a genuine measure of citation instability. This directly invalidates the headline claim about flip rate.Disclosed
Recorded at design review (unresolved): The headline claim attributes X% of citations to 'format-class F,' but the study design does not specify the taxonomy of content formats, how formats are operationally classified from raw citations, or whether classification is done before or after results are observed. A post-hoc grouping chosen to maximize the share of the leading format class would inflate the headline figure and undermine the Wilson CI's nominal coverage.Disclosed
Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If engines are weighted equally regardless of market share, a niche engine that over-cites a particular format or domain could dominate the pooled result, making the X% format share and the flip-rate metric unrepresentative even within the declared inventory.Disclosed
Recorded at design review (unresolved): The variance panel uses only same-day repeats. Intra-day instability conflates genuine algorithmic variance with transient index states (e.g., cache refreshes, A/B tests by the engine). Without multi-day or multi-week repeats, the reported flip rate Y% cannot be attributed to structural citation instability versus ephemeral engine behavior, limiting its interpretive value.Disclosed

Ryan Kings (Founder & CTO) publishes this instrument under AEOForged (founded 2026, Stratford-upon-Avon, UK). See /proof for the self-audit dossier (e.g. case-study AEO score 8→81, entity-chain figures) — directional labels and negatives apply. That dossier is product proof, not this freeze. Proof dossier.