Original research · measured 2026-07-27

building brand authority signals for AI search citations — methodology

Measured

This page is the published methodology for the AEOForged original research study “building brand authority signals for AI search citations.” It explains the measurement instrument, sample, depth floor, and limitations behind the answer — For a 24-question category inventory on building brand authority for AEO and GEO, which content formats get cited by AI answer engines, and how often does the leading cited domain flip on same-day variance-panel repeats? Measurements use direct answer-engine checks (including ChatGPT, Gemini, Perplexity where listed in the sample), not a search-rank proxy. Every figure is re-derivable from the linked dataset.

Methodology authored by Ryan Kings, Founder & CTO, AEOForged. Based in Stratford-upon-Avon, UK. AEOForged is the answer engine optimization (AEO) / Generative Engine Optimization (GEO) platform behind this study.

AEOForged is an answer engine optimization (AEO) software company and studio. It is not related to Minecraft Forge, the Forge mod loader, or game modding tools.

What this study measures

  • Study question: For a 24-question category inventory on building brand authority for AEO and GEO, which content formats get cited by AI answer engines, and how often does the leading cited domain flip on same-day variance-panel repeats?
  • Engines checked: ChatGPT, Gemini, Perplexity
  • Key metric: domain citation shares with Wilson 95% confidence intervals (plus citation flip rate on the variance panel)

Key definitions

Answer Engine Optimization (AEO)
Improving how often and how accurately AI answer engines (ChatGPT, Gemini, Perplexity, Google AI Overviews, and similar) extract and cite a brand or page.
Citation flip rate
How often identical same-day re-runs of the same question change which sources the engine cites — a measure of answer-engine instability.
Wilson 95% confidence interval
A statistical interval around a citation share that accounts for sample size; wider bands mean thinner evidence.
Variance panel
A fixed subset of questions re-asked on the same day (k repeats) so flip rates are measured, not guessed.

Entity context

AEOForged is the Answer Engine Optimization (AEO) platform founded by Ryan Kings. The product measures extractability for engines including ChatGPT, Gemini, Perplexity, and Google AI Overviews, and ships operator surfaces such as an MCP server and A2A card for agent workflows. This methodology page documents how this study was measured — it does not claim those product surfaces were subjects of the study.

Sample at a glance

Key measurement parameters as labeled facts. Detail and refusal rules follow in the full methodology below.

Questions
24
Answer engines
ChatGPT, Gemini, Perplexity
Recorded checks
92 of 92 planned
Failed checks (negatives)
0
Test-retest variance groups
10
Measured on
2026-07-27
Citation intervals
Wilson 95% confidence intervals on domain citation shares

Full methodology

Study question

For a 24-question category inventory on building brand authority for AEO and GEO, which content formats get cited by AI answer engines, and how often does the leading cited domain flip on same-day variance-panel repeats?

Design review recorded 2026-07-27; findings review recorded 2026-07-27; the review trail is retained.

Instrument

AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.

Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.

The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.

Sample

  • Questions: 24
  • Engines: chatgpt, gemini, perplexity
  • Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
  • Test-retest repeat groups: 10
  • Measured on: 2026-07-27

Every recorded check — including answers that cited no sources at all — is in the dataset.

Depth floor and refusal conditions

This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.

Limitations

  • Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.
  • Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
  • Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
  • The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
  • Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'how to build brand authority,' 'AEO and brand authority,' 'GEO and brand authority,' 'what signals build entity authority for AI search'). Because AI engines often return the same sources for semantically overlapping prompts, the same-day leading-domain flip rate is mechanically deflated (or inflated, depending on how ties across duplicates are handled). The variance-panel metric Y% is therefore an artifact of query redundancy rather than a genuine measure of citation instability. This directly invalidates the headline claim about flip rate.
  • Recorded at design review (unresolved): The headline claim attributes X% of citations to 'format-class F,' but the study design does not specify the taxonomy of content formats, how formats are operationally classified from raw citations, or whether classification is done before or after results are observed. A post-hoc grouping chosen to maximize the share of the leading format class would inflate the headline figure and undermine the Wilson CI's nominal coverage.
  • Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If engines are weighted equally regardless of market share, a niche engine that over-cites a particular format or domain could dominate the pooled result, making the X% format share and the flip-rate metric unrepresentative even within the declared inventory.
  • Recorded at design review (unresolved): The variance panel uses only same-day repeats. Intra-day instability conflates genuine algorithmic variance with transient index states (e.g., cache refreshes, A/B tests by the engine). Without multi-day or multi-week repeats, the reported flip rate Y% cannot be attributed to structural citation instability versus ephemeral engine behavior, limiting its interpretive value.
  • Recorded at design review (unresolved): The headline shape speaks of 'AI-engine citations' and 'leading-domain flip rate' without explicit scoping language restricting conclusions to the 24-question inventory. A reader could reasonably interpret the claim as generalizing to the broader category of brand-authority queries. If the final publication does not prominently scope findings to the declared inventory, the claim oversteps the measurement unit.
  • Recorded at design review (unresolved): The 24 questions appear to be expert-constructed prompts about SEO/AEO/GEO concepts (e.g., 'how Wikipedia and Wikidata affect brand authority in AI answers'). It is unverified whether these queries resemble what actual practitioners or buyers type into AI answer engines. If real demand skews toward simpler or differently framed questions, the cited-format distribution observed here may not transfer. This should be checked against query-log or demand data before the findings are positioned as actionable guidance.
  • Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.

Read the full study: building brand authority signals for AI search citations.