Original research · measured 2026-07-27

answer engine optimization and AI search visibility — methodology

Measured

This AEOForged methodology covers 24 category questions asked to ChatGPT, Gemini, Perplexity (92 of 92 planned checks recorded; 0 failed checks kept as negatives). It answers: For the linked sector's 24 active category questions on answer-engine optimization and AI search visibility, which specific content formats and structural patterns distinguish domains that get cited from those that do not, and what does the same-day variance-panel flip rate reveal about common mistakes brands make that cost them AI search citations?

Methodology authored by Ryan Kings, Founder & CTO, AEOForged. Based in Stratford-upon-Avon, UK. AEOForged is the answer engine optimization (AEO) / Generative Engine Optimization (GEO) platform behind this study.

Sample at a glance

Measurement parameters for this freeze. Refusal rules and procedure follow in the methodology body.

Questions
24
Answer engines
ChatGPT, Gemini, Perplexity
Recorded checks
92 of 92 planned
Failed checks (negatives)
0
Test-retest variance groups
10
Measured on
2026-07-27
Citation intervals
Wilson 95% confidence intervals on domain citation shares

Scope

  • Study question: For the linked sector's 24 active category questions on answer-engine optimization and AI search visibility, which specific content formats and structural patterns distinguish domains that get cited from those that do not, and what does the same-day variance-panel flip rate reveal about common mistakes brands make that cost them AI search citations?
  • Engines checked: ChatGPT, Gemini, Perplexity
  • Key metric: domain citation shares with Wilson 95% confidence intervals (plus citation flip rate on the variance panel)

How to read the results

  1. Read Sample at a glance first — those labeled facts are the admission parameters for the freeze.
  2. Treat every failed or unreachable check as data; the study does not drop negatives to inflate rates.
  3. Citation shares (engine lane) carry Wilson 95% confidence intervals — wider bands mean thinner evidence.
  4. Use the linked dataset to re-derive every figure; narrative claims that lack a row are out of scope.

Key terms

Answer Engine Optimization (AEO)
Improving how often and how accurately AI answer engines (ChatGPT, Gemini, Perplexity, Google AI Overviews, and similar) extract and cite a brand or page.
Citation flip rate
How often identical same-day re-runs of the same question change which sources the engine cites — a measure of answer-engine instability.
Wilson 95% confidence interval
A statistical interval around a citation share that accounts for sample size; wider bands mean thinner evidence.
Variance panel
A fixed subset of questions re-asked on the same day (k repeats) so flip rates are measured, not guessed.

Full methodology

Instrument

AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.

Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.

The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.

Sample

  • Questions: 24
  • Engines: chatgpt, gemini, perplexity
  • Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
  • Test-retest repeat groups: 10
  • Measured on: 2026-07-27

Every recorded check — including answers that cited no sources at all — is in the dataset.

Depth floor and refusal conditions

This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.

  • Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
  • Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
  • The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
  • Recorded at design review (unresolved): Several questions in the 24-item inventory are near-paraphrases of each other (e.g., 'what is answer engine optimization' vs. 'what is generative engine optimization'; 'best tools to track ChatGPT citations' vs. 'platforms that measure share of voice in ChatGPT'; 'AEO platforms for marketing agencies' vs. 'what are the best AEO tools'). Because the same engines will likely retrieve overlapping source sets for semantically similar queries, the same-day variance-panel flip rate is mechanically suppressed: domains that appear for one phrasing almost certainly appear for its near-duplicate, making citation presence look more stable than it actually is. Conversely, any domain that is absent from one phrasing is likely absent from its twin, inflating the apparent consistency of non-citation. This directly distorts the headline Z% flip-rate metric and the 'common mistakes' narrative derived from it.
  • Recorded at design review (unresolved): The headline attributes citation presence to 'content-format or structural traits,' but domain authority, backlink profile, brand salience in training data, and pre-existing indexing relationships with each engine are powerful confounders. A domain could be cited despite poor structural formatting simply because it is a high-authority site the engine's retrieval stack preferentially surfaces. Without controlling for these factors, the study cannot isolate content format/structure as the distinguishing variable, yet the headline shape implies a causal or at least uniquely explanatory role for those traits.
  • Recorded at design review (unresolved): Both the citation sweep and the variance panel are conducted on the same day. Answer-engine outputs can shift substantially across days or weeks due to index refreshes, model updates, or ranking algorithm changes. A single-day design cannot distinguish stable structural advantages from transient retrieval artifacts. The headline's prescriptive framing ('what to fix first') implies durable, actionable patterns, but the evidence base is a single temporal cross-section with only intra-day replication.
  • Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If results from ChatGPT, Perplexity, Google AI Overviews, and others are pooled with equal weight, an engine that cites many sources per answer could dominate the Y% citation-presence figure, while an engine that cites few could be drowned out. Without a pre-registered aggregation rule, the headline Wilson CI could shift materially depending on the weighting scheme chosen post-hoc.
  • Recorded at design review (unresolved): The headline shape promises to reveal 'the most common structural mistakes that explain why brands are not cited by AI search engines.' This is a causal and population-level claim ('brands' generally, 'AI search engines' generally) that goes well beyond what a 24-question inventory scoped to one sector can support. The study observes correlational differences in citation presence within a narrow prompt set; it cannot establish that the identified patterns generalize to other sectors, other query types, or other engines not in the panel, nor that absence of a structural trait causally explains non-citation.
  • Recorded at design review (unresolved): The 24 questions originate from a revenue-prompt cluster around 'best AEO and AI visibility tools,' but no external demand data (e.g., search volume, chatbot query logs, or buyer-intent surveys) are presented to verify that these questions reflect what real practitioners or buyers actually ask AI assistants. If the inventory over-represents definitional or tool-comparison queries and under-represents implementation, pricing, or integration queries, the structural patterns found may not generalize to the queries that matter most for purchase decisions. This should be verified against available demand data before the headline is published.
  • Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.

Read the full study: answer engine optimization and AI search visibility.

Limitations (disclosed)

Extracted from the methodology body. Status is always disclosed — never omitted to inflate confidence.

LimitationStatus
Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.Disclosed
Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.Disclosed
Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.Disclosed
The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.Disclosed
Recorded at design review (unresolved): Several questions in the 24-item inventory are near-paraphrases of each other (e.g., 'what is answer engine optimization' vs. 'what is generative engine optimization'; 'best tools to track ChatGPT citations' vs. 'platforms that measure share of voice in ChatGPT'; 'AEO platforms for marketing agencies' vs. 'what are the best AEO tools'). Because the same engines will likely retrieve overlapping source sets for semantically similar queries, the same-day variance-panel flip rate is mechanically suppressed: domains that appear for one phrasing almost certainly appear for its near-duplicate, making citation presence look more stable than it actually is. Conversely, any domain that is absent from one phrasing is likely absent from its twin, inflating the apparent consistency of non-citation. This directly distorts the headline Z% flip-rate metric and the 'common mistakes' narrative derived from it.Disclosed
Recorded at design review (unresolved): The headline attributes citation presence to 'content-format or structural traits,' but domain authority, backlink profile, brand salience in training data, and pre-existing indexing relationships with each engine are powerful confounders. A domain could be cited despite poor structural formatting simply because it is a high-authority site the engine's retrieval stack preferentially surfaces. Without controlling for these factors, the study cannot isolate content format/structure as the distinguishing variable, yet the headline shape implies a causal or at least uniquely explanatory role for those traits.Disclosed
Recorded at design review (unresolved): Both the citation sweep and the variance panel are conducted on the same day. Answer-engine outputs can shift substantially across days or weeks due to index refreshes, model updates, or ranking algorithm changes. A single-day design cannot distinguish stable structural advantages from transient retrieval artifacts. The headline's prescriptive framing ('what to fix first') implies durable, actionable patterns, but the evidence base is a single temporal cross-section with only intra-day replication.Disclosed
Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If results from ChatGPT, Perplexity, Google AI Overviews, and others are pooled with equal weight, an engine that cites many sources per answer could dominate the Y% citation-presence figure, while an engine that cites few could be drowned out. Without a pre-registered aggregation rule, the headline Wilson CI could shift materially depending on the weighting scheme chosen post-hoc.Disclosed

Ryan Kings (Founder & CTO) publishes this instrument under AEOForged (founded 2026, Stratford-upon-Avon, UK). See /proof for the self-audit dossier (e.g. case-study AEO score 8→81, entity-chain figures) — directional labels and negatives apply. That dossier is product proof, not this freeze. Proof dossier.