Study question
For a 24-question category inventory on building brand authority for AEO and GEO, which content formats get cited by AI answer engines, and how often does the leading cited domain flip on same-day variance-panel repeats?
Design review recorded 2026-07-27; findings review recorded 2026-07-27; the review trail is retained.
Instrument
AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.
Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.
The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.
Sample
- Questions: 24
- Engines: chatgpt, gemini, perplexity
- Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
- Test-retest repeat groups: 10
- Measured on: 2026-07-27
Every recorded check — including answers that cited no sources at all — is in the dataset.
Depth floor and refusal conditions
This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.
Limitations
- Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.
- Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
- Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
- The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
- Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'how to build brand authority,' 'AEO and brand authority,' 'GEO and brand authority,' 'what signals build entity authority for AI search'). Because AI engines often return the same sources for semantically overlapping prompts, the same-day leading-domain flip rate is mechanically deflated (or inflated, depending on how ties across duplicates are handled). The variance-panel metric Y% is therefore an artifact of query redundancy rather than a genuine measure of citation instability. This directly invalidates the headline claim about flip rate.
- Recorded at design review (unresolved): The headline claim attributes X% of citations to 'format-class F,' but the study design does not specify the taxonomy of content formats, how formats are operationally classified from raw citations, or whether classification is done before or after results are observed. A post-hoc grouping chosen to maximize the share of the leading format class would inflate the headline figure and undermine the Wilson CI's nominal coverage.
- Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If engines are weighted equally regardless of market share, a niche engine that over-cites a particular format or domain could dominate the pooled result, making the X% format share and the flip-rate metric unrepresentative even within the declared inventory.
- Recorded at design review (unresolved): The variance panel uses only same-day repeats. Intra-day instability conflates genuine algorithmic variance with transient index states (e.g., cache refreshes, A/B tests by the engine). Without multi-day or multi-week repeats, the reported flip rate Y% cannot be attributed to structural citation instability versus ephemeral engine behavior, limiting its interpretive value.
- Recorded at design review (unresolved): The headline shape speaks of 'AI-engine citations' and 'leading-domain flip rate' without explicit scoping language restricting conclusions to the 24-question inventory. A reader could reasonably interpret the claim as generalizing to the broader category of brand-authority queries. If the final publication does not prominently scope findings to the declared inventory, the claim oversteps the measurement unit.
- Recorded at design review (unresolved): The 24 questions appear to be expert-constructed prompts about SEO/AEO/GEO concepts (e.g., 'how Wikipedia and Wikidata affect brand authority in AI answers'). It is unverified whether these queries resemble what actual practitioners or buyers type into AI answer engines. If real demand skews toward simpler or differently framed questions, the cited-format distribution observed here may not transfer. This should be checked against query-log or demand data before the findings are positioned as actionable guidance.
- Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.
Read the full study: building brand authority signals for AI search citations.