Instrument
AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.
Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.
The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.
Sample
- Questions: 24
- Engines: chatgpt, gemini, perplexity
- Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
- Test-retest repeat groups: 10
- Measured on: 2026-07-27
Every recorded check — including answers that cited no sources at all — is in the dataset.
Depth floor and refusal conditions
This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.
- Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
- Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
- The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
- Recorded at design review (unresolved): Several questions in the 24-item inventory are near-paraphrases of each other (e.g., 'what is answer engine optimization' vs. 'what is generative engine optimization'; 'best tools to track ChatGPT citations' vs. 'platforms that measure share of voice in ChatGPT'; 'AEO platforms for marketing agencies' vs. 'what are the best AEO tools'). Because the same engines will likely retrieve overlapping source sets for semantically similar queries, the same-day variance-panel flip rate is mechanically suppressed: domains that appear for one phrasing almost certainly appear for its near-duplicate, making citation presence look more stable than it actually is. Conversely, any domain that is absent from one phrasing is likely absent from its twin, inflating the apparent consistency of non-citation. This directly distorts the headline Z% flip-rate metric and the 'common mistakes' narrative derived from it.
- Recorded at design review (unresolved): The headline attributes citation presence to 'content-format or structural traits,' but domain authority, backlink profile, brand salience in training data, and pre-existing indexing relationships with each engine are powerful confounders. A domain could be cited despite poor structural formatting simply because it is a high-authority site the engine's retrieval stack preferentially surfaces. Without controlling for these factors, the study cannot isolate content format/structure as the distinguishing variable, yet the headline shape implies a causal or at least uniquely explanatory role for those traits.
- Recorded at design review (unresolved): Both the citation sweep and the variance panel are conducted on the same day. Answer-engine outputs can shift substantially across days or weeks due to index refreshes, model updates, or ranking algorithm changes. A single-day design cannot distinguish stable structural advantages from transient retrieval artifacts. The headline's prescriptive framing ('what to fix first') implies durable, actionable patterns, but the evidence base is a single temporal cross-section with only intra-day replication.
- Recorded at design review (unresolved): The design says questions are asked to 'every enabled engine' but does not specify how citations are aggregated across engines. If results from ChatGPT, Perplexity, Google AI Overviews, and others are pooled with equal weight, an engine that cites many sources per answer could dominate the Y% citation-presence figure, while an engine that cites few could be drowned out. Without a pre-registered aggregation rule, the headline Wilson CI could shift materially depending on the weighting scheme chosen post-hoc.
- Recorded at design review (unresolved): The headline shape promises to reveal 'the most common structural mistakes that explain why brands are not cited by AI search engines.' This is a causal and population-level claim ('brands' generally, 'AI search engines' generally) that goes well beyond what a 24-question inventory scoped to one sector can support. The study observes correlational differences in citation presence within a narrow prompt set; it cannot establish that the identified patterns generalize to other sectors, other query types, or other engines not in the panel, nor that absence of a structural trait causally explains non-citation.
- Recorded at design review (unresolved): The 24 questions originate from a revenue-prompt cluster around 'best AEO and AI visibility tools,' but no external demand data (e.g., search volume, chatbot query logs, or buyer-intent surveys) are presented to verify that these questions reflect what real practitioners or buyers actually ask AI assistants. If the inventory over-represents definitional or tool-comparison queries and under-represents implementation, pricing, or integration queries, the structural patterns found may not generalize to the queries that matter most for purchase decisions. This should be verified against available demand data before the headline is published.
- Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.
Read the full study: answer engine optimization and AI search visibility.