Study question
For a 24-question category inventory on how buyers evaluate, vet, and trust AEO/GEO platforms, marketing agencies, and citation-tracking tools, which source domains do answer engines actually cite, and how often does the leading cited domain flip on same-day variance-panel repeats?
Design review recorded 2026-07-27; findings review recorded 2026-07-27; the review trail is retained.
Instrument
AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.
Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.
The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.
Sample
- Questions: 24
- Engines: chatgpt, gemini, perplexity
- Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
- Test-retest repeat groups: 10
- Measured on: 2026-07-27
Every recorded check — including answers that cited no sources at all — is in the dataset.
Depth floor and refusal conditions
This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.
Limitations
- Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.
- Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
- Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
- The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
- Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'best tools to track ChatGPT citations' / 'best tools to track Perplexity citations'; 'best GEO tools for marketers' / 'AEO platforms for marketing agencies'). Because answer engines often return the same top-cited domain for semantically overlapping queries, these clusters mechanically inflate the leading-domain citation share and deflate the same-day flip rate. The headline claim of X% leading-domain share and Y% flip rate would be an artifact of query-cluster redundancy rather than a genuine measure of citation concentration across 24 distinct evaluation topics.
- Recorded at design review (unresolved): The variance panel uses only same-day repeats, so it captures intra-day stochasticity but not the far larger between-day, between-index-update, or between-model-version variance. Reporting the flip rate as 'Y%' without qualifying it as same-day-only would understate true citation instability and mislead readers into treating the metric as a general reliability estimate. If the headline shape does not explicitly restrict the flip rate to same-day variance, the claim overstates measurement stability.
- Recorded at design review (unresolved): The design says '24 questions asked to every enabled engine' but does not specify how results are aggregated across engines. If each engine is weighted equally regardless of its real-world usage share, a niche engine that always cites one domain could swing the leading-domain share and flip rate. Conversely, if only majority-rule across engines is used, minority engines' divergent citation patterns are suppressed. Without a declared, justified aggregation rule, the headline numbers are not reproducible and could be steered by post-hoc engine inclusion/exclusion.
- Recorded at design review (unresolved): The inventory questions read like researcher-composed category prompts (e.g., 'what to look for in an AEO platform', 'red flags when evaluating an AEO agency'). It is unverified whether real buyers phrase evaluation queries this way when interacting with answer engines. If actual buyer queries differ materially in syntax or intent, the citation patterns observed here may not reflect the citation landscape buyers actually encounter. This should be verified against query-demand or log data before the results are positioned as relevant to buyer decision-making.
- Recorded at design review (unresolved): Answer engines surface sources in multiple ways — inline hyperlinks, footnote-style references, 'Sources' drawers, brand mentions without links, and paraphrased attributions. The design does not define which of these count as a 'citation.' Different operationalizations would yield different domain counts and different leading-domain identities, making the headline claim sensitive to an undisclosed measurement decision.
- Recorded at design review (unresolved): At least three of the ten sample questions ('is AEO a scam', 'red flags when evaluating an AEO agency', 'how to vet an AEO or GEO vendor') carry a skeptical or negative valence. If the full 24-question inventory over-represents skepticism-framed queries relative to neutral or positive evaluation queries, the citation distribution will be skewed toward domains that publish cautionary or debunking content, inflating those domains' share and potentially flipping the leading domain away from what a balanced inventory would show. Because the inventory is the declared population, this is a composition disclosure issue rather than a sampling error, but it still biases the headline if the inventory's valence balance is not reported.
- Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.
Read the full study: evaluating and trusting AEO and GEO platforms, agencies, and citation-tracking tools.