Original research · measured 2026-07-27

evaluating and trusting AEO and GEO platforms, agencies, and citation-tracking tools — methodology

Measured

This page is the published methodology for the AEOForged original research study “evaluating and trusting AEO and GEO platforms, agencies, and citation-tracking tools.” It explains the measurement instrument, sample, depth floor, and limitations behind the answer — For a 24-question category inventory on how buyers evaluate, vet, and trust AEO/GEO platforms, marketing agencies, and citation-tracking tools, which source domains do answer engines actually cite, and how often does the leading cited domain flip on same-day variance-panel repeats? Measurements use direct answer-engine checks (including ChatGPT, Gemini, Perplexity where listed in the sample), not a search-rank proxy. Every figure is re-derivable from the linked dataset.

Methodology authored by Ryan Kings, Founder & CTO, AEOForged. Based in Stratford-upon-Avon, UK. AEOForged is the answer engine optimization (AEO) / Generative Engine Optimization (GEO) platform behind this study.

AEOForged is an answer engine optimization (AEO) software company and studio. It is not related to Minecraft Forge, the Forge mod loader, or game modding tools.

What this study measures

  • Study question: For a 24-question category inventory on how buyers evaluate, vet, and trust AEO/GEO platforms, marketing agencies, and citation-tracking tools, which source domains do answer engines actually cite, and how often does the leading cited domain flip on same-day variance-panel repeats?
  • Engines checked: ChatGPT, Gemini, Perplexity
  • Key metric: domain citation shares with Wilson 95% confidence intervals (plus citation flip rate on the variance panel)

Key definitions

Answer Engine Optimization (AEO)
Improving how often and how accurately AI answer engines (ChatGPT, Gemini, Perplexity, Google AI Overviews, and similar) extract and cite a brand or page.
Citation flip rate
How often identical same-day re-runs of the same question change which sources the engine cites — a measure of answer-engine instability.
Wilson 95% confidence interval
A statistical interval around a citation share that accounts for sample size; wider bands mean thinner evidence.
Variance panel
A fixed subset of questions re-asked on the same day (k repeats) so flip rates are measured, not guessed.

Entity context

AEOForged is the Answer Engine Optimization (AEO) platform founded by Ryan Kings. The product measures extractability for engines including ChatGPT, Gemini, Perplexity, and Google AI Overviews, and ships operator surfaces such as an MCP server and A2A card for agent workflows. This methodology page documents how this study was measured — it does not claim those product surfaces were subjects of the study.

Sample at a glance

Key measurement parameters as labeled facts. Detail and refusal rules follow in the full methodology below.

Questions
24
Answer engines
ChatGPT, Gemini, Perplexity
Recorded checks
92 of 92 planned
Failed checks (negatives)
0
Test-retest variance groups
10
Measured on
2026-07-27
Citation intervals
Wilson 95% confidence intervals on domain citation shares

Full methodology

Study question

For a 24-question category inventory on how buyers evaluate, vet, and trust AEO/GEO platforms, marketing agencies, and citation-tracking tools, which source domains do answer engines actually cite, and how often does the leading cited domain flip on same-day variance-panel repeats?

Design review recorded 2026-07-27; findings review recorded 2026-07-27; the review trail is retained.

Instrument

AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.

Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.

The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.

Sample

  • Questions: 24
  • Engines: chatgpt, gemini, perplexity
  • Recorded checks: 92 of 92 planned calls (0 failed — recorded as negatives)
  • Test-retest repeat groups: 10
  • Measured on: 2026-07-27

Every recorded check — including answers that cited no sources at all — is in the dataset.

Depth floor and refusal conditions

This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.

Limitations

  • Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.
  • Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
  • Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
  • The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
  • Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'best tools to track ChatGPT citations' / 'best tools to track Perplexity citations'; 'best GEO tools for marketers' / 'AEO platforms for marketing agencies'). Because answer engines often return the same top-cited domain for semantically overlapping queries, these clusters mechanically inflate the leading-domain citation share and deflate the same-day flip rate. The headline claim of X% leading-domain share and Y% flip rate would be an artifact of query-cluster redundancy rather than a genuine measure of citation concentration across 24 distinct evaluation topics.
  • Recorded at design review (unresolved): The variance panel uses only same-day repeats, so it captures intra-day stochasticity but not the far larger between-day, between-index-update, or between-model-version variance. Reporting the flip rate as 'Y%' without qualifying it as same-day-only would understate true citation instability and mislead readers into treating the metric as a general reliability estimate. If the headline shape does not explicitly restrict the flip rate to same-day variance, the claim overstates measurement stability.
  • Recorded at design review (unresolved): The design says '24 questions asked to every enabled engine' but does not specify how results are aggregated across engines. If each engine is weighted equally regardless of its real-world usage share, a niche engine that always cites one domain could swing the leading-domain share and flip rate. Conversely, if only majority-rule across engines is used, minority engines' divergent citation patterns are suppressed. Without a declared, justified aggregation rule, the headline numbers are not reproducible and could be steered by post-hoc engine inclusion/exclusion.
  • Recorded at design review (unresolved): The inventory questions read like researcher-composed category prompts (e.g., 'what to look for in an AEO platform', 'red flags when evaluating an AEO agency'). It is unverified whether real buyers phrase evaluation queries this way when interacting with answer engines. If actual buyer queries differ materially in syntax or intent, the citation patterns observed here may not reflect the citation landscape buyers actually encounter. This should be verified against query-demand or log data before the results are positioned as relevant to buyer decision-making.
  • Recorded at design review (unresolved): Answer engines surface sources in multiple ways — inline hyperlinks, footnote-style references, 'Sources' drawers, brand mentions without links, and paraphrased attributions. The design does not define which of these count as a 'citation.' Different operationalizations would yield different domain counts and different leading-domain identities, making the headline claim sensitive to an undisclosed measurement decision.
  • Recorded at design review (unresolved): At least three of the ten sample questions ('is AEO a scam', 'red flags when evaluating an AEO agency', 'how to vet an AEO or GEO vendor') carry a skeptical or negative valence. If the full 24-question inventory over-represents skepticism-framed queries relative to neutral or positive evaluation queries, the citation distribution will be skewed toward domains that publish cautionary or debunking content, inflating those domains' share and potentially flipping the leading domain away from what a balanced inventory would show. Because the inventory is the declared population, this is a composition disclosure issue rather than a sampling error, but it still biases the headline if the inventory's valence balance is not reported.
  • Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.

Read the full study: evaluating and trusting AEO and GEO platforms, agencies, and citation-tracking tools.