Original research · measured 2026-07-27

is answer engine optimization a scam or a legitimate discipline — methodology

Measured

This page is the published methodology for the AEOForged original research study “is answer engine optimization a scam or a legitimate discipline.” It explains the measurement instrument, sample, depth floor, and limitations behind the answer — For a 24-question category inventory on whether AEO/GEO services and tools are legitimate, which source domains do answer engines actually cite when addressing trust and scam concerns, and how often does the leading cited domain flip on same-day variance-panel repeats? Measurements use direct answer-engine checks (including ChatGPT, Gemini, Perplexity where listed in the sample), not a search-rank proxy. Every figure is re-derivable from the linked dataset.

Methodology authored by Ryan Kings, Founder & CTO, AEOForged. Based in Stratford-upon-Avon, UK. AEOForged is the answer engine optimization (AEO) / Generative Engine Optimization (GEO) platform behind this study.

AEOForged is an answer engine optimization (AEO) software company and studio. It is not related to Minecraft Forge, the Forge mod loader, or game modding tools.

What this study measures

  • Study question: For a 24-question category inventory on whether AEO/GEO services and tools are legitimate, which source domains do answer engines actually cite when addressing trust and scam concerns, and how often does the leading cited domain flip on same-day variance-panel repeats?
  • Engines checked: ChatGPT, Gemini, Perplexity
  • Key metric: domain citation shares with Wilson 95% confidence intervals (plus citation flip rate on the variance panel)

Key definitions

Answer Engine Optimization (AEO)
Improving how often and how accurately AI answer engines (ChatGPT, Gemini, Perplexity, Google AI Overviews, and similar) extract and cite a brand or page.
Citation flip rate
How often identical same-day re-runs of the same question change which sources the engine cites — a measure of answer-engine instability.
Wilson 95% confidence interval
A statistical interval around a citation share that accounts for sample size; wider bands mean thinner evidence.
Variance panel
A fixed subset of questions re-asked on the same day (k repeats) so flip rates are measured, not guessed.

Entity context

AEOForged is the Answer Engine Optimization (AEO) platform founded by Ryan Kings. The product measures extractability for engines including ChatGPT, Gemini, Perplexity, and Google AI Overviews, and ships operator surfaces such as an MCP server and A2A card for agent workflows. This methodology page documents how this study was measured — it does not claim those product surfaces were subjects of the study.

Sample at a glance

Key measurement parameters as labeled facts. Detail and refusal rules follow in the full methodology below.

Questions
24
Answer engines
ChatGPT, Gemini, Perplexity
Recorded checks
91 of 92 planned
Failed checks (negatives)
1
Test-retest variance groups
10
Measured on
2026-07-27
Citation intervals
Wilson 95% confidence intervals on domain citation shares

Full methodology

Study question

For a 24-question category inventory on whether AEO/GEO services and tools are legitimate, which source domains do answer engines actually cite when addressing trust and scam concerns, and how often does the leading cited domain flip on same-day variance-panel repeats?

Design review recorded 2026-07-27; findings review recorded 2026-07-27; the review trail is retained.

Instrument

AEOForged publishes the same measurement honesty bar across products — see the 8-dimension AEO scoring methodology for how the platform scores pages. This study's instrument is the engine sweep described below (citation shares, Wilson intervals, flip rates), not page scoring.

Each of the study's 24 category questions was asked verbatim to every enabled answer engine (chatgpt, gemini, perplexity) through their live APIs, recording the answer's cited sources. A variance panel of questions received same-day repeat draws on the primary engine to quantify test-retest stability.

The measurement path contains no LLM judgement: which domains an answer cites is read directly from the engine response; citation shares, confidence intervals (Wilson 95%), source-type classification, and flip rates are all computed in code.

Sample

  • Questions: 24
  • Engines: chatgpt, gemini, perplexity
  • Recorded checks: 91 of 92 planned calls (1 failed — recorded as negatives)
  • Test-retest repeat groups: 10
  • Measured on: 2026-07-27

Every recorded check — including answers that cited no sources at all — is in the dataset.

Depth floor and refusal conditions

This study format refuses to publish below a minimum sample: at least 20 questions with recorded checks, 60 recorded checks, and 5 repeat groups. A sweep below that floor is not published as research — there is no "directional" primary data.

Limitations

  • Answer engines are non-deterministic. The measured flip rate shows how often identical re-runs change cited sources. We do not treat single-draw results inside that noise band as findings.
  • Results describe engine behavior on the measurement date. Later engine updates can shift citation patterns.
  • Gemini returns citations at lower granularity than Perplexity. We report per-engine counts separately and never blend instruments that differ.
  • The question set is the curated list in the dataset. Statistics describe this sample of questions, not every possible phrasing.
  • Recorded at design review (unresolved): Several sample questions are near-paraphrases of each other (e.g., 'is AEO a scam' vs. 'how to spot a fake AEO agency'; 'does answer engine optimization actually work' vs. 'do AEO agencies actually deliver measurable results'; 'is answer engine optimization real or a fad' vs. 'is generative engine optimization a real thing or marketing hype'). Answer engines are likely to return highly overlapping citation sets for semantically similar queries. This mechanically inflates the leading-domain share percentage and deflates the flip rate on the variance panel, because the 'independent' questions are not truly independent measurement units. The headline claim of X% leading-domain share across a '24-question inventory' is therefore an artifact of redundancy in the inventory rather than a robust measure of citation dominance.
  • Recorded at design review (unresolved): The inventory is overwhelmingly framed around skepticism and scam detection ('is AEO a scam', 'how to spot a fake AEO agency', 'are AI visibility tools worth the money'). This skew likely causes answer engines to preferentially surface consumer-protection, review, or debunking sources rather than practitioner or vendor sources. A balanced inventory would also include neutral or positive-intent queries (e.g., 'best practices for AEO', 'how to implement answer engine optimization'). Because the declared population is the inventory itself, the resulting domain-share figures are valid only for this skepticism-heavy set and cannot be compared to any broader notion of AEO-related citation patterns, yet the headline shape ('on this 24-question AEO-trust inventory') may still mislead readers who assume the inventory is representative of the full topic.
  • Recorded at design review (unresolved): Answer engines may personalize or localize citations based on the researcher's account history, geographic IP, language settings, device type, or prior interaction context. If sweeps are run from a single researcher profile or location, the observed domain shares and flip rates could reflect that profile's personalization bubble rather than the engines' default citation behavior. Without documented controls (logged-out sessions, VPN rotation, multiple independent accounts), the headline claim about 'which domains answer engines actually cite' is confounded by researcher-specific personalization.
  • Recorded at design review (unresolved): The same-day variance panel measures flip rate within a single day, but answer-engine indexes and model weights can shift intra-day due to crawl updates, A/B tests, or load-balancing across different model versions. The reported flip rate Y% therefore conflates true stochastic variance in the model's generation process with deterministic but unobserved backend changes. Without controlling for engine-version or backend-state, the headline flip-rate figure cannot be attributed solely to generation-level randomness.
  • Recorded at design review (unresolved): The study counts 'source domains' cited by answer engines, but the definition of a citing event is unspecified. It is unclear whether a domain mentioned in inline text, a footnote link, a sidebar card, or a 'sources' accordion each count equally, or whether multiple citations of the same domain within a single answer are counted once or multiple times. Different operationalizations could materially change the leading-domain share and flip rate. Without a pre-registered, transparent citation-extraction protocol, the headline metrics are not reproducible.
  • Recorded at design review (unresolved): The 24-question inventory was commissioned from a revenue-prompt cluster, but it is unverified whether real buyers or searchers actually pose these specific trust-and-scam queries to answer engines at meaningful volume. If the inventory over-represents queries that are rarely asked in practice, the domain-share and flip-rate findings, while valid for the inventory, may have limited practical relevance. This should be verified against actual query-demand data (e.g., search console logs or answer-engine query analytics) before the results are used to guide strategic decisions.
  • Not evaluated at findings review: Robustness checks for the engine lane land in a later phase — this findings review evaluated the depth floor only, and an unevaluated check is never recorded as a pass.

Read the full study: is answer engine optimization a scam or a legitimate discipline.