By Ryan Kings, Founder & CTO at AEOForged · Published August 2026 · 12 min read
How AEOForged Operates: Agent Workflows, Layers, and Citation Data
AEOForged is a measurement-first Answer Engine Optimization platform: agents research, score, ship, and verify; humans own payment and publish judgment. This paper describes the operating model behind that loop — four layers we run in public — and what our own citation instruments recorded for aeoforged.com as of 8 August 2026. It is not a guarantee that the same numbers transfer to your domain. Independent retests are welcome; the methods below are designed to be checked from outside.
What are the four operating layers?
The four operating layers are intelligence, feedback, agent surface, and measurement. Product vision names three technical pillars — intelligence, feedback loop, and platform — and the live product adds a fourth, first-class measurement spine so claims stay instrument-named. Frase’s 2026 AEO guide still splits the category into research, optimisation, and citation tracking; we treat those as jobs inside one stack an agent can call, not three separate vendors.
| Layer | Job | What you see | What stays private |
|---|---|---|---|
| Intelligence | Ground work in retrieved sources; score extractability; diagnose the site | Research packs, 8-dimension scores, audits, schema | Exact detector weights and internal ranking heuristics |
| Feedback | Iterate until the live page clears a pinned bar | Score tips, action queues, verify-page | Model prompts that only restate public acceptance criteria |
| Agent surface | Let Cursor / Claude Code / MCP clients run the loop | Handoff tokens, REST, MCP, task packs | Org secrets, rate cards, operator-only surfaces |
| Measurement | Record real-engine citations and misses | Branded vs generic rates, page vs domain tier, SOV, excerpts | Nothing about a win that we would not show a client |
HelpGuides’ metrics overview lists citation frequency and share of voice as standard AEO campaign metrics. We agree — and we refuse to blend branded prompts into category SOV, because a query that already contains your brand name is not a competitive authority win.
How does the intelligence layer work without becoming a black box?
Intelligence is the research → score → audit path: retrieved sources only, dimensional scores as feedback, and site diagnosis that becomes a claimable queue. Agents call the same tools over REST and MCP; humans see the same artifacts in the dashboard. We publish what is measured (structure, direct answers, entities, schema, crawl access, and related extractability signals) and that scores are deterministic code — not an LLM grade. We do not publish the internal weight tables. Copycats can clone a prompt; they cannot clone months of citation-correlated detector work by reading a marketing page.
Grounding is enforced at the finished boundary: articles that claim facts need research-allow-listed citations or verified primary sources before submit-for-review. That rule is product behaviour, not a slogan. Discovered Labs’ measurement-infrastructure note argues agencies need instrumentation that survives handoff between people and tools; our answer is one score function and one audit snapshot model for both.
Related reading on the public site: AEO scoring dimensions, what is an AEO audit, and what AI engines look for.
How does the feedback loop decide that work is done?
Done is system-granted after re-fetch and re-score against a pinned baseline — not when an agent marks a checkbox. The feedback loop is score → targeted rewrite → ship → verify. On Fix Programme and connected work, verify-page is the only writer that can promote an item to verified. That design exists because self-reported “fixed” is how AEO theatre starts.
Craft beats score-chasing. Stacking FAQ mirrors, key-takeaway blocks, and question-stuffed headings only to move a dimension is a failed fix even if the number rises. AirOps describes AEO as a feedback loop: test prompts as customers would, track change, compound. We agree on the loop; we disagree with any workflow that treats the score as the product.
How do AI agent workflows run day to day?
Day to day, an agent bootstraps a scoped handoff token, loads the next task pack, applies work on the client’s allowed surface (repo, GitHub PR, or CMS), then verifies. Conductor’s Agentic AEO framing treats agents as the continuous operating layer; Novelty SEO’s 2026 platform review notes that MCP-connected citation data can compress the gap-to-remediation loop from weeks toward hours when the interfaces exist. Profound and other vendors also ship MCP surfaces — see Profound’s MCP workflow note — so the category is converging on agent-callable measurement. Our differentiator is not “we have an API”; it is that research, scoring, audit queues, and citation reads share one honesty bar: misses recorded, branded and generic kept separate, prices never invented by a model.
A typical content month on Dominate looks like: pick a gap-tied topic → research (credits) → draft in the repo → score loop → mesh internal links → publish the live URL → distribution kit for humans to post. Payment and brand-profile confirmation stay human-only. For the agent toolkit shape, see AEO tools for AI agents and AEO reports for AI agents.
What does our own citation measurement show in August 2026?
As of the workspace scoreboard generated 8 August 2026, aeoforged.com had 58 tracked queries and 562 real-engine checks across Perplexity, Gemini, and ChatGPT. These are our instruments on our site after we built the content — useful dogfood, not an independent lab study. Numbers below are check-level unless labelled as query-status wins.
Branded standing is strong and expected. On 24 branded queries (201 checks), 168 checks were cited (83.6%). At query status, 24 of 24 branded prompts show a win, with 22 at page tier (a specific URL, not only the domain). When a query names the brand and the site has clear entity pages, engines — especially Perplexity — often cite the primary source. That is plumbing, not category authority.
Generic / goal questions have started, thinly. On 34 generic queries (361 checks), 11 checks were cited (3.0%). At query status, 4 of 34 show a page-tier win — all on Perplexity so far:
- “AI ready content” → /articles/is-your-content-ai-ready
- “MCP for AEO” → /docs
- “sector reports for AI visibility” → /intel/aeo-ai-search-visibility
- “what AEO sites can i connect my agent to?” → /docs
Those queries map closely to pages we deliberately structured. Broader category prompts (“best AEO tools”, “how to get cited in AI answers”, agency discovery queries) remain mostly gaps where competitors such as Profound, Semrush, Otterly, Conductor, and Scrunch still appear.
Share of voice on generic queries is the interesting early signal — and still small in absolute terms. Measured SOV on the generic basis is 6.1% (11 client citation voices vs 170 competitor voices in the tracked set). In the same competitor rollup, Semrush leads at 22.4% of competitor-attributed voices, then Profound (20.6%), Otterly (12.4%), Conductor (11.2%), AEO Engine (8.2%), and Scrunch (7.6%). Sitting near that pack after roughly a month of public dogfood is non-trivial; resting a strategy narrative on 11 generic voices would be theatre. We label it directional.
Answer coverage vs inventory is further ahead than citation wins. Deterministic answer-coverage join (tracked generic query × site inventory + citation evidence) reads 21 of 34 covered (62%). Coverage means we have an extractable page candidate; it does not mean engines cite it yet. New articles take a measurement cycle to show up in both inventory and engines.
Engine skew is real. Blended citation rates on our checks (branded + generic) were highest on Perplexity (41.9%, 95% CI 35.3–48.7), then Gemini (30.3%), then ChatGPT (27.7%). Generic SOV by engine in this window concentrates on Perplexity. Authority that only works on the currently most citation-aggressive engine is incomplete.
Day buckets for the same board (directional, short window): 6 Aug branded 88.9% / generic 2.5%; 7 Aug 83.1% / 3.6%; 8 Aug 75.5% / 3.4% — movement inside a few days is noise until intervals separate.
Where should skepticism still apply?
Skepticism should apply to sample size, selection, engine skew, self-measurement, and timeline — in that order.
- Sample and selection. Four generic page wins and 11/361 generic cited checks are real and non-zero. They are not half the category board. Wins cluster on narrow technical queries our content targets.
- Engine skew. Page-level generic wins here are Perplexity-dominant. ChatGPT and Gemini show more domain-level or partial branded signals. Cross-engine consistency is the harder bar.
- Self-measurement on self. We built the scoring system and the pages, then ran our citation instruments. That is honest internal feedback. It is not yet proof that outsiders should trust the scores without their own retests.
- Timeline. One month of public testing is early. New technical sites often earn niche citations because content is structured, recent, and the topic is underserved. Sustaining and expanding into a broader unbranded set is the actual programme.
For category context beyond our domain, see the sector citation benchmark explainer and our four-study citation-economy synthesis. The earlier on-site score lift is documented separately in the 8→81 case study — score lift and citation standing are related, not identical claims.
How can an independent party verify the claims?
Independent verification means re-asking the same query classes on the same engines, recording misses, and separating branded from generic — not trusting a screenshot. Practical protocol:
- Freeze a query list. Include brand prompts and category prompts. Publish the list.
- Call real engines (or a third-party tracker you trust). Record cited URL, engine, date, and miss rows.
- Classify branded vs generic with a written rule (brand name / domain stem present → branded).
- Prefer page-tier wins when claiming “this article was cited.” Domain mentions are a weaker rung.
- Report absolute counts with rates. “6.1% SOV” without 11/170 voices is not enough.
- Refuse movement inside noise until intervals separate or the sample is large enough to care.
We invite that retest. If your numbers disagree, the disagreement is data — send the query list and engine dates. The product promise on the homepage stays the bar we hold ourselves to: measured, not promised.
Summary
- AEOForged operates as four layers: intelligence, feedback, agent surface, and measurement — with verify-page and citation reads as the honesty brakes.
- Agents run research → score → ship → verify over REST/MCP; humans own payment and final publish judgment.
- On 8 August 2026 instruments: branded 83.6% cited checks (24/24 query wins); generic 3.0% cited checks with 4 page-tier wins; generic-basis SOV 6.1% (11 vs 170 voices).
- Those generic wins are early, Perplexity-skewed, and selection-prone — useful dogfood, not a finished authority claim.
- Independent parties can retest with a frozen query list, real engines, miss rows, and a branded/generic split; we welcome the comparison.