Method
How generative engines pick sources, what we measure, how we work, and what we refuse to sell. Written so that a buyer can check every claim.
How do AI engines choose what to cite?
Every engine runs query fan-out: one prompt becomes several hidden searches, and pages that rank for those sub-queries are fetched and cited. Each engine searches a different index.
Five findings from published research shape the work. Sources are listed at the end.
- Organic rank is a weakening proxy. The share of AI Overview citations taken from Google's top 10 fell from 76% to 38% in eight months. ChatGPT, Gemini and Copilot cite Google's top 10 about 8% of the time.
- ChatGPT cites what it fetches live: pages it fetched were cited 74% of the time, pages seen only as a snippet 7%.
- No AI crawler runs JavaScript except Googlebot. The answer has to be in the server-rendered HTML.
- Third-party sites carry 82–89% of citations. A vendor's own site is cited in about 12% of B2B answers.
- Brand mentions correlate with AI Overview presence at 0.66; backlinks at 0.22. Models pick brands from memory first, then retrieve.
The page that gets cited states the answer in its first 100 words, backs it with a number and a named source, and lives on a site the engine is allowed to fetch.
What content traits are proven to help?
- Statistics, expert quotes and source citations: 28–41% more visibility in generative answers in the Princeton GEO study (KDD 2024).
- A definitional sentence directly under a question-shaped heading.
- The answer early: 44% of cited passages sit in the first 30% of a page.
- Freshness: AI-cited pages are on average 25.7% newer than organic results for the same query.
- Entity consistency: the same name, address, description and founding year on the site, LinkedIn, Wikidata and directories.
How do we measure?
A locked prompt set, the same wording every month, runs by API on ChatGPT, Perplexity, Gemini and Claude, five repeats per prompt, spread over three days, one prompt per fresh session with no memory. Every run is stored as JSON. Each run is classified: brand mentioned, page cited, position, sentiment, accuracy, competitors named, URLs cited.
We report pooled percentages: mention rate, citation rate, share of voice against 3–5 named competitors, lead rate. Per prompt we say never, sometimes or usually. We never report a rank, because the same prompt returns the same brand list under 1% of the time.
Traffic is measured with a GA4 "AI Assistant" channel, set up in your account, and cross-checked with server logs per bot. We say in every report what GA4 misses: in-app traffic and AI Mode clicks that look organic.
API answers differ from the consumer apps. The top 10 prompts are also run by hand, with screenshots, as ground truth. A full re-audit against the baseline runs every 90 days.
What we do not sell
- llms.txt. 97% of llms.txt files are never requested and no engine has confirmed using one. We publish one on our own site because it costs nothing, and we say so.
- "AI schema" as a citation key. A controlled test showed no citation lift from schema. We add schema for entity clarity, and price it as hygiene, not magic.
- Wikipedia page creation. Only if the company is notable by Wikipedia's standard, and never paid for.
- Programmatic content. Hundreds of generated pages without unique data get pruned, not cited.
- Rank or placement promises. There is no rank to promise. We promise the work and the report.
- Blocking training bots as a strategy. Blocking GPTBot, ClaudeBot or Google-Extended does not change citations. Blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot makes a site uncitable. Your decision on training bots is recorded and respected.
Sources
- Google Search Central, AI features and your website.
- Ahrefs, AI Overview citations from the top 10 and brand mentions vs backlinks, 2025–2026.
- Ahrefs, llms.txt study.
- Search Engine Land, how ChatGPT's retrieval stack works and ChatGPT citation content study.
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024, Princeton.
- SparkToro, AI brand recommendation inconsistency; Conductor, recommendation consistency analysis.
- Muck Rack, What is AI reading, May 2026.