The short answer
We classified the sources AI engines cited across 90 days of live tracking for client brands in three unrelated verticals: auto insurance, beauty/skincare, and cannabis retail. The finding that held in every vertical: a brand's own website is only 2 to 6 percent of the sources AI engines cite when its category comes up. Auto insurance came in at 4.0 percent, beauty/skincare at 5.8 percent, cannabis retail at 2.0 percent. The third-party bucket (editorial, UGC, and reference sources combined) ranged from 26 to 59 percent. The rest is corporate sites, competitors, institutional sources, and a small residual. The practical conclusion is blunt: you cannot win AI search from your own site alone, because your own site is a sliver of the citation pool the engines are drawing from.
Key findings
A brand's own website is only 2 to 6 percent of the sources AI engines cite in its category. Auto insurance 4.0 percent, beauty/skincare 5.8 percent, cannabis retail 2.0 percent, an average of roughly 3.9 percent. In no vertical did the brand's own domain reach even a sixteenth of the citation pool.
Third-party surfaces take 26 to 59 percent of citations, several times the brand's own share in every vertical. Editorial, UGC, and reference sources combined reached 34.9 percent in auto insurance, 59.0 percent in beauty/skincare, and 26.0 percent in cannabis retail. These are the surfaces a brand can influence through earned placements but does not own.
The source mix is vertical-specific, so a single AEO playbook does not transfer between industries. Editorial dominates beauty/skincare at 41.4 percent of citations, while corporate sites dominate cannabis retail at 65.6 percent and auto insurance at 52.2 percent. The same engines, tracked the same way, build answers from structurally different source pools depending on the category.
Reference sites are a real citation surface in insurance but nearly absent in beauty. Encyclopedias and documentation-style sources took 13.8 percent of auto insurance citations, against 0.9 percent in beauty/skincare and 1.9 percent in cannabis retail. Regulated, definition-heavy categories lean on reference material; taste-driven consumer categories do not.
Community carries far more weight in consumer categories than in financial ones. UGC sources (forums, social, community sites) took 16.6 percent of beauty/skincare citations versus 4.3 percent in auto insurance and 9.9 percent in cannabis retail. Where buyers ask peers, engines cite peers.
Methodology
We run our client tracking on Peec AI daily. For this study we aggregated the source domains that AI engines (ChatGPT, Perplexity, Gemini, Google AI Overviews, Copilot, and others) cited across the buyer questions we track for live client brands, over a 90-day window from 2026-04-12 to 2026-07-11.
The sample covers three unrelated verticals: auto insurance, beauty/skincare, and cannabis retail. We chose the word "unrelated" deliberately. These categories differ in regulation, purchase journey, price point, and audience, which is exactly why a pattern that holds across all three is worth reporting.
For each project, we took the top roughly 1,000 cited domains and classified every one into seven source types:
- You: the client brand's own website.
- Corporate: any other company site.
- Editorial: news sites, blogs, and magazines.
- Reference: encyclopedias and documentation-style sources.
- UGC: forums, social platforms, and community sites.
- Institutional: government, education, and nonprofit domains.
- Competitor: direct competitor domains.
Domains that did not fit any category are reported as a residual "Other" bucket. Percentages represent each type's share of total inline citations across the top cited domains per project.
A note on why we measured it this way. Most public commentary about AI search visibility is built on one-off screenshots: someone asks an engine a question, notices who gets named, and generalizes. That approach cannot separate signal from noise, because AI answers vary run to run, engine to engine, and week to week. Aggregating inline citations across a full quarter of daily tracking smooths that variance out. It also shifts the question from "who got mentioned today" to the more durable one: what kinds of sources do these engines structurally rely on when they build category answers? That structural pattern is what a marketing team can actually plan against.
Limitations, stated plainly
Research is only useful if you know what it does not show, so here are the limits.
- Three verticals is a sample, not a census. The own-site pattern held across all three, but we make no claim that the specific mix percentages generalize to your industry. The direction is the finding; the decimals are the evidence.
- Clients are anonymized. These are live client brands, so we report vertical-level aggregates rather than named brands. The full anonymized breakdown is available by email (see the Cite this study section).
- This is citation share, not traffic share. We measured how often each source type appears among the domains engines cite, not how many visitors or conversions any source drives. A source can be cited often and clicked rarely.
- The prompts are buyer questions, not brand questions. We track category and purchase-intent prompts. A brand's own site presumably fares better on direct navigational questions about itself; that is not what this study measures.
The full source mix, vertical by vertical
Here is the complete seven-way classification for each vertical, plus the residual bucket. Every number is the share of total inline citations across the top roughly 1,000 cited domains for that project.
| Source type | Auto insurance | Beauty/skincare | Cannabis retail |
|---|---|---|---|
| Corporate | 52.2% | 25.5% | 65.6% |
| Editorial | 16.8% | 41.4% | 14.1% |
| Reference | 13.8% | 0.9% | 1.9% |
| UGC | 4.3% | 16.6% | 9.9% |
| You (brand's own site) | 4.0% | 5.8% | 2.0% |
| Competitor | 6.2% | not broken out | 1.1% |
| Institutional | 2.0% | 2.3% | 2.6% |
| Other | 0.8% | 7.3% | 2.9% |
Competitor share was not broken out as a separate class in the beauty/skincare project; competitor domains there fall inside the corporate and other buckets.
Three patterns in this table deserve a closer look.
Corporate is the largest single class everywhere except beauty. In auto insurance (52.2 percent) and cannabis retail (65.6 percent), engines lean heavily on company websites: not the brand's own, but the wider population of vendor, aggregator, and industry company domains. In beauty/skincare, editorial takes that crown at 41.4 percent, with corporate at 25.5 percent.
The You row is the smallest meaningful class in every vertical. At 4.0, 5.8, and 2.0 percent, the brand's own site is out-cited by editorial in all three verticals, by UGC in all three, and by the corporate bucket in all three. If you have read our piece on why a brand goes missing from AI search, this is the data behind the third-party citation gap we call the most underestimated cause.
Reference and UGC trade places depending on the category. Insurance buyers get answers built on reference material (13.8 percent); beauty buyers get answers built on community discussion (16.6 percent). The engines mirror where trust lives in each category. That mirroring is worth sitting with, because it means the engines are not imposing a uniform sourcing standard on the web. They are amplifying whatever evidence ecosystem already surrounds a purchase decision, and a brand's earned-media strategy has to meet that ecosystem where it is rather than where the brand wishes it were.
What this means if you run marketing
The instinct when a brand is invisible in AI answers is to fix the website: more content, better structure, more schema. That work matters, and we walk through it in how to rank in ChatGPT. But this data says the website is, at best, a 2 to 6 percent play. The other 94-plus percent of the citation pool is territory you influence through earned presence, or not at all.
Three implications follow directly from the numbers.
The battleground is third-party. Editorial, UGC, and reference sources took 26 to 59 percent of citations in our sample, several times the own-site share in every vertical. Getting your brand into the roundups, comparison articles, community threads, and reference pages the engines already cite moves you into the pool that actually gets quoted. This lines up with the broader industry evidence: Ahrefs found that roughly 43.8 percent of pages cited by ChatGPT are listicles ("best X" and "top X" formats), based on over 1 billion data points. If almost half of ChatGPT's citation diet is listicles, being absent from the listicles in your category is a structural handicap no amount of on-site optimization repairs.
The buyers are already there. This is not a future problem. G2 reports that roughly 48 percent of buyers now use AI in the buying process. The sources AI engines cite are, increasingly, the shortlist your buyers see. Citation share in those sources is market presence.
We can see this in our own analytics. Over the 33 days ending 2026-07-29, chatgpt.com was the second largest traffic source to aeolabs.ai at 169 sessions, nearly double the 90 sessions Google search sent us, and 47 percent of those ChatGPT visits were engaged sessions. A site that practices AEO gets found through AI answers. That is the mechanism this study describes, visible in our own referral logs.
Measure citation share, not just rankings or traffic. None of this shows up in a normal analytics stack. AI-influenced buyers tend to arrive as direct or branded visits, long after the answer that shaped their shortlist, and your Google rankings say little about which sources the engines quote. The metric that maps to this study is citation share: of the sources cited when your buyer prompts run, what fraction mention or belong to you, and how is that trending against competitors? That is a number you can baseline, report monthly, and hold a program accountable to, the same way search teams once held themselves to rankings.
Your playbook must be vertical-specific. A beauty brand should be fighting for editorial (41.4 percent of citations) and community presence (16.6 percent). An insurance brand needs reference and editorial coverage on top of a strong corporate footprint. A cannabis retailer operates in a pool where corporate domains take 65.6 percent, which makes directory-style and industry corporate placements disproportionately valuable. Copying another industry's AEO strategy means fighting on the wrong surface.
The operational loop we run for clients is the same one we used to produce this study: track the prompts, classify the citations, find the gap between where the engines look and where the brand appears, then close it surface by surface. If you want to run the tracking side yourself, start with our guide to the best AEO tracking tools. If you want the baseline done for you, our AI visibility audit measures your citation share across engines and shows you exactly which third-party surfaces are citing your competitors instead of you.
For a primer on the discipline behind all of this, this short Ahrefs explainer is a solid starting point.
Key takeaways
- Across 90 days and three unrelated verticals, a brand's own website was only 2 to 6 percent of the sources AI engines cited for its category (auto insurance 4.0 percent, beauty/skincare 5.8 percent, cannabis retail 2.0 percent).
- Third-party editorial, UGC, and reference sources took 26 to 59 percent of citations, several times the own-site share in every vertical.
- The mix is category-dependent: editorial leads beauty (41.4 percent), corporate leads cannabis (65.6 percent) and insurance (52.2 percent), reference matters in insurance (13.8 percent) but barely registers in beauty (0.9 percent).
- The method: live Peec AI tracking, top roughly 1,000 cited domains per project, classified into seven source types plus a residual.
- The strategy that follows: your site makes you eligible; third-party citations get you named. Budget accordingly.
Cite this study
You are welcome to reference these findings in articles, newsletters, decks, and roundups. All we ask is attribution to AEO Labs and a link to this page so readers can check the method and the limits for themselves.
Suggested citation: AEO Labs, "Where AI Engines Get Their Answers: We Classified the Sources Behind 90 Days of AI Citations" (2026), aeolabs.ai/blog/ai-citation-sources-study.
Journalists and researchers who want the full anonymized breakdown behind the vertical percentages can email aidan@aeolabs.ai and we will share it. If your brand's own citation picture looks like the one in this data, our piece on why a brand goes missing from AI search is the diagnostic to run next.