Table of Contents
- The Brands LLMs Actually Cite
- The Disconnect Index
- Where Traditional Authority Outpaces AI Recall
- The AI Overperformers
- Most Brands Exist in Only One AI's World
- The Sector Breakdown
- What This Means for Brand Strategy
When a user asks ChatGPT for “the best fintech platforms for my startup,” the model returns three to five names. It doesn’t hedge. It doesn’t list ten options with caveats. It picks a handful and explains each one. That answer is a brand choice the user never made and never compared against, and it happens millions of times a day across ChatGPT, Gemini, and Claude.
We wanted to know which brands are winning that slot, which ones are losing it, and how well it tracks the traditional measures marketers have spent the last decade optimizing. We prompted GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6 with 96 sector-specific prompts across 8 industries, ran each one 15 times, and extracted over 8,500 distinct brand mentions across 4,320 model responses. We then matched those brands against Ahrefs data on Domain Rating, organic traffic, referring domains, and keyword volume. The gap between the two is the story.
Key Takeaways
- AI visibility is not all-encompassing. Fewer than 1 in 8 brands are cited by all three models we analyzed.
- Digital PR is the lever for LLM recall, not owned SEO. The brands that overperform are the ones that live inside other people’s content (i.e., roundups, comparisons, and expert lists).
- Insurance shows the widest gap between traditional and AI visibility. Insurtech challengers like Lemonade and Root have replaced legacy carriers in model output.
- Education institutions punch above their weight. Stanford, MIT, and Khan Academy lead the sector despite smaller search footprints than commercial platforms.
- The default-answer slot is still claimable in most categories, but it’s consolidating fast. Once a model settles on a category leader, dislodging it gets structurally harder.
The Brands LLMs Actually Cite
Every industry we tested has a short list of “default-answer” brands, which are names that surface in nearly every model, in nearly every run, regardless of how the prompt is phrased. The list is shorter than the traditional leaderboard in that sector. Models don’t hedge. They name a handful of brands, explain each one, and move on.
A few patterns are immediately visible:
- Travel was the most concentrated category. Booking.com (285 mentions), Airbnb (227), and Expedia (215) accounted for roughly 20% of the sector’s entire mention volume between them. Three brands, one out of every five recommendations.
- HealthTech had a runaway leader. Teladoc (275) led Amwell (220), the runner-up, by roughly 25%. The models have settled.
- Wellness had the deepest bench. Peloton, Headspace, Calm, Whoop, and Oura all cleared 168 mentions. No single brand dominates, which means the category-leader slot is still up for grabs.
- Insurance broke the way industry watchers should have expected. Lemonade (213) ranked above State Farm (172). Root Insurance (165) ranked above Progressive (114). The models cited the digital-native carriers first and treated the legacy ones as the alternatives, not the default.
- Lifestyle was the loudest signal of the methodology working. Patagonia, Allbirds, Eileen Fisher, Everlane (sustainability-coded direct-to-consumer brands) outranked Sephora, Samsung, and Whirlpool. Whether that’s a function of how product roundup articles get written or a deeper bias in training data, the result is the same: the models prefer the brand story that gets told in third-party content.
The pattern across all eight sectors was consistent. The models have already crowned a handful of brands in each category. Sometimes those crowns go where you’d expect (Stripe in FinTech, Notion in SaaS, Coursera in Education). Often they don’t.
Within each category, the field of competitive brands the models will surface is far narrower than the field of brands competing for Google rankings. If you’re not in the top 5-10 in your category for LLM recall, you’re effectively not in the consideration set.
The Disconnect Index
This is the finding that anchors the rest of the study. We took every brand in the dataset that had measurable presence in both LLM citation and traditional search authority (over 8,500 brands across 8 industries) and scored each one on both axes. The vast majority sit roughly where you’d expect. But the brands at the tails tell a more interesting story, and they break in opposite directions.

Over 9 in 10 brands (91%) were aligned. Their AI visibility roughly tracked their traditional search authority. If they ranked well on Google in their category, the models cited them at a comparable rate. This is the expected outcome and it’s the majority case.
The other 9% is where the strategic insight lives.
Five percent (471 brands) were underrepresented in AI. These were brands with strong Domain Ratings, heavy organic traffic, large keyword portfolios, and the models barely mentioned them. They have the SEO infrastructure of a category leader and the LLM presence of a startup nobody has written about.
Four percent (377 brands) were AI overperformers. Modest traditional signals, outsized LLM recall. The opposite problem. The models cited them aggressively in their categories even though Ahrefs would tell you they don’t deserve the placement.
The 9% in the tails tracked one signal more strongly than any other: how often a brand appeared in third-party content that wasn’t written by the brand itself. Overperformers showed up repeatedly in roundups, expert lists, and comparison articles. That kind of coverage gets ingested broadly into model training. Brands with high traditional search authority but light third-party coverage tend to land in the underrepresented group, regardless of category.
The pattern points one direction. LLM recall isn’t driven by what you publish on your own domain. It’s driven by what other people publish about you.

The visualization shows what the headline numbers describe. The diagonal is the world most marketers assume they’re operating in: do well on Google, do well in models. The two off-diagonal clusters are the world they actually have to plan for.
Where Traditional Authority Outpaces AI Recall
The brands underrepresented in AI are, in most cases, the blue chips of their categories on every traditional measure. Domain Ratings above 80. Millions of monthly organic visits. Hundreds of thousands of keywords ranked. And the models barely cite them in their own sectors.

A few of the patterns worth pulling out:
- Microsoft and Spotify topped the list in FinTech. They’re the platonic giants of traditional visibility, but the models didn’t categorize them as fintechs when prompted with fintech queries. That’s not a bug. Microsoft is heavily cited when you ask about productivity software or cloud platforms. It’s a categorization gap. The models have decided what counts as a fintech, and Microsoft wasn’t on the list.
- Legacy insurance carriers got displaced. Aetna, Cigna, Humana, Liberty Mutual, and Trustpilot all sat at DR 80+ with millions of monthly visits. The models barely cited them. Lemonade and Root sat higher in the same prompts despite a fraction of the traditional footprint.
- Legacy retail brands got displaced. Sephora, Samsung, Whirlpool, GE Appliances. The models preferred Patagonia, Allbirds, and Eileen Fisher when the prompt asked for lifestyle recommendations. This is the most visible signal that the training data heavily weights direct-to-consumer brand narratives over legacy retail.
- Healthcare giants were under-indexed. Medtronic, GoodRx, Humana, 23andMe. The models cited telehealth-native brands like Teladoc, Amwell, and Doxy.me instead.
The note on Microsoft applies broadly. Low recall in one category doesn’t mean a brand has zero LLM presence overall. It means the brand isn’t part of the models’ default mental map for that category. That’s a more specific and more solvable problem than “we’re invisible.” It also means the underrepresentation isn’t random. It’s categorical. The models have decided where each brand belongs, and if your brand isn’t where you think it is, your AI strategy needs to address the categorization before it addresses the visibility.
The AI Overperformers
The opposite tail is more useful for marketers, because it tells you what worked for the brands that beat their traditional footprint. Modest Ahrefs profiles. Outsized LLM recall. The 377 brands in this group share one common trait: they appear repeatedly in the third-party content the models were trained on.

The list tells the story by itself:
- Monday.com was the #1 overperformer in the dataset. AI visibility score: 0.71. The note in the visualization matters here: Monday.com has substantial traditional traffic in absolute terms (close to 1M organic visits and 50k+ keywords). Its overperformer status reflected how aggressively the models cited it relative to its sector peers. Among SaaS comparison content, Monday.com gets pulled into nearly every roundup.
- Root Insurance won insurance outright. More AI mentions than any traditional carrier, even though its DR sits well below most of them.
- Six of the top 15 overperformers were in Education. Stanford, MIT, Khan Academy, LinkedIn Learning, IBM Data Science, Google Career Certificates. These are institutions and certificate programs, not high-DR commercial platforms. The models cited them for credibility, not for traffic.
- Smaller players showed the same pattern. Doxy.me, Nike Training Club, and Google Flights all shared strong product-roundup coverage, heavy expert-list inclusion, and outsized LLM recall.
The trait these brands share isn’t a content strategy. It’s a coverage profile. They show up in the kind of third-party content that gets ingested into model training: product comparison articles, expert roundups, review content, category lists. That coverage doesn’t come from publishing more on your own blog. It comes from what other people write about. That’s the surface area of digital PR, not owned SEO.
If you’re trying to understand why a competitor with half your organic traffic is showing up in the model’s answer and you aren’t, the question to ask isn’t, “what’s their SEO strategy?” It’s, “what’s their press coverage and how did they earn it?”
Most Brands Exist in Only One AI’s World
The other major finding in the study cuts against the way most marketing teams currently think about AI visibility as a single metric. Cross-model agreement is the exception, not the rule.

Only 900 brands (11%) were cited by all three models. Seventy-seven percent of brands (6,264) were mentioned by just one model. The remaining 12% were split between two models.
The implication is unambiguous: if you optimize for one model, you’re statistically likely to be invisible in the others. A brand winning ChatGPT might be missing from Gemini, and that’s a different kind of problem than not winning at all. It means there’s slack in the system. The default-answer slot in each model is being contested separately, and a brand can hold one of those slots without holding the others.
The models have distinct fingerprints, too. Claude over-indexed on SaaS and insurance, while Notion, Linear, and Lemonade got cited at higher rates in Claude than the other two. Gemini over-indexed on travel and healthcare, while Booking.com, Teladoc, and Livongo showed up disproportionately in its responses.
GPT-4o is the most consensus-driven model; its top brands had the highest overlap with the other two. These aren’t huge swings, but they’re directionally consistent across the dataset, and they suggest the models have inherited different training-data biases that show up in which categories they have more confident defaults in.

The heatmap shows how unevenly distributed even consensus brands are. Notion was cited 17 times more often by Claude than by Gemini. Whoop almost four times more often by Claude than Gemini. State Farm got almost half its mentions from GPT-4o and only a quarter from Gemini. Booking.com was the rare brand that’s roughly evenly distributed, and even there, GPT-4o cited it more often.
For a marketing team, the strategic translation is direct. An aggregate “AI visibility” score is fiction. Per-model scores are the real picture. A brand winning one model isn’t the same as winning all three, and the work to close the gap on one model isn’t the same as the work to close it on another.
The Sector Breakdown
The disconnect rate isn’t evenly distributed across industries either. Some sectors have been almost entirely re-mapped by the models; others still look the way Google sees them.
The pattern breaks down differently by sector:
- Insurance had the highest underrepresentation rate at 8%. Legacy carriers like Aetna, Cigna, Humana, and Liberty Mutual, had been displaced by digital-native carriers in model output. If you’re an established insurance brand, you have the widest gap in the dataset to close.
- Education had the highest overperformer rate at 7%. Stanford, MIT, Khan Academy, IBM Data Science, Google Career Certificates, LinkedIn Learning. Institutions and professional certificates dominate, and they do it without the SEO infrastructure of commercial platforms. Education is the sector where credibility coverage outweighs traffic coverage most clearly.
- HealthTech was roughly balanced (5% under, 6% over). The shift has been from pharma and device incumbents to consumer-facing digital health. Both ends of the trade are visible.
- Travel had the lowest disconnect overall (4%/4%). LLM recall and traditional visibility track tightly here. Booking.com, Expedia, Airbnb are top of mind for both Google and the models. The category leaders are stable.
- SaaS and FinTech sit at 6% under/3-4% over. These are categories where the displacement is more about which incumbent the models picked (Stripe over PayPal, Notion over Microsoft) than about new entrants replacing old ones wholesale.
The pattern across sectors mirrors the broader finding. Where the legacy authority of a category was built primarily on SEO and traffic, the models have re-sorted. Where the category was built around brand stories that circulated in third-party media, the alignment is tighter.
What This Means for Brand Strategy
Here are five things to take from the study, in roughly the order you’d act on them:
- Traditional authority does not equal AI authority. This is the load-bearing finding. DR 95 brands sat below DR 60 brands in LLM recall on a regular basis. The traditional visibility playbook tells you whether Google ranks you. It doesn’t tell you whether the models remember you. Treat them as two distinct metrics with two distinct optimization paths.
- Digital PR is the LLM lever. The overperformers (Monday.com, Root, Stanford, Doxy.me, Khan Academy) shared one thing. They appeared repeatedly in third-party content. Product roundups, expert lists, comparison articles, category coverage. That coverage gets ingested into model training broadly and durably. Owned content on your own domain doesn’t carry the same weight in model recall because the models weigh what other people say about a brand more heavily than what the brand says about itself. That’s the structural argument for digital PR as the AI-era brand lever.
- Measure per-model, not in aggregate. Only 11% of brands were cited by all three models. The rest is model-specific. Score your AI visibility against GPT-4o, Gemini, and Claude separately. A win on one isn’t a win on the others, and the work to win each is different.
- The default-answer window is closing. Once an LLM has settled on a category leader, dislodging it gets structurally harder. The training data gets reused. The categorization sticks. Categories that still have a wide-open field, like Wellness, Lifestyle, and some pockets of HealthTech, are the ones to move on first. Categories where the default has consolidated (Travel, parts of FinTech, parts of SaaS) are harder targets and require more sustained third-party coverage to shift.
- Categorization matters as much as visibility. Microsoft had more LLM mentions than most brands in this dataset. It just didn’t have them in fintech. If your AI visibility is low in your own category but high overall, your problem isn’t recall — it’s categorization. That’s a different intervention. It’s about the kind of content you appear in, not how often you appear.
The brands that win the AI-era discovery surface won’t be the ones with the biggest content libraries or the most aggressive SEO programs. They’ll be the ones whose names show up most often in the writing other people do about their category. That’s a different motion than the one most marketing teams have been built around, and the data here suggests the gap between the two motions is widening.
Methodology
This study analyzed how three major large language models, including GPT-4o (gpt-4o-2024-11-20), Gemini 2.5 Flash, and Claude Sonnet 4.6, cite brands across 8 industries (Education, FinTech, HealthTech, Insurance, Lifestyle, SaaS, Travel, Wellness).
We developed 96 sector-specific prompts (12 per industry × 4 intent types: recommendation, comparison, trend, problem-solving) and ran each prompt 15 times per model at temperature 0.7, generating 4,320 total responses for statistical stability. Brand extraction used a hybrid pipeline combining GPT-4o-mini JSON extraction, spaCy named-entity recognition, and fuzzy deduplication with 50+ alias resolutions (Alphabet → Google, Meta Platforms → Meta, etc.) and stopword filtering to remove category noise (AI, ML, API, SaaS, etc.). The pipeline identified over 8,500 distinct brands across all responses.
Traditional Visibility Scores were built from Ahrefs Batch Analysis for the top 200 brands per industry, using a weighted blend of Domain Rating (30%), Organic Traffic (25%), Referring Domains (25%), and Organic Keywords (20%), with each metric percentile-ranked within its industry to normalize for category-level differences in search volume.
A note on snapshot data: these results reflect a single scrape period in 2026. AI citation behavior shifts as models are updated, and outputs vary by user, query phrasing, and session context. We recommend treating these findings as a directional benchmark rather than a definitive ranking.
About Fractl Marketing
Fractl is a growth marketing agency that helps brands earn attention and authority through data-driven content, digital PR, and AI-powered strategies. Our work has been featured in The New York Times, Forbes, and Harvard Business Review. Fractl is also the team behind Fractl Agents, including the AI Brand Visibility Agent used to power this study.
Fair Use Statement
Fractl encourages journalists and publishers to share findings from this study with proper attribution. Please link back to this page when referencing or citing our research so readers can access the full context and explore the complete dataset behind the insights.







