Image

The AI Visibility Index: Which Major Brands Are Vanishing From ChatGPT, Gemini, and Claude

A study of how over 8,500 large brands show up inside three major LLMs, and why traditional visibility signals no longer predict who the models cite.

Avatar of Kelsey Libert

Kelsey Libert

Cofounder
Avatar of Aditya Sachdeva

Aditya Sachdeva

Head of Data Journalism
Icon

15 min read

The AI Visibility Index: Which Major Brands Are Vanishing From ChatGPT, Gemini, and Claude

Table of Contents

When a user asks ChatGPT for “the best fintech platforms for my startup,” the model returns three to five names. It doesn’t hedge. It doesn’t list ten options with caveats. It picks a handful and explains each one. That answer is a brand choice the user never made and never compared against, and it happens millions of times a day across ChatGPT, Gemini, and Claude.

We wanted to know which brands are winning that slot, which ones are losing it, and how well it tracks the traditional measures marketers have spent the last decade optimizing. We prompted GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6 with 96 sector-specific prompts across 8 industries, ran each one 15 times, and extracted over 8,500 distinct brand mentions across 4,320 model responses. We then matched those brands against Ahrefs data on Domain Rating, organic traffic, referring domains, and keyword volume. The gap between the two is the story.

Key Takeaways

  • AI visibility is not all-encompassing. Fewer than 1 in 8 brands are cited by all three models we analyzed.
  • Digital PR is the lever for LLM recall, not owned SEO. The brands that overperform are the ones that live inside other people’s content (i.e., roundups, comparisons, and expert lists).
  • Insurance shows the widest gap between traditional and AI visibility. Insurtech challengers like Lemonade and Root have replaced legacy carriers in model output.
  • Education institutions punch above their weight. Stanford, MIT, and Khan Academy lead the sector despite smaller search footprints than commercial platforms.
  • The default-answer slot is still claimable in most categories, but it’s consolidating fast. Once a model settles on a category leader, dislodging it gets structurally harder.

Every industry we tested has a short list of “default-answer” brands, which are names that surface in nearly every model, in nearly every run, regardless of how the prompt is phrased. The list is shorter than the traditional leaderboard in that sector. Models don’t hedge. They name a handful of brands, explain each one, and move on.

Infographic comparing AI search visibility and traditional search authority across brands, showing that 91% of brands have aligned performance between AI search results and traditional SEO visibility, while a small percentage significantly overperform or underperform in AI-generated search experiences.

A few patterns are immediately visible:

  • Travel was the most concentrated category. Booking.com (285 mentions), Airbnb (227), and Expedia (215) accounted for roughly 20% of the sector’s entire mention volume between them. Three brands, one out of every five recommendations.
  • HealthTech had a runaway leader. Teladoc (275) led Amwell (220), the runner-up, by roughly 25%. The models have settled.
  • Wellness had the deepest bench. Peloton, Headspace, Calm, Whoop, and Oura all cleared 168 mentions. No single brand dominates, which means the category-leader slot is still up for grabs.
  • Insurance broke the way industry watchers should have expected. Lemonade (213) ranked above State Farm (172). Root Insurance (165) ranked above Progressive (114). The models cited the digital-native carriers first and treated the legacy ones as the alternatives, not the default.
  • Lifestyle was the loudest signal of the methodology working. Patagonia, Allbirds, Eileen Fisher, Everlane (sustainability-coded direct-to-consumer brands) outranked Sephora, Samsung, and Whirlpool. Whether that’s a function of how product roundup articles get written or a deeper bias in training data, the result is the same: the models prefer the brand story that gets told in third-party content.

The pattern across all eight sectors was consistent. The models have already crowned a handful of brands in each category. Sometimes those crowns go where you’d expect (Stripe in FinTech, Notion in SaaS, Coursera in Education). Often they don’t.

Within each category, the field of competitive brands the models will surface is far narrower than the field of brands competing for Google rankings. If you’re not in the top 5-10 in your category for LLM recall, you’re effectively not in the consideration set.

This is the finding that anchors the rest of the study. We took every brand in the dataset that had measurable presence in both LLM citation and traditional search authority (over 8,500 brands across 8 industries) and scored each one on both axes. The vast majority sit roughly where you’d expect. But the brands at the tails tell a more interesting story, and they break in opposite directions.

Infographic visualizing where AI recall and traditional search visibility diverge across brands, featuring a scatter plot that compares LLM mention frequency with traditional SEO authority and highlights brands that overperform or underperform in AI search results relative to their organic search presence.

Over 9 in 10 brands (91%) were aligned. Their AI visibility roughly tracked their traditional search authority. If they ranked well on Google in their category, the models cited them at a comparable rate. This is the expected outcome and it’s the majority case.

The other 9% is where the strategic insight lives.

Five percent (471 brands) were underrepresented in AI. These were brands with strong Domain Ratings, heavy organic traffic, large keyword portfolios, and the models barely mentioned them. They have the SEO infrastructure of a category leader and the LLM presence of a startup nobody has written about.

Four percent (377 brands) were AI overperformers. Modest traditional signals, outsized LLM recall. The opposite problem. The models cited them aggressively in their categories even though Ahrefs would tell you they don’t deserve the placement.

The 9% in the tails tracked one signal more strongly than any other: how often a brand appeared in third-party content that wasn’t written by the brand itself. Overperformers showed up repeatedly in roundups, expert lists, and comparison articles. That kind of coverage gets ingested broadly into model training. Brands with high traditional search authority but light third-party coverage tend to land in the underrepresented group, regardless of category.

The pattern points one direction. LLM recall isn’t driven by what you publish on your own domain. It’s driven by what other people publish about you.

Infographic identifying brands whose traditional search authority significantly exceeds their AI search visibility, comparing AI recall and traditional SEO performance for companies including Microsoft, Spotify, eBay, Sephora, Mayo Clinic, Notion, GoodRx, Samsung, and Typeform across multiple industries.

The visualization shows what the headline numbers describe. The diagonal is the world most marketers assume they’re operating in: do well on Google, do well in models. The two off-diagonal clusters are the world they actually have to plan for.

The brands underrepresented in AI are, in most cases, the blue chips of their categories on every traditional measure. Domain Ratings above 80. Millions of monthly organic visits. Hundreds of thousands of keywords ranked. And the models barely cite them in their own sectors.

Infographic highlighting brands that outperform their traditional search footprint in AI search results, showing examples such as Root Insurance, Stanford, MIT, Google Flights, Khan Academy, LinkedIn Learning, and HubSpot CRM that achieve disproportionately strong AI visibility relative to their organic search authority.

A few of the patterns worth pulling out:

  • Microsoft and Spotify topped the list in FinTech. They’re the platonic giants of traditional visibility, but the models didn’t categorize them as fintechs when prompted with fintech queries. That’s not a bug. Microsoft is heavily cited when you ask about productivity software or cloud platforms. It’s a categorization gap. The models have decided what counts as a fintech, and Microsoft wasn’t on the list.
  • Legacy insurance carriers got displaced. Aetna, Cigna, Humana, Liberty Mutual, and Trustpilot all sat at DR 80+ with millions of monthly visits. The models barely cited them. Lemonade and Root sat higher in the same prompts despite a fraction of the traditional footprint.
  • Legacy retail brands got displaced. Sephora, Samsung, Whirlpool, GE Appliances. The models preferred Patagonia, Allbirds, and Eileen Fisher when the prompt asked for lifestyle recommendations. This is the most visible signal that the training data heavily weights direct-to-consumer brand narratives over legacy retail.
  • Healthcare giants were under-indexed. Medtronic, GoodRx, Humana, 23andMe. The models cited telehealth-native brands like Teladoc, Amwell, and Doxy.me instead.

The note on Microsoft applies broadly. Low recall in one category doesn’t mean a brand has zero LLM presence overall. It means the brand isn’t part of the models’ default mental map for that category. That’s a more specific and more solvable problem than “we’re invisible.” It also means the underrepresentation isn’t random. It’s categorical. The models have decided where each brand belongs, and if your brand isn’t where you think it is, your AI strategy needs to address the categorization before it addresses the visibility.

The opposite tail is more useful for marketers, because it tells you what worked for the brands that beat their traditional footprint. Modest Ahrefs profiles. Outsized LLM recall. The 377 brands in this group share one common trait: they appear repeatedly in the third-party content the models were trained on.

Infographic analyzing brand recognition across GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6, showing that most AI-cited brands appear in only one model while relatively few brands are consistently cited across all major AI search and answer engines.

The list tells the story by itself:

  • Monday.com was the #1 overperformer in the dataset. AI visibility score: 0.71. The note in the visualization matters here: Monday.com has substantial traditional traffic in absolute terms (close to 1M organic visits and 50k+ keywords). Its overperformer status reflected how aggressively the models cited it relative to its sector peers. Among SaaS comparison content, Monday.com gets pulled into nearly every roundup.
  • Root Insurance won insurance outright. More AI mentions than any traditional carrier, even though its DR sits well below most of them.
  • Six of the top 15 overperformers were in Education. Stanford, MIT, Khan Academy, LinkedIn Learning, IBM Data Science, Google Career Certificates. These are institutions and certificate programs, not high-DR commercial platforms. The models cited them for credibility, not for traffic. 
  • Smaller players showed the same pattern. Doxy.me, Nike Training Club, and Google Flights all shared strong product-roundup coverage, heavy expert-list inclusion, and outsized LLM recall. 

The trait these brands share isn’t a content strategy. It’s a coverage profile. They show up in the kind of third-party content that gets ingested into model training: product comparison articles, expert roundups, review content, category lists. That coverage doesn’t come from publishing more on your own blog. It comes from what other people write about. That’s the surface area of digital PR, not owned SEO.

If you’re trying to understand why a competitor with half your organic traffic is showing up in the model’s answer and you aren’t, the question to ask isn’t, “what’s their SEO strategy?” It’s, “what’s their press coverage and how did they earn it?”

The other major finding in the study cuts against the way most marketing teams currently think about AI visibility as a single metric. Cross-model agreement is the exception, not the rule.

Infographic comparing the largest gaps between AI search visibility and traditional SEO authority across industries including Insurance, SaaS, FinTech, Wellness, Lifestyle, Education, HealthTech, and Travel, highlighting the percentage of brands that are underrepresented or overperforming in AI search results.

Only 900 brands (11%) were cited by all three models. Seventy-seven percent of brands (6,264) were mentioned by just one model. The remaining 12% were split between two models.

The implication is unambiguous: if you optimize for one model, you’re statistically likely to be invisible in the others. A brand winning ChatGPT might be missing from Gemini, and that’s a different kind of problem than not winning at all. It means there’s slack in the system. The default-answer slot in each model is being contested separately, and a brand can hold one of those slots without holding the others.

The models have distinct fingerprints, too. Claude over-indexed on SaaS and insurance, while Notion, Linear, and Lemonade got cited at higher rates in Claude than the other two. Gemini over-indexed on travel and healthcare, while Booking.com, Teladoc, and Livongo showed up disproportionately in its responses. 

GPT-4o is the most consensus-driven model; its top brands had the highest overlap with the other two. These aren’t huge swings, but they’re directionally consistent across the dataset, and they suggest the models have inherited different training-data biases that show up in which categories they have more confident defaults in.

Infographic showing brands with the greatest variation in visibility across GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6, featuring a heat map of AI citation rates for brands such as HubSpot, Notion, Google Flights, Khan Academy, Airbnb, Stripe, and LinkedIn Learning across multiple industries.

The heatmap shows how unevenly distributed even consensus brands are. Notion was cited 17 times more often by Claude than by Gemini. Whoop almost four times more often by Claude than Gemini. State Farm got almost half its mentions from GPT-4o and only a quarter from Gemini. Booking.com was the rare brand that’s roughly evenly distributed, and even there, GPT-4o cited it more often.

For a marketing team, the strategic translation is direct. An aggregate “AI visibility” score is fiction. Per-model scores are the real picture. A brand winning one model isn’t the same as winning all three, and the work to close the gap on one model isn’t the same as the work to close it on another.

The disconnect rate isn’t evenly distributed across industries either. Some sectors have been almost entirely re-mapped by the models; others still look the way Google sees them.

Infographic ranking the most-cited brands in AI-generated responses across Education, FinTech, HealthTech, Insurance, Lifestyle, SaaS, Travel, and Wellness industries, comparing citation frequency across GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6 to reveal leaders in AI search visibility.

The pattern breaks down differently by sector:

  • Insurance had the highest underrepresentation rate at 8%. Legacy carriers like Aetna, Cigna, Humana, and Liberty Mutual, had been displaced by digital-native carriers in model output. If you’re an established insurance brand, you have the widest gap in the dataset to close.
  • Education had the highest overperformer rate at 7%. Stanford, MIT, Khan Academy, IBM Data Science, Google Career Certificates, LinkedIn Learning. Institutions and professional certificates dominate, and they do it without the SEO infrastructure of commercial platforms. Education is the sector where credibility coverage outweighs traffic coverage most clearly.
  • HealthTech was roughly balanced (5% under, 6% over). The shift has been from pharma and device incumbents to consumer-facing digital health. Both ends of the trade are visible.
  • Travel had the lowest disconnect overall (4%/4%). LLM recall and traditional visibility track tightly here. Booking.com, Expedia, Airbnb are top of mind for both Google and the models. The category leaders are stable.
  • SaaS and FinTech sit at 6% under/3-4% over. These are categories where the displacement is more about which incumbent the models picked (Stripe over PayPal, Notion over Microsoft) than about new entrants replacing old ones wholesale.

The pattern across sectors mirrors the broader finding. Where the legacy authority of a category was built primarily on SEO and traffic, the models have re-sorted. Where the category was built around brand stories that circulated in third-party media, the alignment is tighter.

Here are five things to take from the study, in roughly the order you’d act on them:

  1. Traditional authority does not equal AI authority. This is the load-bearing finding. DR 95 brands sat below DR 60 brands in LLM recall on a regular basis. The traditional visibility playbook tells you whether Google ranks you. It doesn’t tell you whether the models remember you. Treat them as two distinct metrics with two distinct optimization paths.
  2. Digital PR is the LLM lever. The overperformers (Monday.com, Root, Stanford, Doxy.me, Khan Academy) shared one thing. They appeared repeatedly in third-party content. Product roundups, expert lists, comparison articles, category coverage. That coverage gets ingested into model training broadly and durably. Owned content on your own domain doesn’t carry the same weight in model recall because the models weigh what other people say about a brand more heavily than what the brand says about itself. That’s the structural argument for digital PR as the AI-era brand lever.
  3. Measure per-model, not in aggregate. Only 11% of brands were cited by all three models. The rest is model-specific. Score your AI visibility against GPT-4o, Gemini, and Claude separately. A win on one isn’t a win on the others, and the work to win each is different.
  4. The default-answer window is closing. Once an LLM has settled on a category leader, dislodging it gets structurally harder. The training data gets reused. The categorization sticks. Categories that still have a wide-open field, like Wellness, Lifestyle, and some pockets of HealthTech, are the ones to move on first. Categories where the default has consolidated (Travel, parts of FinTech, parts of SaaS) are harder targets and require more sustained third-party coverage to shift.
  5. Categorization matters as much as visibility. Microsoft had more LLM mentions than most brands in this dataset. It just didn’t have them in fintech. If your AI visibility is low in your own category but high overall, your problem isn’t recall — it’s categorization. That’s a different intervention. It’s about the kind of content you appear in, not how often you appear.

The brands that win the AI-era discovery surface won’t be the ones with the biggest content libraries or the most aggressive SEO programs. They’ll be the ones whose names show up most often in the writing other people do about their category. That’s a different motion than the one most marketing teams have been built around, and the data here suggests the gap between the two motions is widening.

Methodology

This study analyzed how three major large language models, including GPT-4o (gpt-4o-2024-11-20), Gemini 2.5 Flash, and Claude Sonnet 4.6, cite brands across 8 industries (Education, FinTech, HealthTech, Insurance, Lifestyle, SaaS, Travel, Wellness).

We developed 96 sector-specific prompts (12 per industry × 4 intent types: recommendation, comparison, trend, problem-solving) and ran each prompt 15 times per model at temperature 0.7, generating 4,320 total responses for statistical stability. Brand extraction used a hybrid pipeline combining GPT-4o-mini JSON extraction, spaCy named-entity recognition, and fuzzy deduplication with 50+ alias resolutions (Alphabet → Google, Meta Platforms → Meta, etc.) and stopword filtering to remove category noise (AI, ML, API, SaaS, etc.). The pipeline identified over 8,500 distinct brands across all responses.

Traditional Visibility Scores were built from Ahrefs Batch Analysis for the top 200 brands per industry, using a weighted blend of Domain Rating (30%), Organic Traffic (25%), Referring Domains (25%), and Organic Keywords (20%), with each metric percentile-ranked within its industry to normalize for category-level differences in search volume.

A note on snapshot data: these results reflect a single scrape period in 2026. AI citation behavior shifts as models are updated, and outputs vary by user, query phrasing, and session context. We recommend treating these findings as a directional benchmark rather than a definitive ranking.

About Fractl Marketing

Fractl is a growth marketing agency that helps brands earn attention and authority through data-driven content, digital PR, and AI-powered strategies. Our work has been featured in The New York Times, Forbes, and Harvard Business Review. Fractl is also the team behind Fractl Agents, including the AI Brand Visibility Agent used to power this study.

Fair Use Statement

Fractl encourages journalists and publishers to share findings from this study with proper attribution. Please link back to this page when referencing or citing our research so readers can access the full context and explore the complete dataset behind the insights.

Avatar of Kelsey Libert

Kelsey Libert

Cofounder

Kelsey Libert is a cofounder of Fractl, a top-ranked content marketing and digital PR agency recognized on "Clutch’s Leaders Matrix" among 30,000+ firms. She has helped lead 5,000+ campaigns for brands including Adobe, Discover, and Paychex, earning coverage in The New York Times, USA Today, Vice, CNET, and other top publishers. Her industry research has appeared in Harvard Business Review, Search Engine Land, and Inc., and she has spoken at MozCon, Pubcon, SMX Advanced, and BrightonSEO.

Avatar of Aditya Sachdeva

Aditya Sachdeva

Head of Data Journalism

Aditya Sachdeva is the Head of Data Journalism at Fractl, where he has spent over five years producing data-driven content campaigns that earn coverage in publications like The New York Times, Forbes, USA Today, CNN, and Hypebeast. He turns original research and survey data into stories that place in top-tier outlets. More recently, he's been exploring how generative search is changing SEO, including GEO, LLM ranking, and AI-driven approaches to organic visibility.