Table of Contents

AI Guide for Winning B2B Visibility in ChatGPT, Claude & Gemini

There is a number on a dashboard somewhere in your marketing stack right now that is quietly lying to you. It says you “rank #3 in ChatGPT” for your category, or that your “AI visibility position” improved from 5.2 to 4.7 last month. It looks precise. It looks trackable. It looks exactly like the SERP-position metrics your team has chased for fifteen years. And that is precisely the problem – because it is measuring something that does not exist.

We work with B2B and SaaS teams every week who have wired up AI-visibility dashboards the same way they once wired up rank trackers, and then made budget decisions based on a “position” that changes every single time the prompt is run. Here is the uncomfortable truth that the best primary research now confirms: large language models do not rank. They shortlist. They assemble a consideration set – a pool of brands they consider acceptable answers to a question – and then they pull from that pool in an order that is, for all practical purposes, random.

If you internalize one idea from this article, make it this one: in AI search, the game is getting into the set and staying in the set, not climbing to an imaginary “position one.” Below, we will walk through the research that proves the consideration-set model is real, explain why traditional rank-tracking is the wrong instrument entirely, and then give you a concrete, three-vector playbook – with a 90-day implementation plan – for getting your brand shortlisted in ChatGPT, Claude, Google’s AI Mode, Gemini, and Perplexity.

The Study That Should End Rank-Tracking in AI

In January 2026, SparkToro founder Rand Fishkin and Gumshoe‘s Patrick O’Donnell published a study that ought to be taped to the wall of every marketing team experimenting with AI visibility. Titled, with characteristic bluntness, “New Research: AIs Are Highly Inconsistent When Recommending Brands or Products”, the project recruited 600 volunteers to run 12 prompts through ChatGPT, Claude, and Google’s AI a combined 2,961 times.

The findings dismantle the entire premise of “AI ranking”:

  • There was a less than 1 in 100 chance that two runs of the same prompt would return the same list of brands.
  • There was a less than 1 in 1,000 chance that two runs would return that list in the same order.
  • The most-recommended brands in a given category (for headphones: Bose, Sony, Sennheiser, Apple) showed up in only 55% to 77% of responses – never 100%, even for category leaders.
  • A real-world example, City of Hope hospital, appeared in 97% of ChatGPT runs (69 of 71) for a relevant query – but was the top mention only 25 times. Visibility was nearly universal; “position” was a coin flip.
  • The digital agency Smartsites appeared in 85 of 95 runs – high inclusion, scattered placement.
  • Crucially, the semantic similarity between how different real users phrased “the same” prompt averaged just 0.081 – meaning humans ask in wildly different ways, and each phrasing reshuffles the deck.

Fishkin’s verdict on the metric most AI-visibility tools sell you was characteristically direct: tracking a visibility percentage across many prompts run many times (he suggests 60-100 repetitions) is reasonable; tracking a “ranking position in AI” is, in his words, “full of baloney.”

Sit with the City of Hope number for a second, because it is the whole thesis in miniature. A brand can be in the answer 97% of the time and still “rank first” only a third of the time it appears. If you were optimizing for position, you would conclude you were failing. If you were optimizing for inclusion, you would recognize you had already essentially won.

From “Rank” to “Consideration Set”: The Mental Model You Actually Need

Marketers borrowed the word “rank” from the world of ten blue links, where it genuinely meant something. In classic Google, position one received a wildly disproportionate share of clicks, position two less, and so on down a stable, repeatable ladder. Run the query twice, get the same ladder. Optimize, climb a rung, capture more clicks. The model was deterministic enough that an entire industry built billion-dollar tooling on top of it.

Generative engines do not work this way, and pretending they do is the single most expensive misunderstanding in AI marketing today. A language model answering “what’s the best customer data platform for a mid-market SaaS company?” is not consulting a sorted index. It is doing something closer to what a knowledgeable human does when a friend asks for a recommendation: it calls to mind a handful of brands it considers credible answers, and then it names some of them – not always all of them, not always in the same order, often shaped by how the question was phrased.

That handful is the consideration set. And the marketing implications are profound:

  1. The boundary that matters is the edge of the set, not the top of it. The decisive question is binary – are you in or out? A brand mentioned 60% of the time is winning. A brand mentioned 4% of the time is, functionally, invisible, regardless of where it “ranks” on the rare occasions it appears.
  2. Order within the set is noise you should largely ignore. Because two runs almost never produce the same order, “moving from third to second” is not a real achievement you can engineer or defend. It is sampling variance dressed up as progress.
  3. The right success metric is share of voice – your inclusion rate across many runs of many phrasings. This is the only number that is both meaningful and stable enough to optimize against.

This is not a fringe interpretation. It maps directly onto how the major engines actually behave. Google’s AI Mode, as we will see, fans a single question out into many sub-queries and aggregates sources across all of them – a process that structurally produces a set rather than a ranking. The consideration-set model is not a metaphor we invented for clarity; it is the closest plain-English description of the machinery.

The reframe in one line: Stop asking “What’s my rank in ChatGPT?” Start asking “What percentage of relevant answers include my brand?”

Proof the Old Map Doesn’t Fit – AI Citations Are Decoupled From Search Rankings

If AI visibility were just SEO with a chatbot skin, your existing Google rankings would predict your AI inclusion. They don’t – and the data here is overwhelming and consistent across independent researchers.

In February 2026, Moz published a landmark analysis of nearly 40,000 search queries (using data from STAT) examining what Google’s AI Mode actually cites. The “AI Mode Citations” study found:

  • 88% of AI Mode citations are not in the organic SERP for that exact query.
  • Only 12% of citations were a strict URL match to the organic top 10.
  • Even at the looser domain level, only about 1 in 5 (≈20%) citations came from a site appearing in the top 10.
  • YouTube was the #2 most-cited external source.
  • 96% of responses included at least one citation, and 91% cited 10 or more sources.
  • The mechanism is “query fan-out”: one user question silently spawns many related sub-queries, and citations are aggregated across all of them.

Read that again: nearly nine in ten things AI Mode cites are not even on the first page of Google for the question being asked. The two systems are drawing from different wells.

The pattern repeats on ChatGPT. Ahrefs found just a 6.82% URL overlap between the pages ChatGPT cites and Google’s top 10 for the same queries – fewer than 7 in 100 (Ahrefs, September 2025). Profound analyzed over 650 ChatGPT queries and measured an 8-12% URL overlap, with product queries showing a correlation of roughly r ≈ −0.98 between ChatGPT citation frequency and Google rank (Profound, 2025) – a near-perfect inverse relationship for some query types. And the divergence isn’t just Google-vs-AI; 89% of the domains cited differ between ChatGPT and Perplexity as well, meaning even two AI engines barely agree with each other.

The takeaway for B2B and SaaS teams is blunt: optimizing exclusively for search rankings does not build AI citation authority. You can own page one of Google and still be missing from the consideration set. They are separate games with separate rulebooks, and you now need a strategy for both.

How Brands Actually Get Into the Set – Entity Resolution and Earned Authority

If rank-climbing is the wrong job, what is the right one? Getting included. And inclusion has two preconditions that, once you see them, reorganize your entire content strategy.

Precondition 1: The model has to know you exist as an entity

BrightEdge analyzed tens of thousands of prompts across ChatGPT, Perplexity, and Google’s AI systems and surfaced a deceptively important concept: brand resolution (BrightEdge, October 2025). ChatGPT included brand mentions in 99.3% of eCommerce responses – but only for brands it had already resolved. A resolved brand is one the model has encountered across enough credible, independent sources to construct a stable internal “entity” – a confident understanding of who you are, what you do, and who you serve.

A brand that lives only on its own domain, with no recognized third-party coverage, is not a resolved entity. It is, in BrightEdge’s framing, simply absent from the citation layer – regardless of its domain authority or its Google position. This is why the same study found that 48-77% of citations across AI engines come from “specialized sites”: industry publications, niche experts, and topic-specific resources. The long tail of earned placement is where AI goes to confirm you are real.

Precondition 2: Most of what AI cites is not your website

This is where the numbers get genuinely jarring for teams that have poured years into owned content. McKinsey’s October 2025 report, “New front door to the internet: Winning in the age of AI search” (authored by Elizabeth Silliman, Julien Boudet, and Kelsey Robinson; usefully summarized by IndexLab if the original is gated), reports that brand-owned pages typically make up only ~5-10% of the sources AI uses to construct answers. Publishers, affiliates, review platforms, and user-generated content dominate. McKinsey also projects that 20-50% of traditional search traffic is at risk and estimates roughly $750 billion in spend will move through AI search by 2028, with about half of consumers already using AI-powered search intentionally as of its August 2025 survey.

The earned-media skew is corroborated everywhere you look:

  • Muck Rack‘s May 2026 “What Is AI Reading?” analysis of more than 25 million links from AI responses across ChatGPT, Claude, and Gemini found that earned media accounts for 84% of all AI citations, while paid and advertorial content accounts for just 0.3% (Muck Rack, May 2026). Journalism alone comprised 25-27% of cited sources, spanning over 20,000 distinct outlets.
  • A 5WPR study independently found 85.5% of AI citations reference earned media, and that brands appearing on 4 or more third-party platforms were 2.8x more likely to be cited in ChatGPT (5WPR, 2026).
  • xFunnel.ai analyzed 40,000 AI responses containing 250,000 citations across ChatGPT, Gemini, and Perplexity and found earned content was the single largest citation category across all three platforms and all stages of the buyer journey (xFunnel.ai, 2025).

Put the two preconditions together and the strategy writes itself. You get into the consideration set by becoming a resolved entity that is corroborated across the earned, third-party sources AI trusts – not by polishing one more landing page. Your website still matters (more on that below), but it is the smallest slice of the pie that decides whether you’re in or out.

A note on how each engine behaves

The same Muck Rack data shows the engines have distinct citation personalities, which is worth knowing because your set membership is engine-specific:

EngineCites in responsesAvg. citations per responseNotable top domain
ChatGPT96%~5Wikipedia
Gemini82%~8Reddit
Claude55%~13PubMed Central

ChatGPT cites almost every answer but pulls from a narrow set; Claude cites less often but, when it does, casts a wide net of sources. The practical implication: a “win” in one engine does not guarantee a win in another, which is exactly why share of voice must be tracked per engine.

Ready to
turn these
insights
into a
measurablepipeline?
Our one-time SEO service gives you everything you need in 30 days – technical audit, content strategy, and implementation-ready recommendations – with no retainer or long-term commitment.

For B2B brands ready to scale paid acquisition, our B2B performance services manage campaigns for companies like Miro, Navan, and GoCardless.

For broader performance marketing needs, our performance team scales companies to their edge of potential.

Trusted by:

The 3-Vector Consideration-Set Strategy

Once you accept that the objective is inclusion and share of voice, the work organizes cleanly into three vectors. Think of them as three doors into the set, each pushing on a different lever the engines use to decide who belongs. The strongest B2B and SaaS brands in AI search are not the ones who picked one – they are the ones running all three in parallel.

Vector 1: Structured Presence – Win the Formats AI Loves to Quote

AI engines disproportionately pull from a specific shape of content: the list, the comparison, the ranked roundup. Analysis by Barchart found that listicles and comparison articles account for 59.5% of AI citations – well over half. There’s a clean logic to it: a “Top 10 CDPs for Mid-Market SaaS” article is already pre-structured as a consideration set. The model barely has to do any synthesis; it can lift the candidates directly.

Your structured-presence work, in priority order:

  1. Get listed on the third-party “best-of” and comparison pages that already rank for your category. These are the raw material AI reshuffles into its answers. If the top “best [your category] tools” listicles don’t mention you, you are starting the set-inclusion race from behind. Pursue inclusion the way you’d pursue a backlink – outreach, data, a compelling reason to be added.
  2. Maintain a strong, current presence on review platforms (G2, Capterra, TrustRadius, Gartner Peer Insights for enterprise). These are entity-corroboration and structured-comparison sources simultaneously.
  3. Build your own genuinely useful comparison and “alternatives” content – honest “X vs. Y” and “best tools for [use case]” pages – structured with clear headings, tables, and extractable summaries. You won’t be the only source, but you reinforce the entity and occasionally become the cited one.

Vector 2: Earned Authority – Distribute Beyond Your Own Domain

If owned content is ~5-10% of what AI cites, then the highest-leverage move available to most B2B teams is getting your story, data, and expertise onto third-party properties AI already trusts. The evidence that distribution moves the needle is now quantified and statistically robust.

Stacker‘s March 2026 study – the largest controlled GEO experiment published to date – analyzed 87 stories across 30 brands, querying more than 2,600 prompts across 8 AI platforms, and found a median citation lift of 239% from earned-media distribution, with 97% of distributed stories earning at least one AI citation and 64% of all citations coming from third-party publisher sources (p < 0.006) (Stacker, March 2026). An earlier Stacker and Scrunch pilot (December 2025) found an even larger 325% lift – moving citation rates from 7.6% to ~34% – and, strikingly, that in 19.2% of answers AI cited the third-party version and did not cite the original brand piece at all.

Distribution doesn’t just raise inclusion; it makes it durable. Stacker’s source-decay research, analyzing over 3 million citation events across 120,000+ domains over 26 weeks, found that distributed content maintains citation authority 2.1x longer than non-distributed content – roughly 10 weeks of half-life versus 4.5. Non-distributed content fades from AI answers in about a month; distributed content stays relevant for most of a quarter.

Your earned-authority playbook:

  1. Produce original, citable data. Surveys, benchmark reports, and proprietary statistics are catnip for journalists and, by extension, for AI. A single well-covered data study can seed dozens of third-party citations.
  2. Pursue real earned media and expert commentary in the trade and industry publications your buyers (and the models) already read. Recall 5WPR’s finding: being on 4+ third-party platforms made a brand 2.8x more likely to be cited.
  3. Distribute and syndicate strategically so your narrative reaches multiple trusted editorial sources, multiplying the touchpoints AI has to encounter – and re-encounter – your brand.
  4. Show up where the engines clearly lean. Reddit is Gemini’s top domain; YouTube is the #2 cited external source in Google’s AI Mode (and Ahrefs has measured a 0.737 correlation between YouTube presence and AI visibility). A credible community and video presence is now an AI-visibility channel, not a vanity one.

Vector 3: Entity & Extraction Signals – Make Yourself Easy to Resolve and Easy to Quote

The first two vectors get you mentioned across the web; the third makes sure that when AI encounters you, it can confidently identify you and cleanly extract what you said.

The foundational evidence here is the Princeton and Georgia Tech “GEO: Generative Engine Optimization” paper (Aggarwal et al., presented at SIGKDD 2024), which tested 10,000 queries and found that adding statistics, quotations, and source citations each produced a 30-40% improvement in AI visibility – while keyword stuffing decreased visibility by 10%. AI engines treat attributed data and verifiable claims as reliability signals, and they reward content built that way independent of the source domain.

Your entity-and-extraction checklist:

  1. Lead with data and cite your sources. Pack your content with specific statistics, direct quotes, and linked references. You’re not just informing readers; you’re sending the exact reliability signals the Princeton GEO research proved move the needle.
  2. Keep your entity consistent everywhere. Identical brand name, description, category, and key facts across your site, your “about” pages, Wikipedia/Wikidata where eligible, LinkedIn, Crunchbase, and review platforms. Consistency is what lets a model resolve you with confidence.
  3. Structure for extraction. Clear H2/H3 headings phrased as the questions buyers ask, short TL;DR summaries, tables, and direct-answer paragraphs near the top of each section. The easier you are to quote, the more quotable you become.
  4. Implement structured data (Organization, Product, FAQ, Article schema) so machines can parse your entity unambiguously. Schema alone won’t get you into the set – but it removes friction from resolution and extraction.

Measure the Right Thing – Share of Voice, Not Rank

You cannot manage what you mismeasure. If your team’s AI dashboard reports a “position,” you are tracking variance and calling it performance. Here is how to measure the consideration set the way the research says it actually behaves.

1. Track inclusion rate (share of voice), per engine. For each priority query, run it many times – Fishkin’s SparkToro study suggests 60 to 100 repetitions – and record the percentage of responses that mention your brand. That percentage is your metric. Aim to move City-of-Hope-style toward the 55-97% band the category leaders occupy.

2. Vary the phrasing on purpose. Because real-user prompt similarity averages just 0.081, a single phrasing tells you almost nothing. Build a prompt set that mirrors how actual buyers ask – different words, different framings, different stages of intent – and measure inclusion across the whole set.

3. Track per engine, separately. With 89% of cited domains differing between ChatGPT and Perplexity, and each engine showing distinct citation behavior, a blended “AI visibility score” hides more than it reveals. Maintain separate share-of-voice numbers for ChatGPT, Claude, Gemini/Google AI Mode, and Perplexity.

4. Adopt the new KPIs McKinsey names. McKinsey is explicit that classic KPIs (organic clicks, SERP positions) are incomplete and that you need GEO KPIs: AI citation share, share of voice in AI summaries, and AI sentiment. Wire these into your monthly reporting and tie them to commercial outcomes – and remember the 5WPR finding that AI-referred visitors converted at 14.2% versus 2.8% from Google, which is why this share of voice is worth fighting for.

5. Ignore “position” as a primary KPI. Watch it, if you must, as a soft secondary signal – but never set targets, budgets, or bonuses against an “AI rank” that has a less-than-1-in-1,000 chance of repeating.

The Main Facts for the B2B and SaaS Teams

The brands winning AI search in 2026 are not the ones who “rank first” in ChatGPT – because, as the research now decisively shows, that position is a statistical mirage that won’t survive a second run of the prompt. The winners are the ones who understood the deeper shift: AI doesn’t rank you, it shortlists you, and the entire game is earning a permanent seat in that shortlist.

That seat is earned in a specific way. You become a resolved entity the model is confident about. You get corroborated across the earned, third-party sources that make up the overwhelming majority of what AI cites. You show up in the structured, comparison-shaped content AI loves to quote. And you make yourself trivially easy to extract with data, citations, and clean structure. Then you measure the only number that means anything – your share of voice across many runs, per engine – and you optimize that, patiently, over quarters.

Stop chasing a rank that doesn’t exist. Start engineering your way into the set. That is where B2B and SaaS visibility – and the disproportionately high-converting traffic that comes with it – is actually won.

Ready to
turn these
insights
into a
measurablepipeline?
Our one-time SEO service gives you everything you need in 30 days – technical audit, content strategy, and implementation-ready recommendations – with no retainer or long-term commitment.

For B2B brands ready to scale paid acquisition, our B2B performance services manage campaigns for companies like Miro, Navan, and GoCardless.

For broader performance marketing needs, our performance team scales companies to their edge of potential.

Trusted by:

We value your privacy. By using our website, you agree to our use of cookies and our Privacy Policy.