Key Takeaways

  • The Growthwaves study of 24,067 AI search prompts shows that AI engines pull disproportionately from third-party listicles, review roundups, and comparison content rather than vendor .com pages, which changes where SaaS marketers should be investing.
  • A useful working benchmark for AI search visibility for SaaS is a 20 to 30 percent citation rate across a tracked prompt set that mirrors your buyer's real questions; below 10 percent means you're effectively invisible in the answer layer.
  • Prompt-level tracking beats keyword-level tracking, because AI answers are assembled per-query from a shifting mix of sources, and one prompt phrasing can flip which vendors get named.
  • The fastest wins sit off your own site, in the third-party sources the models already trust, and in making your own pages easier for a model to quote in one clean paragraph.
  • Treat AI visibility as a category-share metric, not a ranking. Track share of answer against a defined competitor set and a defined prompt set, and re-baseline monthly as the models change.

Most of the AI search advice floating around right now is vibes. Someone asks ChatGPT about their category, doesn't see their brand, and writes a LinkedIn post declaring SEO over. That's not data, that's a mood. So when Growthwaves published a study that actually ran 24,067 prompts through AI search engines and looked at which SaaS brands got cited and why, I read it twice and then made our team read it. It's the first thing I've seen at that scale that gives a CMO something to benchmark against instead of guess against.

What I want to do here is walk through what the study actually says, what I think it means for how a B2B SaaS marketing team should spend the next two quarters, and where I think the conventional read of it is wrong. The headline isn't that AI search is important, you already know that. The headline is that the mechanics of getting cited look almost nothing like the mechanics of ranking on Google, and if you port your SEO playbook over one-for-one you'll spend a lot of money on the wrong things.

The one idea I want you to leave with: getting cited by AI engines is a category-share game played mostly on other people's websites, and you win it by being the easiest company for a model to quote, not the loudest.

What the 24,000-prompt study actually found

The Growthwaves study ran a very large prompt set across the major AI search surfaces and coded what got cited. The pattern that jumped out at me: the sources the models lean on are overwhelmingly third-party. Listicles, review sites, comparison pages, industry roundups, Reddit threads, community posts. Vendor .com pages show up, but they show up less than a marketing team would want, and they show up mostly when the prompt is already navigational ("what does Notion do") rather than evaluative ("best tool for X").

That second bucket, evaluative prompts, is the one that matters, because that's where buying happens. A CMO trying to get their product into consideration doesn't care that the model can describe their homepage. They care whether the model names them when a buyer asks "what are the best options for onboarding B2B customers under 500 seats." And in that bucket, the model is mostly reading G2, Capterra, a handful of respected blogs, a Reddit thread, and two or three listicles from publications the model has learned to trust. Your homepage is not in that stack.

The other finding worth sitting with is how unstable the citations are. Rephrase the prompt slightly and the cited set shifts. This is why measuring AI search visibility as a share metric, share of answer across a defined prompt set, is more honest than measuring it as a ranking. There is no rank. There is a probability that you get named, and that probability moves.

Why the old SEO instinct misfires here

That instability is exactly where marketing leaders are about to waste money. The reflex, when you learn that AI engines pull from third-party content, is to publish more of your own third-party-looking content. More blog posts. More comparison pages on your own domain. More "Us vs Competitor" pages. That is the SEO instinct, and it is mostly wrong for this problem.

The models are not choosing your comparison page over G2's comparison page. They're choosing G2. The trust signal the model has learned isn't "this page is well-optimized," it's "this domain has been cited a lot by other trusted sources when the topic comes up." The work sits off your site, in the sources the model already trusts. That's a PR and partnerships motion dressed up as an SEO motion, and most SaaS marketing teams are not staffed for it.

The second misfire is treating this as a keyword problem. The Martech piece on ranking for SaaS in AI search gets this right: you have to think in prompts, not keywords. A keyword is a stem. A prompt is a fully formed question with context, constraints, and intent baked in. "Onboarding software" is a keyword. "What onboarding tool works best for a Series B B2B SaaS with a PLG motion and under 50 CS people" is a prompt. The second one is what your buyer actually types, and it's what the model actually answers.

Your prompt set, the tracked list of questions you care about, is now the artifact that matters, more than your keyword list ever did.

A working definition of AI search visibility for SaaS

That prompt set is also the thing you have to measure against, which means you need a number. The definition I've landed on, borrowing from a few sources including Averi's guide and SimpleTiger's benchmarks, is this: your AI search visibility is the percentage of prompts, in a defined prompt set relevant to your category, where your brand is named or your pages are cited by a major AI engine.

Two things about that definition are load-bearing. First, "defined prompt set", someone on your team has to sit down and write the 100 to 300 questions your buyer actually asks, across the stages of their journey, and commit to them as the tracked set. Second, "major AI engine", plural. ChatGPT, Claude, Gemini, Perplexity, and increasingly Copilot inside Microsoft's stack. Each one has different source preferences. Being strong in one and weak in three is a real situation and worth knowing.

The benchmarks I've seen converge in a useful range. Under 10 percent means you're not in the conversation. 20 to 30 percent is the threshold where AI search is meaningfully contributing to pipeline. 30 to 50 percent puts you in category-leader territory. Those numbers won't be identical for every category, a five-vendor category and a fifty-vendor category are different math problems, but as a starting frame they're honest.

What I'd actually do in the next 90 days

If I were a VP of Marketing at a mid-market SaaS company reading this study on a Monday morning, I'd start with the artifact everything else is measured against. Get your product, sales, and CS leads in a room and write down the 150 questions a real buyer asks across awareness, evaluation, and selection, in full sentences with context, and commit to them as the tracked set.

Then run those prompts through ChatGPT, Claude, Gemini, and Perplexity and log which brands get named and which sources get cited. You can do this manually for a first pass, or use one of the AI search visibility platforms now on the market. The goal of the baseline isn't to fix anything yet. It's to know the truth about where you actually stand, engine by engine, prompt by prompt, before you spend a dollar reacting to it.

Then comes the real game, which is the source work, not the site work. Look at the third-party pages the models cited when your competitors got named and you didn't; if it's a G2 category page, your G2 presence needs work, if it's a specific listicle on a specific publication, that publication needs a relationship, and if it's a Reddit thread, someone on your team needs to be a real, credible participant in that community rather than a spammer.

Only after that would I go back to your own pages and make them more quotable, because models like to lift clean, self-contained paragraphs that answer a question directly, and if your product pages bury the answer three scrolls down inside a hero animation, the model won't quote you even when it lands there. Rewrite the key pages so the answer to the obvious question sits in the first paragraph, in one sentence, in plain language. Then re-baseline.

When this is the wrong thing to focus on

I want to be honest about when I'd tell a CMO to not spend cycles on this yet. If your ICP doesn't use AI search to evaluate vendors, and there are still categories where that's true, particularly in regulated enterprise buying with heavy procurement involvement, you're solving a problem your buyer doesn't have.

Talk to twenty customers. Ask them, specifically, whether they used ChatGPT or Perplexity anywhere in their last vendor evaluation. If the answer is mostly no, park this and revisit in two quarters.

Second, if your fundamentals are broken, your positioning is unclear, your site can't convert the traffic you already have, your sales team can't articulate why you win, AI visibility is a distraction. Getting named by ChatGPT in a category where you can't close deals just accelerates disappointment. Fix the machine first, then feed it more air.

And third, if you're at a stage where hiring an agency of any kind, ours included, would eat a quarter of your marketing budget for one channel bet, don't. This is a program you can start in-house with a spreadsheet and a prompt list. The tooling gets useful at scale, not at the start.

What to do about it

The move I'd push for is treating the tracked prompt set as a first-class marketing artifact, the same way you treat the ICP document or the messaging house. Version it, review it monthly, and check every content brief against it. And swap the question on every brief from "will this rank" to "will a model quote this paragraph when a buyer asks this prompt." That's a small reframe and it changes almost everything about how a team writes.

The study doesn't tell you what to do. It tells you you can finally stop guessing about whether it matters.

Frequently Asked Questions

What is a good AI search visibility benchmark for a B2B SaaS company?

A citation rate of 20 to 30 percent across your tracked prompt set is the working threshold where AI search starts contributing meaningfully to pipeline. Below 10 percent, you're effectively invisible. Category leaders often sit at 30 to 50 percent. These numbers assume you've defined a prompt set that reflects real buyer questions, not just brand searches.

How is AI search visibility different from SEO ranking?

SEO ranking is per-keyword and per-page. AI search visibility is per-prompt and per-answer, and the answer is assembled from many sources at once. There is no single rank. You're measuring the probability your brand gets named when a buyer asks a specific question, which means the unit of tracking is the prompt rather than the keyword, and the unit of optimization is often a third-party source rather than your own page.

Where do AI engines actually pull their citations from?

The Growthwaves study and other prompt-level research consistently show AI engines lean heavily on third-party sources: review platforms like G2 and Capterra, respected industry publications, comparison listicles, and community threads like Reddit. Vendor .com pages show up more for navigational prompts than for evaluative ones, which means most buying-intent answers are being assembled from sources you don't control.

Do I need an AI visibility tool, or can I track this manually?

For the first baseline, you can track it manually with a spreadsheet, a prompt list, and an afternoon per engine. Purpose-built AI visibility tools become worth it when you're tracking hundreds of prompts across four or five engines on a weekly cadence, or when you need to compare share of answer against a defined competitor set over time. Start manual, upgrade when the manual work stops scaling.

How often should I re-baseline my AI search visibility?

Monthly is a reasonable cadence for most SaaS teams. The underlying models change, the source preferences shift, and your own actions take a few weeks to show up. Weekly tracking is overkill unless you're actively running a campaign to influence a specific set of prompts. Quarterly is too slow, you'll miss shifts that matter.

Sources