Key Takeaways
- Showing up in an AI answer is not the same as being recommended inside it. Most AI visibility dashboards count the first and quietly ignore the second, which is the one that actually moves pipeline.
- Foundation Marketing ran 21 prompts twice across six AI surfaces to test ClickUp and found meaningful drift between runs, between models, and between mentioned and chosen. The variance is the finding.
- A useful AI visibility tracking tool has to separate three things: presence (are we in the answer), position (are we the recommendation or a footnote), and persistence (does that hold across reruns and surfaces).
- Treat your dashboard like a research instrument. Rerun the same prompts, log the deltas, and read the shape of the variance before you brief the team on what to fix.
- If your category is small or your buyers don't ask AI about it yet, chasing a visibility score is the wrong spend. Fix the source material the models are reading first.
Foundation Marketing published a study last week that is worth sitting with for a few minutes before you log back into whatever visibility dashboard you bought this year. They ran 21 prompts, twice, across ChatGPT, Claude, Gemini, Perplexity, Google AI Overview, and AI Mode, with ClickUp as the test brand. The headline finding is simple and a little uncomfortable: showing up in an AI answer and being recommended by that answer are two very different outcomes, and most of the tracking tools in market collapse them into one number.
I think that collapse is the real story. We are at the point in the cycle where every CMO I talk to has either bought an AI visibility tracker or is about to, and the pitch deck always shows a tidy score climbing up and to the right. The Foundation data suggests the score is often measuring the wrong thing, or measuring the right thing with enough noise that week-over-week movement is mostly drift. If you are going to spend budget here, you want to know what the instrument is actually telling you.
So here is what I want to walk through: what the ClickUp study actually shows, why mention-versus-recommendation is the split that matters, and how I would set up an AI visibility tracking tool so the number on the dashboard is one a marketing leader can act on instead of just report on.
What the ClickUp study actually found
Foundation's setup was deliberately plain. 21 prompts a buyer might realistically type, run twice across six AI surfaces, logging where ClickUp appeared and in what role. Not a thousand prompts, not a scraped corpus, not a synthetic benchmark. Just the kind of test a careful marketing team could run in an afternoon.
The thing that jumped out at me reading it (the full write-up is here) is that the same prompt, run twice, same day, same model, returned different brand sets. ClickUp would be the lead recommendation in one run and a mid-list mention in the next. Across models the drift was larger. ChatGPT and Claude did not agree on who the shortlist was. Perplexity and Google's AI Mode had their own shapes. And inside each surface, there was a meaningful gap between the prompts where ClickUp was named as the answer and the prompts where ClickUp was listed among options.
That gap is the whole game. A buyer asking "what's the best project management tool for a 20-person agency" and getting ClickUp as the answer is a very different outcome than getting a list of eight tools with ClickUp at position six. Both count as a mention. Only one counts as a recommendation.
Why most of these dashboards hide the split
Walk the category and you will see the same pattern across the tools marketers are evaluating right now, from free checkers like Ahrefs' to platforms like Profound and Rankscale, and you'll see the same pattern. Most lead with some version of a share-of-voice or visibility score. Most of the rest lead with some version of a share-of-voice or visibility score. The number is clean. The number is comparable across competitors. The number fits in a board slide.
The number is also doing a lot of work it probably should not. Under the hood, most of these tools are counting whether your brand string appeared in the model's response to a prompt. Some weight by position. A few are starting to tag sentiment or citation link. Very few are splitting "we were the recommendation" from "we were in the list," and almost none are reporting the variance between reruns of the same prompt, which is the signal Foundation's study is really surfacing.
If you only see the mean and not the variance, you cannot tell the difference between a real position change and model drift. You will brief your team to go fix something that was never broken, or celebrate a jump that will reverse itself on the next run. This is the Marketo-dashboard problem from ten years ago in a new costume: a tidy number that abstracts away the thing you actually needed to see.
The three things a useful tracker has to separate
My read on what to demand from an AI visibility tracking tool, after sitting with the Foundation data, is that it needs to report three distinct measures and never let them blur into each other.
Presence. Did our brand appear in the response at all? This is the easy one and the one every tool already does. It is table stakes and it is the least interesting signal.
Position. When we appeared, were we the recommendation, a shortlisted option, or a footnote citation? These are three different jobs in the buyer's head and they deserve three different counts. A recommendation is a sales lead. A shortlist mention is a consideration-set win. A footnote citation is an SEO-adjacent trust signal. Rolling them into one score is the equivalent of rolling MQLs, SQLs, and newsletter subscribers into one "engagement" number.
Persistence. If we rerun the same prompt in an hour, a day, a week, does the result hold? This is the measure Foundation's methodology implicitly argues for and the one almost no vendor is showing cleanly. Persistence is what tells you whether a win is structural or whether you got lucky on a sampling run.
A dashboard that gives you presence, position, and persistence as three columns, per prompt, per model, is a dashboard you can actually brief against. A dashboard that gives you a single visibility score is a dashboard that will get you in trouble.
How I'd actually set this up
If I were standing up this practice inside a B2B tech marketing team this quarter, I would not start with the vendor evaluation. I would start with the prompt set.
Sit with sales and customer success and write down the 20 to 40 real questions a buyer asks on a discovery call before they have made up their mind, in the shape of "what's the best X for a Y-sized Z." These are your test prompts and they are more important than the tool you buy to run them. Foundation used 21. That is about right.
Then pick two or three AI surfaces that your buyers actually use. For most B2B tech categories right now that is ChatGPT, Perplexity, and Google's AI Overview or AI Mode. Claude and Gemini matter but matter less for buyer research in most segments I see. Do not try to track all six at once in month one.
Run the prompts twice. On different days. Log four things per run: did we appear, in what role (recommendation, shortlist, citation), who else appeared, and what the model said about us. That last column is the one humans need to read, not a dashboard. The language the model uses about your brand is the actual GEO work surfaced in plain English, it tells you what source material the model is leaning on and what it got wrong.
Pick the tool after you have done this once by hand. You will know what you actually need the tool to automate, and you will be a much harder customer for a vendor to oversell.
When this is the wrong thing to spend on
Honestly, the whole category is in hype mode and someone should say it: if your buyers are not using AI tools to shortlist in your category yet, an AI visibility tracking tool is a premature spend. Enterprise security buyers, regulated-industry buyers, and most deeply technical infrastructure buyers are still running RFPs and analyst calls, not asking ChatGPT who to buy. Check with your sales team. If "how did you first hear about us" never comes back as an AI answer, you have time.
Secondly, if your source material is thin, if your site, your docs, your G2 presence, your third-party coverage are all weak, a tracker will just tell you repeatedly that you do not show up. You do not need a dashboard for that. You need to go write the pages, earn the citations, and get into the comparison posts the models are reading. Measurement sits on top of that work. It does not substitute for it.
The piece worth repeating to clients is that the AI surfaces are reading the open web and a handful of trusted sources. If you are not in those sources with clear, specific, well-structured answers to buyer questions, no amount of visibility tracking will change your score. If you want a walkthrough of how we sequence that work, our growth blueprint lays out the order of operations we use.
What the Foundation study really changes
The useful shift from Foundation's piece is not "AI visibility tools are bad." It is that the tool is a research instrument, not a scoreboard. Expect noise. Rerun your prompts. Look at variance, not just the mean. Split mention from recommendation every time you report. Pick the handful of prompts where being the chosen answer actually matters to pipeline, and go deep on those instead of boiling the ocean across hundreds of keywords.
That is a less sexy dashboard. It is also a dashboard that will not embarrass you in six months when the board asks why the score went up and the pipeline did not.
The move for most teams right now is to rebuild internal reporting on this basis and give it a quarter before trusting the trend line. If you want a place to start Monday morning, pick the single prompt where being the recommended answer would most clearly show up in pipeline, and rerun it three times before you touch anything else.
Frequently Asked Questions
What is the difference between an AI visibility tool and an SEO rank tracker?
An SEO rank tracker measures where your pages appear in a search results list for a keyword. An AI visibility tool measures whether and how your brand appears inside a generated answer from a model like ChatGPT or Perplexity for a prompt. The jobs are related but not the same, one tracks document ranking, the other tracks whether a model picks you up and recommends you.
Which AI surfaces should a B2B tech brand track first?
Start with ChatGPT, Perplexity, and Google's AI Overview or AI Mode. These are the surfaces most B2B buyers are actively using for research right now. Claude and Gemini are worth adding once you have a rhythm, but you do not need all six on day one. Pick the ones your buyers are actually in.
How often should I rerun my prompt set?
Weekly is a reasonable cadence for an established program. Monthly is fine if you are just starting. The important thing is to run each prompt at least twice per cycle so you can see variance, not just a single-point reading. Foundation's study shows how much the same prompt can drift between runs on the same day.
Can I do this without buying a tool?
Yes, for a small prompt set. Twenty prompts across three surfaces, run twice, is a few hours of work with a spreadsheet. That is actually how I would recommend starting, because it forces you to look at the actual language the models use about your brand instead of just a number. You will know what to buy, and why, once you have done it manually.
Does being recommended by an AI model actually drive pipeline?
Yes, in categories where buyers are using AI for shortlist research — early signal is real. Sales teams are starting to hear "ChatGPT suggested you" on discovery calls. The volume varies wildly by segment, so check with your sales team before you size the opportunity. If no one is hearing it yet, the pipeline impact is still ahead of you.
Sources
- Why AI visibility scores hide whether brands get chosen, Foundation Marketing
- The 9 best AI visibility tools in 2026, Zapier
- 18 Best AI visibility tools for marketing agencies, Profound
- The Best AI Visibility Tracking Tools, Compared, Brainlabs
- 5 AI Visibility Tools to Track Your Brand Across LLMs, Backlinko
- Free AI Visibility Checker, Ahrefs
- AI Visibility Tracking Tools, Lumar
- Rankscale: AI Visibility Platform
- AI Search Visibility Tool, SE Ranking
- Best AI Search Monitoring Tools for Marketers in 2026, Nightwatch