This website uses cookies

Read our Privacy policy and Terms of use for more information.

🛤️ Why prompt tracking is crucial in understanding what your ICP sees in AI

A stream of scattered AI answers converging into a single, trackable signal.

Which approach should you take for prompt tracking?

Back in January, Rand Fishkin and Patrick O'Donnell from Gumshoe.ai published new research testing how consistent AI answers actually are. 600 volunteers ran 2,961 AI queries across ChatGPT, Claude and Google's AI Overview, using 12 different product recommendation prompts. The point was to see how consistent AI answers actually are.

According to Fishkin, not very. There's less than a 1 in 100 chance ChatGPT or Google's AI gives you the same list twice if you ask it repeatedly. Get two lists in the same order and you're looking at more like 1 in 1,000. Ask an AI who the best running shoe brands are, ask again five minutes later, and you'll probably (well, you will) get a different answer both times.

But are the brands included the same? Take Butternut Box, a brand we track but don't work with. Across the 25 prompts in their "dog food subscription" topic, tracked daily, they've held visibility around 82%.

A single day's read on 25 prompts carries a wide margin of error, easily ±15 percentage points either way. But we're not taking a single reading, we're tracking the same 25 prompts every day, and each extra day adds to the sample. By the time you're looking at a month of the same 25 prompts, that margin tightens to under ±3 points. That's what the chart below shows: not a snapshot, a number converging on something real, holding in a tight band rather than swinging from 80% one day to near zero and back.

Butternut Box, Different Dog and Tails.com's daily AI visibility over a month. Each holds its own band rather than swinging randomly.

Fishkin's actual conclusion from this was more sensible. His words: "any tool that gives a ranking position in AI is full of baloney." But he also said "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric."

In other words, stop trying to track where you sit on any single prompt, a one-off "you're #2 today" is just as unreliable as a single yes or no on visibility. What actually holds up is average prominence, where you tend to sit once you've run it enough times to see the real pattern, not a single snapshot dressed up as a fact. Track how often you show up across a broad set of them too.

That's not where it stopped, though. A follow-up piece built on top of that same research, using worst-case statistical assumptions to work out how many prompts you'd need to run before the numbers were trustworthy. The answer it landed on was thousands. That's the version that got forwarded to me by more than one prospect, basically as a challenge: prove your numbers aren't nonsense.

A sensible statistical assessment

So I put this to Mike, my co-founder, who's a statistician by training and knows this stuff better than anyone else I know.

His response, in short: the maths behind "you need thousands of prompts" wasn't wrong, but the assumptions feeding it were needlessly harsh for a commercial context. Worst-case probability and clinical-trial confidence intervals are the right call if you're testing a drug.

They're overkill for tracking whether a brand shows up in ChatGPT. Dial those down to something a marketer would actually use, and the number of prompts required drops from the thousands into the hundreds, especially once you account for tracking as a rolling time series rather than a single snapshot. Mike wrote an article on the optimal number of prompts you should be tracking to understand what your customer could see in AI search.

Margin of error against number of prompts tracked, using Obsero's own 90% confidence framework. It flattens hard past a couple hundred prompts.

Worth flagging: a few months later Fishkin backtracked himself, calling broad prompt tracking a reasonable metric he's "converted" on, even if rankings are still "baloney" to him.

Round two: Evertune vs Profound

Then last week it flared up again. Evertune's CEO, Brian Stempeck, took a public swing at Profound, a much bigger rival in the same space, and it landed on almost exactly the same answer.

He posted a breakdown arguing that Profound's sampling approach produces unacceptable noise. His numbers: run a single prompt around 30 times and you're looking at an 11-point margin of error. Run it 100 times and you're still sitting at roughly 6 points. Fishkin turned up in the comments backing this approach and dunking on Profound stating “they’re bad news in more ways than one”, whatever that means.

Profound's response was the interesting part. They're not hammering a single prompt over and over in one day. Each prompt still runs daily, building its own day-by-day time series, but the scale comes from the portfolio, hundreds or thousands of distinct prompts, not from repeating any one of them within a session.

Their own test across 753 prompts found only a 0.25 percentage point difference between running a given prompt once a day versus ten times a day, so that extra repetition buys you almost nothing.

Significance vs diversity

I asked Mike what he made of it.

"You're better having a diverse prompt set within one topic, capturing more of the query fanout that's relevant to what you're trying to sell, and running those multiple prompts once a day, than running the same prompt 100 times a day. The cost is the same either way. But with the same prompt repeated, you're learning less about the breadth of questions people actually ask, and the breadth of answers you might be mentioned in.

I'm 100% on Profound's side of this one. You learn more that way, and you're less reliant on guessing exactly how people talk to LLMs, which you don't actually know. Putting in a wide range of prompts reduces the odds you've guessed wrong about how your customers phrase things. Evertune's take isn't wrong exactly, it's just coming from someone thinking like a statistician rather than someone who's sat with a customer and watched how they actually use this stuff."

Mike Logue

Mike Logue · Co-founder, Obsero

Fishkin gets quoted as the guy who thinks AI visibility tracking is basically snake oil. Read the actual research and that's not his point at all. His point is: don't trust a single prompt, and don't trust anyone selling you a "ranking position." Track broadly, track repeatedly, and you've got something real. That's Profound's approach, that's Mike's approach and ultimately, Obsero’s approach too.

Statistical significance tells you whether you can trust a number. Prompt diversity tells you whether that number reflects anything real about your customers. You need both, but diversity is the one people keep ignoring to only then argue about sample sizes.

Track the same prompt 100 times a day and you'll get a very precise answer to a question nobody's actually asking. Track a genuinely varied set of prompts across a topic, even a smaller number of them, and you start to see what your customers are actually seeing when they ask an AI about your category.

That's the whole point of prompt tracking. Not "are we technically visible", but "what does a real customer, typing a real question, actually get told about us" at scale.

It doesn’t start and end with AI visibility

But here's the bit I want to be honest about. Visibility is one layer of a much bigger GEO playbook, the ongoing work of understanding, measuring and actually influencing how customers discover your brand through ChatGPT, Claude, Google's AI mode and Gemini.

All that visibility % tells you is: are you being mentioned. That's it. It doesn't tell you how you're being mentioned. Are you the first brand recommended, or are you buried three sentences down as an afterthought. Personalisation muddies this day to day, sure, but with the right persona analytics you can still get a handle on what a given user is likely to see, based on who they are and what they're asking.

This isn't a new problem, it's the old keyword research problem dressed up, one I was wrestling with back in the early 2010s. Someone searching "mascara" could be a 20 year old, a 40 year old or a 60 year old, and Google was never going to hand back the same intent or the same right answer for all three.

Who's actually asking, and what do they need, was always the real job. It still is. It comes down to how well you understand your customers, and whether you've got a partner who can build the touch points and the narrative that actually reaches them.

Do you know whether your ICP is using these tools logged in or logged out? That changes whether the answer they get is personalised to their history or a generic public one. Do you talk to your customers about their journey to discover your brand? Do you run surveys? When someone converts, do you ask them where they last heard about you, and do you give them "an AI" as an option rather than lumping it into "other."

Visibility tracking answers one narrow question well. Everything else on that list is what tells you whether the answer means anything.

Loving this newsletter? Why not create your own?

One of the best things I've ever done was start a newsletter.

I learnt a new skill. I write about my passion (AI search) and it's become the best business development channel I have. Prospects reply to editions. Clients mention stories weeks later. It's put me in rooms I'd never have got into cold.

Beehiiv's the only platform I've used, but it really works for me. I think it's brilliant. They've also just had a big couple of weeks: last week they launched Community, a built-in space to actually talk with your readers, an AI copilot that handles the boring bits, plus a proper drag-and-drop visual editor and native podcast hosting, all in one place.

Get paid for something you love doing. Sign up with my referral link and get 20% off your first three months.

Redesigning Obsero’s prompt window

I'm super pleased with how our new prompt window looks. If you're an Obsero customer, jump in and take a look today.

The redesigned prompt window, showing which AI answer components are present and the query fan-out behind them.

We know how much the individual components inside an AI answer matter, so we now show whether a shopping carousel, a map pack or local business information is present in the result.

We now capture query fan-out on the right-hand side too, along with a clear read on what's enabled and what isn't. That gives you a proper answer to whether you're actually visible inside the map pack, or whether there's work to do to get there.

Built for marketers by experienced marketers

For anyone who's followed Obsero's journey for a while, this is a big step up from what we had before. I'm honestly loving being able to work on stuff like this every day, with a team that's got the drive, energy and experience to build it properly.

A customer's take on Obsero as a whole, and exactly the kind of feedback that reminds me why I started this journey.

I was delighted to get that feedback from a new customer earlier in the week. This is exactly why you do it, helping customers understand the category better than they could on their own, and being a partner, not just a platform.

How I actually write this newsletter

I use Wispr Flow for everything: emails, prompts, even this newsletter. It gets my thoughts down properly, and I can then go back and place them logically within whatever I'm writing. It's saved me so much time.

Recommended, and I genuinely love it. Give Wispr Flow a try.

That is it for this week.

Until next time.

Andy

🗣️ This week’s stories

Five stories from the week in AI search.

Reply

Avatar

or to participate

Keep Reading