This website uses cookies

Read our Privacy policy and Terms of use for more information.

Sponsored by

AI agents built to get work done

Skydive agents take on work for you, using the same tools your team already uses.

Talk to your agent on the web, Slack, email, iMessage, or right from your terminal. Wherever you pick up the conversation, your agent keeps the context.

Put agents to work across customer support, sales, marketing, engineering, ops, and more.

Give your agent a job. They’ll take it from there.

How to run your own AI search audit

There is one question most brands still cannot answer with any confidence, which is what an AI assistant says about them when a customer asks. Plenty have typed a few prompts into ChatGPT and had a look, and that is a very different thing from a baseline you can measure against and come back to.

Here is how to find out.

Five steps, and you can run a rough version of it yourself.

Step one. Decide what you are measuring against

Before you touch anything, get clear on what you want to achieve with the audit. There are three dimensions worth covering and most brands only ever get to two.

  • Your brand, which is what an assistant says when someone asks about you directly.

  • Your category, which is whether you come up at all when someone asks about the kind of thing you sell.

  • And your competitors, which is what gets said about the brands you are losing deals to.

Pick three to five of those competitors, and pick the ones your sales team genuinely loses to. Aspirational comparisons will hand you a flattering number and nothing you can act on.

This matters because a visibility score means nothing on its own. Appearing in 40% of answers could be excellent or dreadful, and the only way to know which is the rival number next to it.

If they are on 70% while you are on 40%, you are being left out of the shortlist, and the work is third-party presence and category content rather than anything on your own product pages.

Step two. Build the prompt set

In search you track keywords. In AI search you track prompts, which are the actual questions people ask when they are working something out.

This step decides whether the rest of the audit is worth anything, and it is where most of them fall over. People pick the prompts that feel important internally, the ones lifted straight out of the positioning deck, and they end up with prompts that flatter them.

If you want a head start rather than a blank page, I have put twenty of them into a free set you can take and adapt. Ten category, five branded and five competitor.

AI visibility audit prompt set
AI visibility audit prompt set
Twenty ready-made prompts for finding out what AI assistants actually say about your brand, your category and the competitors you lose deals to. Paste them straight into ChatGPT, Claude, Perplexity...
$0.00 usd

Build across three types. Category prompts are the broad exploratory ones, along the lines of "what is the best travel insurance for pre-existing conditions". Branded prompts are about you specifically, so "is [brand] any good for [use case]". Competitor prompts name the rivals you picked in step one, so "how does [brand] compare to [rival]".

Fifty is plenty to start with, and depth beats breadth here.

One practical thing. Write them the way a person actually talks to an assistant, in full sentences, because "best project management tool agencies" and "what is the best project management tool for a small agency" will not come back with the same answer.

Pro tip. Use Apify to scrape People Also Ask data and you will get the questions people are genuinely asking around a topic, which is a far better starting point for a prompt set than a brainstorm in a meeting room.

A branching tree of nine People Also Ask questions for the seed term macbook pro in the UK, shown across three levels of expansion, with five marked as objections.

People Also Ask for the seed term macbook pro in the UK, three levels deep. Five of the nine questions are objections. All pulled using Apify.

Step three. Run the set across several platforms

ChatGPT, Perplexity, Google AI Mode and Gemini do not agree with each other. They pull from different sources, they give different brands prominence and the sentiment can swing between them on the same question. Audit one platform and you have audited one platform.

Run your full set across at least three of them and record four things for every answer: whether you appear at all, where you appear in the response, the actual language used about you and which sources got cited.

Resist interpreting any of it yet. This step is collection and nothing else, and the urge to start drawing conclusions off the first five answers is exactly how these things go wrong.

The platforms report in Obsero, showing how a brand performs across different AI search platforms alongside its competitors.

The platforms report from Obsero allows you to view a breakdown of your prompts across different AI search platforms to assess where you are strong or weak vs. competitors.

Step four. Read visibility, prominence and sentiment together

Now you interpret it. Three things matter and most audits stop after the first.

  • Visibility is how often you turn up across the set. That is your headline number and the closest thing AI search has to share of voice.

  • Prominence is where you turn up when you do. Being mentioned in an answer and being the answer are different commercial outcomes, and the second one is what converts. A brand sitting fifth in a hedging paragraph at the bottom of a response is technically visible and practically absent.

  • Sentiment is what it says about you, and whether any of it is accurate. A visible brand with consistently lukewarm framing has a completely different problem to an invisible one, and the fixes have nothing in common.

Look at patterns across the whole set as you go. Individual answers move around between runs, so one bad response is noise and a consistent gap in a single topic area is a finding.

Pro tip. Use Obsero's free trial to get set up with fifty free prompts. It will show you the questions people are asking about your brand and category, and identify the competitors that keep turning up in AI answers. You might be surprised at who is recommended 👀

Step five. Map the citation sources

This is the step that turns a diagnosis into a list of things to do. Citations are the sources an assistant leans on when it builds an answer, and three questions get you everything useful out of them.

What types of sources are being cited? If your category is being explained by review platforms and Reddit threads, that tells you where the assistant is forming its view of you, and it is probably not your homepage.

Are your own pages in there? If your content is getting cited it is being read and used. If it is absent while a competitor's is not, the reason is usually content that does not directly answer the question being asked.

Where are competitors cited that you are not? This is the most useful thing in the whole exercise. A specific publication, a specific review site or a specific type of page where they show up and you do not is a targeted list you can hand straight to whoever owns PR and content.

Three panels showing citation source types, own-site citation share, and a publisher gap chart for a US denim brand.

Where the answers come from for a US denim brand. Its own site is 3.4% of citations, and 0.7% in answers it does not appear in.

What you end up with

Prompt-level detail does not survive the trip into a board paper, so the audit needs an output that travels. I would keep it to a red, amber, green across three things: category visibility, sentiment and citation coverage.

That format does two jobs. It gives you an honest read of where you stand, and it gives leadership a sentence they can act on. A denim brand I looked at this week comes out green on visibility and green on sentiment, and it is named in under 8% of answers about what to buy this season. "We win the category and lose the trend" is a far easier conversation than asking for GEO budget in the abstract.

Worth being straight about the measurement itself. Analytics will not show you AI referrals cleanly, Search Console does not track answer appearances and the tooling across the board is still maturing. That is an argument for starting now with an imperfect baseline, because the brands that are ahead in a year will be the ones already watching their own trend line.

A baseline audit dashboard showing category visibility at 57 percent, sentiment at 65 to 75 percent, citation coverage at 4 percent, and visibility broken down by topic.

The baseline for a US denim brand. Green on visibility and sentiment, amber on citation coverage, and under 8% on what to buy this season.

What to do today.

  • Write down your three to five competitors and 50 questions that your ICP asks about your brand and category. Your best approach is to speak with customers, use Apify to scrape PAA or Reddit threads and analyse call logs if you have access to these.

  • Run them through the assistants that matter to you and your category, whether that is Google AI Mode, Gemini or ChatGPT. Doing this by hand is slow work. Obsero's free trial tracks fifty prompts for you, which takes the manual part out of it.

  • Do it again on a regular basis. A prompt is a sample, not a fact. The same question asked twice will come back differently, and that is sampling rather than the tool being unreliable, which is why how consistently you appear is the actual measure, and you only get that from repeat runs.

Pro tip. Different models give different answers, so it is worth knowing which one your ICP is actually in. Free users are on GPT-5.6 Luna. Plus subscribers are on GPT-5.6 Sol, and only see GPT-6 Astra inside Work and Codex. Personalisation shifts things again, so audience traits matter as much as model choice.

I'm speaking in Dublin on the 29th

Speaker card for Andy Francos, co-founder and Chief Product Officer at Obsero, on Panel 1, Building AI on AWS, in Dublin on 29 September.

Panel 1, Building AI on AWS, at the Guinness Enterprise Centre on Tuesday 29 September.

If you are anywhere near Dublin on Tuesday 29 September, I am on a panel at the Guinness Enterprise Centre on building AI on AWS, alongside Christian Saam from VoiceTune AI and Amy Neale from Delta Partners. MCS Group and Kimber are running the morning.

There is also a cybersecurity keynote from Niranjan Kunwar, CTO at Genese Solution, and a second panel on hiring in the age of AI with Anthony Brew from Zapier, Sergey Volkodav from Darkhunt AI and Dee Coakley from Payoneer.

It runs 09.30 to 13.00, it is invitation only and places are limited. Reply to this email if you want to come and I will put you in touch with the organisers.

🗣️ This week's stories

Illustration marking the weekly roundup of AI search stories.

What caught my eye in AI search this week.

  • Profound raised $180m at a $1.8bn valuation, seven months after its last round. Sequoia and Kleiner Perkins led it, revenue is up 3x in six months and they are past 1,000 enterprise customers including Comcast, Estée Lauder and Walmart. Credit where it is due. Profound has done more than anyone to build this category, and a raise at that size is the clearest validation yet of something a lot of us have been arguing for two years, which is that brands need to know what AI is saying about them. That is the job I started Obsero to do, and the market has just told everyone it is a real one. TechCrunch

  • OpenAI is testing ads that open a conversation instead of your website. Announced on Wednesday. Click a ChatGPT ad and you can start a labelled chat with a brand's sponsored agent, running in a session separate from your own. Advertisers also get campaign management by prompt, copy and imagery drafted from their landing page, and an opt-in setting that adapts ad copy and headlines to the user. Shopify's ChatGPT Ads app goes international on the 23rd. OpenAI

  • A federal judge spared Google's ad tech business, with AI reshaping the market in the background. Judge Brinkema's full opinion landed on Wednesday, ordering interoperability with rivals and stopping short of divestiture. She put the remedy period at six years when the DOJ had pushed for fifteen, and both she and Judge Mehta in the search case noted that AI is reshaping the markets in front of them. Nearly four years of litigation produced a code of conduct. Digiday

  • 61% of US consumers have used an AI assistant for shopping research in the past three months. That comes from a Dept survey of 2,600 shoppers, and insurance and automotive marketers are already quoting numbers like it in budget conversations. Which makes AI search a media planning problem as much as an organic one. Digiday

  • Google is paying publishers based on measured usage of their content. A pay-per-value licensing scheme, which critics in the piece call a legal fig leaf more than a meaningful payout. Axios is working the same seam from the other end, building feeds sold directly to AI models and agents as a 2027 revenue line. Digiday

  • Creators have started pitching brands on AEO. They are now selling citation share alongside audience size, with the new line in the pitch being that they got cited. It follows Perplexity going public earlier in the week with a serious push to court creators, who mostly still see AI as a threat to their living. Digiday

That is it for this week, until next time.

Have a good one.

Andy

Reply

Avatar

or to participate