ai search audit tools

AI Search Audit Tools: How Agencies Audit Client Visibility in AI Answers

By Brian Shelton — Founder of GrowPredictably.com

TL;DR: An AI search audit checks three things separately: whether an assistant mentions your client, whether it cites your client’s own domain as the source, and which third-party pages it cites instead. Most tools report only the first, so most audits hide the gap that actually matters.

Key Takeaways

  • Mention and citation are different measurements. A brand can be named in an answer while a competitor’s page supplies the source, and a single visibility score conceals that entirely.
  • The classic deliverable is worth less than it was. An AI Overview correlates with a materially lower clickthrough rate for the top-ranking page, so a number one ranking converts to less traffic than it did.
  • Assistants differ enough that per-platform reporting is mandatory. One cites many sources per answer, another cites very few.
  • The audit is a procedure, not a tool purchase. Build the question set, record mention and citation separately, then cluster the third-party sources the answers depend on.
  • It packages cleanly as a fixed-scope diagnostic, which is the practical reason to build it now rather than later.

Most agencies answering a client question about AI search reach for a visibility tool, run it, and report the score it produces. That is faster than the alternative and it reproduces the original problem at higher cost, because the score conflates two things the client needs kept apart.

What follows is the procedure rather than the pitch: what to measure, in what order, with which tools doing which part, and how to scope the whole thing as something you can sell. It assumes you already run technical audits competently and are not looking for another crawler.

What is an AI search audit, and why does it pay now?

An AI search audit is a structured check of whether a brand appears in AI-generated answers to its buyers’ real questions, whether its own domain is cited as the supporting source, and which third-party sources those answers draw from instead.

It is a question-level measurement, not a site-level crawl, and the two answer different things.

It pays now because the deliverable most agencies have always sold is quietly depreciating. Ahrefs compared Search Console clickthrough data across 300,000 keywords, 150,000 with AI Overviews present and 150,000 informational keywords without, and found that the presence of an AI Overview correlates with a 58 percent lower average clickthrough rate for the top-ranking page.

Sit with what that means commercially. You win the ranking. You report the win honestly. And you hand your client materially less traffic than the same win produced two years ago. Nothing in your reporting stack shows you this, because your stack measures position and the loss happens above it.

The ranking win is worth less than it was

Agencies feel this before they can prove it, which is roughly where the anxiety in the market comes from. SparkToro’s second annual survey of agency owners found that 53 percent now agree AI poses a significant threat to the agency business model, up from 44 percent the year before.

My take is that this is a measurement problem rather than an existential one. The threat is not that AI writes faster. It is that the thing you sell is now valued against an outcome you cannot see. You cannot argue with a client about a number neither of you is measuring.

One boundary before the method. This is not a classic technical audit. Crawl health, indexation and Core Web Vitals remain prerequisite infrastructure, and a site that fails those will fail this too. They are simply not what this measures.

There is one infrastructure check worth running before anything else, because it can explain an entire result on its own. Confirm the client is not blocking the AI crawlers in robots.txt. GPTBot, ClaudeBot, PerplexityBot and Google-Extended are the ones to look for, and any of them can be sitting in a file nobody has opened in a year.

A blocked crawler produces exactly the same audit output as bad content, so check it first. You do not want to bill for a diagnosis that a single line in a text file already explains.

Why is being mentioned not the same as being cited?

Because an assistant can name your client in its answer while linking a competitor, a directory or a review site as the source it actually relied on. Mention is a brand outcome. Citation is an owned-asset outcome. Only the second is something your client’s own site can be optimized toward, and only the second sends anyone anywhere.

The gap is the norm rather than an edge case. Semrush analyzed 126 million United States AI search prompts between January and April 2026 and reported that being mentioned in an AI-generated answer does not necessarily mean a brand’s own website is cited as the supporting source.

On Gemini specifically, the overlap between mentioned brands and cited domains reached only 30 percent.

The platforms also behave differently enough that averaging them misleads. In the same Semrush analysis, ChatGPT averaged 15 sources per response. Gemini averaged 3. An audit that blends those into one number reports an artifact of platform mix rather than anything true about your client. Keep them apart.

I’ve watched marketing leaders hire and fire agencies for fifteen years, from the agency side and from inside the buying organization, and the agencies that lose accounts are rarely the ones with worse numbers. They are the ones reporting numbers the client has stopped believing map to anything.

A visibility score with no citation breakdown is the next candidate for that treatment.

So record the two separately from the first run. Retrofitting the distinction later means running everything again, and when I audit a stalled reporting setup this is the single most common thing nobody thought to capture at the start.

How do you run the audit, step by step?

By building a question set, running it across the assistants your client’s buyers actually use, and recording three fields per question rather than one score. The procedure below takes a competent practitioner a day or two on a first pass and is repeatable per client after that.

Build the question set first

Start from the questions the client’s buyers ask, not from the client’s keyword list. Assistants answer questions, and a keyword is not a question. Pull them from sales call notes, the objections list, the support inbox, and the queries the client’s own team gets asked in demos.

Twenty to forty questions is enough for a first pass. Weight them toward the early stage. That is where a buyer has a problem and no shortlist yet, and where being absent costs the most. I tell agency owners to write the questions before they open any tool, because the tool will happily measure the wrong forty.

Record mention and citation separately

The first pass is a manual job, and that is deliberate. Open each assistant, ask the question, and note what came back. The tools further down exist to scale repeat runs, not to gate your first one. For every question, on every assistant, record three fields:

FieldWhat you are recording
MentionedIs the client’s brand named anywhere in the answer?
CitedIs the client’s own domain linked as a source?
Cited insteadWhich domains were cited, if not the client’s?

That third column is the one that turns an audit into a plan. Run it and the same handful of third-party domains will keep reappearing: a review site, a listicle, a community thread, a competitor’s comparison page. Those are the pages the answers are built from.

Cluster them, then split remediation into two piles. Things fixable on the client’s own site, which are yours to do. And placements you have to earn on someone else’s, which run on a different timeline entirely. Be explicit about that second pile before you promise anything.

I’ve seen the confusion between those two piles turn a good audit into a missed expectation inside a month.

A first pass stretches when the client sells into several distinct buyer types, since each needs its own question set. That is the main thing that turns two days into a week, and it is worth knowing before you quote.

Which tools actually measure this?

Tools in this category monitor whether and how a brand appears in AI-generated answers, and the better ones record which sources those answers cited. The established names are Profound, Peec AI, Otterly, Scrunch AI and SE Ranking’s AI Visibility Tracker.

A note on how this list was made. Capabilities were checked against each vendor’s own product pages rather than third-party roundups, and where a claim was not documented it is not repeated here. There was no hands-on bake-off across all five, so treat these as the category to evaluate rather than a ranked verdict.

What to check before you buy

Three questions separate them for agency use, and none is about the interface.

Does it report cited sources, or only mentions? This is the whole distinction above, rendered as a purchasing decision. A tool that returns a visibility score with no source-level detail cannot produce the third column of your table, which means you will be doing that part manually regardless.

How many assistants does it cover, and can you see them separately? Given how differently the platforms cite, a blended figure is close to useless for client reporting.

Can you export a per-question, per-assistant record? An agency deliverable has to survive being put in front of a client who was not in the room. A dashboard nobody can turn into a document is a tool for you, not a tool for the engagement.

One honest caution. This category is young and consolidating, so a long contract is a bet on a market that may look different by renewal. Buy a short commitment and expect to re-evaluate.

For the wider question of picking tools across your workflows rather than for this one job, choosing tools by workflow covers the selection logic, and buying AI tools against the job covers the same discipline from the client’s side of the table.

How do you package this as a client offering?

As a fixed-scope diagnostic with a defined deliverable, priced on the size of the question set and the number of assistants covered rather than on hours. The work scales with questions and platforms, so pricing it by time punishes you for getting faster at it.

The deliverable is four things: the question set you built, the per-assistant mention and citation record, the source-cluster map showing which third-party domains the answers depend on, and a prioritized list split into own-site fixes and third-party placements. That last split is what makes it a plan rather than a report.

Commercially it does the job a diagnostic is supposed to do. It is small enough to say yes to, specific enough to prove competence, and it ends with work only you are positioned to do. That matters more than usual right now, given the same SparkToro survey found only 14 percent of agencies describe their current sales pipeline as healthy.

One caveat worth being straight about with yourself. The running of this audit automates well, and the interpretation does not.

As Paddy Moogan, founder of Founder Focus and author of that survey write-up, put it:

The collection is the repetitive part. Deciding which of forty questions actually matters to a buying decision, and which third-party placement is worth chasing, is the part clients pay you for.

Price accordingly. Building versus buying the capability covers the staffing question underneath it.

What would you find if you ran this on your own site first?

The audit is a diagnostic, and the honest first test is your own agency rather than a client’s. You sell visibility for a living. So the question of whether an assistant actually recommends you when a buyer asks for an agency in your specialty is not a rhetorical one, and most principals have never checked.

Build a ten-question set on your own brand this week and run it before you sell it to anyone. If you are mentioned but never cited, you have just found your own first project, and you will sell the service better for having done it.

Want to know which stage of your growth is actually capped before you build a new offering around it? Find your growth gap.

Frequently Asked Questions

What is an AI search visibility audit?

A structured check of whether a brand appears in AI-generated answers to its buyers’ real questions, whether its own domain is cited as the supporting source, and which third-party sources those answers draw from instead. It is a question-level measurement rather than a site-level crawl, so it complements a technical audit rather than replacing one.

Is being mentioned in an AI answer the same as being cited?

No, and conflating them is the most common error. An assistant can name a brand while linking a competitor, a directory or a review site as its source. Semrush’s 2026 index found the overlap between mentioned brands and cited domains reached only 30 percent on Gemini, so the gap is normal rather than exceptional.

Can ChatGPT do an SEO audit?

It can summarize a page, but it cannot tell you whether assistants cite your client for the questions their buyers actually ask. That requires running a fixed question set across the assistants and recording which sources each answer returned. The measurement is the work, and a single chat session is not it.

How long does an AI search audit take?

A competent practitioner can complete a first pass in a day or two for a single buyer type, and it is faster on repeat runs once the question set exists. The main thing that stretches it is a client selling into several distinct buyer types, since each needs its own question set.

How much should an agency charge for an AI visibility audit?

Price it on the size of the question set and the number of assistants covered rather than on hours, because the work scales with questions and platforms and pricing by time penalizes you for getting faster. Treat it as a fixed-scope diagnostic that opens a larger remediation conversation.

Do I still need a classic technical SEO audit?

Yes. Crawl health, indexation and Core Web Vitals are prerequisite infrastructure, and a site failing those will fail this too. They are simply measuring something different, so run them as separate deliverables rather than merging them into one confusing report.

Similar Posts