Home/Blog/measure AI search visibility

How to Measure AI Search Visibility Without Fooling Yourself

AI answers are non-deterministic, so the obvious measurement method produces confident nonsense. Here is a four-layer approach that survives the noise.

Updated August 2026
How to Measure AI Search Visibility Without Fooling Yourself
The short answer

Measure four layers: presence (mention rate across a fixed prompt set), referral (traffic from AI sources), demand (branded search volume), and outcome (qualified inquiries). Run a fixed prompt set monthly and track the rate over time. Never draw conclusions from a single query, because the same prompt produces different answers between sessions.

Someone asks ChatGPT whether it recommends their business, gets a good answer, and concludes their AEO is working. The next week they ask again, get nothing, and conclude it has broken.

Neither conclusion is supported. They ran a sample size of one against a non-deterministic system, twice.

This is the central measurement problem in AEO and most of the tooling and advice glosses over it. Rankings were stable enough that checking once told you something. Generated answers are not. Same prompt, same day, different session, different answer.

Why this is genuinely harder than rank tracking

Three properties make it awkward, and they compound.

Non-determinism. These systems sample from probability distributions. Variation between runs is inherent, not a bug to be controlled for.

Personalization and context. Answers vary by account history, location, and conversation context. There is no single canonical result to observe.

No position. There is no equivalent of "we rank fourth." You are either mentioned or you are not, which is a binary outcome per run and only becomes informative in aggregate.

The consequence is that everything must be measured as a rate over repeated trials, not as a state.

The Four-Layer AEO Measurement Model

Each layer is noisier and more distant from revenue as you go up, and more actionable as you come down. Track all four, weight the lower ones.

LayerWhat it measuresSignal qualityEffort
1. PresenceMention rate across a fixed prompt setNoisy, directionalManual, monthly
2. ReferralSessions arriving from AI sourcesReliable but undercountedAnalytics
3. DemandBranded search volumeReliable, laggingSearch Console
4. OutcomeQualified inquiriesThe only one that paysCRM

Layer 1: presence

The direct measure, and the least trustworthy in isolation.

Build a fixed set of prompts, run them on a schedule, and record whether you were mentioned. The number that matters is the rate: mentioned in 6 of 20 prompts this month against 4 of 20 last month.

Single results tell you nothing. A rate across twenty prompts, tracked monthly, tells you something.

Building a prompt set that is worth anything

This is the part that decides whether Layer 1 is useful or theatre.

Write fifteen to twenty prompts a real buyer would type. Not "best SEO agency," which nobody asks conversationally. More like "how do I find someone to fix my law firm's website" or "what should I pay for monthly SEO for a small business."

Cover the buying journey. Some problem-stage, some solution-stage, some vendor-stage. You will likely appear in different ones at different times, and which ones is informative.

Include your competitors' territory. If a prompt returns three competitors and never you, that is the most useful data point on the sheet.

Freeze the set. The moment you edit prompts, your time series breaks. Add new ones as a separate cohort rather than changing existing ones.

Record the whole answer, not just yes or no. Who else appeared, and in what terms, is often more useful than your own presence.

Running it without deceiving yourself

Use a clean session each time, logged out where possible, so you are not measuring your own history.

Run the same set against the two or three assistants your buyers actually use. Do not average across them, since they behave differently. Track separately.

Monthly is the right cadence for most businesses. Weekly produces noise you will over-interpret. Quarterly loses the thread.

Expect the number to bounce even when nothing has changed. Two consecutive months of movement in the same direction is weak evidence. Four is worth acting on.

Layer 2: referral

More reliable, and systematically undercounted.

Some assistants pass a referrer when a user clicks through, so sessions from AI sources are identifiable in analytics. Set up a segment or channel grouping for them and watch the trend.

The undercount is structural and worth understanding. Much AI visibility produces no click at all, because the answer was the point. And plenty of downstream visits arrive as direct traffic later, after someone read about you in an answer and searched your name a day afterwards. So treat referral as a floor, not a total.

Layer 3: demand

The most underrated signal in this entire area.

If off-site presence and AI mentions are growing, people begin searching your brand name. Branded search volume in Search Console is reliable, free, and lags the cause by weeks rather than quarters.

It is also hard to fake, which makes it a good corrective when Layer 1 is bouncing around and you cannot tell whether anything real is happening.

Segment branded from non-branded queries and chart branded impressions and clicks monthly. A steady rise is the clearest evidence that awareness work is landing.

Layer 4: outcome

The layer everything else exists to serve.

Ask new inquiries how they found you, and record the answer in a field you can report on. It is imperfect, people misremember, and attribution in a multi-touch world is genuinely hard. It is still worth more than the three layers above it combined, because it is the only one denominated in money.

The specific thing to watch for is a change in inquiry quality. Several practitioners have observed that leads arriving via AI-mediated discovery tend to arrive better informed and further along, presumably because the assistant did the education. If your close rate rises while volume stays flat, that is worth knowing and no dashboard will show it to you.

How to Measure AI Search Visibility Without Fooling Yourself

What about paid AI visibility tools

They exist, they are proliferating, and they mostly automate Layer 1.

That has real value once your prompt set grows or you want more runs per prompt than you can do by hand. What they cannot do is fix the underlying problem, which is that the thing being measured is inherently variable. A tool reporting a precise visibility score is applying a confident number to a noisy process, and the precision is presentational.

My recommendation is to start manual. Twenty prompts in a spreadsheet, once a month, takes under an hour. Do that for a quarter, learn what normal variation looks like for your business, and then buy tooling if the manual version has become the constraint. Buying first means you will not know how to read the output.

What not to measure

A single query result. Sample size of one against a stochastic system.

Whether you appear for your own brand name. You almost certainly do, and it tells you nothing about discovery.

Composite visibility scores you cannot decompose. If you cannot explain what moved the number, you cannot act on it.

Anything weekly. The variance will exceed the signal and you will chase noise.

Screenshots of good answers. These circulate constantly and they are the least reliable evidence in the field. A screenshot shows one sampled output from one session, chosen because it was favourable, which is selection bias with a picture attached. If a provider sends you one as proof that their work is succeeding, ask what the mention rate across the full prompt set was that month. The honest version of that evidence is a number with a denominator, not an image.

Your own repeated checking. Running the same query at your desk five times in a week feels like measurement and functions as reassurance-seeking. It also pollutes your own results if you are logged in, since context and history influence what you see. Leave the measurement to the scheduled run.

What the sheet looks like after six months

The abstract description of a measurement model is less useful than knowing what you are actually going to be staring at, so here is the shape of it.

The sheet has one row per prompt and one column per month, with the cell recording whether you were mentioned and, ideally, who else was. Twenty rows, six columns. At the bottom, a single summary row: mentions out of twenty, per month.

What you will see in that summary row is not a clean upward line. It is more likely to read something like 1, 0, 2, 1, 3, 2. Six months of genuine effort producing a series that could plausibly be random noise. This is the moment where most people either abandon the measurement or start telling themselves a story about the good months.

The correct reading is that the trend is probably mildly positive and you cannot yet distinguish it from chance, which is an honest and useful thing to know. It is also why the other three layers exist. If branded search impressions have risen steadily across the same six months, and referral sessions from AI sources have gone from zero to a small but consistent trickle, then the noisy presence data is corroborated by two cleaner signals and you can believe the direction.

The column that turns out to be most valuable is the one recording who else appeared. After six months you will have a list of the competitors the assistants consistently name, and that list is usually shorter than you expect and remarkably stable. Those are the firms the models already know. Studying what is true of them, where they are mentioned, what they have published, which communities they participate in, is more actionable than any amount of analysis of your own site.

You will also notice that certain prompts never return anyone useful, and others return the same three names every time. Prompts in the first category are probably badly written or genuinely have no consensus answer. Prompts in the second are the contested ground and deserve the most attention.

One habit worth building: write a one-line note next to each month recording what you actually did that month. Restructured six pages. Got listed in two directories. Published the benchmark piece. Six months later, that note column is the only thing that lets you connect cause to effect, and without it you have a time series with no explanatory variables and no ability to learn anything from it.

The whole apparatus takes under an hour a month once it is built. The discipline is not the effort, it is resisting the urge to conclude anything from the first three columns.

Frequently asked questions

How many prompts do I need for a reliable rate?

Fifteen to twenty is a workable minimum for a small business. Below about ten, a single prompt changing state swings the rate too much to interpret. If you can sustain forty, the number gets meaningfully steadier.

Should I run each prompt more than once?

Ideally yes, three times per prompt, which is how you see the variability directly rather than assuming it away. Most small businesses will not sustain that, and one run per prompt across a larger set is a reasonable trade.

Why does my competitor appear and I do not?

Most often because the model already knows them, which is a corroboration and presence problem rather than an on-page one. It usually means the work is off your site.

Can I track AI Overviews specifically?

Partially. Some third-party tools report AI Overview presence for tracked keywords, with varying reliability. Search Console does not separate AI Overview impressions from ordinary ones, so you cannot isolate the effect cleanly in first-party data.

What does good look like after six months?

Honestly, for a small business starting from no presence: a presence rate that has moved from near zero to something detectable, branded search rising, and a handful of inquiries mentioning AI. Anyone promising a specific visibility percentage is quoting a number they cannot control.

The discipline here is resisting the confident conclusion. A good month is not proof and a bad month is not failure. Only the trend across several months, across several layers, means anything.

If you want a prompt set built for your business and a baseline run, email adphconsulting@gmail.com. ADPH is a US-facing consultancy delivering its work from the Philippines.