Teams that need cleaner marketing decisions tied to inquiries, revenue, and useful visibility signals.
The measurement model is only as good as the prompts you feed it. Here is how to write a set that produces a usable signal instead of a spreadsheet of noise.
Write 15 to 20 prompts a real buyer would type, spread across problem, solution and vendor stages, including territory your competitors own. Freeze the wording, run them monthly against each assistant separately, and record who else appeared. The rate across the set is the signal; any single result is noise.
Tracking AI visibility fails in a specific way. Someone writes down ten prompts, runs them once, gets a mixed result, and concludes the tracking does not work.
Usually the model is fine and the prompts are the problem. Most first attempts are written by the business about itself, which is not how buyers search, so the set measures something nobody was ever going to ask.
This is how to build one that gives you a real signal. It takes about an hour, and it only has to be done once.
Write prompts a buyer would type, not a marketer
The most common failure. A set full of "best SEO agency" and "top digital marketing consultant" measures a query almost nobody types conversationally.
People talk to assistants in full sentences, with context and a problem attached. "Our law firm's website gets traffic but no calls, what should we look at" is a real prompt. "Best law firm SEO" is a keyword pasted into the wrong tool.
The test: read the prompt aloud. If it sounds like a search box rather than a question you would ask a knowledgeable friend, rewrite it.
Cover all three buying stages
Sets that only ask vendor questions miss most of the picture, because you can be present at one stage and invisible at another, and which one tells you what to fix.
Problem stage. The buyer knows something is wrong and not what. "Why has our website traffic dropped this year." Appearing here means your educational content is doing its job.
Solution stage. They know the category and are working out what it involves. "How much should a small business pay for SEO." This is where being cited is most valuable, because the buyer is forming their criteria.
Vendor stage. They are ready to shortlist. "Who should I hire to fix my nonprofit's website." Rarest, hardest, and closest to revenue.
Roughly a third each is a reasonable starting split.
Include territory you do not own
The instinct is to write prompts you might win. Do the opposite for part of the set.
Include several where you fully expect competitors to appear. Those are the most informative rows on the sheet, because the column recording who else was mentioned becomes a list of the firms these systems already know. After a few months that list is short, stable, and more actionable than anything about your own presence.
If you never appear and the same three names always do, that is a finding: go and look at what is true of them off-site.
Freeze the wording
Once written, do not edit prompts. Changing the wording breaks the time series, and you will not be able to tell whether a change in results came from your work or your edit.
If you want to add prompts, add them as a new batch with its own start date and track them separately. Resist the urge to improve the originals.
Keep the set small enough to sustain
Fifteen to twenty is the practical range. Below about ten, one prompt flipping state swings the rate too much to read. Above thirty, most people stop running it by month three, and an abandoned set is worth nothing.
The discipline of running the same modest set every month beats a comprehensive one you do twice.
How to run it
Use a clean session. Logged out where possible, so you are measuring the system rather than your own history.
Run each assistant separately and keep the columns apart. They behave differently and averaging them destroys the most useful signal, which is divergence between them.
Record more than yes or no. Whether you were mentioned, who else was, and roughly in what terms. The competitor column is the one you will actually use.
Monthly. Weekly produces variance you will over-read. Quarterly loses the thread.
Note what you did. One line per month recording what shipped. Without it you have a time series with no explanatory variable and no way to learn from it.
What the results will look like
Set expectations now, because the first few months are discouraging.
A realistic six-month summary row reads something like 1, 0, 2, 1, 3, 2 out of twenty. That is not failure and it is not proof of success. It is a mildly positive trend indistinguishable from chance at this sample size.
That is exactly why the other measurement layers exist. If branded search impressions are rising over the same period and AI referral sessions have gone from zero to a trickle, the noisy presence data is corroborated and you can trust the direction.
Two consecutive months moving the same way is weak evidence. Four is worth acting on. One good month is worth nothing at all, and screenshots of favourable answers are the least reliable evidence in this field.
A starting template
Adapt the shape rather than the wording:
Three problem-stage prompts describing symptoms your buyers actually report. Three solution-stage prompts about what the fix involves and what it costs. Three vendor-stage prompts about who to hire, phrased conversationally. Three naming your niche and location specifically. Four covering territory where competitors are strong. Two or three about adjacent problems you also solve.
That is eighteen, spans the journey, and takes about twenty minutes a month to run.
Frequently asked questions
How many prompts do I actually need?
Fifteen to twenty for a small business. Fewer and single flips distort the rate; more and most people stop running it.
Should I run each prompt more than once?
Ideally three times, which shows you the variability directly rather than assuming it away. Most people will not sustain that, and one run across a larger set is a reasonable trade.
Can I automate this?
Tools exist and mostly automate this layer. Start manual for a quarter so you learn what normal variation looks like, then buy tooling if the manual version has become the bottleneck. Buying first means you will not know how to read the output.
What if I never appear at all?
Common at the start, and it usually points off-site rather than on-page. If competitors appear consistently for the same prompts, the difference is that models already know them.
The prompt set is the cheapest piece of measurement infrastructure you can build, and the one that determines whether everything above it is meaningful. An hour to write, twenty minutes a month to run.
For the full four-layer model this feeds into, read how to measure AI search visibility without fooling yourself. If you want a set built for your business and a baseline run, email allan@adph-consulting.com.
