Why does AI name a different business every time you ask?
Because the answer is generated, not looked up. Ask the same question four times and you often get four overlapping but different lists. That variation is not a bug to be explained away -- it is the reason a single answer tells you almost nothing.
How much does it actually vary?
A lot at the bottom, much less at the top. In our own scans the businesses named most often come back consistently, while the tail of one-off mentions turns over almost completely between runs.
That pattern is the useful part. The leaders are a stable fact about the market; the tail is noise. Which means a business that appeared once in twelve answers has learned something quite different from one that appeared nine times.
So what should be measured instead?
A rate, not a result. Ask the same question many times and count how often each business is named. That share of answers is stable enough to compare, to track over time, and to act on -- and it comes with an honest margin of error.
| One answer | Twelve answers | |
|---|---|---|
| What you learn | Who was named that time | How often you are named, roughly |
| Repeatable? | No | Yes, within a stated range |
| Comparable to rivals? | Not meaningfully | Yes -- same questions, same day |
| Can it show change? | No | Yes, if the change is bigger than the noise |
Why do you show error bars on the number?
Because twelve answers cannot tell you much precisely, and rounding that into a confident-looking score would be dishonest. A rate of 25% from twelve asks might really be anywhere from about 7% to 52%.
We would rather show that range and explain that a larger scan narrows it than quote a single figure chosen to look authoritative. It is also self-protective: a number presented without its uncertainty will eventually move on its own, and then the whole report looks unreliable.
How would anyone prove the work made a difference?
Ask the same questions again, with enough repeats to detect the size of change you are aiming at, and compare the two rates with the uncertainty carried through. If the interval on the difference includes zero, nothing has been shown yet.
This is the part most reporting skips. Going from 25% to 33% sounds like progress and, on a small scan, is well within what random variation produces on its own. Detecting a genuine twenty-point shift usually takes something on the order of forty to sixty asks per question set.
Common questions
Does that mean AI visibility cannot be measured?
It can, and it has to be measured as a rate rather than an outcome. What cannot be done honestly is claiming a result from one answer, or from a small scan that moved a few points.
Why not just ask once and check?
By all means ask once -- it is a good way to see the problem. Just do not make a decision on it. The list changes.
Sources
Where a figure here comes from one study we say so, because one study is one study. Where something rests on practice rather than published work, we say that instead.
If this answered one question, these answer the next ones.
How does an AI assistant decide which businesses to name?
It searches, reads a handful of pages, and writes an answer from what those pages say. So the question is not 'how do I rank' but 'which pages does it read, and is my business on them'. Usually those pages are directories, reviews and round-ups rather than your own site.
How can I check whether AI recommends my business?
Open ChatGPT, ask the questions your customers would ask, and write down who gets named. Do it several times, because the list changes. Ten minutes gets you an honest picture of whether this is a problem worth spending money on.
What are the ways to get this done, and what does each cost?
Four realistic options: do it yourself, buy a monitoring tool, ask your existing SEO agency, or hire someone who does this specifically. They differ mainly in who does the work and whether anybody measures whether it helped.