This post explains what an AI visibility score can and cannot prove, so consultants, lawyers, and other expertise-led business owners can tell a real measurement from a marketing number before paying for one.
On 13 August someone put a question to a room of SEO professionals on Reddit. Is anyone actually paying three hundred dollars a month for AI visibility tracking. The thread ran to fifty-one comments and the answer was mostly no.
The word people kept using was snake oil.
One commenter described a company that fired its SEO, bought a forty-five hundred dollar a month AI visibility platform, and lost ground on keywords, organic traffic and AI presence inside a year. Another described building a client report by screenshotting a vendor dashboard and feeding the screenshots into Claude, because the tool had no usable export.
Five days later a small business owner made the same point without the anger. An AI visibility score is not evidence yet.
I sell an AI visibility report. The score is the least interesting page in it.
I mean that. When I run a check on a business, the number is a summary. The finding is the gap between what an AI model already says about that business before it has seen the website, and what the website actually says about it. That gap is specific. It points at something the owner can go and fix this week. The score is a way of writing the gap down, and it is the part I would delete first if I had to lose a page.
So I have some sympathy for the snake oil crowd, and some for the people they are shouting at.
Here is what most tools in this category are doing. The tool writes a set of prompts it believes your buyers might use. It runs them a number of times. It records whether you appeared. Then it averages that into a percentage and puts the percentage on a dashboard.
That is an approximation, and the people building these tools mostly say so. The large models do not publish citation data, so there is no ground truth to check the estimate against.
The estimate also moves for reasons that have nothing to do with your business. Change one word in the prompt and the answer changes. Run it through the API instead of the chat window and the answer changes. The user's location changes it. Whether they are signed in changes it. Whether they are on a free account changes it.
A score is not evidence. It is one answer, on one day, to a question you picked.
None of that makes the underlying worry wrong. Business owners are finding out, one at a time, that someone asked ChatGPT for a business like theirs and got a competitor instead. That is real, and it is happening to real revenue.
So the question is not whether to care. The question is what counts as evidence.
Start with a fixed set of questions, written down. Not questions a tool generated for you. The actual sentences a buyer would type when they do not yet know your name. Twenty is plenty. They stay the same every time, because a question set that quietly changes between reports is how a decline gets reported as growth.
Run them on dates you set in advance, more than once. A single run tells you what happened that afternoon.
Run them across more than one model. The models disagree with each other constantly. A report built on one of them is a report about one company's product.
Record what you can actually see. Some tools and some chat modes show their sources, and where you get them, write down which page of yours was used. Where you do not get them, write that down too, instead of treating the number underneath as precise. Record who showed up instead of you. And when you drop out of an answer, find out what replaced you before you file it as lost ground, because a lot of the time the answer was reorganised rather than reassigned.
Then connect it to something that happened. A booked call. A form fill. A person who said they found you through ChatGPT. Those are the numbers I report back to clients when the work is done, and they are the only ones that survive a year. More mentions can mean better visibility without meaning better business, and if nobody is checking which one you got, the dashboard is decoration.
That method is free and you should run it. It will tell you whether you are showing up.
It will not tell you why. For that you need the other half, which is what the models already believe about you before they go anywhere near your site, and how that compares to a competitor being asked the same questions. That comparison is the work I do inside an AI Visibility Check. It is the part a dashboard cannot reach, because a dashboard only ever sees the answer, never the picture the model started with.
If someone is quoting you a monthly price, ask three things before you sign.
Ask which questions they are running, and whether you get to see the list. If the list is proprietary, you are buying a number you cannot audit.
Ask how many models they run, and whether they use the API or the chat interface. If it is one model through an API, they are measuring something your buyers will never see.
Ask what they will show you when the score moves. If the answer is a bigger number, keep your money. If the answer names which page got cited, which competitor took your place, and what changed on the site in between, that is a report.
You do not need a subscription to start. You need twenty questions, a calendar reminder, and the discipline to write down what you actually saw.
If you would rather see where you stand before you build any of that, the AI Reality Check is free and takes a few minutes.