What Is a Good "AI Visibility Score," and Should You Trust the One Your AEO Tool Gives You?

There is no good AI visibility score, because there is no such thing as an AI visibility score. Not one that OpenAI, Anthropic, Google, or Perplexity publishes, recognizes, or audits. Every number in every AEO dashboard is a vendor construction: a private blend of mention rate, list position, citation rate, sentiment, and share of voice, weighted by rules the vendor picked and does not show you. Trust it as a rough directional read on your own prompt set. Do not put it on a board slide as if it means something outside the tool that produced it.
Why did the industry start scoring AI visibility at all?
Because scores sell. This movie already ran once, when Moz introduced Domain Authority around 2010 as a model of how a site might perform. Google said, repeatedly and on the record, that it does not use Domain Authority. It did not matter. DA became currency. Agencies sold DA packages. Link brokers priced their inventory by DA. A modeled number invented by a vendor turned into the thing the industry optimized for, and the vendor selling the score also sold the fix.
AI visibility scores are following the same arc, faster. A single 0 to 100 figure is easy to sell, easy to screenshot, and easy to move without changing anything about your business. Add fifteen long-tail prompts you already win and the score climbs. Nothing happened. That is Goodhart's Law with a subscription attached.
Why does my AI visibility score change when nothing changed?
Three things move it, and none of them are you.
Large language models are not deterministic, even at temperature zero. Batch-size variance and floating-point behavior in GPU inference mean the same prompt can return different answers on different runs. Brands sitting outside the clear top three are the most volatile, so a mid-pack company can appear, vanish, and reappear across identical asks in the same hour. Most day-over-day score movement is noise wearing a decimal point.
Share of voice is relative. Your score can fall because a competitor gained ground, not because you lost any. A composite hides which of those two happened, and they call for completely different work.
Then there is personalization. Location, memory, account type, and API versus product UI all change what an engine says. A score measured from a fixed datacenter IP with no buyer context describes an answer no real prospect ever saw. This is why we put a short buyer persona in front of every question Clark asks, and a location line for local and nationwide businesses. The question text stays as the buyer would type it. The context around it stops being a robot's.
What should you measure instead of a composite score?
Numbers that trace back to a specific answer you can open and read.
- Mention rate, kept separate from citation rate. Being named in an answer and having your domain cited as a source are different events with different fixes. Answers routinely name eight brands while linking three sources, often a review site or a Reddit thread rather than any vendor. Blend them into one figure and you cannot tell which lever is broken. This is where most tools over-report: if your URL shows up as a source but your name never appears in the answer text, that is a citation, not a mention, and it should be counted as one.
- Average position when you are mentioned. Third in a list is a different business outcome than eighth.
- How many answers were analyzed, and how many companies appeared. Sample size is context. Twenty-five prompts on a scheduled rotation, each one coming back around roughly weekly, is directionally honest. Ten prompts checked by hand once a month is a coin flip.
- The full answer text and its source links. Receipts beat aggregates. Reading which third-party pages the engines lean on tells you where the work actually is, and a lot of that work is off-site.
- AI sessions, leads, and revenue by engine and by page. This is the only tier that survives a CFO question.
Clarity ships no composite score on purpose. Clark records mention rate, average position, answers analyzed, and companies seen, and stores every full answer plus up to twenty deduplicated source links, so any number can be traced back to the text that produced it. We would rather hand you four traceable figures than one confident invention. If you want the competitor side of that, here is how to find who is winning citations in your category.
So what counts as "good"?
Good is your mention rate on a prompt set you did not rig, moving up over a rolling 30-day window, on questions your buyers actually type. The bands people circulate, under 10 percent means invisible and 30 percent and up means strong, only mean anything relative to your own prompt list and your category structure. In a consolidated category the top three brands can hold half the share of voice. In fragmented B2B SaaS, 7 percent can be first place.
The better test is downstream. A modest mention rate built on high-intent comparison questions will out-earn a fat score built on informational trivia, every time. So the number that ends the argument is not visibility at all. It is a prospect who read a Perplexity answer, clicked the post named in it, landed on your pricing page, and closed six weeks later, tied back to that engine and that post. We compared the ways people try to prove that in closed-loop attribution versus dark-traffic estimation versus rank-only tracking.
How do you report this without overselling it?
Two tiers, never substituted for each other. Leading indicators on the left: mention rate, citation rate, prompts won, posts published. Lagging indicators on the right: AI sessions, leads, pipeline, closed revenue. Show both, side by side, with the tiers labeled. The failure mode in a 90-day proof window is presenting a leading indicator as a result, and a composite score is the most tempting way to do it, because it looks like an outcome and costs nothing to produce.
Before you trust any vendor's number, ask which prompts it runs, how often it runs them, and how the pieces are weighted. If the weighting is proprietary, the score is a vibe with a chart around it.
Start with the diagnosis rather than the dashboard. Our free AI visibility check shows you the questions your buyers ask and which competitors get named instead of you. No score. Just the answers, and who is in them.
Written by the Clarity Search AI team.
Is AI recommending you?
See where your brand shows up across ChatGPT, Claude, Gemini, and Perplexity, and win back the customers AI is sending to your competitors.
Check my visibility




