How to Measure AI Visibility (2026 Guide)
Measure AI visibility by asking a fixed set of buyer questions on a fixed schedule and scoring how often you are mentioned, cited and recommended.

By the Alphaa team — we build an AI agent that checks what ChatGPT, Gemini, Claude and Perplexity say about local businesses, so these are the metrics we run on our own customers every week.
Short answer: To measure AI visibility, write 20 to 40 buyer questions, ask them on every engine that matters on a fixed schedule, and score four things: how often you are mentioned, how often your site is cited, where you sit in the recommended list, and how you are described. Treat any single answer as noise.
This is a different measurement problem from SEO, and the instinct to reuse rank tracking is where most teams go wrong. There is no stable ranked list to monitor, the same question returns different answers an hour apart, and the metric that pays — being recommended by name — is not in any analytics tool you already own. Here is a method that survives contact with that reality.
What does AI visibility actually mean?
AI visibility is how often an AI assistant names, cites or recommends you when someone asks a question you should win. It is a property of the answer, not of a results page. Three distinct things hide inside the phrase, and conflating them is why vendor numbers disagree so wildly.
- Mention. The answer says your business name, with or without a link. This is the loosest and most common form of visibility.
- Citation. The answer links to a page on your site as a source. This is what sends traffic and what shows up in your logs.
- Recommendation. The answer actually puts you forward as a choice, usually in a short list. This is the one that produces calls.
A business can be mentioned constantly and recommended never — described as "also in the area" while three competitors get the shortlist. If you only track mentions you will not see that, so measure all three separately.
Which four metrics are worth tracking?
Four metrics cover almost every useful question: mention rate, citation rate, recommendation rank and sentiment. Compute each one across your whole prompt set rather than per answer.
| Metric | How to compute it | What a change tells you |
|---|---|---|
| Mention rate | Prompts where your name appears ÷ all prompts asked | Whether the engines know you exist in this category at all |
| Citation rate | Prompts where your domain is cited ÷ all prompts asked | Whether your own pages, rather than directories, are the evidence |
| Recommendation rank | Your average position in the named list, counting misses as unranked | Whether you are a candidate or just context |
| Sentiment and attributes | The adjectives and facts attached to you across answers | Whether the engines have the right story, price and specialisms |
Share of voice — your mentions as a percentage of all brand mentions across the set — is a useful fifth if you have a tight competitor list. We define that metric on its own page: what AI citation share is.
How do you build a prompt set that is worth measuring?
Build the prompt set from how customers actually ask, not from your keyword list, and then freeze it. Twenty to forty prompts is enough for a single-location business; a multi-location or multi-service brand needs a set per market. Work through it in this order.
- Start with the money questions. "Best emergency plumber in Leeds", "who repairs Bosch dishwashers near me", "cheapest MOT in Bristol". Full sentences, the way someone types into a chat box, not two-word keywords.
- Add the qualifying questions. Questions asked just before the decision: open on Sundays, do they take Medicaid, how much does it usually cost, do they handle insurance claims.
- Add the comparison questions. "X vs Y" and "alternatives to X" for each real competitor. These answers are where you find out how you are positioned.
- Add a brand-check block. "What do you know about your business?", "is your business any good?". This is where you catch wrong hours, a closed location or a merged entity. Fixing those is covered in how to fix wrong AI information about your business.
- Freeze the wording and the schedule. Changing a prompt resets its history. Re-ask weekly, on the same day, with no personalisation, no memory and no prior chat context.
One discipline matters more than the rest: sample each prompt several times per run. Answers are generated fresh each time and vary even with identical input, which we explain in why AI answers change every time. Three samples per prompt turns an anecdote into a rate.
What does Google Search Console now show about AI visibility?
Search Console now has a dedicated generative AI performance report, and it shows impressions only. Google announced the reports in June 2026 and noted on the same page that, as of 31 August 2026, the insights had rolled out to all websites worldwide. The report gives you impressions inside generative AI features on Search — AI Overviews and AI Mode — plus generative AI features in Discover.
What you get, per Google's documentation, is impressions broken down by page, country, device, date and search type. What you do not get is clicks, click-through rate, position or query data, and the usual Search Console caveats still apply, including the 1,000-row limit and preliminary recent data. Google has said it expects to add metrics over time.
So treat it as one input, not the scoreboard. It is free, it is first-party, and it is the only place Google tells you anything about AI Overviews and AI Mode. It is also silent on ChatGPT, Claude and Perplexity, which is most of the question for a lot of businesses.
How do you measure the traffic side of AI visibility?
Measure AI traffic by isolating assistant hostnames as referral sources, and expect it to undercount. In GA4 there is no built-in AI channel, so clicks from a citation inside ChatGPT land in Referral alongside newsletters and forums; filtering session source for chatgpt.com, perplexity.ai, gemini.google.com and claude.ai pulls them out, and promoting that into a custom channel group makes it permanent. The GA4 dimension reference lists the fields involved, and we walk through the setup in how to track AI traffic in Google Analytics.
Your server logs are the other half, and they answer a different question: is anyone fetching your pages on behalf of these engines at all? Both OpenAI and Perplexity publish their crawler user agents. If you see no retrieval fetches for a page you want quoted, no amount of prompt tracking will fix it — that is a crawling problem, and it is the first thing to rule out.
How often should you measure, and what counts as a real change?
Weekly measurement with a monthly read is the right rhythm for almost everyone, and a real change is one that holds for three consecutive runs. Answers drift day to day for reasons that have nothing to do with you — a model update, a different retrieval set, a competitor's new page — so a single week's improvement is not evidence.
Two things do deserve a scheduled review rather than continuous watching. Refresh your competitor list quarterly, because the names the engines put next to yours change as new pages and new entrants get indexed, and a stale list makes your share numbers quietly wrong. And re-read your prompt set against seasonality twice a year: a roofer's winter questions are not the summer ones, and measuring the wrong half of the year is a common reason a flat chart hides real movement.
Pair the series with a change log. Every time you publish a page, fix a listing, add schema or collect a batch of reviews, write the date next to it. Without that, you will have a chart and no idea which of five things moved it. And keep expectations calibrated: engines need to recrawl and then choose your page during retrieval, so weeks is normal — see how long AEO takes.
How do you turn the numbers into a priority list?
Rank the gaps by how commercial the prompt is and how cheap the fix is, and work the cheap commercial ones first. The same four metrics point at four different kinds of problem, so read them as a diagnosis rather than a scoreboard.
- Wrong facts about you. Always first. A closed location, old hours or a merged entity costs you answers across the whole set, and the fix is listing and schema work, not content.
- Zero mentions on a money prompt. The engines do not associate you with that service or that place. This needs a page that is unmistakably about it, plus corroboration elsewhere.
- Mentioned but never cited. The engines know you through directories rather than your own site. The fix is a page that answers the prompt better than the directory does.
- Cited but never recommended. You read as reference material, not as a candidate. This usually means missing the comparison and "best X for Y" questions, which we cover in comparison pages for AI search.
- Described wrongly. Right name, wrong story — too expensive, wrong specialism, wrong area served. Fix it where the engines read it: your own pages, then your listings and reviews.
What can no AI visibility measurement tell you?
No measurement can prove why an engine changed its answer, and none can promise it will keep the new one. You are observing a system you cannot inspect: nobody outside the labs sees the retrieval set, the ranking, or the model revision behind today's answer. Google's own guidance on AI features is explicit that there is no special markup that buys inclusion, and the same is true of every other engine.
That is an argument for honest measurement, not for giving up on it. Knowing that ChatGPT names three competitors and not you for your best question is genuinely actionable, even without knowing why. If you want the baseline without building the harness yourself, the free 60-second check at our scan page asks ChatGPT, Gemini, Claude and Perplexity about your business and shows you what came back.
What else do people ask about measuring AI visibility?
Can I just use rank tracking for AI visibility?
No, because there is no stable ranked list to track. AI answers are generated per request and vary between runs, so the unit of measurement is a rate across many samples of a prompt, not a position.
How many prompts do I need to track?
Twenty to forty prompts covers a single-location business, split across money questions, qualifying questions, comparisons and brand checks. Add a set per market or per service line if you operate in several, as covered in multi-location AI visibility.
Does Google Search Console show ChatGPT data?
No. The generative AI performance report covers Google's own surfaces — AI Overviews, AI Mode and generative features in Discover — and reports impressions only. ChatGPT, Claude and Perplexity are not included anywhere in Search Console.
Why does my AI visibility score differ between tools?
Because every vendor picks its own prompts, sample count, engines and definition of a mention. A score is only comparable with itself over time, so pick one method and keep it rather than reconciling two dashboards.
What is a good mention rate to aim for?
There is no universal benchmark, and any vendor quoting one is guessing. The useful target is your own baseline plus a direction of travel, measured against the same prompt set and the same competitors over several months.
Should I measure before or after I start fixing things?
Before, always. A baseline taken after you have changed five things is worthless for attribution, and the first measurement run usually surfaces the wrong-information problems that are cheapest to fix.
Sources
- Generative AI performance report (Search) — Google Search Console Help
- Introducing Search Generative AI performance reports in Search Console — Google Search Central Blog
- AI features and your website — Google Search Central
- Overview of OpenAI Crawlers — OpenAI
- Perplexity Crawlers — Perplexity
- Analytics dimensions and metrics — Google Help