From Ritu Raj | Product & Market Analysis
Build an AI Visibility Dashboard in a Day, Before You Buy One
On this page
Three leading models were asked the same 250 category questions, five times each. They agreed on the single top recommended brand only 41.6% of the time. That number decides how an AI visibility dashboard has to be built. One query against one assistant is not a measurement, it is an anecdote, and most brand monitoring stops there.
Key takeaways
- Cross-model agreement on the top brand was 41.6%. Across 3,750 responses covering 50 brands and five industries, a leading position on one model did not reliably hold on another. Tracking one assistant measures one assistant.
- Two of the four layers are already free inside tools you own. Google Analytics 4 now carries a native AI Assistant channel and Search Console carries a generative AI impressions report. Both have documented gaps you patch by hand.
- Crawling is the leading indicator, referrals are the lagging one. Cloudflare attributed 80% of AI crawler requests to training and 18% to search over the 12 months to July 2025. Your access log sees movement long before analytics does.
- The commercial tools are worth buying second, not first. Profound lists its Growth tier at $399 a month for 100 prompts across three engines. A 25-prompt daily panel on one engine costs $18.75 a month in published request fees.
What an AI visibility dashboard actually measures
An AI visibility dashboard is not one chart. It is four separate measurements that people collapse into a single word, and the collapse is where most builds go wrong.
Four events, not one metric
A crawler fetches your page. A model names your brand in an answer. That answer carries a link. Someone clicks the link. Those are four events with four different data sources, and each one moves while the others sit still.
Treat them as one number and your dashboard reports a rise when the only thing that changed was a crawler schedule. Separate them and you can say which part moved.
Fix the vocabulary before you build anything. A mention is your brand named in an answer. A citation is a link inside that answer. A referral is a session that arrived from an assistant. A crawl is a request from a platform's bot. Four words, four columns, no overlap.
Visibility is not traffic, and the gap is measurable
Similarweb's panel put the share of AI responses carrying a citation at 1.6% in June 2025 and 6.8% in May 2026. The rate more than quadrupled and it is still under one response in fourteen.
That single ratio explains why a mentions dashboard and a referrals dashboard disagree. Most answers name brands without linking to them. A brand can be recommended constantly and see almost nothing land in analytics.
The industry spread matters more than the average. Similarweb put travel and hospitality near 23% and professional services under 4%. If you sell B2B services, your ceiling on referral volume is set before you write a word, which is the case for measuring mentions directly rather than inferring them from traffic you can attribute.
| Layer | Question it answers | Source | Cost to stand up |
|---|---|---|---|
| Crawl | Are AI platforms reading this page at all? | Your own access or edge logs | Free, already collected. |
| Mention | When a buyer asks, does the model say our name? | A prompt panel you run yourself | Roughly $19 a month per engine. |
| Citation | Which of our pages get linked or surfaced? | Search Console generative AI report | Free where available. |
| Referral | Who arrived, and did they do anything? | GA4 AI Assistant channel, patched | Free, one hour of setup. |
Costs are incremental, assuming you already run analytics and keep server logs. The mention figure is derived later from published API prices.
Layer 1: the traffic you can already see
Start with the two layers you already pay for. Together they take about an hour, and they set the baseline everything else gets compared against.
GA4 names five assistants and excludes two surfaces
Google's Default Channel Group documentation now defines an AI Assistant channel as the route by which users arrive from sources like ChatGPT, Gemini, Deepseek, Copilot or Grok. Assignment happens when the medium exactly matches ai-assistant, with the campaign stamped as (ai-assistant).
Read the exclusions in the same document, because they are the important part. The channel does not cover Google's own AI Overviews or AI Mode. Those stay inside organic search, so the largest generative surface on the web is absent from the channel named after generative assistants.
Two patches, both worth doing on the first morning. Build a regex custom channel group that catches the assistants outside Google's named list, then write down the date your property first recorded the channel. Processed sessions are not reclassified, so your history starts on that date.
Search Console shows impressions and no clicks
Google launched generative AI performance reports in Search Console on 3 June 2026, covering AI Overviews, AI Mode and generative features in Discover. The reports carry five dimensions: impressions, pages, countries, devices and dates.
They do not carry clicks, click-through rate, average position or queries. Google said it would introduce additional metrics over time and gave no date. The rollout began with a subset of UK site owners, with global expansion stated as an intention.
Do not wait for access to build the slot. Pull the pages dimension weekly through the Search Console API into the same table as everything else, and the day access arrives you have a citation history rather than a starting line. The relationship between those impressions and classical rankings is examined in the analysis of whether AI citations follow search position, and the wider evidence is weighed in the comparison of what GEO can and cannot claim over SEO.
Layer 2: the prompt panel nobody gives you free
This is the layer no vendor hands over without a contract, and it is the only one that answers what a founder actually asks. When a buyer asks an assistant to recommend something in your category, how often does your name come back?
Pick 25 prompts a buyer would type
Twenty-five is where the commercial tools start, so it is a defensible panel size. Semrush lists its AI Visibility Toolkit at $99 a month per domain for 25 custom prompts with daily rankings, drawing mentions from ChatGPT, Google AI, Gemini and Perplexity.
Write your prompts from sales calls, not from a keyword tool. The prompts worth tracking are category questions with no brand in them: best tool for a named job, alternatives to a named incumbent, how a specific team should solve a specific problem. Brand-free queries are what the research uses, and they are what a buyer at the top of the funnel actually types.
Then freeze the wording. A panel whose prompts get edited mid-quarter produces a trend line that measures your edits. Keep a change log and treat any wording change as the start of a new series.
Run each prompt five times
The strongest published argument for repetition is the June 2026 arXiv preprint on brand category ownership. Its author ran 250 brand-free category queries five times each across three models under what the paper calls a dice-roll stability protocol, producing 3,750 responses across 50 brands and five industries.
Two findings shape the build. Cross-model agreement on the top recommended brand was 41.6%, and concentration inside a category was moderate rather than winner-take-all, with a mean Gini of 0.28 and a 95% confidence interval of 0.16 to 0.41.
State the caveat with the finding. It is a preprint, it is not peer reviewed, and 50 brands across five industries is a small map. The design implication survives anyway: identical prompts return different answers, so a single-shot query cannot characterise anything.
Five repetitions is a floor rather than a target. It is enough to distinguish a brand mentioned in most answers from one mentioned in a minority. It is not enough to call a two-point week-on-week move, and you should resist reporting one.
Three engines is the floor
If a top position on one model does not hold on another, then one engine measures one engine. Profound's Starter tier at $99 a month tracks ChatGPT alone, which is why the Growth tier at $399 exists and covers ChatGPT, Perplexity and Google AI Overviews.
Running the panel yourself makes engine coverage a budget line rather than a plan change. Perplexity publishes sonar at $1 per million input and output tokens plus $5 per 1,000 requests at low search context. Add an engine, add a line to the config, pay the difference.
The reason to run more than one is not completeness. It is that the engines disagree, and the shape of the disagreement is the finding, a pattern also visible in the collapse of overlap between AI citations and classical rankings.
Layer 3: the crawler log you already have
Your access log is the earliest signal in this stack. It is free, first-party, complete, and almost nobody reads it for this purpose.
The user agents worth matching
OpenAI documents three crawlers with distinct jobs. GPTBot fetches content for model training, OAI-SearchBot surfaces sites inside ChatGPT search features, and ChatGPT-User handles fetches triggered by a person in a live conversation. Each has a published IP range file, which is how you verify that a request claiming to be a named bot really is one.
The distinction is the whole point of the layer. A spike in training crawls tells you the next model generation may know about your page. A spike in search-bot or user-triggered fetches tells you people are asking questions your page answers, right now.
Cloudflare's split shows why the layers move independently. Over the 12 months to July 2025 it attributed 80% of AI crawler requests to training, 18% to search and 2% to user actions. Four requests in five are not looking for a reader to send you.
Crawl-to-refer is the ratio to compute on your own data, and Cloudflare publishes the method. Divide HTML requests from a platform's crawler user agents by HTML requests whose referer header carries that platform's hostname. On its own network in July 2025 it measured Anthropic near 38,066 to 1, OpenAI near 1,091 to 1 and Perplexity near 195 to 1.
Your number will differ, and that is fine. The network figure is a benchmark, yours is the diagnostic. A ratio moving the wrong way while mentions rise is the clearest early sign that visibility is decoupling from traffic, and mentions you do not control feed the same problem, as set out in the piece on community mentions.
The one-day build order
The order below is deliberate. Every step writes into the same table, so the last hour is joining rather than rebuilding.
| Hour | Task | Output |
|---|---|---|
| 1 | Define the four columns and create one table: date, layer, platform, page or prompt, value | The schema everything writes into. |
| 2 | Verify the GA4 AI Assistant channel is populating, add a regex custom channel group for assistants outside the named list | Referral layer, live. |
| 3 | Connect the Search Console API, back-fill the pages dimension for every date available | Citation layer, where enabled. |
| 4 | Grep 30 days of access logs for the documented AI user agents, verify a sample against the published IP files | Crawl layer, plus a 30-day baseline. |
| 5 | Write the 25 prompts. Do this from call notes, not from a keyword list | The panel, frozen and version-controlled. |
| 6 | Script the panel: for each prompt, five calls per engine, store the full response text | Mention layer, first run. |
| 7 | Parse responses for your name and three competitors, plus any URL on your domain | Mention rate and citation rate per prompt. |
| 8 | One page with four charts, and a scheduled daily run | The dashboard. |
Store the full response text, not just a boolean. The parsed flag answers this quarter's question and the raw text answers next quarter's, when someone asks what the model actually said about your pricing. Storage is the cheapest part of this build by an order of magnitude.
What this costs against buying it
Here is the arithmetic, using published prices rather than a vendor's comparison table.
Twenty-five prompts, five repetitions, once a day, is 125 requests a day and 3,750 a month per engine. At Perplexity's published $5 per 1,000 requests for low search context, request fees are $18.75 a month. Token charges sit on top at $1 per million in each direction, which on typical answer lengths adds single-digit dollars. That token component is an estimate from published unit prices, not a measured bill.
| Option | Monthly price | Prompts and engines | What you also get |
|---|---|---|---|
| Self-built panel, one engine | About $19 in request fees. | 25 prompts, 5 repetitions, daily | Raw response text you own. |
| Semrush AI Visibility Toolkit | $99 per domain, billed annually. | 25 custom prompts, four named AI sources | Prompt research, competitor analysis, a site audit. |
| Profound Starter | $99, billed yearly. | 50 prompts, ChatGPT only, 1,500 responses | Agent credits and a maintained collector. |
| Profound Growth | $399, billed yearly. | 100 prompts, three engines, 9,000 responses | Perplexity and Google AI Overviews coverage. |
Vendor prices are list prices from each company's own pricing page, consulted 31 August 2026. The self-built figure excludes engineering time, which is the largest real cost in that row.
The honest reading of that table is not that the tools are overpriced. It is that they are priced against a job you have not defined yet. Buy a tool to remove a constraint you have hit, not to answer a question you have not asked.
My position, stated plainly: run the free layers and a one-engine panel for a month first. If at the end of it you can name the limit you hit, more engines, more prompts, a competitor view you cannot build, then the $99 or $399 is easy money. If you cannot name it, the tool becomes another unused seat in the sprawl you are already trying to cut.
Four numbers that belong on the front page
Most dashboards fail by showing everything. Four numbers, each tied to one layer, is the version people read on a Monday.
Mention rate. The share of your panel responses that name your brand, computed across all repetitions and engines. This is your headline. Report it as a band across the week, not a point, because the underlying process is stochastic.
Citation rate. The share of responses that both name you and link to you. Against a category average near 6.8%, a low number here is normal rather than alarming, and the trend matters more than the level.
Share of voice against three named competitors. Your mentions divided by mentions of the four brands combined. This is the number that survives a board meeting, because it does not depend on knowing the true size of assistant traffic.
Crawl-to-refer, computed on your log. One line per platform, monthly. Rising crawls with flat referrals is a warning worth acting on, and the content dynamics behind it are set out in the analysis of saturation and search visibility.
Where this dashboard is weakest
Four limits, listed because a dashboard whose limits are undocumented gets over-read within a fortnight.
The panel is your hypothesis, not the market's queries. No assistant publishes real user query logs. You are measuring 25 questions you guessed, and the guess carries every bias your sales calls carry. Search volume data for classical queries is a weak proxy, because people phrase things differently in a chat window.
The sample is smaller than it looks. A hundred and twenty-five observations a day sounds substantial. It is 25 prompts observed five times, so the effective independent sample is 25. Week-on-week wobble of several points is expected, and treating it as signal is the most common way this build gets misused.
API answers are not consumer answers. This is the one that should worry you most. The model behind an API call is not the logged-in consumer product with memory, personalisation, a different retrieval stack and, increasingly, advertising. A self-built panel measures a nearby thing rather than the exact thing, and this is where the commercial tools have a real edge.
There is still no click data anywhere. Search Console's generative AI report ships impressions only. GA4 catches the sessions that arrive with a referrer and misses the rest. Nobody, including the paid tools, closes that loop end to end, so any conversion number attached to AI visibility is a modelled estimate rather than a measurement.
The counter-case deserves a hearing too. If your category is one where assistants cite heavily, a commercial tool with nine engine integrations beats a homemade panel on coverage from day one, and the $399 buys back the maintenance of a scraper that breaks whenever a provider changes its output format. That is a real argument, and it applies to a minority of teams who mostly already know who they are.
Frequently asked questions
How do I track brand mentions in ChatGPT?
Run a fixed prompt panel. Pick 20 to 30 questions a buyer would type, run each one five times against every assistant you care about, and record whether your brand appears and whether it is linked. Repetition matters because identical prompts return different answers on every run. A single ad hoc check tells you what one sampled response said on one day, which is not something you can trend.
Does Google Analytics track AI traffic?
Partly. Google's Default Channel Group documentation defines an AI Assistant channel for arrivals from sources such as ChatGPT, Gemini, Deepseek, Copilot and Grok, assigned when the medium matches ai-assistant. The same documentation states the channel excludes Google's AI Overviews and AI Mode, which continue to land in organic search. Assistants outside that named list still fall into referral or direct, so a custom channel group is worth adding.
What is a crawl-to-refer ratio?
It is the number of pages an AI platform crawls from your site for every visitor it sends back. Cloudflare calculates it by dividing HTML requests from a platform's crawler user agents by HTML requests carrying that platform's hostname in the referer header. In July 2025 it measured Anthropic near 38,066 to 1, OpenAI near 1,091 to 1 and Perplexity near 195 to 1 across its network.
How many prompts do I need to track AI visibility?
Start at 25 and hold it there for a quarter. Semrush's entry tier tracks 25 custom prompts daily and Profound's starter tier tracks 50, so that band is where the commercial tools sit. What matters more than the count is repetition and stability. Running 25 prompts five times each gives 125 observations a day, and changing the wording mid-quarter destroys the comparison you were building.
Are AI visibility tools worth the money?
Buy one once you have run the free layers for a month and hit a specific limit. Published list prices are $99 a month for Semrush's toolkit on one domain and $399 for Profound's Growth tier across three engines. Both buy engine coverage, competitor panels and a maintained collector. Neither tells you anything your own panel cannot, until you need more engines or more prompts than you want to run yourself.
Can I see which queries triggered an AI Overview citation?
No. Google's generative AI performance reports in Search Console launched on 3 June 2026 with impressions, pages, countries, devices and dates. Click data, click-through rate, average position and query breakdowns are all absent, and Google has said further metrics will follow without giving a date. You can see which of your pages were cited, but not what was asked to surface them.
Where to start this week
Block a Friday afternoon and a Monday morning. They do different jobs and splitting them is the point.
On Friday, do the free layers. Confirm the AI Assistant channel is populating in GA4, note the date it started, and grep 30 days of access logs for GPTBot, OAI-SearchBot and ChatGPT-User. That is one afternoon and it gives you a crawl baseline before you have spent anything.
On Monday, write the 25 prompts with someone who takes sales calls in the room. That conversation is the part of this build that cannot be automated, and it decides whether the dashboard measures your market or your assumptions. Script the panel after the prompts exist, never before.
Related analysis
Once the dashboard runs, the next question is what the traffic is worth. Start with what AI referral sessions convert at, then read the evidence behind GEO claims.
References
- D. Zatuchin, Who Owns the AI Recommendation? arXiv preprint 2606.23057, June 2026. Used for the 41.6% cross-model agreement figure, the Gini coefficient and the repetition protocol. Preprint, not peer reviewed.
- Similarweb, AI search stats, June 2025 to May 2026. Used for citation rates overall and by industry.
- Google Search Central, Search Generative AI performance reports, June 2026, with rollout and metric detail as reported by Search Engine Land, 3 June 2026.
- Google Analytics Help, Default channel group, consulted 31 August 2026. Used for the AI Assistant channel definition and its stated exclusions.
- Cloudflare, The crawl-to-click gap, 2025. Used for the purpose split and the July 2025 crawl-to-refer ratios.
- Cloudflare, The crawl before the fall of referrals, 2025. Used for the calculation method.
- OpenAI, Bots documentation, consulted 31 August 2026. Used for the crawler definitions and IP range files.
- Vendor list prices, consulted 31 August 2026: Profound, Semrush and Perplexity.
The weakest part of this source base is the panel research. The 41.6% figure comes from a single unrefereed preprint covering 50 brands, with no independent replication. Cloudflare's figures describe its own network, and the platform ratios moved by an order of magnitude within seven months.
Related reading