Measuring AI Visibility: What to Track – and What the Management Report Should Look Like
Measuring AI visibility in 2026: why prompts, not keywords, are the unit of measurement, which KPIs belong in quarterly reporting and how to build a reliable measurement setup – with an education-provider example.
AI visibility is not measured with keywords but with prompts: a fixed set of realistic user questions is run regularly against ChatGPT, Google AI Overviews, Perplexity and Gemini, and for each answer you record whether your brand is mentioned, your website is cited and how much share your competitors capture. This produces five metrics - mention rate, citation rate, share of voice, position, sentiment - which, combined with AI referral traffic and leads, form the quarterly report for management.
The trigger for this article is a real search query that reached us: an education provider has to report quarterly to management on how its AI visibility is developing - and doesn't know what to track in the first place: prompts? Brands? Keywords? It's a fair question, because the market is young and the terminology is fuzzy. This article lays out the measurement setup we use ourselves.
Why does AI visibility belong in management reporting in 2026?
Because a growing share of purchase decisions is shaped inside AI answers before your website is ever visited. ChatGPT counts around 900 million weekly users, 68% of US Google searches end without a click according to SparkToro/Similarweb, and in Switzerland the IGEM Digimonitor 2025 reports that 60% of the population already uses AI tools - 79% among 15- to 34-year-olds.
For education providers the shift is especially measurable. In the US, according to an EAB study (February 2026), 46% of prospective students already use AI tools like ChatGPT in their college search - almost twice as many as six months earlier. 18% removed an institution from their shortlist because of AI-generated answers. Among adult learners, UPCEA found that half use AI-powered search at least weekly, and 56% are more likely to trust institutions cited in AI summaries. If you don't appear there, you lose prospects who never show up in any analytics report.
At the same time, McKinsey shows that only 16% of companies systematically measure their AI search performance. That is exactly the opportunity - whoever builds a clean measurement setup now sees shifts before they hit revenue.
Prompts, brands or keywords - what is the right unit of measurement?
The prompt is the unit of measurement. Brand and source are the metrics. Keywords remain an SEO instrument. The confusion arises because all three terms appear in the same report - but on different levels:
| Level | Role in the setup | Example |
|---|---|---|
| Prompt | Unit of measurement: the question posed to the AI | "Which project management further-education course is worth it in Zurich?" |
| Brand | Metric: is our name mentioned in the answer? | Mention, position, sentiment |
| Source | Metric: is our website linked as evidence? | Citation |
| Keyword | SEO instrument, not an AI measurement unit | "project management course" |
People don't type keywords into AI systems - they ask situational questions. A good prompt set therefore mirrors real decision situations of your audience, structured by funnel stage:
- Unbranded (awareness): "How do I become a certified project manager in Switzerland?" - this is where it's decided whether you are considered at all.
- Category (consideration): "Compare the best providers for a CAS in digital marketing in German-speaking Switzerland" - this is where the shortlist is built.
- Branded (decision): "Is [your institute] reputable? What do graduates say?" - this is where the AI audits your reputation.
In practice, 50-150 carefully curated prompts covering personas, regions and product lines are enough. Representativeness beats volume: 30 precise prompts along real decision paths deliver better steering signals than 500 generic ones. What matters is that the set stays stable across quarters - that is the only way comparable trends emerge.
Which metrics belong in the report?
Five platform metrics, measured across the prompt set, plus two business metrics from your own website:
| Metric | What it measures | Why it matters |
|---|---|---|
| Mention rate | Share of answers naming your brand | An AI recommendation - works even without a click |
| Citation rate | Share of answers linking your website as a source | Your content is the evidence base |
| Share of voice | Your mentions relative to competitors | The management metric: market share inside AI answers |
| Position | Where in the answer you appear (first recommendation or footnote) | Order shapes perception |
| Sentiment | How the AI talks about you | Catch false or negative portrayals early |
| AI referral traffic | Sessions from ChatGPT, Perplexity, Gemini, Copilot | The link to website reality |
| Leads / conversions from AI | Inquiries and sign-ups from AI traffic | The business case for management |
Two things get mixed up regularly. First: a mention is not a citation. The AI can recommend your brand without linking your website - and cite your website without recommending you. Both must be tracked separately, because both call for different actions. Second: visibility alone is not a result. A sound report therefore connects the platform level (are we named?) with the business level (does it generate business?).
The data makes clear why the business level matters: according to Semrush, visitors from AI search convert at 4.4 times the rate of traditional organic visitors, and Adobe Analytics has measured since March 2026 that AI traffic in retail converts 42% better than non-AI traffic. Fewer visitors, but far more qualified - exactly what a quarterly report has to make visible.
Why is a single ChatGPT query not a measurement?
Because AI answers are non-deterministic: the same question yields different answers on different days. Profound analyzed over 240 million ChatGPT citations: 40-60% of cited domains change from month to month for identical queries. SE Ranking ran 10,000 queries through Google's AI Mode three times - only 9.2% of cited URLs were identical across all three runs.
For your measurement setup this means:
- Measure repeatedly: Run the same prompt set several times a week, run critical prompts multiple times in a row, and work with majority values.
- Report trends, not snapshots: Only values aggregated over four to six weeks are reliable. The quarterly comparison smooths out the noise.
- Break out models separately: ChatGPT, AI Overviews, Perplexity and Gemini cite different sources - an average across all systems hides where you are winning or losing.
How do you measure AI traffic on your own website?
Since May 2026, GA4 has a dedicated "AI Assistant" channel that automatically groups visits from ChatGPT, Gemini, Claude, Copilot, DeepSeek and Grok. Three gaps remain: Perplexity is missing and still lands in Referral, the classification is not retroactive, and clicks from AI apps without a referrer end up in Direct. A custom channel group with a regex over the known AI domains (chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, etc.) therefore remains best practice - especially since ChatGPT automatically appends utm_source=chatgpt.com to its source links.
In June 2026, Google Search Console introduced dedicated reports for generative AI: impressions in AI Overviews and AI Mode are now visible separately. Clicks, CTR and queries are still missing there - so for impact measurement, the combination of prompt monitoring and GA4 remains essential.
What does a quarterly report to management look like?
One page, five blocks - nothing more is needed:
- Share-of-voice trend across the last four quarters, per AI system, against your three most important competitors.
- Gains and losses: On which prompt topics did we gain, where did we lose - and to whom?
- Source analysis: Which websites does the AI cite most often in our category? (Often directories, comparison portals and trade media - that is where the next action list comes from.)
- Business impact: AI referral sessions, inquiries generated from them, conversion rate compared to the rest of organic traffic.
- Actions: What was implemented, what is planned, what do we expect from it?
The last block is the most important one. Measurement without derived actions is reporting theater. The Princeton GEO study shows that targeted content optimization - statistics, source references, citable structure - can lift visibility in AI answers by up to 40%. We described how that works in practice in From SEO to GEO.
Buy a tool or have the setup managed?
The market for AI visibility tools exploded in 2026 - Adobe acquired Semrush for 1.9 billion dollars, and category pioneer Profound was valued at over a billion. There are now dozens of SaaS monitors at every price point. The problem is not access to data, but what tools don't deliver: a well-designed prompt set for your audience, statistically sound interpretation of fluctuating values, the connection to GA4 and CRM data - and above all the actions that follow from the numbers. A dashboard does not answer the question of why a competitor ranks ahead of you in Perplexity and what to do about it.
We solved this problem for ourselves before solving it for clients: Hierarchy runs its own in-house AI visibility tracking system that continuously measures our brand and our clients' brands across ChatGPT, Google AI Overviews, AI Mode, Perplexity and Gemini - with individually built prompt sets, repeated sampling against answer noise and competitive benchmarks. Clients receive the tracking as a managed setup: we build the prompt set together, measure continuously, deliver the quarterly management report and implement the derived GEO actions directly - as part of our SEO & content service. Measurement and execution from one team, instead of yet another dashboard nobody reads.
Conclusion
The answer to the original question - "prompts, brands or keywords?" - is: prompts as the unit of measurement, brand mentions and citations as the metrics, share of voice as the management KPI, connected to AI traffic and leads from GA4. Measure continuously and repeatedly, report quarterly as a trend. Build the setup properly once and you turn a vague management requirement into a steerable metric - and spot market shifts before they reach your enrollment numbers.
You need to report quarterly on AI visibility and want a measurement setup that lasts? Let's talk.
Sources
- TechCrunch - ChatGPT reaches 900M weekly active users (February 2026)
- Search Engine Land - SparkToro/Similarweb zero-click study 2026 - 68% of US Google searches end without a click.
- IGEM Digimonitor 2025 - 60% of the Swiss population uses AI tools.
- EAB - AI in College Search Survey (February 2026) - 46% use AI in their college search; 18% removed an institution based on AI answers.
- UPCEA & Search Influence - Adult Learner Study - 50% weekly AI search, 56% higher trust when cited.
- McKinsey via MarketingTech - only 16% of brands systematically track AI search.
- Profound - AI Search Volatility - 40-60% of cited domains change monthly.
- SE Ranking - AI Mode Research - 9.2% identical URLs across three runs.
- Semrush - ChatGPT Search Insights - AI visitors convert at 4.4x.
- Adobe Analytics via Digital Applied - AI traffic converts 42% better (March 2026).
- Google - Generative AI performance reports in Search Console (June 2026)
- TechWyse - GA4 "AI Assistant" channel (May 2026)
- Aggarwal et al. - GEO: Generative Engine Optimization (KDD 2024) - up to 40% visibility lift through content optimization.
Frequently asked questions
- What should we track: prompts, brands or keywords?
- The unit of measurement for AI visibility is the prompt, not the keyword. A curated set of 50-150 realistic user questions is run regularly against ChatGPT, Google AI Overviews, Perplexity and Gemini. For each answer you then record whether your brand is named (mention), whether your website is used as a source (citation) and how your share compares to competitors (share of voice). Keywords remain relevant for classic SEO but are not a suitable basis for AI measurement.
- Which KPIs belong in an AI visibility report?
- Five platform metrics: mention rate, citation rate, share of voice, average position within the answer and sentiment. Plus two business metrics from your own website: AI referral traffic (sessions from ChatGPT, Perplexity and others) and the leads or conversions they generate. Only the combination shows management whether visibility translates into business results.
- Why isn't a single ChatGPT query enough?
- Because AI answers are non-deterministic. According to Profound, 40-60% of cited domains change month to month for identical queries; SE Ranking found only 9.2% identical URLs across three runs of the same queries. Reliable statements only emerge from repeated measurement of the same prompt set over weeks - single queries are anecdotes, not data.
- Can we see AI traffic in Google Analytics?
- Yes. Since May 2026, GA4 automatically classifies visits from ChatGPT, Gemini, Claude and Copilot in the new 'AI Assistant' channel. However, Perplexity is missing and still lands in Referral, and app clicks often end up in Direct. A custom channel group with a regex filter therefore remains the cleaner approach. Since June 2026, Google Search Console additionally shows impressions from AI Overviews and AI Mode - but no clicks yet.
- How often should we report to management?
- Measure continuously (weekly or daily runs of the prompt set), report quarterly. AI answers fluctuate heavily in the short term; only the quarter-over-quarter comparison separates real trends from noise. A good quarterly report fits on one page: share-of-voice trend, biggest gains and losses, source analysis, AI traffic and leads, and the actions derived from them.
