Updated By Tenzin Langdun

Measuring AI Visibility: What to Track – and What the Management Report Should Look Like

Measuring AI visibility properly: prompts not keywords, KPIs for quarterly reporting, ChatGPT traffic in GA4 – with our own measurement as the example.

AI VisibilityGEOAI Search

Quarterly dashboard with a rising AI visibility curve, share-of-voice bars and chips for ChatGPT, Perplexity and AI Overviews
AI visibility is measured per prompt and reported quarterly as a trend.

AI visibility is not measured with keywords but with prompts: a fixed set of realistic user questions is run regularly against ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude, and for each answer you record whether your brand was named, your website cited or merely retrieved, and how much share your competitors capture. This produces five metrics - mention rate, citation rate, share of voice, position, sentiment - which, combined with AI referral traffic and enquiries, form the quarterly report for management.

The trigger for this article is a real search query that reached us: an education provider has to report quarterly to management on how its AI visibility is developing, and doesn't know what to track in the first place: prompts? Brands? Keywords? It's a fair question, because the market is young and the terminology is fuzzy. This article lays out the measurement setup we use ourselves, including one of our own measurements from September 2026 as the example.

Why does AI visibility belong in management reporting in 2026?

Because a growing share of purchase decisions is shaped inside AI answers before your website is ever visited. ChatGPT counts around 900 million weekly users, 68% of US Google searches end without a click according to SparkToro/Similarweb, and in Switzerland the IGEM Digimonitor 2025 reports that 60% of the population already uses AI tools, 79% among 15- to 34-year-olds.

Add a finding from our own AI visibility study of 1,515 AI answers and 14,135 traced mentions across twelve Zurich industries: around 77% of the firms the assistants recommended are not in Google's top 10, even for the exact searches the assistants themselves ran. Anyone reporting only rankings is reporting on a different game.

For education providers the shift is especially measurable. In the US, according to an EAB study (February 2026), 46% of prospective students already use AI tools like ChatGPT in their college search, almost twice as many as six months earlier. 18% removed an institution from their shortlist because of AI-generated answers. Among adult learners, UPCEA found that half use AI-powered search at least weekly, and 56% are more likely to trust institutions cited in AI summaries. If you don't appear there, you lose prospects who never show up in any analytics report.

At the same time, McKinsey shows that only 16% of companies systematically measure their AI search performance. That is exactly the opportunity. Whoever builds a clean measurement setup now sees shifts before they hit revenue.

Prompts, brands or keywords - what is the right unit of measurement?

The prompt is the unit of measurement. Brand and source are the metrics. Keywords remain an SEO instrument. The confusion arises because all three terms appear in the same report, but on different levels:

LevelRole in the setupExample
PromptUnit of measurement: the question posed to the AI"Which project management further-education course is worth it in Zurich?"
BrandMetric: is our name mentioned in the answer?Mention, position, sentiment
SourceMetric: is our website linked as evidence?Citation
KeywordSEO instrument, not an AI measurement unit"project management course"

People don't type keywords into AI systems. They ask situational questions. A good prompt set therefore mirrors real decision situations of your audience, structured by funnel stage:

  • Unbranded (awareness): "How do I become a certified project manager in Switzerland?" This is where it's decided whether you are considered at all.
  • Category (consideration): "Compare the best providers for a CAS in digital marketing in German-speaking Switzerland." This is where the shortlist is built.
  • Branded (decision): "Is [your institute] reputable? What do graduates say?" This is where the AI audits your reputation.

In practice, 20 to 150 carefully curated prompts covering personas, regions and product lines are enough. Representativeness beats volume: 30 precise prompts along real decision paths deliver better steering signals than 500 generic ones. What matters is that the set stays stable across quarters; that is the only way comparable trends emerge. To see what an evaluated prompt set looks like for a whole industry, read our analyses for dental practices, physiotherapy practices, accounting firms and architecture firms in Zurich.

Since June 2026, Google Search Console has been a goldmine for real phrasings: its report on generative AI features lists the long, spoken-style queries for which your website appeared in AI Overviews and AI Mode. Ours contains sentences like "Our best new client this month found us through a ChatGPT recommendation. Which agencies specialise in this?". Queries like that belong in the prompt set word for word.

Which metrics belong in the report?

Five platform metrics, measured across the prompt set, plus two business metrics from your own website:

MetricWhat it measuresWhy it matters
Mention rateShare of answers naming your brandAn AI recommendation, works even without a click
Citation rateShare of answers linking your website as a sourceYour content is the evidence base
Share of voiceYour mentions relative to competitorsThe management metric: market share inside AI answers
PositionWhere in the answer you appear (first recommendation or footnote)Order shapes perception
SentimentHow the AI talks about youCatch false or outdated portrayals early
AI referral trafficSessions from ChatGPT, Perplexity, Gemini, Claude, CopilotThe link to website reality
Enquiries / deals from AIEnquiries and deals from AI trafficThe business case for management

Two things get mixed up regularly. First: a mention is not a citation. The AI can recommend your brand without linking your website, and cite your website without recommending you. Between the two sits a third stage, visible only on systems that expose their search results: retrieved but not used. The assistant saw your page and chose another. All three stages must be tracked separately, because all three call for different actions. Second: visibility alone is not a result. A sound report therefore connects the platform level (are we named?) with the business level (does it generate business?).

The data makes clear why the business level matters: according to Semrush, visitors from AI search convert at 4.4 times the rate of traditional organic visitors, and Adobe Analytics has measured since March 2026 that AI traffic in retail converts 42% better than non-AI traffic. Fewer visitors, but far more qualified. Exactly what a quarterly report has to make visible.

How do you measure AI visibility properly?

With three rules that all follow from the same property: AI answers are non-deterministic. The same question yields different answers on different days. Profound analysed over 240 million ChatGPT citations: 40-60% of cited domains change from month to month for identical queries. SE Ranking ran 10,000 queries through Google's AI Mode three times; only 9.2% of cited URLs were identical across all three runs.

  1. Measure repeatedly: Run the same prompt set several times a week, run critical prompts multiple times in a row, and work with majority values.
  2. Report trends, not snapshots: Only values aggregated over four to six weeks are reliable. The quarterly comparison smooths out the noise.
  3. Break out systems separately: ChatGPT, AI Overviews, Perplexity, Gemini and Claude cite different sources. An average across all systems hides where you are winning or losing.

There is a fourth rule that is often forgotten: measure through the APIs with web search switched on, not through screenshots from the chat window. Only then are the search queries, the retrieved pages and the citations machine-readable, and only then can the same measurement be repeated exactly next quarter.

What a measurement actually shows: our own case

Our website has been live since June 2026. On 8 September 2026 we put 14 questions about AI workshops in Zurich ("Which providers of AI workshops for SMEs in Zurich can you recommend?", "We are looking for an in-house ChatGPT training for our team", "What does an AI training for 15 people cost?") to ChatGPT with web search, Gemini and Claude through their APIs, with Zurich as the location. The result, broken down by the three stages above:

SystemNamedCited but not namedAbsent
ChatGPT (with web search)11 of 1403
Gemini6 of 14not evaluated8
Claude2 of 1448

Four observations that are worth more in a report than the hit rate:

  • The description comes from your own website. All three systems lifted phrasings from our workshop page almost verbatim: formats, process, review after 30 days, price. What is on the page is in the answer. What is missing is missing there too.
  • The three misses on ChatGPT were answers without a search. For all three questions the model answered from memory and never triggered a web search. That is a different gap from "searched and not found", and it calls for different actions: mentions in third-party sources that feed into training, rather than changes to your own page.
  • Cited is not named. Claude used our page as a source in four answers without recommending us. Counting only citations overstates visibility; counting only mentions overlooks that the page is already being read.
  • Google looks different. For the same topics our pages sat at positions 50 to 90 in Google search, with impressions but no clicks. Rankings would not have shown this visibility.

And the most important caveat: this is one sample from one day. Whether 11 of 14 is a position or a good day will only be known after the repeat in October. That is exactly why the prompt set is fixed and run again every month.

How do you set up ChatGPT traffic tracking in GA4?

Since May 2026, GA4 has a dedicated "AI Assistant" channel that automatically groups visits from ChatGPT, Gemini, Claude, Copilot, DeepSeek and Grok. Three gaps remain: Perplexity is missing and still lands in Referral, the classification is not retroactive, and clicks from AI apps without a referrer end up in Direct. A custom channel group therefore remains best practice, especially since ChatGPT appends utm_source=chatgpt.com to its source links. Here is how to set it up:

  1. Create the channel group: Admin › Data display › Channel groups › Create new channel group. The default group cannot be edited, a copy can.
  2. Define the "AI assistants" channel: Add a new channel with the condition "Source matches regex": chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com. Place the channel above "Referral" and "Organic Search" in the order, otherwise the wrong rule fires first.
  3. Mark key events: Form submission, clicks on the email address and phone number, appointment bookings. Without key events the "Conversions" column stays at zero for every AI session, and the report cannot attribute any enquiry.
  4. Exclude internal traffic: Admin › Data streams › Tag settings › Internal traffic, then activate the filter under Data filters. Otherwise your own test visits, Tag Assistant and the local development server count as sessions. For us, in summer 2026, that was more sessions than the entire AI channel.
  5. Close the no-referrer gap: A "How did you find us?" field on the contact form with "ChatGPT / AI assistant" as an option. Many people type the name from the answer instead of clicking. Those visits land in Direct and are otherwise invisible.
  6. Save a comparison: In the "Acquisition › Traffic acquisition" report, choose your custom channel group as the dimension and save a comparison against the previous month.

What comes out of it is shown by our own GA4 for 1 June to 8 September 2026: 464 sessions in total, 32 of them in the AI channel, 28 from chatgpt.com, 2 each from claude.ai and perplexity.ai. Google organic brought 97 sessions in the same window. ChatGPT is therefore almost the entire AI channel and brings about a third as many sessions as Google organic. The website had been online for three months at that point. Most AI visits landed on the homepage, not on a service page; the homepage is therefore the page that has to be citable first.

In June 2026, Google Search Console introduced dedicated reports for generative AI. As of September 2026 the report shows impressions and the queries behind them; for us, 525 impressions in three months and several hundred long, spoken-style queries. Clicks per query are still rare there and barely usable. For impact measurement, the combination of prompt monitoring, GA4 and the form question therefore remains essential.

What does a quarterly report to management look like?

One page, five blocks, nothing more is needed:

  1. Share-of-voice trend across the last four quarters, per AI system, against your three most important competitors.
  2. Gains and losses: On which prompt topics did we gain, where did we lose, and to whom?
  3. Source analysis: Which websites does the AI cite most often in our category? Often directories, comparison portals and trade media. In our study, around 40% of mentions had no source at all, the assistant already knew the firm; for the remaining 60% it opened a page, and being on that page was worth 8.2 times as much as being the stronger firm on it. This list is where the next action list comes from.
  4. Business impact: AI referral sessions, enquiries generated from them, conversion rate compared to the rest of organic traffic.
  5. Actions: What was implemented, what is planned, what do we expect from it?

The last block is the most important one. Measurement without derived actions is reporting theatre. The Princeton GEO study shows that targeted content optimisation - statistics, source references, citable structure - can lift visibility in AI answers by up to 40%. We described how that works in practice in From SEO to GEO.

Which arguments convince the board?

Five sentences, each with a number behind it, are enough for budget approval:

  1. The channel already exists. 60% of the Swiss population uses AI tools, ChatGPT has 900 million weekly users. Your customers are asking there for providers today, whether you measure or not.
  2. Rankings don't show it. 77% of the firms recommended in our study are not in Google's top 10. The existing SEO report is blind to this channel.
  3. Competitors aren't measuring yet. Only 16% of companies systematically track their AI search performance. Whoever starts now owns the baseline the others will be looking for in a year.
  4. The visitors are better. According to Semrush, AI visitors convert at 4.4 times the rate of organic visitors. Few sessions, but the right ones.
  5. It costs less than a tool without actions. A managed measurement setup with implementation starts with us at CHF 600 per month, monitoring included. What other Zurich agencies charge is in our agency comparison.

And the argument that usually lands in the end: without measurement you don't find out when you disappear. A competitor who gets onto the three list pages ChatGPT reads in your industry pushes you out of the answer, and in analytics all that drops is a channel nobody reports separately.

Buy a tool or have the setup managed?

The market for AI visibility tools exploded in 2026: Adobe acquired Semrush for 1.9 billion dollars (announced November 2025, closed April 2026), and category pioneer Profound was valued at one billion dollars in its February 2026 Series C. There are now dozens of SaaS monitors at every price point. The problem is not access to data, but what tools don't deliver: a well-designed prompt set for your audience, statistically sound interpretation of fluctuating values, the connection to GA4 and CRM data, and above all the actions that follow from the numbers. A dashboard does not answer the question of why a competitor ranks ahead of you in Perplexity and what to do about it.

We solved this problem for ourselves before solving it for clients: Hierarchy runs its own in-house AI visibility tracking system that continuously measures our brand and our clients' brands across ChatGPT, Google AI Overviews, AI Mode, Perplexity, Gemini and Claude, with individually built prompt sets, repeated sampling against answer noise and competitive benchmarks. The same method sits behind our study and behind the example above. Clients receive the tracking as a managed setup: we build the prompt set together, measure continuously, deliver the quarterly management report and implement the derived actions directly, as part of our SEO and AI visibility service. Measurement and execution from one team, instead of yet another dashboard nobody reads.

Conclusion

The answer to the original question, "prompts, brands or keywords?", is: prompts as the unit of measurement, brand mentions and citations as the metrics, share of voice as the management KPI, connected to AI traffic and enquiries from GA4. Measure continuously and repeatedly through the APIs, report quarterly as a trend. Build the setup properly once and you turn a vague management requirement into a steerable metric, and spot market shifts before they reach your enrolment numbers.

Want to know where you stand first? The AI visibility check measures for free whether ChatGPT, Google AI Overviews, Perplexity and Gemini name your brand and which sources they cite instead. You need to report quarterly on AI visibility and want a measurement setup that lasts? Let's talk.

Sources

Frequently asked questions

What should we track: prompts, brands or keywords?
The unit of measurement for AI visibility is the prompt, not the keyword. A curated set of 20 to 150 realistic user questions is run regularly against ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude. For each answer you record whether your brand is named (mention), whether your website is used as a source (citation) and how your share compares to competitors (share of voice). Keywords remain relevant for classic SEO but are not a suitable basis for AI measurement.
How do you measure AI visibility properly?
With a fixed prompt set that stays unchanged across quarters, with repeated runs of the same prompt because AI answers fluctuate, and with separate evaluation per system. For each answer you record whether the brand was named, the website cited or merely retrieved, at which position and with which description. You report the trend over weeks, never a single answer.
Which KPIs belong in an AI visibility report?
Five platform metrics: mention rate, citation rate, share of voice, average position within the answer and sentiment. Plus two business metrics from your own website: AI referral traffic (sessions from ChatGPT, Perplexity and others) and the enquiries or deals they generate. Only the combination shows management whether visibility translates into business results.
Why isn't a single ChatGPT query enough?
Because AI answers are non-deterministic. According to Profound, 40-60% of cited domains change month to month for identical queries; SE Ranking found only 9.2% identical URLs across three runs of the same queries. Reliable statements only emerge from repeated measurement of the same prompt set over weeks. Single queries are anecdotes, not data.
How do I set up ChatGPT traffic tracking in GA4?
In GA4 go to Admin › Data display › Channel groups, create a custom channel group and define a channel "AI assistants" with the condition "Source matches regex chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com". Then mark form submissions, email and phone clicks as key events, exclude internal traffic and ask on the contact form how the person found you. The full step-by-step guide is in the article.
Can we see AI traffic in Google Analytics?
Yes. Since May 2026, GA4 automatically classifies visits from ChatGPT, Gemini, Claude and Copilot in the "AI Assistant" channel. However, Perplexity is missing and still lands in Referral, and app clicks often end up in Direct. A custom channel group with a regex filter therefore remains the cleaner approach. Since June 2026, Google Search Console also shows impressions and queries from AI Overviews and AI Mode; clicks per query are still barely usable there.
What does AI visibility cost for an SME?
A managed measurement setup with a compact prompt set is included in our Base package from CHF 600 per month, together with technical optimisation, one to two content pieces and two meetings. The Growth package from CHF 1,200 per month adds the extended prompt set, the competitive benchmark and the work in third-party sources. Pure monitoring tools cost between roughly CHF 100 and several thousand francs per month depending on the vendor, but deliver no actions. A one-off assessment is free with the AI visibility check.
How often should we report to management?
Measure continuously (weekly or daily runs of the prompt set), report quarterly. AI answers fluctuate heavily in the short term; only the quarter-over-quarter comparison separates real trends from noise. A good quarterly report fits on one page: share-of-voice trend, biggest gains and losses, source analysis, AI traffic and leads, and the actions derived from them.
Tenzin Langdun

About the author

Tenzin Langdun

AI Expert & Marketing Lead at Hierarchy

Tenzin is an AI expert and marketing lead with an MSc in Artificial Intelligence from the University of Bath and over 10 years of marketing experience across strategy, paid acquisition, and SEO. He has held roles at leading organisations including KPMG, EY, Siemens, and Adnovum — with expertise in AI, cybersecurity, audit and consulting, and the insurance sector. Together with Martin Oswald, he co-authored an award-winning research paper on AI-based cancer detection, published in Nature and recognised with the National Siemens Excellence Award and the Lab Sciences Award.