Launching soon. Preferium for agencies is in private preview — partner registration isn’t open yet. Join the waitlist →

All articles

How to actually measure AI visibility (hint: ask the engines)

Estimated 'AI visibility scores' are guesswork. The only honest measurement is asking ChatGPT, Claude, Gemini and Perplexity real questions — with web search on — and recording who gets cited.

Photograph: two colleagues reviewing results together on a monitor, one pointing at the screen.

TL;DR: There are two ways to put a number on AI visibility. One estimates it from proxies — domain authority, content scores, gut feeling. The other measures it: send real questions to the assistants buyers actually use, with web search enabled, in the client’s market and language, and record whether the brand is cited, how it’s described, and who appears instead. Only the second holds up in front of a skeptical client.

Measurement, not estimation

An honest AI-visibility metric has to answer a concrete question: when a potential customer asks ChatGPT for recommendations in my category, do I appear? No proxy metric answers that. The only method that does is running the question — repeatedly, systematically, across engines.

That’s how Preferium’s citation tracking works: a panel of real prompts per brand, executed weekly against all four engines — ChatGPT, Claude, Perplexity and Gemini — with web grounding enabled, localized to the client’s market. Each response is parsed for brand mentions, the actual cited source URLs, competitor appearances, and sentiment (how the brand is characterized, not just whether it appears). Google’s answer layer gets its own tracking — AI Overviews and AI Mode citations per keyword.

The metrics that matter

  • Citation rate — of your tracked prompts, how many cite the client at all.
  • Share of voice — the client’s citations versus each competitor’s, on the same prompts. Single most persuasive chart in an agency report.
  • New and lost mentions — week-over-week deltas. A lost mention is actionable the way a lost ranking is.
  • Cited sourceswhich pages the engines pull from. When a competitor’s comparison page is the answer’s source, that page is your content brief.
  • Sentiment — being mentioned as the caveat (“some users report issues with…”) is not a win. Track how, not just whether.

One more, unique to the moment: sponsored placements inside AI answers. Ads are appearing inside assistant responses; knowing who is paying to appear on your client’s topics is competitive intelligence you can bill for.

Why “web-grounded” is non-negotiable

Asking a model from its training data measures the past. Buyers use assistants with search enabled — so measurement must too. Web-grounded runs return the sources the engine actually consulted, which turns visibility from a mystery into a supply chain: retrieval → source pages → citation. Each link in that chain is optimizable (the retrieval link runs through indexes — including Bing’s).

Key takeaways

  • Measure by asking the four engines real questions with web search on — everything else is estimation.
  • Share-of-voice against named competitors is the metric clients understand instantly.
  • Track cited sources, not just mentions: they tell you exactly what content to build.
  • Sentiment separates “mentioned” from “recommended” — only one of those grows revenue.
More articles Become a partner