You measure AI visibility by establishing a baseline before you change anything: ask the engines the questions your customers actually ask, record who gets named and cited, then corroborate the picture with referral data in analytics and crawler activity in your server logs. Without that baseline, "improving AI visibility" is guesswork with extra steps.
Why measure before you optimize?
Because AI visibility work is full of plausible-sounding tactics, and without a baseline you can't tell which ones did anything. If you add schema, rewrite pages into answer-shaped content, and publish an llms.txt file all in the same month, then notice ChatGPT mentioning you — was it the schema? The rewrite? Something a competitor broke? A baseline turns "we feel more visible" into "we're now cited for six of our twenty money questions, up from two." One of those statements can run a marketing meeting. The other is a mood.
What are your money questions?
Money questions are the queries a ready-to-spend customer asks an assistant: "who's the best family dentist in town," "how much does water heater replacement cost near me," "emergency vet open right now." Not brand searches — those tell you little, because engines handle a direct question about your business very differently from a category question where they must choose whom to recommend.
Build a list of ten to twenty. Pull them from what customers say on the phone, the searches that already convert in your SEO data, and the services with the best margins. Write them the way people talk, not the way keyword tools do — assistants get asked full questions in plain language.
How do you run a baseline citation audit?
Ask every money question in each engine that matters for your market: ChatGPT with web search active, Perplexity, Google's AI Overviews and AI Mode, and Gemini. For each answer, record four things:
- Mentioned: does your business appear at all?
- Cited: is your site linked as a source, or are you name-checked without a link?
- Accurate: are the facts stated about you — services, hours, pricing posture — actually right?
- Instead: who is recommended when you aren't, and which sources does the engine lean on — directories, review sites, a competitor's service pages?
One honest caveat: these answers are not deterministic. The same question asked twice can produce different recommendations. Run each question more than once, on different days, and read for patterns rather than treating any single answer as truth. A business cited most of the time is visible; a business cited once, once, is an anecdote.
How do you track AI referrals in analytics?
In GA4 or your analytics tool of choice, look for referral traffic from AI domains — chatgpt.com, perplexity.ai, gemini.google.com and their relatives — and group them into their own channel so you can watch the trend. Two warnings. First, this undercounts badly: a large share of AI answers are consumed without any click, so referral traffic is the visible tip of influence you mostly can't see. Second, treat it as a direction rather than a total — rising AI referrals alongside improving citation-audit numbers is a story you can trust. Watch branded search volume too; people often hear about you from an assistant and then Google you to make sure you're real.
What do your server logs tell you?
Logs measure the supply side of visibility: whether AI crawlers can and do read you at all. Search your access logs for GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot. A healthy pattern looks like regular fetches across your important pages with 200 responses. The warning signs are no AI bot traffic at all, or bots requesting pages and receiving 403s — often a firewall or CDN blocking crawlers by default. If the bots can't read you, no amount of content work will surface you, which is why this check comes first in any diagnosis.
How often should you re-measure?
Monthly is a sensible cadence for a single business — same question set, same engines, same scorecard, so the numbers stay comparable. Re-run after anything significant: a site redesign, a CDN or security-plugin change, a competitor suddenly showing up everywhere. The manual version of all this is genuinely doable, and worth doing at least once for the feel of it. It's also tedious enough that most people quietly stop after the second month — which is the honest reason monitoring platforms like Speak Local exist: they track citations, crawler access, and share of voice continuously so the baseline keeps itself current. However you do it, measure first. You can't improve what you never saw.