Skip to content
Ante
Deciding what to doAugust 28, 2026

How to measure AEO

Rankings and traffic cannot tell you whether an engine named you. Here is what can, and how to keep the numbers honest.

Table of contents

Key takeaways

  • Presence is binary and the most important: were you named at all.
  • Sentiment and accuracy matter as much as presence. Being described wrongly can be worse than being absent.
  • Never change the question set mid-quarter. A trend on a shifting set is not a trend.
  • Match names by word boundary, not substring. Careless matching inflates results badly and silently.
  • Engines are non-deterministic. Ask more than once and treat a single answer as a sample, not a fact.

Why traffic and rankings do not answer the question

An AI answer often ends the search. The buyer asks, reads three or four names, and forms a shortlist without visiting anyone. If you were named and not clicked, your analytics show nothing at all, and the most valuable event in the funnel is invisible to every tool you already own.

Rankings have the same blind spot from the other side. You can hold position three on a query and never be quoted in the AI answer above it, because the answer is assembled from passages and sources, not from the ranked list.

So the instrument has to be the engines themselves. You ask them what a buyer would ask, and record what comes back.

The five things worth recording

  1. Presence. Were you named, yes or no. Binary, unglamorous, and the metric everything else is conditional on.
  2. Position within the answer. First name mentioned carries more weight than fourth. Order is not random and it moves.
  3. Sentiment. How you are characterised. "The most established option" and "a smaller newer entrant" are both presence and they are not the same outcome.
  4. Accuracy. Is what the engine says about you currently true. Stale pricing and outdated positioning are extremely common and directly cost deals.
  5. Sources cited. Which URLs the answer drew on. This is the most actionable of the five, because it tells you exactly which pages to influence next.

Building a question set that holds up

This is where most measurement goes wrong, and it goes wrong quietly.

  • Write questions a buyer would type, not queries a marketer would target. "What software do med spas use for patient intake" beats "patient intake software".
  • Cover the awareness range. Some questions where the buyer does not know any vendor, some where they name a competitor, some where they name you. These behave differently and mixing them tells you more.
  • Twenty to forty is enough. Fewer and noise dominates. More and nobody maintains it.
  • Then freeze it. Add questions only at a version boundary you record, and when you do, restate the history so nobody compares a new set against an old baseline.

The mistakes that corrupt the numbers

Every one of these produces confidently wrong reporting, which is worse than no reporting.

  • Substring matching on your brand name. Searching for a short brand name inside answer text will match it inside unrelated words and inflate your results substantially. Match on word boundaries, and read the sentences rather than trusting a count.
  • Substring matching on domains. A review-site URL that contains your domain as a path segment is not a citation of your site. Compare hostnames, not string fragments.
  • Asking once. These models are non-deterministic. A single answer is a sample. Repeat and look at frequency.
  • Changing the question set to look better. Usually unintentional, always fatal to the trend.
  • Reporting presence without accuracy. Being named in a wrong description is a problem your presence metric will report as a win.

What a good report actually shows you

Ask for the evidence, not the score. A report that gives you a single visibility number is measuring in a way you cannot check or act on.

  • The question asked, verbatim.
  • The engine and the date.
  • The answer text, or enough of it to see the context you appear in.
  • Where your name appeared, or a clear statement that it did not.
  • The sources the answer cited, as URLs you can go and look at.

Connecting it to something commercial

Citation counts are a leading indicator, not a business result, and it is worth being honest about that gap rather than papering over it.

The most practical bridge is a qualitative one: ask new inbound leads how they found you and whether they used an AI assistant while researching. It is imperfect and it is far better than nothing. A reasonable early commercial target is modest and specific, something like one genuinely qualified inbound conversation per week attributable to AI-assisted research. That is a real bar, it is checkable, and it does not pretend to a precision this channel cannot yet deliver.

The minimum viable version

Twenty buyer questions in a spreadsheet. Asked across ChatGPT, Claude, Perplexity, Gemini and Google AI Mode. Once a week. Five columns: named, position, sentiment, accurate, sources. That is it, and it beats most paid dashboards because you can see the evidence behind every row.

Ante runs this weekly across five engines on every account, and the free AEO report is the same measurement run once so you can see where you stand before deciding anything.

Frequently asked questions

How do you measure AI search visibility?

Ask a fixed set of real buyer questions across the engines on a regular cadence, and record whether you were named, in what position, how you were described, whether that description is accurate, and which sources the answer cited. Traffic and rankings cannot answer it because AI answers often end the search.

What is a good AEO metric?

Presence against a fixed question set is the primary one, with sentiment and accuracy alongside it. Sources cited is the most actionable, since it tells you which pages to influence. Any single composite visibility score should be treated with suspicion unless you can see the evidence behind it.

Can Google Analytics track AI citations?

No. If an engine names you and the buyer does not click, no analytics tool records anything. Some referral traffic from AI products is visible, but that undercounts the event that matters, which is being named while forming a shortlist.

How often should we measure?

Weekly is enough to see movement without drowning in noise, and monthly is the minimum for a trend you can act on. What matters more than frequency is that the question set stays fixed between readings.

Why do we get different answers each time we ask?

These models are non-deterministic, so the same question can produce different names on different runs. That is why single answers are samples rather than facts, and why frequency across repeated asks is more trustworthy than any one result.

Book a strategy call