Back to Blog
AI VisibilityCitationSeptember 17, 2026

How to Build an AI Citation Benchmark: A Measurement Method

Define the scope, counting unit, and data hygiene for an AI citation benchmark. Reproduce competitor comparisons with a transparent method.

Prefer Maya AI in Google

Highlight our stories in Search, AI Mode & AI Overviews.

An AI citation benchmark compares how brands or sites are cited within a defined set of questions and answers. To be trustworthy, it must be clear not only who is being compared but also which answers are being counted. A ranking that mixes different time periods, platforms, or question groups can be misleading even when it looks precise.

This guide does not announce a research result; it offers an applicable method for building your own comparison. Before finding the largest citation count, its goal is to define what that count represents. That way the report's reader understands the scope of the finding, and the team can reproduce the same calculation.

Narrow the research question

“Which brands does AI like?” is not a measurable starting point. Instead, write a question with clear boundaries, such as “Which sites were cited in the new answers to selected Turkish category questions over the past four weeks?”

This definition sets the language, topic, time period, and unit of measurement. If the platforms were not run the same number of times, report the results separately or choose a shared scope. A platform producing more answers can cause it to be represented more heavily in the overall ranking.

Record the competitor list in advance. Avoid selecting, after seeing the results, only the brands that give you the comparison you want. You can add newly noticed competitors later, but show the list change in the method note.

Define the citation unit

The same answer can cite a URL three times. Another answer can cite four different pages from the same domain. Which of these you count changes the result.

To start, you can use the URL–answer pair as a unit: if the same URL appears again in the same answer, it counts once. If you are preparing a report at the domain level, produce a separate count that combines all pages in the same answer into a single site view. Keep the two counts under different names.

As a representative example, suppose 20 of 100 valid answers contain a link to a site. That site's answer-based citation share is 20%. Across those same 20 answers, the site may have a total of 35 different page–answer pairs. These two numbers do not contradict each other; they measure different units.

Record your cleaning decisions

Redirects, tracking parameters, and canonical addresses can cause the same content to appear under different URLs. Before merging URLs, define which transformation you applied. Do not indiscriminately strip parameters that genuinely change the language, product variant, or content page.

Explain how you handle erroneous answers, incomplete runs, and results that come back from the cache. Counting missing results as “brand did not appear” can artificially lower the rate. Counting the same answer over and over can make some sources look stronger than they are.

Also classify the content type of the source page in a way that can be manually reviewed: product, price, guide, comparison, news, forum, or other. The word “blog” appearing in the URL can be a helpful signal; on its own it does not confirm the type of content.

Do not declare page features the cause of success

Seeing a table, a current price, or an implementation example on frequently cited pages is a useful editorial observation. But it does not prove those pages were cited because of these features. Brand awareness, topic, source network, and the questions measured can also relate to the outcome.

Explain your sample selection too. You cannot generalize to the whole web by reading only pages that receive high citation counts. Comparing them with lower-visibility pages can improve the research; even so, drawing firm causal conclusions from observational data requires care.

Google's helpful content guidance recommends offering original information, analysis, and added value. Rather than copying the format of competitor pages, aim for a contribution that improves the reader's decision. Google content evaluation guidance.

Publish a method card alongside the report

The following fields can form the backbone of a method card:

FieldInformation to disclose
ScopeDate, language, country, platform, and question group
SamplePlanned and valid counts of new answers
CountingURL–answer or domain–answer unit
CleaningDeduplication, error, and cache rules
LimitsClients, questions, and platforms not represented
ReproductionData version, calculation method, and correction date

When preparing a public report, remove the client questions and raw answers you do not have the right to share. Explaining the method does not require publishing all of the private data. Where needed, present examples explicitly as representative.

You can use Maya's data sources hub to review its research approach, and its competitive benchmarking page to assess your measurement needs. A good benchmark does not only say who is ahead; it also shows on which questions, under which conditions, and with which uncertainties they are ahead.

About the author

Ahmet Vefa Akbacı

GEO researcher at Maya. Focused on measuring AI visibility — the metrics, benchmarks, and methods behind tracking how brands appear in AI answers.

Ready to improve your AI visibility?

See how your brand appears across ChatGPT, Claude, Gemini, and other AI assistants.

THE NEXT ANSWER COULD BE YOURS.

Get your brand
mentioned in AI Search.

Let’s make it happen