ChatGPT brand visibility: a 100-prompt benchmark setup
How to build a fixed 100-prompt benchmark for measuring your brand’s visibility in AI answers: set design, run protocol and scoring rules.
By Roozbeh Nazari · CEO
The question of visibility in AI answers gets answered, in most companies, with a one-off experiment: a few questions are asked, and if the brand does not come up there is worry, and if it does there is relief. The problem with that method is that asking the same question the next day produces a different answer. What it takes to become measurement is not more questions but a fixed set of questions and an unchanging protocol.
This article explains how to build such a set. We are not publishing measurement results or citation rates: figures of that kind are meaningful only for a specific brand, a specific date and a specific language, and they mislead when carried to another brand. For what the measurement means, our article on prompt-level measurement is the companion; for the implementation side, our AI search visibility service.
Why a hundred prompts, and why a fixed set
There is no magic in the number itself; a hundred is where two practical constraints intersect. Below it, the number of questions per layer drops far enough that a single answer can move the table. Above it, reviewing the set by hand and repeating it on every run becomes a load most teams cannot sustain.
What actually matters is not the number but the fixity of the set. A benchmark’s value comes not from an absolute rate but from how the same set changes over time. If the set is updated on every run, what is being measured is not the brand’s visibility but the change in the questions asked. The rule is this: the set is updated once a quarter, with written reasoning and a version number; between updates not a single prompt changes.
Designing the prompt set: four layers
A hundred prompts are not chosen at random; they are divided into layers so as to represent the different contexts in which the brand should appear. The distribution that works in practice is this:
- Category questions (around forty prompts): questions asked about a service or a problem, with no brand name in them. This is the layer where visibility is genuinely tested, because the model has to choose the brand itself.
- Comparison questions (around twenty-five prompts): questions asking for a choice between two approaches, two methods or two cities. This layer shows the context in which the brand gets mentioned.
- Brand questions (around twenty prompts): questions with the brand name in them directly. These measure accuracy rather than visibility: does the model describe the brand correctly, does it produce wrong information?
- Long-tail questions (around fifteen prompts): narrow, specific questions written in real customer language. The highest visibility rate usually appears here, and this is the fastest-improving layer.
For brands operating in several languages the set is built separately per language; translation is not enough, because the way category questions get asked differs by language. For a clinic working in four languages that means four separate sets of a hundred, and that cost has to be accepted at the outset.
The run protocol
The purpose of the protocol is to be sure the difference between two runs comes from the brand. Four rules achieve that. The first is repetition: each prompt is asked at least three times within the same run, because there is natural variability between answers to the same question and a single answer is not a measurement. The second is session hygiene: each prompt is asked in a clean session, otherwise the previous question influences the next one’s answer.
The third is recording: the full text of each answer, its date, and the interface version used are kept. That record is the only thing that settles a later argument about how it "was different last month". The fourth is a fixed date window: the run is completed within a single day rather than spread over a week. If model-side updates fall inside the measurement window, the result is a mixture of two different systems.
Scoring: what counts and what does not
The scoring rules are written before the run and are not changed during it. Three separate things are recorded for each prompt: whether the brand is mentioned in the answer text, whether a link is given as a source, and whether the context of the mention is positive, neutral or wrong. These are three separate measures and should not be summed into a single "visibility score"; when they are, the critical difference between being mentioned without attribution and receiving a link disappears.
What will not count is defined explicitly too. Where the brand was placed in the question by the user, the model repeating it is not visibility. A similarly named other organisation being mentioned does not count either — this happens more often than expected in multilingual sets because of transliteration. A brand appearing in only one of the three repetitions is marked "unstable" and does not count in full.
Mistakes to avoid when interpreting results
The most common mistake is reading the result of a single run as a level. The first run is not a level but only a baseline; the meaning arrives with the second run. The second mistake is comparing your rate against a competitor’s: unless the competitor’s set is identical to yours, the two rates are not comparable.
The third and most expensive mistake is moving from a result to a commitment. A brand being cited in AI answers is an outcome that can be influenced on the content side but cannot be guaranteed; which source gets quoted is not under the publisher’s control. The function of the benchmark is to see the direction of the changes made — not to promise a target. On the access side, how the relevant crawlers are handled in robots.txt should be decided deliberately before the measurement.
Who runs the set, and what it costs
Running a hundred prompts with three repetitions means three hundred question-and-answer pairs for a single language. Done by hand, that takes an experienced person roughly a full day; for a brand operating in four languages it comes to four days per quarter. When that cost is not accepted at the outset, what happens is predictable: the first run is done, the second is postponed, and the benchmark reverts to a one-off experiment.
The right way to reduce the cost is not to shrink the set but to systematise the run. Tying record-keeping to a standard template, having the scoring rules written in advance, and gathering the repetitions into a single session all shorten the time noticeably. If the set has to be shrunk, it should be shrunk with the layer ratios preserved; keeping only the category questions and dropping the brand questions loses the accuracy dimension of the measurement entirely.
Who runs it affects the result too. The same person running it every quarter reduces protocol drift. If the person is going to change, having one run done in parallel by two people before the handover and comparing the results shows whether the scoring rules are genuinely shared — a comparison that, on most teams, ends with the rules being rewritten.
What the benchmark changes, and what it does not
A benchmark is not a visibility tool but a feedback loop. What it changes is that the direction of decisions taken on the content and technical sides becomes measurable: after a particular set of pages is strengthened, you see how the mention rate in the relevant layer moves. This does not isolate the effect of a single page, but it shows direction, and for most decisions that is enough.
What it does not change is the outcome itself. Which source the model quotes is not under the publisher’s control, and that is independent of how well the measurement is set up. The honest use of a benchmark is to observe the effect of the work done, not to set a target. A provider committing to a particular mention rate on the basis of benchmark results inverts the purpose of the measurement.