GEO 2026: measuring visibility at prompt level
The unit of GEO measurement is not the keyword, it is the prompt. A practical framework for building a fixed prompt set, logging answers and checking access.
By Roozbeh Nazari · CEO
Most of the debate around GEO (Generative Engine Optimization) is still stuck at the definition stage: is it a new discipline, an extension of SEO, or a passing fashion. That argument is not much use in practice, because the real problem teams run into is not definition, it is measurement.
Here is what it looks like on the ground: management asks "are we showing up in ChatGPT?", someone on the team types a question, and if the brand appears in the answer the verdict is "yes". The following week the same question gets asked and a different answer comes back. There is nothing comparable between the two observations, and in that state it is not data, it is anecdote.
This piece describes the shift that makes GEO measurable, and the practical system you can build around it. There is no such thing as a ranking guarantee or a guarantee of being cited; the only thing you can measure is visibility itself and how it moves over time.
The unit is the prompt, not the keyword
In classic SEO the unit is the query. A query is a fixed string, the results page is largely repeatable, and position is a single number.
In generative systems all three of those break:
- The input is not a fixed string. The same need can be phrased in dozens of ways, and the phrasing visibly changes the answer.
- The output is not repeatable. Ask the same prompt twice and you may get two different answers. A single observation proves nothing.
- There is no such thing as a position. The brand either appears or it does not; if it appears, it appears inside a context and sometimes with a source link.
The fix for all three breaks is the same: move the unit from a single prompt up to a prompt set, and measure the rate across repeated runs instead of a single observation. The answerable version of "are we showing up in ChatGPT" looks like this: "across our defined set of [N] prompts, averaged over [K] repetitions, our brand appears at [rate]; last period it was [rate]." The brackets are deliberate: your own measurement fills in the numbers, and nobody else's benchmark can be your baseline.
How to build the prompt set
The prompt set is the backbone of GEO work. A badly built set does not measure the thing you think it measures.
Three principles:
First, prompts should be derived from the way real people phrase things. Turning your own service page heading into a prompt means you are testing yourself. Questions that reach the sales team, support tickets and long-tail queries in Search Console are far better sources.
Second, the set should reflect a range of intents. Include at least four types: the category question ("how do you choose an agency for X"), the comparison question, the problem question ("why does Y happen") and the direct brand question. These four behave differently, and melting them into a single average destroys the information.
Third, the set has to stay fixed. If you change the prompt set every month, what you are measuring is the set, not visibility. When you do need to change it, run the old and the new set in parallel for one period.
Preparing that set is work that resembles classic keyword research but produces something different. In our AI search visibility engagements, building this set is usually the first step.
What you record: a four-field log
Logging only "appeared / did not appear" on each run throws away most of what you have in front of you. Keep four fields at minimum:
- Appearance: was the brand mentioned in the answer.
- Context: how it was mentioned — as one option among several, as the single example, or inside an unfavourable comparison. This field is far more informative than the raw appearance rate.
- Source: if the answer linked out, which page it linked to. Your own site, a directory, a news outlet. This field steers content strategy directly.
- Co-mentions: which other brands showed up in the same answer. This is where you see the competitive landscape, and it may not match your classic SERP competitors.
The source field is the biggest surprise for most teams: in answers where the brand is mentioned, the link often goes not to your own site but to a third-party page. That tells you part of the work sits not on your pages but on the pages that talk about you.
Run the log three times and report the rate. One run is noise; three runs are the beginning of a signal.
The access layer: can they see you at all
Before you get to measurement there is a precondition to verify: can these systems reach your site. This is the layer most often skipped in GEO work, and the cheapest one to fix.
What to check:
- Know exactly which user agents you are allowing in robots.txt. OpenAI's own documentation defines separate agents for separate purposes: one for search surfaces, one for model training, one for user-triggered visits. Blocking them all in a single line also blocks search visibility.
- Do not assume how a rule will be interpreted. Google's robots.txt specification document shows that matching and precedence rules can work differently from what you expect.
- Your content should be present in the first HTML response. Content rendered only on the client is not handled the same way by every system. If you have any doubt on that side, there is a setup worth reviewing on the technical SEO side.
- Note also that Google offers no separate "optimize for AI" layer: its own documentation says there is no extra requirement for appearing on AI surfaces, and that the basic condition is the page being eligible to appear with a snippet in normal search.
What changes on the content side
Once measurement is in place, content changes stop being guesswork; the source and context fields in the log tell you what to work on.
The direction that generally works: answer the question in one place, completely. An answer that arrives after a long preamble does not produce a quotable unit. Completing the answer in two or three sentences right under the heading works better both for the reader and for citation.
The second direction is sourcing your claims. Unsourced general statements do not form a quotable unit. Structuring pages this way is the core of what we cover on the content SEO side.
The third is terminological consistency: calling the same concept by different names across pages scatters the chance of being associated with that concept.
Finally, expectation management: the output of this work is not a guarantee, it is a curve. Nobody can commit to a brand appearing in a generative answer; what you can measure is the direction of your appearance rate over time within a defined set. If that direction is up, the work is doing something — and that is useful enough to know without any guarantee attached.
Who you report to, and how
One last practical point: the way this measurement is reported determines whether the work continues. Presenting the raw appearance rate as a single number generates needless panic whenever it wobbles.
A healthier presentation looks like this: show the four prompt types on separate rows, give the average of three repetitions and the change from the previous period on each row, and keep the distribution of the source field in its own table. That way, instead of an uninterpretable sentence like "we went down overall", you get a workable one: "we slipped on comparison prompts, we are flat on category prompts".
Also keep the limits of the measurement written into every report: results are not repeatable, a single observation is not evidence, and the rates depend on the defined set. That note stops the numbers from being taken too seriously in the months ahead.