A practical framework for measuring AI search visibility
How to move from occasional prompt checks to a documented baseline that distinguishes mentions, recommendations, citations and representation.

AI-search visibility is easy to demonstrate and surprisingly difficult to measure well. A person asks one familiar question, sees a brand in the answer and captures a screenshot. That observation can be useful, but it is not yet a baseline.
A defensible measurement programme needs defined journeys, a stable prompt sample, recorded test conditions and separate measures for different outcomes. The purpose is not to manufacture one impressive score. It is to produce evidence that helps a team decide what to investigate and improve.
Define the business question before collecting prompts
Begin with the decisions people make, the audiences making them and the products or entities being considered. A prompt set for brand discovery will look different from one designed around comparison, validation, troubleshooting or purchase.
Write down what the organisation needs to learn. It may be whether the brand enters a consideration set, how accurately an important product is described, which competitors are recommended, or what sources repeatedly support an answer. This definition stops the project becoming a tour of interesting screenshots.
- Audience and market
- Journey or decision being represented
- Brand, product, service or topic under investigation
- Platforms and environments included
- What would cause the team to take action
Build a prompt sample, not a bag of questions
Prompts should be grouped by a documented logic. Include core questions, meaningful variations and the stages of the journey that matter. Avoid selecting only phrases that already produce favourable answers.
Keep a stable core sample for comparison over time. A separate exploratory sample can change as products, competitors and audience behaviour develop. This creates continuity without preventing the research from learning.
Measure distinct outcomes separately
A mention, recommendation and citation are not interchangeable. Neither is a positive description the same as accurate representation. Record each outcome separately before deciding whether a combined summary is useful.
- Presence: whether the organisation appears at all
- Prominence: where and how strongly it appears in the answer
- Recommendation: whether it is presented as an option for the user
- Citation: whether an owned or third-party source is linked or referenced
- Accuracy: whether important claims are correct
- Message: which attributes, strengths or concerns are associated with the brand
- Competition: which alternatives appear and what evidence supports them
Record the conditions and the limitations
AI outputs can vary by model, interface, location, time, account state and wording. Record the date, platform, prompt, relevant settings and answer. Where the programme permits, repeat a sample to understand how stable the observation is.
A baseline represents the defined test, not every possible answer a person could receive. Reporting should make that boundary clear. False precision is more dangerous than an honest limitation because it encourages teams to act on noise.
Investigate the sources behind the pattern
When a pattern is material, examine the evidence environment. Important information may come from owned pages, reviews, publications, communities, reference sources, product feeds and other third parties. Compare the brand with competitors across those spaces.
Technical fundamentals still matter. Google states that pages need to be indexed and eligible to appear with a snippet to be considered as supporting links in its AI features, and that normal search best practices continue to apply. Important information should therefore be crawlable, indexable, internally discoverable and available in text.
Turn the baseline into a repeatable operating rhythm
Agree which prompts remain stable, how often the sample will be rerun and which movements deserve investigation. Connect the observations to technical, content, digital PR, product and brand owners rather than leaving the work inside an SEO dashboard.
The useful output is a history of evidence and decisions: what changed, why the team acted, and whether the later sample moved in the expected direction. Referral traffic and conversions should be considered alongside visibility observations when those data are available.
Key takeaways
- Define the decision and audience before writing prompts.
- Keep a stable core sample and a separate exploratory set.
- Separate presence, recommendation, citation, accuracy and message.
- Record test conditions and report limitations openly.
- Use the baseline to assign action and compare later measurements.