When you track AI visibility, the headline number is a ratio: answers that mention your brand, divided by all answers collected for the prompts you chose. Change the prompts and the ratio changes, even if nothing changed in how ChatGPT or Gemini sees your company. That makes the prompt list the most important setting in any tracker, and the one teams spend the least time on.
The denominator problem
Take a brand that appears in 95% of answers to prompts that name it (“Is Brand X reliable?”) and in 10% of answers to category prompts (“best project management tool for a 5-person agency”). These are illustrative numbers, but the effect is arithmetic:
| Prompt set | Reported visibility |
|---|---|
| 10 branded + 10 category prompts | 52.5% |
| 20 category prompts | 10% |
| 18 category + 2 branded prompts | about 19% |
Same brand, same week, three different scores. Branded prompts are useful for checking how AI describes you, but they belong in a separate report. Mixed into the main score, they inflate it and hide the number that matters: how often you show up when a buyer has not heard of you yet.
Real people don’t type your keywords
SEO keyword lists don’t translate directly into prompts. In research published by SparkToro and Gumshoe in January 2026, 142 people each wrote a prompt for the same task: choosing headphones for a family member who travels. The average semantic similarity between any two prompts was 0.081. People added budgets, ages, flight lengths and brand dislikes. Almost no two prompts looked alike.
The useful finding came next. Across 994 answers to those 142 prompts, the same brands kept appearing: Bose, Sony, Sennheiser and Apple showed up in 55–77% of responses. When the intent changed to gaming or podcasting headphones, the brand set changed too. AI tools respond to intent, not wording. So the job is to cover intents, and to phrase each one the way a buyer would, with the constraints they actually mention.
Category width sets the ceiling
The same research shows why a visibility score cannot be compared across industries. In narrow categories, the top brands reached 90–100% visibility. For a broad request, brand design agencies for a coffee shop, the leaders sat in the 30–40% range, because the model had far more candidates to choose from.
A 35% score can be excellent in one market and weak in another. The meaningful comparison is your share against competitors on the same prompt set, measured over the same period.
One prompt, many searches
When AI search tools answer a prompt, they often run their own searches first. Google’s Search Central documentation says AI Overviews and AI Mode may use “query fan-out”, issuing several related searches across subtopics to build one response. The prompt you track is therefore not one keyword. It is a bundle of hidden queries.
That explains why classic rank tracking is a poor proxy. Moz analyzed nearly 40,000 queries and found that 88% of AI Mode citations did not match the URLs ranking in the organic top 10 for the same query. Ranking for the head term does not mean you are cited for the prompt.
Answers move; your prompt list shouldn’t
AI answers change constantly. Ahrefs tracked AI Overviews for 43,000 keywords and found the content changed every 2.15 days on average, with 45.5% of citations replaced at each update. The overall meaning stayed close (0.95 semantic similarity), but the sources and brands rotated.
This is why a trend needs a stable prompt set run repeatedly. If you rewrite half the list every month, you can no longer tell whether visibility moved because the AI changed or because you changed the question.
How to build a prompt set worth tracking
- Collect real wording: question-style queries from Google Search Console, recurring questions from sales calls and support tickets, and threads where people ask for recommendations in your category.
- Group prompts by intent. Typical groups are category discovery (“best X for Y”), comparison (“X vs Y”), problem-led (“how do I stop Z”), local (“near me”, city names) and branded. Report the branded group separately.
- Keep the buyer’s constraints. Budget, company size, location, industry and use case change which brands appear, and a prompt without them tests a buyer who doesn’t exist.
- Hold a stable core for at least a quarter. Add new prompts in batches and note the date, so a jump in the chart has an explanation.
- Match each intent group to a page, on your site or a third-party site, that ought to be cited. If no such page exists, that is your first content task.
What to check in your tracker
Before you trust a visibility score, confirm the tool lets you edit prompts, tag them by intent, filter out branded prompts, and see results for each prompt on each platform. Also check whether it logs when prompts were added or removed.
A visibility score without its prompt list is just a number. With a clean, stable, intent-based list, it becomes a measurement you can act on and explain to the rest of the business.





