Skip to content

Prompt tracking is bullsh*t: why questions beat prompts

More prompts don't mean better intelligence. Why the buyer's question, not the prompt, is the right unit for AI visibility strategy.

A checklist of three buyer questions at the centre of a purple gradient, connected by dotted paths to question mark badges

Prompt tracking was a reasonable first move. Buyers started researching products in ChatGPT, Gemini, Perplexity, and marketers wanted to know what those systems were saying about them. So tools appeared that built lists of prompts and tracked the answers.

Then the tool became the strategy. Vendors started selling prompt volume as coverage, and buyers started repeating the number back. Five hundred became five thousand. Nobody stopped to ask whether tracking more prompts told you anything more about your market.

Buyers do not think in prompt libraries. They have something they need to figure out, and they ask it in whatever words make sense in the moment. That is the problem with building an AEO strategy around prompts. It starts with the wording instead of the intent behind it.

More prompts don't mean better intelligence

We covered the first problem in our last post: Your buyers don't write prompts, they ask questions. Responses to prompts are inherently non-deterministic. The model, search behavior, available sources, context, and timing all affect the responses, so you can ask the same question many times and get different answers back each time.

Adding more prompts does not solve that. More repeated probes can improve the estimate for a question set you have already chosen. They cannot tell you whether that set covers the decisions your buyers are actually making, or whether it repeats one intent in dozens of phrasings and leaves the rest of the journey unmeasured.

It makes an already questionable metric much more expensive to maintain and you can never, ever fix an irrelevant question set with 'more prompts'.

Every prompt has to run across models, often repeatedly. When search is enabled, each test may also require retrieval. Then the results have to be stored, compared, monitored, and interpreted.

The costs grow quickly, but the useful information does not necessarily grow with them.

We've seen companies actioning 5,000 to 7,000 prompts without learning materially more than teams working with a fraction of that number. A large prompt set can easily contain hundreds of variations of the same underlying question while missing an entire stage of the buyer journey altogether. Tracking more prompts may improve the estimate for the intent you already had. It does not tell you that you are covering more of what buyers need to know. Often, it means you are paying to monitor more ways of saying the same thing.

Prompt volume has become a proxy for sophistication. The real judgement here is whether the additional data changes what you know or what you should do? If it does not, it is just expensive data collection.

The prompt isn't even fixed

There is another problem with treating a prompt as the unit of measurement. The answer engine may not use it as written.

OpenAI says ChatGPT Search typically rewrites a user's question into one or more targeted search queries. It can search, review what comes back, and search again before producing the final answer.

By that point, the exact sentence the buyer typed into a prompt is no longer the important part. The system is trying to understand what the person actually wants to know, their real question.

Most prompt-tracking systems still treat that original string almost like a keyword in traditional search. But answer engines increasingly interpret context, expand queries, and follow the meaning behind what someone asked. Buyer behavior is moving in the same direction.

Google says queries in AI Mode are about two to three times longer than traditional searches and often lead to follow-up questions. People are no longer forced to compress a complicated need into a few keywords.


Traditional search might have looked like this:

enterprise project management software


Now someone can say:

We have 600 employees. Engineering lives in Jira, but the rest of the company hates it. I need something to run cross-functional projects without forcing every team into the same workflow. What should I look at?


There is no practical prompt library that can anticipate every version of that conversation. And even if you tracked that exact sentence, you would be measuring a string the engine already threw away.

Go back to the job

AEO has introduced plenty of new language, but the marketing job has not changed very much. You need to understand what buyers need to know and make sure your company has a credible answer. The difference is that buyers now have far more freedom in how they ask.

Instead of starting with a spreadsheet of prompts, start by mapping the important questions in your market.

  • What are buyers trying to understand before they know they have a problem?
  • What do they need to learn before they consider your category?
  • What comes up when they compare you with someone else?
  • What information do they need before they are comfortable making a decision?

Those questions are much more useful than thousands of possible ways to phrase them. Map them and you get a structured view of what buyers need to know, and where your company is strong, weak, or absent.

That is question intelligence, and GrowthX built it.

Build the map

Start with the market you already have in front of you. Look at your own site and your competitors' sites. Look at what prospects ask sales, what people search for, and what comes up repeatedly in customer conversations. Look at the third-party sources buyers rely on when they are trying to understand your category.

Then work backward to the questions behind that information. Different wording can roll up to the same underlying need, and twenty slightly different searches might all come back to one question your buyer needs answered well.

Once those questions are mapped, you can see the gaps. Some will already have strong answers on your site. Others will be buried in weak content, answered only partially, or missing altogether. In some cases, a competitor will have a much stronger answer to a question your company should own.

The GrowthX platform makes this distinction visible in a Topics section. Coverage is measured as questions answered out of questions found, not pages published. A topic can have 22 pages behind it and still leave all four of the questions buyers are asking unanswered.

Then you take that same question set to the engines. The AI Visibility Segments section has a different job. It runs a maintained panel of buyer questions across answer engines and records whether the company was cited, who appeared instead, and which sources shaped the answer.

The order is what makes it different. The questions come from your market. The engines tell you how you are doing against them. Flip it and you are back to tracking a list of sentences somebody invented.

A gap might call for a new page. It might mean rewriting an existing one. It could expose weak positioning, stale documentation, poor proof, or a comparison experience that is losing the buyer before they ever talk to sales. The system shows you where the gap is. You decide what to do about it, and the answer is not always another piece of content.

Then use prompts to test the answer

Prompt tracking is still useful when it has a specific job to do. Once you know a question matters, you can test how answer engines respond to it. If you repeatedly monitor it under comparable conditions, it can provide reinforcement. You can try different wording, look at whether your company is cited, see how it is described, and understand which sources are shaping the answer.

Then read it honestly. A range or a persistent pattern is more useful than a weekly movement presented to one decimal place. When a pattern emerges, you'll know what to change, what to watch, or even why no action is warranted.

That is a much better role for the prompt. It is a test against a buyer question you already know matters, not the foundation for understanding the market itself.

The goal is not to maintain the largest possible library of sentences somebody might type into ChatGPT. It is to know what your buyers are trying to understand, make sure your company has a strong answer, and see whether that answer is actually getting through. Prompt tracking belongs downstream of the strategy, because the question comes first.

We're going deeper on this in a live workshop: what 2.6 million AI responses showed us about which questions actually earn citations, and why different engines reward different content. Free, and everyone who registers gets the replay. Register here