Back to Learn
#AEO

How to choose prompts for AI tracking

Learn a six-step method to select high-intent prompts for monitoring brand citations, mentions and share of voice across AI search engines.

Scored AI tracking prompts grouped by intent and journey stage

Choosing prompts for AI tracking is a six-step procedure: pick the metric, source real demand, score candidates, bucket by intent, size to budget, and refresh on a cadence. Keyword research gives you strings with volume. This method gives you a monitored set where every prompt feeds citations, mentions, or share of voice.

Prompt tracking is not the go-to. It is a diagnostic you run on buyer questions that already matter, not a library you grow until the dashboard looks thorough. Marcel Santilli is walking through that live on October 22 at https://luma.com/growthx-ai-workshop. If you miss it, join the community at https://ailedgrowth.com to catch the replay.

How do you choose prompts for AI tracking in one method?

Run the steps in this order, because each one constrains the next.

  1. Define the metric. Pick citations or mentions. If you track share of voice, write down the denominator and name the engines and competitive set you will count across.
  2. Source candidates. Pull question queries from Search Console, forum threads, keyword-tool SERP filters, and your own sales calls and support tickets.
  3. Filter with a score. Rate each candidate on commercial intent, answer likelihood, competitor presence and search demand, weighted, and cut below a line you set.
  4. Bucket by type and stage. Tag survivors as branded, unbranded, comparison, use case or transactional, and map each to the awareness stage it serves.
  5. Size to budget. Fix the prompt count per topic and engine against the run count you can afford, and set the branded share.
  6. Refresh on schedule. Review monthly, re-source quarterly, and retire prompts on defined triggers while keeping a locked core for baseline continuity.

What should you start measuring first?

The metric you report determines which prompts earn a slot in the set. A mention is your brand named in the generated answer, while a citation is a page the engine retrieved to build that answer, and the two move independently. Share of voice adds a denominator problem on top: Peec divides your mentions by all brand mentions, while Search Engine Land’s Share of Model Voice divides your appearances by total answers in the set. One preprint restricts the denominator to competitors that recur or resolve to a real domain so one-off hallucinated names do not inflate it. Pick one definition, write it at the top of the tracking sheet, and never compare numbers built on different ones. For the tracking layer around those numbers, see https://ailedgrowth.com/learn/track-brand-ai-search.

You need a different prompt shape for each metric. Across 160 queries, Google AI Overviews triggered on 98% of informational queries, 40% of commercial, 35% of navigational and 0% of transactional, so a citation program built on transactional prompts will show nothing on that surface. In a 1,000-prompt study across four engines, researchers found that AI search skewed toward earned media over brand-owned sources. Use prompts your pages answer directly when you want to measure own-domain citations rather than relying on category lists.

Researchers found another pattern in transactional and comparison prompts. A 1,000-query study found Google’s sourcing ran 41% earned, 34% social and 26% brand, while the AI systems raised brand-owned citations to 52–68% on transactional intent. And comparison, table, list and ranking prompts surfaced about 20% more brands on average in a Peec-based study reported by Search Engine Journal. Use list-shaped prompts for mentions and share of voice. Use transactional and how-to prompts for own-domain citations, and know which metric you are buying before you source a single candidate.

Where do you source candidate prompts?

Candidates come from two kinds of source: public demand signals every competitor can also pull, and internal data only you hold.

How do you mine search and community demand?

Search Console is the fastest start because your own question queries are already phrased the way people ask assistants. In the Performance report, click + Add filter, choose Queries, choose Custom (regex), paste a pattern, leave Matches regex selected and click Apply, as Google’s filter documentation lays out. Google’s own example pattern is what|how|when|why, published in its 2021 post introducing negative regex matching. A tighter anchored pattern from Search Engine Land catches more question forms: ^(who|what|where|when|why|how|which|whose|whom|is|are|was|were|do|does|did|can|could|will|would|should|has|have|had)\b.

Search Console uses RE2, which has no lookahead, lookbehind or backreferences, so keep patterns to alternation and anchors. For the branded split, Google added a branded and non-branded query filter in November 2025 that runs on Google’s own classification rather than regex, though it is unavailable for low-impression sites.

Keyword tools add two things Search Console cannot: competitor demand and SERP-feature context. Nikki Lam’s keyword-tool workflow filters Semrush keyword lists to the “Discussions and forums” SERP feature and runs Keyword Gap against reddit.com, then checks upvotes and sentiment on the threads behind them. Pull People Also Ask questions for your top 20 queries while you are there. Copy the related questions Perplexity shows under each answer as well. Those are the engine’s own follow-up turns, and they belong in your multi-turn pool later.

Forum threads are the closest public proxy for how people talk to an LLM. Search Engine Land recommends site:reddit.com "best project management tools" and scanning titles and comments for openers like “I’m struggling with”, “How do I” and “What’s the best way to”. Appending &udm=18 to a Google search URL restricts results to the Forums view. Google does not document the parameter, SerpApi catalogued it as the forums mode, and one vendor reported Google sometimes stripping or rewriting udm values, so treat it as a shortcut rather than infrastructure. The payoff is phrasing. Joanna Lambadjieva’s example of a thread asking whether anyone has moved off HubSpot to something lighter for a small agency reads almost exactly like a ChatGPT prompt, constraint and all.

How do you mine internal data competitors cannot copy?

Sales call transcripts contain the prompts your actual buyers would type, including the objections. Kyle Poyar’s playbook describes pulling roughly 200 keywords out of discovery-call transcripts, and Webflow’s Brett Domeny suggests auditing call transcripts for the questions reps could not answer, since those are the questions a prospect will take to an assistant instead. Support tickets and on-site search logs round out the list. Add customer service records to the same pool. Search Engine Land’s prompt-level measurement guide pairs synthetic benchmark prompts with real ones from sales, support, communities and site search for exactly this reason.

This is the step most teams skip, because exporting call transcripts feels like a detour from the tracking tool, and it is why their prompt sets end up identical to every competitor’s. Keep the set as a versioned sheet with source, score, bucket and metric per row. The tracking platform is one place you run that sheet. It should never be the only place the set exists.

How do you filter candidates with a scoring matrix?

Four weighted criteria turn a pile of candidates into a ranked list you can cut.

  • Commercial intent (weight 3): Does the person asking fit your ICP and sit close to a purchase decision? This carries the most weight because a mention on a prompt nobody with budget asks is a vanity number.
  • Answer likelihood (weight 2): Will the engine name vendors at all? Transactional prompts draw no AI Overview, and list-shaped prompts surface more brands, so score the shape, not the topic.
  • Competitor presence (weight 2): Do named competitors appear in the answer today? Without them the share-of-voice denominator is empty and you measure nothing competitive with the prompt.
  • Search demand proxy (weight 1): Google volume for the nearest keyword, the same proxy Ahrefs uses to estimate AI impressions. It gets the lowest weight because conversational prompts have no volume of their own.

Score each criterion 1 to 5, multiply by the weight, and total out of 40. The cut line is yours. In the illustrative example below, a project management tool sold to agencies tracks everything at 28 or above, which keeps roughly the top half:

[@portabletext/react] Unknown block type "table", specify a component for it in the `components.types` prop

The scores are an operator’s judgment, and the table is the artifact you argue about with your team. Two cuts show why the weights matter. The freelancer prompt fails at 27 despite near-certain answers and plenty of competitors, because freelancers sit outside this ICP and intent is weighted triple. The kickoff-meeting prompt fails at 19 even with the highest demand score, because engines answer it every time and name almost no vendors. The operating rule that follows is specific: treat a prompt that engines answer without naming brands as a citation prompt. It belongs in the set only when own-domain citations are the metric you defined in step one.

How do you bucket prompts by type and journey stage?

Tagging survivors by type and awareness stage tells you which metric each prompt feeds and where the gaps in your coverage are. Five buckets cover most B2B SaaS sets.

  • Branded (product-aware through purchase): Phrasings like “is [your brand] worth it” and “[your brand] pricing”. Add “[your brand] reviews” to the same bucket. Use these for accuracy and sentiment checks, along with own-domain citations. They trigger AI Overviews on only about a third of navigational queries, so expect blanks on Google surfaces.
  • Unbranded category (solution-aware): “best project management tools for small agencies” or “which project management tools do agencies use”. Use these for mention rate and share of voice, since competitors tend to appear here.
  • Comparison (product-aware): “[your brand] vs Asana for agency teams” and “Asana alternatives for agencies under 20 people”. Engines surface 20% more brands from the list shape in this bucket, which makes it a high-yield bucket for share of voice.
  • Use case (problem-aware): “project management tool with built-in time tracking and client portals”. Search Engine Journal recommends tracking constraint variants for integration, team size and budget, and this is where they go.
  • Transactional (purchase): “project management software that integrates with QuickBooks and Slack” or “[your brand] free trial”. Researchers found the highest brand-owned citation rate here and the lowest AI Overview rate, so prioritize these on ChatGPT and Perplexity over Google AI Overviews.

Problem-unaware prompts like “how to run a project kickoff meeting” sit upstream of all five. Add one or two only if citations are your metric, and label them so nobody reads a zero mention rate as a problem. For the overall mix, the Peec-based study points to 25% top-of-funnel, 50% mid-funnel and 25% bottom-of-funnel, and Kyle Poyar runs roughly a third core, a third competitive and a third experimental. Search Engine Land’s intent list adds objections, validation and implementation as distinct categories. Objections and implementation are the two most sets leave out entirely.

How many prompts should you track, and what branded vs unbranded split?

Start at 40 to 50 prompts across four or five topics, and run the identical set on every engine you report. Hold branded prompts to 20–30% of the total. Practitioners derive the topic math from sets that cluster three to five topic areas with eight to ten prompts each, and Anna Crowe’s GEO audit template logs 50 prompts weekly with engine, date, question, inclusion, snippet and other brands per row. Per-engine comparison only works if the set is identical across engines, so never let one platform’s default prompts stand in for a second engine’s set.

With fifty prompts, you get a trend rather than a precise point estimate. For a 95% confidence interval spanning five percentage points on citation share, one preprint estimated Gemini needs roughly 40–50 queries per topic, Perplexity about 100 and SearchGPT 150 or more. Treat those as the ceiling you grow toward on your two or three highest-stakes topics, after your first month of results shows which topics move.

No study documents an ideal branded ratio. The one data point is single-site vendor data from Bing Webmaster Tools, where branded queries were 29.7% of grounding queries but drew 75 citations each against 35 for non-branded, which is why branded prompts can stay a minority and still anchor your own-domain citation count. The 20–30% figure is a starting default drawn from that pattern and the one-third core share above. Move it toward 15% once sentiment and accuracy are stable.

Tracking everything is the common failure. Five hundred prompts run seven times a day across four engines is 420,000 calls a month by straight multiplication, and at OpenAI’s $10 per 1,000 web-search calls that is $4,200 a month in search fees before a single token is billed. More prompts also means more noise to explain to stakeholders and less attention per prompt that matters.

Why is prompt selection not keyword research?

A keyword is a fixed string with a measurable volume. A prompt is a conversational, multi-variable request that the engine rewrites before it ever hits an index, and the rewrite is what your content has to match. OpenAI’s help documentation describes ChatGPT search rewriting a request into one or more targeted queries, then sending more specific follow-ups after reviewing initial results. Google Search Central defines query fan-out as a set of concurrent, related queries the model generates to fetch additional results. Nobody publishes an official average, but independent measurements put ChatGPT at 2.17 searches per prompt with a search triggered on 31% of prompts, and a 501-prompt agency test found Gemini 3 via the API averaging 10.7 fan-out queries per prompt, ranging from 3 to 28 with grounding forced on.

The second difference is determinism. A keyword rank is one observation of one result page. In SparkToro’s volunteer study, ChatGPT and AI Overviews each returned the same brand list on fewer than 1% of repeated runs, and Anthropic’s own documentation states that temperature zero does not make Claude fully deterministic. A prompt is a distribution you sample, and the selection question is which distributions are worth sampling.

That changes what “good prompt” means. You do not pick the phrasing with the highest volume. You pick the parent question whose predictable sub-questions your content can answer, which is why Search Engine Land advises mapping each parent query to the sub-questions it reliably spawns. That parent-query work is a different job from keyword mapping. See https://ailedgrowth.com/learn/prompt-mapping-vs-keyword-mapping. The GSC regex in the sourcing step is the bridge: it takes keyword-era data and filters it down to the question forms that survive the rewrite.

How do you account for variability across runs, engines, and locations?

A single run of a prompt is a snapshot, and researchers recommend seven runs as the floor for a trustworthy daily read. The GEO measurement preprint found within-day source overlap between runs of 0.32–0.43 on Jaccard and recommends at least 7 runs per prompt per day for brand visibility and 8 for source-level coverage, which gives a 95% interval of about ±0.158 on per-brand detection at 7 runs, aggregated over a rolling 2–4 weeks. Mike Sonders’ 100-run test on logged-out ChatGPT found that only about five brands appeared in 80% or more of answers to a B2B software prompt, and he suggests around five runs per key bottom-of-funnel prompt as the practical heuristic when you cannot afford the statistical ideal. Whether API calls vary the way manual sessions do is an open question SparkToro flags, so log the collection method alongside the result.

Engines often disagree with each other. For well-known brands, the four-engine GEO study measured ChatGPT search at 93.5% earned sources with 0% social, Perplexity at 67.4% earned and 23.8% social, Gemini at 63.4% earned and 25.1% brand-owned, and Claude at 87.3% earned. Across 50 U.S. metros, Gemini and ChatGPT cited the same domains only 8% of the time. On Google’s own surface, AI Overviews shared only 18% of linked pages between runs, against 45% for organic results.

Copilot is the thinnest-evidenced engine: consumer Copilot has no public API, and a first-party signal is the AI Performance report Microsoft added to Bing Webmaster Tools in February 2026. GPT-4o search and Gemini 2.5 Flash, the model versions behind studies here, have since rotated. Sonar has as well, so confirm the current default on each engine when you log a run.

Multi-turn and localized variants expand the set fast, so ration them. Search Engine Land recommends evaluating whole conversation paths rather than isolated prompts, and the cheapest version is adding one follow-up turn to your five highest-scoring comparison prompts, something like “which of those is cheapest for a 15-person team”.

For location, the US-versus-Germany study found AI Overviews appearing on 81% of US queries and 65% of German ones with no significant difference in source diversity, while a vendor study of 56,223 citations found Perplexity drawing 56.5% of citations from non-global sources against 5.3% for Gemini. Localize only the prompts where geography changes the answer, and set the location through the user_location parameter on OpenAI’s web search tool or Perplexity’s rather than duplicating prompt text per city. Variants live in the experimental third of the set, and the core two-thirds stays untouched.

What does it cost to track 50, 200, or 500 prompts?

Polling APIs yourself costs tens of dollars a month per engine at one run a day, while third-party platforms charge by prompt count and bundle the engines. The figures below are the published rates on each vendor’s page at the time of writing. The API rows assume one call per prompt per day for 30 days on one engine, with tokens extra unless noted. Pricing on every one of these pages has changed within the past year, so check the linked page before you commit a budget:

[@portabletext/react] Unknown block type "table", specify a component for it in the `components.types` prop

The Otterly 200 and 500 figures are computed from its published extra-prompt rate, and the OpenAI and Perplexity all-in rows name specific models that may have rotated since the pricing pages were captured. Two more platforms matter for sizing even though they do not fit the grid: Semrush’s AI Visibility Toolkit is $99 a month per domain for 25 custom prompts, and Ahrefs Brand Radar runs $199 a month for one platform or $699 for all platforms with 2,500 custom prompt checks. Profound’s free trial covers 50 prompts a day for seven days across three engines, with no prompt customization, which makes it a sanity check rather than a tracking plan.

Adjust the table to turn it into a real budget. API rows are per engine, so three engines triple them. The seven-runs-a-day guidance multiplies them again: 50 prompts at 7 runs for 30 days is 10,500 calls, which is about $105 in OpenAI search fees or about $52.50 on Perplexity’s Search API by straight multiplication from the rates above.

Gemini bills per search query rather than per prompt, so the 10.7-query fan-out average means a Gemini set can cost several times the per-prompt row. Platform plans avoid those multipliers by fixing one daily run and bundling engines, which is the trade you are making when you pay $95 for 50 prompts instead of $15.

When should you refresh, retire, or add prompts?

Review the set monthly and re-source candidates quarterly. Reset baselines whenever an engine changes its default model. Vendors update models on their own schedules: OpenAI’s policy is that a ChatGPT model stays available for about 90 days after its successor ships, and Google made Gemini 3 the global default for AI Overviews on January 27, 2026 after moving AI Mode to it the previous month. No vendor publishes a research-based rerun interval, and the only quantified stability guidance is the 2–4 week rolling aggregation from the GEO measurement preprint. Search Engine Land’s tracking guide covers the practical defense: record model name, version and test date on every run, so a step change in your chart lines up with a release note instead of a content push.

Retire a prompt on one of four triggers.

  • Answer likelihood collapsed: No engine has named a vendor in the response for four consecutive weeks. Treat the prompt as citation-only, and keep it only if citations are the metric.
  • ICP or product moved: You dropped a segment, killed a feature, or repositioned, and the prompt’s constraint no longer describes a buyer you want.
  • Duplicate signal: Two prompts return near-identical brand sets week after week. Keep the higher-scoring one and spend the slot on an objection or implementation prompt.
  • Competitor set changed: A named competitor in a comparison prompt was acquired, renamed or exited the category, so the comparison denominator is no longer valid.

Add prompts from the quarterly re-source, and give new ones a denser sampling window before they count. Scrunch’s own collection cadence is a useful signal here: it polls new prompts daily for 14 days, then drops to a 72-hour refresh. Swap no more than the experimental third in any one quarter, because rotating the whole set prevents baseline comparison. One preprint found citation volatility was predominantly structural and system-level rather than driven by content updates, which supports a locked core set. Without stable prompts, you cannot separate the engine’s drift from your own work.

How do you put the method to work this week?

You can build a working set in one week if you hold the steps in order.

  • Monday: Pull 16 months of question queries from Search Console with the anchored regex and split them with the branded filter.
  • Tuesday: Build the public pool from ten Reddit threads via site:reddit.com and the People Also Ask questions on your top 20 queries. Add Perplexity’s related questions.
  • Wednesday: Block the morning for the step teams skip. Export the last 20 discovery calls and the last quarter of support tickets, then lift every question a prospect asked in their own words.
  • Thursday: Score the pool in the matrix and cut below your line. Tag buckets and stages, then check the mix against the 25/50/25 funnel split and the 20–30% branded share.
  • Friday: Run the full set on every engine you report, at least five times for the bottom-of-funnel prompts. Log engine, model version, date, inclusion, snippet and other brands per row.

Store the sheet in a shared, versioned workspace and keep it out of one person’s chat history. For a weekly working example of this kind of operational prompt work, subscribe to the newsletter.

Frequently Asked Questions

Related Content