Back to Learn
#AEO

How to Run an AI Search Content Audit

A practical audit for measuring AI citations, diagnosing retrieval gaps, and prioritizing the pages most likely to improve AI search visibility.

Four stacked wireframe blocks representing stages in an AI search content audit

An AI search content audit measures whether answer engines can retrieve, understand, and cite your pages for real buyer questions. Build a prompt set, record citations by engine and URL, diagnose misses across access, relevance, evidence, and passage structure, then prioritize pages as defend, strengthen, consolidate, or create. Traffic alone cannot make those decisions.

What does an AI search content audit measure?

An AI search content audit measures citation coverage, not only page performance. Its basic unit is a buyer question answered by an engine. For each question, record whether your brand appears, whether your domain is cited, which URL earns the citation, which passage supports the answer, and which competing sources appear instead.

That creates three distinct measures. Brand mention rate shows how often an engine names you. Citation rate shows how often it links to your domain. Prompt coverage shows how many questions produce at least one useful appearance. Track each engine separately because a page can appear in Google AI Overviews and remain absent from ChatGPT or Perplexity.

Clicks still matter, but they cannot stand in for citation visibility. A Pew Research panel covering 68,879 Google searches found that users clicked a traditional result in 8% of visits when an AI summary appeared, compared with 15% when one did not. Citation data adds the visibility signal that click data misses.

How is this different from a traditional content audit?

A traditional audit asks whether a URL earns traffic, links, conversions, or another documented business outcome. It usually ends in keep, update, consolidate, or delete. The existing SEO recovery audit covers that workflow.

An AI search audit asks a narrower set of questions. Can an engine access the page? Does the page answer the subquestions generated around a buyer prompt? Can a useful passage stand alone? Does the passage carry enough evidence for an engine to reuse it? The output is a citation-opportunity queue, not a sitewide deletion list.

The two audits should share inventory data but retain separate verdicts. A page can earn healthy organic traffic and still need stronger answer passages. Another can receive AI citations while contributing little business value. Neither signal should erase the other.

What data should you collect before scoring pages?

Start with the URLs that map to buyer questions, category education, comparisons, workflows, and definitions. Utility pages and expired campaign pages rarely deserve passage-level analysis. Join one row per canonical URL to these fields:

  • Search demand: Top queries, clicks, impressions, and average position from Search Console.
  • Business value: Key events, assisted pipeline, sales use, or another documented outcome.
  • Citation coverage: Mentions, citations, cited URL, engine, prompt, and observation date.
  • Source competition: Domains and pages cited when yours is absent.
  • Passage readiness: Question heading, direct answer, standalone context, current evidence, and source links.
  • Access status: Indexability, canonical target, rendered text, robots rules, and server response.

Do not classify an empty analytics row as zero demand until you confirm the export is complete. The Search Console interface returns 1,000 representative rows, while the Search Analytics API supports up to 25,000 rows per request and pagination. Large sites need the API or bulk export before joining query data to URLs.

Use a crawl to find canonical conflicts and orphan pages before testing citations. The orphan-page workflow combines crawl results with analytics and sitemap URLs, which catches pages engines may discover even when your internal navigation does not.

How do you build a useful prompt set?

Build prompts from actual buyer language, then group them by job. Search Console queries, sales calls, support tickets, community discussions, and customer interviews are stronger inputs than a brainstormed list of head terms. Each prompt should represent a question a buyer would plausibly ask without naming your brand.

Use four prompt groups:

  • Category questions: What is the problem, method, or category?
  • Evaluation questions: Which approaches, products, or tradeoffs fit a stated situation?
  • Implementation questions: How does someone complete a specific workflow?
  • Risk questions: What fails, what should be avoided, and what evidence changes the decision?

Keep the first benchmark small enough to review manually. Twenty to fifty prompts per priority topic is usually enough to expose repeated gaps without turning the audit into a monitoring project. Record the exact prompt, engine, date, answer, cited domains, cited URLs, and your brand's role in the answer.

Do not combine variations that reveal different intent. "What is content operations?" and "How do I build a content operations workflow?" may cite different pages because one asks for a definition and the other asks for a process. The audit should preserve that distinction.

How do you establish a citation baseline?

Run every prompt across the engines that matter to your buyers using the same evaluation window. Record results by engine rather than averaging them into one score. An aggregate rate can hide that one engine cites your domain consistently while another never retrieves it.

Calculate citation rate as cited prompts divided by tested prompts. Then calculate citation share among all cited domains. The first measure tells you whether you appear. The second shows how much of the available source visibility you capture against competitors.

Add a passage field whenever the engine exposes one. A 2026 analysis of 15.7 million AI Mode citations found that 47.7% used text-fragment highlights and that the median highlighted passage was 117 words. The analysis is observational, but it makes passage-level review more useful than a page-level cited or not-cited flag.

Save screenshots or answer text for the first benchmark. Answer engines change, citations move, and a naked percentage will not tell you whether the system changed its answer or simply switched sources.

How do you diagnose why a page is not cited?

Test failures in sequence. Access comes first because editorial work cannot repair a page the engine cannot retrieve. OpenAI separates OAI-SearchBot, used for ChatGPT search visibility, from GPTBot, used for model training, in its crawler documentation. Audit the search crawler rule you intend to support instead of treating every AI user agent as the same system.

Check robots rules, canonical tags, status codes, rendered text, and internal links. Robots.txt controls crawler access but does not guarantee deindexing, as the robots.txt documentation explains. A blocked script or contradictory canonical can leave a page technically live but hard to understand.

If access is clean, test relevance. Compare the prompt with the page's actual question, not its keyword list. A 2026 analysis of 863,000 keywords found that only 38% of AI Overview citations came from pages ranking in the top 10 for the exact query. The result suggests that subquery coverage can matter even when the source does not rank for the prompt verbatim.

Then inspect the passage. A strong candidate answers one question immediately, makes sense outside the page, names concrete entities, and supports factual claims. If the section depends on "the process above" or delays the answer for several paragraphs, an engine has to reconstruct context before it can cite the passage.

Evidence is the final test. The original GEO benchmark covered 10,000 queries and found that adding citations, quotations, and statistics improved source visibility in its experimental setup. Treat that as evidence for clearer sourcing, not a promise that adding numbers will create citations in every live engine.

How should you score each page?

Score each target page from zero to two across four dimensions. Zero means the requirement is absent or broken. One means it is partial. Two means it is complete and verifiable.

  • Access: The priority crawlers can fetch the canonical page and read the main answer content.
  • Question fit: The page directly serves at least one prompt cluster and covers the subquestions that cluster produces.
  • Passage quality: Important sections use question-led headings, answer immediately, stand alone, and preserve enough context to quote accurately.
  • Evidence: Claims use current primary or established sources, examples are specific, and dates appear beside time-sensitive numbers.

An eight-point score is a prioritization device, not a universal ranking formula. Record the reason beside every score so a second reviewer can reproduce it. "Passage quality equals one because three of six priority sections refer to earlier context" is useful. "Needs AEO work" is not.

Weight the dimensions only when the business case requires it. A documentation library may weight access and question fit highest. A research-led category page may weight evidence and passage quality. Publish the weighting before reviewers see individual scores so the rubric does not bend around favorite pages.

Which action should each page receive?

Use actions that describe the AI search job instead of copying a traditional deletion framework:

  • Defend: The page receives useful citations for priority prompts. Preserve its successful passages, monitor source accuracy, and avoid broad rewrites.
  • Strengthen: The page fits the prompt but has weak access, incomplete subquery coverage, buried answers, or thin evidence. Update the smallest failing layer.
  • Consolidate: Several pages answer the same prompt cluster and split authority or contradict one another. Select one canonical source, merge unique evidence, and redirect only after checking the wider SEO audit.
  • Create: Competitors receive citations for a buyer question your site does not answer. Brief a new page only when no existing URL can serve that intent without losing its current job.

Retirement belongs in the traditional content audit because citation absence alone is not a deletion signal. A page may support customers, sales, links, or conversions even when no answer engine cites it. Route weak pages back to the broader audit before removing them.

Prioritize strengthen rows that already rank, convert, or attract links. They have demonstrated demand and authority, so a focused passage or evidence update carries less risk than creating another URL for the same topic.

How do you strengthen a page for citation readiness?

Rewrite at the passage level. Start each priority section with the answer, then explain the mechanism, boundary, or evidence. A section should remain accurate when copied out of the page, including the subject and necessary qualifiers.

Use this review sequence:

  • Turn label headings into questions buyers ask.
  • Put a direct answer in the first sentence below each priority heading.
  • Replace backward references with the specific noun or process they describe.
  • Add dates and sources to claims that can expire.
  • Remove unsupported comparisons and generic authority language.
  • Merge repeated answers so one passage owns each subquestion.
  • Add internal links that establish the page's role in the topic cluster.

Do not pad sections to hit a word count. The aim is complete context, not a longer paragraph. The practical AI search optimization workflow covers the rewrite mechanics once the audit identifies the passages worth changing.

Keep useful nuance. A short answer that strips away conditions may be easy to extract and wrong in practice. The helpful-content guidance still asks whether content demonstrates first-hand expertise and leaves readers satisfied. Citation readiness cannot compensate for an answer that fails the reader.

How do you validate the audit after making changes?

Record the implementation date and retest the same prompt set after the engines have had time to recrawl. Compare citation rate, cited URL, cited passage, competitor share, organic performance, and business outcomes. A citation gain paired with declining qualified traffic is not automatically a win.

Review results at the prompt-cluster level. One new citation can be noise. Several related prompts switching to the same improved passage is stronger evidence that the update changed how the page is retrieved or selected.

Protect passages that already perform. If an engine repeatedly cites one section, preserve its answer, evidence, and URL during adjacent updates. Broad rewrites can erase the exact wording and context the engine had learned to reuse.

The baseline also makes the next audit faster. Instead of crawling the whole site and guessing, rerun priority prompts, inspect changed citations, and open only the pages connected to those changes.

For a weekly operating lesson on AI search visibility and content systems, subscribe to The Messy Middle. You will get practical criteria you can use to keep the audit rubric current as citation behavior changes.

Frequently Asked Questions

Related Content