Back to Learn
#AEO

How to optimize comparison pages for LLMs and AI answer engines

Build comparison pages that LLMs cite using a verdict-first structure, flat tables, self-contained claims, schema, and a monthly measurement loop.

Verdict-first comparison page with a flat product table and dated source column

Optimize a comparison page for LLMs by writing it passage by passage: a verdict that names both products, a flat table with one unit-labelled fact per cell, claims that stand alone when extracted, schema that matches the visible text, and a monthly log of which URL each engine cites. Answer engines retrieve spans, not pages.

How do LLMs retrieve and cite comparison content?

Answer engines cite passages, not pages. The Perplexity search architecture breaks each document into self-contained spans that are retrieved and ranked individually. It scores results at both the document and sub-document level, then compresses the surviving sentences into the snippet the model reasons over. The cited URL is the page that held the span answering the sub-question, and that span competes on its own, with the rest of your page out of view.

That retrieval unit is the cause behind every formatting rule in this article. The working target for a citable passage is 40–60 words that name the subject and state the fact with its unit or date inside the same sentence. Neither OpenAI nor Perplexity publishes the chunk size its web retriever applies to third-party pages, so treat 40–60 words as the writing convention for the opening passage of each section rather than as a published parameter.

On a vs page, open each H2 section with a passage that answers that section’s question outright and names both products. A pronoun (“it costs more”) or a back-reference (“as noted above”) turns into noise when the retrieval system extracts the span alone. The system can use a passage that makes sense in isolation more easily than one that only made sense in sequence.

For the category layer this sits in, start with what AEO is. For how those passages should be structured, see how to structure content for AI citations. For the measurement loop after you ship, see how to track your brand in AI search.

How is this different from traditional SEO for comparison keywords?

AI retrieval systems evaluate passage extractability when selecting citations. Search ranking systems determine blue-link placement. Across 11,203 keywords checked at three points between August 2024 and January 2025, 40% of the pages cited in AI Overviews sat outside the organic top 10. For comparison keywords this gap matters more than for most, because the “best X” and “X vs Y” results are held by review platforms and listicle publishers whose domain authority you will not out-build this quarter.

  • The unit of competition: In traditional SEO you optimize one URL against one keyword and win or lose the ranking as a whole. In AI answers you compete section by section, because the retrieval system selects the span that answers the sub-question and ignores the rest of the URL. Your pricing section can be cited in the same answer where your integrations section loses to a competitor’s.
  • The outcome you optimize: An AI answer synthesizes across sources and attaches a short list of cited links. You are writing for the citation and the recommendation the model attaches to your product in the prose, with or without a click. You shift the target from ranking the page to supplying the passage the model quotes when it explains how the two products differ.

What vs page template should you use from verdict to who should pick which?

Models tend to cite pages that put the recommendation first and give each criterion its own section. In a test of six direct ATS matchups, Venture Harbour’s Workable-versus-Greenhouse page appeared in 13 answers across seven prompts and five surfaces, and its structure was a recommendation up front followed by six criteria: pricing, sourcing, structured hiring, setup, reporting, and integrations. Lever’s Greenhouse-versus-Ashby guide, built the same way, appeared in 27 answers across 14 prompts. Those are small single-run counts, but they describe the shape every cited page in that study shared.

Copy this skeleton and keep the wording of the headings, since the heading is the question the model matches against:

# [Product A] vs [Product B] ([Month Year] comparison)

Verdict: 40-60 words. Both names. One measurable reason each. No pronouns.
[Product A] is the better fit for [buyer type] because [one measurable reason with unit].
[Product B] is the better fit for [buyer type] because [one measurable reason with unit].
Last verified [YYYY-MM-DD] against [source type].

## How we compared [Product A] and [Product B]
[Criteria used. Data sources with dates. Who wrote this page and your relationship to either product.]

## [Product A] vs [Product B] at a glance
| Criterion | [Product A] | [Product B] | Source and date verified |
| --- | --- | --- | --- |
| [Criterion 1 (unit)] | [one value] | [one value] | [URL, YYYY-MM-DD] |
| [Criterion 2 (unit)] | [one value] | [one value] | [URL, YYYY-MM-DD] |

## How [Product A] and [Product B] compare on [criterion 1]
[Opening passage, 40-60 words, both names, the fact, the unit, the date.]
[Product A]: [what it does for this criterion, its limit, the unit, the source and date].
[Product B]: [what it does for this criterion, its limit, the unit, the source and date].

## Who should pick [Product A]
[2-3 sentences naming the buyer, the workflow, and the criterion that decides it.]

## Who should pick [Product B]
[2-3 sentences. Written with the same care as the block above.]

## What changed on this page
| Date | Row | Old value | New value | Source |

A compact filled sample, using two reporting tools whose documentation is public, shows how the slots read once the placeholders are gone:

# Google Search Console vs Bing Webmaster Tools for AI search reporting (2026 comparison)

Bing Webmaster Tools is the better fit for a content team that needs to know which prompts triggered a citation, because its AI Performance report lists grounding queries alongside cited pages. Google Search Console is the better fit for teams measuring Google surfaces, because its Search Generative AI report breaks out impressions by page, country, device and date, with no query data. Last verified against Google documentation dated June and August 2026 and Bing documentation dated February 2026.

## Google Search Console vs Bing Webmaster Tools at a glance
| Criterion | Google Search Console | Bing Webmaster Tools | Source and date verified |
| --- | --- | --- | --- |
| Query-level data for AI citations | No | Yes (grounding queries) | Google gen-AI performance reports, 2026-06 and Bing AI Performance preview, 2026-02 |
| Click data for AI surfaces | No (impressions only in the AI report) | Not stated | Same Google source, 2026-06 |
| Isolate AI Mode traffic alone | No | Not applicable | Search Engine Land, Google AI Mode traffic data |

## Who should pick Bing Webmaster Tools
Pick Bing Webmaster Tools when the question you are answering is which prompt produced this citation, because the AI Performance preview reports cited pages with the grounding queries behind them.

The who-should-pick block for the product you are not selling is the part most teams skip, and it is why their first-party pages read as sales copy to a retriever. Write that block with the same specificity as your own. The next section covers the table rules the sample follows.

How do you format a comparison table so models can parse it?

The table is the densest set of facts on the page, and a retriever can only use what it can read as text. Five rules keep every cell extractable:

  • One fact per cell: W3C table guidance says each separate piece of data gets its own cell, and line breaks must never be used to fake extra rows. If a criterion has two values (a monthly cap and an annual cap), it gets two rows.
  • No merged or split cells: a merged header leaves the data cells beneath it without a programmatically determinable label. Flatten the table or split it into two tables when the topic changes.
  • No icon-only cells: a checkmark glyph carries no words a chunker can retrieve. Write “Yes”, “No”, “Not published”, or the actual value.
  • Units and currency spelled out: put the unit in the column header in brackets (“Price (USD per month)”, “Export limit (rows per month)”) and keep cells numeric. Never leave a model to guess whether 500 means seats, rows, or dollars.
  • Real table markup: use table, caption, and header cells with column and row scope rather than CSS-aligned div grids, and never use a table for layout.

Then restate the table in prose. No vendor documents how its web retriever treats HTML table markup, and the compression Perplexity describes operates at sentence level. The rows that decide your verdict should therefore also exist as full sentences directly below the table. Each sentence names both products and includes the value with its unit and date: “As of August 2026, Google Search Console’s Search Generative AI report shows impressions by page, country and device but no query data, while Bing Webmaster Tools’ AI Performance preview lists the grounding queries behind each cited page.”

The reader question underneath all of this, table or prose, has a direct answer: use both for different jobs. The table is where discrete values live and get maintained. The prose restatement provides complete sentences that an answer engine can evaluate alongside the table, without assuming that a particular retriever omits HTML tables.

How do you write self-contained, neutral claims that survive extraction?

A claim survives extraction when it names both products and carries a number with its unit and date. It also avoids adjectives that depend on who is speaking. Keep prices inside the dated table rather than scattering them through prose, and let dates and counts carry the argument.

The before-and-after pattern looks like this on the sample page:

  • Before: “Google’s AI reporting is limited and frustrating, so most teams end up blind.” No product named on the other side, no number, two adjectives, and nothing a retriever can verify.
  • After: “Google Search Console’s Search Generative AI report, released June 2026, shows impressions by page, country, device and date but no click data and no query data, so a team cannot see which prompt produced a citation.” Both the product and the limitation are stated as facts a model can lift intact.

Use this rule from the example: every sentence on the page should still be true and specific if it appears alone inside an AI answer with your logo nowhere in sight. It should also be attributable. “Industry-leading accuracy” fails that test. “A 9% bounce rate acknowledged in the vendor’s own support documentation, checked on this date” passes it.

How do you write a comparison of your own product without LLMs flagging bias?

Publish the competitor’s strengths and a when-to-choose-them line, because models may extract facts from a page that ranks its own product first while sourcing the recommendation elsewhere. In a study of 100 B2B “best software” queries checked three times between April and June 2026, 80 triggered an AI Overview, self-promotional listicles were cited 323 times, and in 224 of those cases Google cited a brand’s own page without recommending that brand. Citation and recommendation are separate outcomes, and a self-ranked page tends to earn only the first.

Two edits move a first-party vs page from sales asset to source:

  • State every criterion you lose: if the competitor’s export limit is higher or its onboarding faster, write that in the cell. State the loss in the criterion passage using the same neutral form you use for your own wins.
  • Document the comparison and disclose your role: link the competitor’s pricing page or documentation, including release notes, for each competitor value and record the date you checked it. In the “How we compared” block, name who wrote the page and state that one of the two products is yours. Place the block above the table, not in a footer.

When you publish the competitor’s strengths on your URL with sources attached, the model can use one page for the full comparison rather than assembling each half from a different source. Without those strengths, the model may quote a review platform for the competitor and your page for your product, leaving the third party to frame the recommendation.

What schema markup belongs on a comparison page?

Google’s AI guidance states that no special structured data is needed to appear in AI Overviews or AI Mode, and that structured data should match the visible text. No published controlled test shows schema alone lifting AI citations. Schema’s job on a vs page is therefore narrow: restate, in machine-readable form, what the visible table already says, so a parser doesn’t have to infer which price belongs to which product.

Two constraints determine what you can mark up:

  • Product scope: Google’s product rich-result guidance only supports pages focused on a single product, so a two-product page should not expect a product rich result. Each Product node needs a name plus at least one of review, aggregateRating, or offers.
  • FAQPage support: FAQPage guidance says FAQ rich results stopped appearing in Google Search on May 7, 2026, so remove FAQPage markup from existing vs pages rather than adding it to new ones.

A generic schema.org pattern for representing the two products is an ItemList wrapping one Product node per product. Google does not document this pattern as a comparison-page rich-result feature:

{
"@context": "https://schema.org",
"@type": "ItemList",
"name": "[Product A] vs [Product B]",
"itemListElement": [
{
"@type": "ListItem",
"position": 1,
"item": {
"@type": "Product",
"name": "[Product A]",
"description": "[one sentence matching the table]",
"offers": {
"@type": "Offer",
"price": "[number identical to the table cell]",
"priceCurrency": "USD",
"priceValidUntil": "[published date through which the price is valid]"
}
}
},
{
"@type": "ListItem",
"position": 2,
"item": {
"@type": "Product",
"name": "[Product B]",
"description": "[one sentence matching the table]",
"offers": {
"@type": "Offer",
"price": "[number identical to the table cell]",
"priceCurrency": "USD",
"priceValidUntil": "[published date through which the price is valid]"
}
}
}
]
}

Use priceValidUntil only when the offer has a published validity date. Your verification schedule belongs in the visible change log instead. Any value in the JSON-LD that differs from the visible cell conflicts with Google’s guidance, so generate the schema from the same source the table renders from rather than maintaining it by hand.

How do you give AI crawlers access?

Allow the search-indexing bots by name, because each vendor documents them separately from its training crawler. OpenAI crawler guidance recommends allowing OAI-SearchBot and says sites that opt out won’t appear in ChatGPT search answers, while GPTBot governs training and is controlled independently. Robots.txt changes take about 24 hours to propagate. Perplexity crawler guidance recommends allowing PerplexityBot, which surfaces sites in its search and is not used for foundation-model training. Its documentation also says that Perplexity-User, the on-demand fetcher, generally ignores robots.txt.

The minimum configuration for a vs page you want cited:

User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

# Training opt-out is a separate decision. Blocking GPTBot does not block OAI-SearchBot.
# User-agent: GPTBot
# Disallow: /

Apply the operational guidance in three areas:

  • Separate search from training: You make separate decisions about blocking a training token and blocking a search token, with different consequences. A blanket User-agent: * disallow also applies to OAI-SearchBot and can remove your pages from ChatGPT search. Audit the file for that mistake before you touch the page.
  • Prioritize robots.txt over llms.txt: No measurable evidence shows that llms.txt improves citations. Gary Illyes said Google rejects llms.txt and will not crawl the files. OpenAI and Anthropic crawler documentation does not mention the convention. Perplexity’s documentation does not mention it either. Publish one if a stakeholder insists, but put the effort into robots.txt.
  • Render core content in HTML: Perplexity-User and ChatGPT-User perform user-triggered fetches, and their vendor documentation says robots.txt rules may not apply. The vendors do not publish a guarantee that these fetchers will execute every client-side JavaScript dependency, so ship the verdict and table in the initial server response. Include the prose restatement there as well.

How do you keep pricing and feature data fresh enough to be trusted?

For recency-demanding queries, stale prices can get a page rated as stale. Google’s Search Quality Rater Guidelines include freshness guidance that treats pages with old product models and prices as stale for recency-demanding queries and assigns them low Needs Met ratings. Comparison queries involving current products often demand recency because the teams behind both products ship changes. No AI vendor publishes a freshness penalty, so the operational reason to stay current is accuracy: an answer engine that quotes your outdated competitor price has you on record being wrong about a competitor.

Track freshness at the row level rather than relying only on a page-level date. Every table row carries its own verified date, because your product’s price and the competitor’s integration count change on different days. Editors should not move a page-level “updated” date unless a row changed. Change the date only when a value changed, and record what changed in the change log the template reserves:

| Date verified | Row | Old value | New value | Source |
| --- | --- | --- | --- | --- |
| [YYYY-MM-DD] | [Criterion (unit)] | [old] | [new] | [competitor release note or pricing page URL] |

Tie the cadence to the competitor’s output rather than to your calendar:

  • Monitor changes: Subscribe to their release notes and status page, then diff their pricing page on a schedule.
  • Verify the data: Re-check every affected row the week a change lands and run a full-table pass quarterly, because some changes arrive without a release note.

Do Reddit and review sites beat your own vs page in AI answers?

By volume, yes. Across 9 million AI answers covering 400-plus enterprise brands, 82% of citations on bottom-of-funnel prompts went to third parties and owned pages received 3%. Your vs page is competing for a minority share from the start, and the share varies by engine: in the ATS matchup test, comparison pages and listicles were 89.4% of sources on Gemini and 72% on Google AI Overviews, but only 3.3% of ChatGPT’s visible inline sources, where product, docs and help pages took 77.3%.

Treat off-page validation as a parallel channel that carries the same facts:

  • Review profiles: Put the dated, unit-labelled claims from your vs page on your G2 and TrustRadius profiles.
  • Community and partner pages: Use the same facts in Reddit threads about the exact matchup and in relevant partner roundups.

When a model triangulates, agreement across sources gives it corroborating evidence. A competitor price that matches on your page and on a review platform has support from more than one source.

Use the engine split to allocate effort. For Gemini, AI Overviews and Perplexity, prioritize the rebuilt vs page. For ChatGPT, the ATS test gives more weight to product documentation and help-center content, so the per-product definition blocks on the vs page should link to the doc page that states the same limit. That doc page should state it in the same words.

How do you measure citation share across ChatGPT, Perplexity, Claude, and AI Overviews?

Build the prompt panel before you rebuild the page, or you will have no baseline. Use 30 to 50 tagged core prompts across the journey, covering branded intent and non-branded or competitor intent. Run each prompt 3 to 5 times within 24 to 72 hours in a clean logged-out session. Record the platform and market alongside location, language, account state and model version. For a vs page, weight the panel toward “[A] vs [B]”, “[A] alternatives”, “is [A] or [B] better for [use case]”, and “[B] pricing compared to [A]”.

Log one row per run per engine. Each row records:

  • The date and engine, including the model version
  • The prompt and its intent tag
  • Whether your URL was cited
  • Every cited URL and competitor URL
  • Which product the answer recommended
  • A screenshot

Citation share for an engine in a month is the number of runs where your URL appeared in the citations divided by the total runs on that engine. Track recommendation share as a separate column, because the Lily Ray data earlier shows the two diverge.

Expect the before-and-after to be slower and smaller than the case-study genre suggests. Treat the template as retrieval infrastructure rather than proof that formatting alone earns citations. Run the panel monthly, hold the prompt set fixed for at least two cycles, and read a change only when it holds across the repeated runs.

How do you put this into practice this month?

Rebuild one page in two stages:

  • Baseline and rebuild: Start with the single vs page whose competitor prompts your sales team hears most, and baseline it with the 30–50 prompt panel across the four engines before you change a word. Then rebuild it on the template with a verdict naming both products and a flat table containing one unit-labelled fact per cell. Add a 40–60 word opening passage under each criterion H2 and write the competitor’s who-should-pick block as carefully as your own. Include the “How we compared” disclosure, link every competitor value to the competitor’s own documentation with the date you checked it, and fill the change log with the first entry.
  • Publish and measure: Generate the ItemList and Product schema from the same data the table renders from, then strip any FAQPage markup. Confirm robots.txt allows the named search bots, including OAI-SearchBot and PerplexityBot, as well as Claude-SearchBot. Re-run the panel a month later and then quarterly. Re-verify rows the week a competitor ships a change. Budget for one test or measurement the competitor’s page cannot offer.

Keep the rebuilt page and the prompt panel in a shared, versioned workspace so the monthly rerun uses the same inputs. For a weekly working example of this kind of AEO measurement, subscribe to the newsletter.

Frequently Asked Questions

Related Content