Back to Learn
#AEO

How to measure AI search pipeline and attribution

Capture AI referrer sources in first-party cookies, store them in CRM fields, and reconcile detected vs. declared sessions into auditable pipeline figures.

CRM pipeline view with an AI first-touch source field on a closed-won deal

AI search pipeline attribution is a stamp, store, reconcile procedure. Capture the AI referrer (chatgpt.com, perplexity.ai, gemini.google.com, claude.ai) at first touch in a first-party cookie, write it to custom CRM fields that persist from lead through opportunity to closed-won, add a mandatory self-reported source question on every demo form, then reconcile detected against declared into one multiplier-adjusted pipeline figure.

The symptom is a Direct traffic line that grew faster than any campaign explains. In a 446,405-visit sample, 70.6% of AI-referred visits landed in GA4 as Direct, and the buyers who read a ChatGPT answer without clicking anything left no trace at all. Your CFO is looking at a channel with no name and a pipeline with no source.

What is the short answer: stamp, store, reconcile?

Stamp the source at the first pageview. Configure a script on every page to read document.referrer and source parameter, match them against the four assistant domains, and write the match to a first-party cookie named ai_first_touch. Store the raw referrer with an ISO timestamp and the landing path. Make the cookie write-once so a Google click three days later never overwrites the ChatGPT click that started the journey.

Store it where the deal lives. Configure four hidden inputs (ai_source, ai_referrer_raw, ai_first_touch_at, ai_landing_path) to copy the cookie into every form submission. Set one workflow to write those values onto the contact or lead only when the target is blank. Set a second workflow to copy them onto the deal or opportunity when it is created, allowing the source to remain on the record through closed-won without a rep re-keying it.

Reconcile detected against declared. Add a mandatory “How did you first hear about us?” picklist to every demo form, name the assistants, and store the answer in self_reported_source. At quarter close, count leads with an AI value in ai_source as detected and leads with an AI value in self_reported_source as declared. Then calculate the union. Declared divided by detected is your multiplier, while the union is the pipeline you report and the detected subset is the hard floor.

Why does AI search traffic disappear from your reports?

Referrer stripping erases the click before your analytics ever sees it. The loss happens through four mechanisms:

  • ChatGPT web policy: chatgpt.com serves a Referrer-Policy of strict-origin-when-cross-origin, the browser default, so a click that does pass a referrer passes only https://chatgpt.com/ with no conversation path.
  • Link-level suppression: In-content links for paid ChatGPT accounts add rel="noreferrer", which suppresses the header entirely.
  • Native app boundaries: The native iOS and Android apps for ChatGPT, Perplexity, Gemini and Claude open links across an OS boundary that never sends a Referer. Most of those sessions arrive as (direct) / (none), though a minority of Gemini app sessions route to google / organic via Google App routing instead.
  • Incomplete UTM coverage: OpenAI appends a chatgpt.com source parameter to citation links in ChatGPT Search, but not to conversational inline links or mobile app links, and it never appends a medium parameter.

Campaign tags therefore recover less than the announcement suggested. A session carrying only a source parameter lands in GA4 Unassigned rather than Referral. Perplexity appends nothing, and neither do Gemini or Claude. GA4 native AI Assistant channel only classifies sessions with a detectable referrer, and it files AI Overviews and AI Mode under Organic Search, so Google’s own AI surfaces are inseparable from an ordinary organic click.

Session-level GA4 setup is covered in how to measure AI search attribution in GA4.

The dark funnel is larger than the stripping problem. 51% of B2B software buyers start research with an AI chatbot more often than Google, and chatbots influenced 69% of shortlists in a survey of 1,076 respondents who said they could see chatbots influencing pipeline but could not measure it cleanly. A buyer who asks Perplexity for the three best vendors in your category, reads the answer and types your URL into a browser a week later is an AI-sourced lead your analytics will file as Direct.

How does each engine pass or strip referrer data?

The five surfaces you care about pass different signals, and no two land in the same GA4 bucket. Behaviour by engine as of September 2026:

[@portabletext/react] Unknown block type "table", specify a component for it in the `components.types` prop

Your detection regex has to key on the source string rather than the channel, because chatgpt.com shows up under three channel names. GA4 cannot separate Google AI surfaces from ordinary organic traffic, which makes the self-reported field in Step 3 the only record-level detection in this setup for AI Overviews.

How do you capture the source in a first-party cookie and hidden form fields?

The capture script runs on every page, checks for an existing stamp, and writes one only on the first qualifying visit. Load it in the head so it fires before any consent banner or tag manager rewrites the referrer. Read the document referrer hostname and the campaign source parameter, match them against chatgpt.com, perplexity.ai, gemini.google.com, and claude.ai, and write a JSON cookie named ai_first_touch with source, raw referrer, ISO timestamp, and landing path. First touch wins and never overwrites.

Matching the campaign source parameter against the assistant hostnames is what catches ChatGPT Search clicks that arrive with a chatgpt.com source tag and no referrer. Without it you lose the entire Unassigned slice.

Every form on the site carries four hidden inputs named ai_source, ai_referrer_raw, ai_first_touch_at, and ai_landing_path. On page load, a second script reads the cookie and copies those four values into the inputs before submit.

The script requests a 90-day expiry. In Safari, ITP deletes all script-writable storage after seven days without user interaction, and a UTM-decorated landing from a classified domain can cap the cookie at 24 hours. For an enterprise cycle that runs a quarter or longer, mirror the stamp into a server-set HttpOnly cookie from the same origin and IP subnet as your site, because cookies set from a CNAME-cloaked subdomain get the same seven-day cap. Without this server-set mirror, Safari-heavy executive buyers can show up as Direct even after the script is live.

For the analytics side, build a GA4 custom channel group with one rule ordered above Referral, Unassigned and Direct, matching Source against chatgpt.com, perplexity.ai, gemini.google.com, and claude.ai.

If you match on medium instead of source, the ChatGPT rule catches AI Assistant sessions and misses the Referral and Unassigned slices, so your channel undercounts the one engine that sends most of the volume. Match on the source string and the three fragments collapse into one channel.

The same logic in BigQuery against the GA4 export filters page_view rows where source, manual source, or page referrer matches those four hosts, then counts distinct sessions by date and source.

Neither query recovers the referrer-less sessions. They count what the browser passed, which is why the CRM stamp and the self-reported field carry the rest of the load.

How do you map the fields into HubSpot, Salesforce, or Pipedrive so they survive to closed-won?

Create four custom properties on the person-level object from the four hidden inputs, then configure the CRM to copy those values to the deal-level object at creation. Use the custom properties as the reporting system of record because each CRM places different write, transfer, or editing constraints on its native source fields.

  • HubSpot: Original Traffic Source is set at the first known session and deals mirror the value of the associated contact with the earliest activity, which sounds like the job is done. It is not, because read-only properties such as Original source cannot be selected as the target of a workflow copy, so you cannot correct a Direct stamp or enrich it with the engine name. Create four single-line text contact properties matching the hidden input names. Build a contact-creation workflow that copies ai_source into a first-touch property only when that property is blank, with re-enrollment off. Build a deal-creation workflow that copies the four contact values onto four matching deal properties. Your closed-won report then groups on the deal-level ai_source.
  • Salesforce: The hidden inputs ride the Web-to-Lead form as custom field IDs into custom Lead fields. At conversion Salesforce creates the Account, Contact and Opportunity, and custom fields only reach the Opportunity if you map them under Map Lead Fields. Map all four, then add a Flow on Opportunity creation that sets the first-touch fields only if blank. Account Engagement’s campaign association is set once from the first asset a visitor touches, which gives you a second first-touch signal, but it is editable and campaign-shaped rather than engine-shaped, so keep the custom fields as your reporting key.
  • Pipedrive: source_origin is system-populated and not editable, while source_channel is admin-editable and any edit applies retroactively to every existing record. An admin renaming a channel value rewrites your historical attribution in one click, so never hold the AI stamp there. Create custom fields on Person and Deal, write them through the API or a webhook (the Web Forms documentation does not cover campaign-parameter capture), and use an automation to copy Person fields to the Deal at creation.

You can apply the same field logic across these CRMs with a standard webhook flow. Configure the form submission to post to an endpoint or automation tool, then have that endpoint upsert the person record with the four fields while writing only where blank. Configure the CRM trigger on deal or opportunity creation to copy the fields down. At closed-won, RevOps or finance can group the revenue report on the deal-level ai_source without requiring manual changes upstream.

How do you add self-reported attribution and calculate a declared-to-detected multiplier?

The self-reported field catches the buyer who researched you in an assistant and arrived by typing the URL. Your marketing ops team reads analytics as the record of where demand was captured and the form response as the record of where it was created, so the two numbers measure different things and should not be expected to match.

Make the question mandatory on the demo or contact form and word it as “How did you first hear about us?” with a picklist: ChatGPT, Perplexity, Gemini, Claude or another AI assistant, Google search, LinkedIn, Podcast or newsletter, Colleague or peer referral, and Other with a free-text follow-up. Store the picklist in a self_reported_source property and the free text in self_reported_source_detail. Run the free-text responses through an LLM classifier monthly so “I asked GPT for options” gets coded to ChatGPT rather than Other.

Docebo runs a mandatory field on its demo-booking page and found 12.7% of high-intent leads named AI tools as their discovery source in Q3 2025, five times the Q3 2024 level. The rule Docebo’s setup makes concrete is that the question has to be mandatory and placed on the highest-intent form, because an optional field on a newsletter signup gives you a response rate too thin to reconcile against.

The reconciliation math uses illustrative numbers that you replace with your own:

  • Detected and declared: Suppose your CRM holds 100 demo requests where ai_source is stamped and 240 where self_reported_source names an assistant, with 60 records appearing in both sets.
  • Union: The union is 280 AI-sourced leads.
  • Multiplier and floor: Your declared-to-detected multiplier is 240 divided by 100, or 2.4. If the 100 detected leads produced 20 opportunities at a $50,000 average, the detected floor is $1,000,000 of pipeline.
  • Reported pipeline: If the 280 union leads produced 52 opportunities at the same average, the reported figure is $2,600,000 of AI-sourced pipeline with a $1,000,000 floor and a 2.4 multiplier disclosed alongside it.

Report all three numbers. Your CFO can audit the floor record by record, while the union reflects buyer behaviour. The multiplier tells you how much of the channel your analytics stack is blind to, and it should shrink as you harden the cookie and fix the channel group.

How do you triangulate the rest with branded search, Direct traffic, and prompt visibility?

Zero-click influence never produces a record you can stamp, so you triangulate it from three signals that should move together if citations are working. Treat the three as corroboration, never as a source of truth, and report them as leading indicators next to the stamped pipeline.

Prompt-panel scoring is covered in how to track ChatGPT citations and measure AI search visibility.

  • Branded search lift: Pull weekly branded query impressions from Search Console and plot them against your citation share. A buyer who reads your name in a Perplexity answer searches for you by name, and that search is the only footprint the assistant session leaves in Google’s data.
  • Direct traffic trends: Segment Direct by new users landing on non-homepage URLs, especially comparison, pricing and integration pages. A bookmark rarely lands on a “vendor A vs vendor B” page, so growth in that segment is the assistant traffic the referrer stripping hid.
  • Prompt visibility: Run your buyer prompt set weekly across the engines and track two rates separately, the share of answers where you are mentioned and the share where you are cited with a link. Mentions move branded search, citations move the stamped cookie, and the gap between them is your zero-click exposure.

When all three rise in the same window and the stamped pipeline lags by roughly one sales cycle, treat the result as a pattern consistent with a chatbot-formed consideration set. When citation share rises and the other two stay flat, test prompt selection before changing the content because the tracked prompts may not reflect the prompts your buyers run.

Which attribution model fits AI search, first-touch, multi-touch or mix modelling?

The model choice follows the failure mode of each option:

  • Recommendation: Run record-level first-touch stamping with the declared field patching its blind spot. Every other model should wait until that layer has been live for several quarters. That is a position no ranking page will commit to.
  • First-touch stamping: This model fails on the zero-click journey. It records the engine only when a click happened, and the assistant sessions that never click are invisible to it. Its strength is that every stamped record is auditable. A CFO can open the deal and see the source, while the declared field fills the zero-click gap with a second auditable value on the same record.
  • Multi-touch attribution: Your attribution team uses algorithmic MTA to distribute credit across the touches the system observed, so an AI touch that arrived with no referrer gets zero weight. The team’s model then assigns that credit to whatever organic or direct session came after it. The output looks precise and undercounts the channel you were trying to measure. Enterprise MTA also tends to sit behind the most expensive tier of your marketing platform, so you pay for a model that structurally ignores your question.
  • Marketing mix modelling: MMM needs a variable that moves week to week, and AI search has no spend series or impression feed you control. Citation share on a fixed prompt set can eventually serve as that variable, which is one more reason to run the prompt set on a schedule now, but a regression on eight data points will fit noise. Stamp first, declare second, and give MMM the citation-share series once it is long enough to model.

How do you choose between visibility trackers and revenue attribution platforms?

Most teams buy a citation tracker and then ask it a CRM question it was never built to answer. External tools cover three signal tiers, but your owned attribution layer is what preserves the source through closed-won.

  • Visibility tools: These tools run synthetic prompts against the engines and return mention share, citation share and sentiment. They receive no session or CRM-object data. Use one to run the prompt set in the triangulation section and to build the competitive citation map your board keeps asking for.
  • Demand tools: Your analytics team uses GA4 with the custom channel group, or the BigQuery export, to identify clicks that passed a referrer or a UTM. The team can determine which pages assistant traffic lands on and whether it converts to a form fill. This data does not cover the zero-click majority or anything past the form.
  • Revenue tools: A CRM-connected attribution platform ingests leads, opportunities and deals, then writes attribution properties back. Dreamdata maps AI sources from ChatGPT, Gemini and Perplexity UTMs to an Organic LLM channel and reports influenced MQL, SQL and new business value against HubSpot and Salesforce records. The platform still receives only what the browser passed, so it inherits the same referrer gap as GA4 unless your own stamp feeds the CRM fields it reads.
  • Owned attribution layer: The stamp-store-reconcile layer in Steps 1 through 3 is the revenue tier you own regardless of which vendor you add on top. Buy a visibility tool for the prompt set. Do not expect it to produce a closed-won number, because it has never received a deal record.

Do AI-sourced leads convert better than organic search leads?

The only peer-reviewed evidence says organic search converts better, and it comes from ecommerce rather than B2B. Across 973 ecommerce websites and more than 50,000 LLM-referred transactions, organic search converted 13% higher than organic LLM sessions, and revenue per session from LLM referrals came in below every traditional channel except paid social. Vendor and practitioner datasets mostly point the other way, but they use last-click attribution, define conversion anywhere from a signup to a Stripe payment, and draw on self-selected SMB cohorts, so they cannot be pooled with each other or with the peer-reviewed result.

For B2B pipeline the field publishes nothing usable. No analyst firm has released an AI search versus organic search comparison on win rate, average contract value or sales cycle length, and no study reviewed reports lead-to-opportunity rate by channel. Any number you see for enterprise deal size by AI source is an estimate, and you should label it one in your own reporting.

Build the comparison from your own CRM instead. Once the deal-level ai_source field has been live for a full cycle, group closed-won and closed-lost by ai_source versus organic, and compute three ratios per group: opportunity-to-close rate, median contract value and median days from first touch to close. Report the record counts next to each ratio, mark every figure as a first-party estimate, and rerun it quarterly so the sample grows. Segment the AI group by declared versus detected as well, because the buyer who typed your URL after a Perplexity session may behave differently from the one who clicked a ChatGPT citation.

How do you report AI search pipeline to a CFO or board?

Report three lines and one leading indicator, and state the blind spots in the same slide. The three lines are detected AI-sourced pipeline (the auditable floor), declared AI-sourced pipeline (the self-reported layer) and the reconciled union with the multiplier beside it. The leading indicator is citation share on your buyer prompt set, plotted against the pipeline lines with a one-cycle lag. The disclosure states which engines you cannot detect, which today means AI Overviews and AI Mode on every surface and Claude on its desktop and mobile apps.

Citation share is the number you optimise because it is the one that moves first. AEO work changes which chunks the engines retrieve, citation share moves within weeks, branded search and non-homepage Direct move next, and stamped pipeline arrives one sales cycle later. Presenting the chain in that order gives the board a sequence they can check against the next quarter rather than a single number they have to take on faith.

The consideration-set framing explains why the influence metrics belong on the slide at all. The Ehrenberg-Bass Institute’s 95-5 heuristic holds that up to 95% of business buyers are not in market for a given category in a given quarter, a figure the authors derive from interpurchase time (a five-year replacement cycle yields about 5% in market per quarter) and describe as a heuristic to be recalculated per category rather than a fixed rule. Combine that with the G2 finding above, that most buyers now form their shortlist inside a chatbot before visiting a vendor site, and the implication for the board is direct. The small in-market slice arrives already shortlisted, and whether you are in that shortlist was decided by citation share months before a form was filled.

Reporting this every quarter means building the stamp, the CRM fields and the reconciliation as a persistent measurement system rather than a one-off analysis, and the same discipline applies to the content that earns the citations in the first place. For more operator notes on this measurement stack, subscribe to the AI-Led Growth newsletter.

Frequently Asked Questions

Related Content