Conversational Engagement Signals as AI Ad Performance Indicators
Conversational exchanges, not clicks, reveal whether an ad truly persuaded a customer.

Click-through rate, cost per mille, and impression count were built to measure discrete, static events: a page loads, an ad appears, a user clicks or doesn't. A conversational AI session does not resolve into any single one of those events, because intent evolves across multiple exchanges, the response itself is generated in real time, and the user's relationship to a product recommendation shifts from one turn to the next.
The clearest evidence of the mismatch is what happens to attribution when a user asks follow-up questions inside a chat, leaves, and converts somewhere else later. Last-click models have no mechanism for crediting the multi-turn dialogue that did the actual persuading, so they undercount the channel by design. Treating those two numbers as equivalent is one of the most common mistakes advertisers make when they bring search-era habits into a conversational environment, and it is a structural error, not a rounding one. Click-based metrics reward the tap itself rather than the quality of the intent that produced it. The measurement system is optimizing for the wrong moment in the interaction.
How conversational AI ad targeting works
Conversational AI advertising does not run on keyword bidding. The primary targeting control is a context hint: a freeform, natural-language description of the kinds of conversations where a given ad belongs, which the system then matches against the live conversation as it unfolds. That matching process has to account for more than the sentence a user just typed. It has to weigh what was asked several exchanges earlier, how the AI responded, what follow-up questions surfaced in between, and where the conversation seems to be heading next.
Microsoft's Copilot shows this architecture. That is a real departure from keyword-match logic, where a single term can trigger a placement regardless of what came before it. In Copilot's system, the entire arc of the conversation determines which advertisers are even eligible to appear.
The reason this works mechanically is that the transformer models underneath ChatGPT already process the full conversation history as context on every inference call, because that history is what lets the model generate a coherent reply. Advertising systems built on top of these models draw on the same contextual representations to judge ad relevance. So when an ad system decides whether a placement fits, it reads the same signal that is shaping the AI's own next sentence.
Research is pushing this further. Both directions assume the same premise: relevance in this environment is a property of the conversation's trajectory, not of a matched term.
That has a direct consequence for measurement. Every conversation with a tool like ChatGPT moves through phases, an opening query that establishes initial intent, then subsequent exchanges that sharpen it into something specific, and a placement that lands in an early phase means something different than the same placement landing after three rounds of narrowing. Because targeting is governed by context rather than keyword match, the behaviors that follow a placement, a follow-up question, a deepened topic, a pivot away, are legible responses to how well the ad fit the conversation. The architecture that places the ad is the same architecture that produces the signal worth reading afterward.
What conversational engagement signals measure
Conversational engagement signals are behaviors that happen inside the dialogue itself, not around it: follow-up questions, topic progression, response depth, and whether the session continues. Each one carries a distinct claim about where a user sits on the path toward a decision, and they do not all carry equal weight.
Follow-up questions immediately after an ad-adjacent response are the strongest signal available. A user who asks how something compares to an alternative, or where to find it, has moved from passive exposure to active inquiry, a shift in kind that no impression count or click can register. Topic progression carries a different kind of information: a conversation that narrows from general research toward specific product or decision criteria tracks movement along a purchase funnel in real time, within a single session. Response depth measures something subtler still. A user who elaborates more in the next turn, volunteering detail rather than giving one-word replies, is co-constructing the conversation.
Session continuation functions as a check on the other three. Whether a user keeps talking after an ad-relevant moment is a reasonable proxy for whether that placement felt native to the conversation or intrusive to it. An abrupt topic change or an ended session is the conversational equivalent of a banner ad getting scrolled past without registering, and it belongs in the same measurement framework as the positive signals, because a framework that only counts the good signals is not actually measuring anything.
What separates these from legacy metrics is that each one reads differently depending on the conversational phase in which it occurs. A follow-up question at the opening of a conversation reveals something different than the same kind of question three exchanges in, after the user's need has already sharpened into specifics. Keywords were never built to capture that distinction, because a keyword match type treats a query as a static object, not a point along a path. The follow-up questions inside a real conversation are a more precise intent signal than any match type a search platform has ever offered.
Why these signals measure intent trajectory
The decisive advantage of these signals is that they are sequential. A click is a binary outcome recorded at one instant: it happened or it didn't. A series of deepening follow-up questions across three conversational turns is a trajectory, and trajectory is what separates a browser from a buyer in a way a single timestamp never could.
Consider how a conversation that opens as a simple restaurant recommendation can evolve into trip planning, then budget discussion, then transportation logistics. Each of those phases represents a distinct intent stage, and a single-click metric would flatten all four into one undifferentiated event, losing the very information that makes the interaction valuable to an advertiser. Meta's systems already operate on this premise at scale: the company's algorithm processes natural-language expressions of purchase intent from more than a billion monthly active users across WhatsApp, Messenger, Instagram, and Facebook, catching moments like a user asking Meta AI to compare running shoes or plan a travel itinerary. That scale is itself evidence that conversational intent signals are not a niche curiosity confined to one chatbot but a pattern recurring across surfaces.
Because these signals derive from the conversation happening right now rather than from a historical profile, they reflect present intent rather than a past behavioral pattern projected forward, which is the thing cookie-based targeting was actually measuring all along.
What the performance data shows with matched conversational context
The available performance data shows an inversion: the metric that looks weaker by legacy standards often marks stronger intent, while the metric that looks strong by legacy standards can mark almost nothing. Microsoft reports that search ads inside Copilot deliver stronger CTR and higher conversion rates than traditional search placements, and Copilot is also the platform where session-level "ad voice" targeting is most developed. That pairing supports the broader argument: richer conversational context produces a better signal, not just a different one.
The same pattern is visible downstream of the click. ChatGPT-referred e-commerce traffic converts at a multiple of what Google organic traffic converts at, because visitors arriving from an AI conversation arrive already pre-sold: they arrived after a recommendation, not before a comparison. That conversion premium is the measurable outcome of a trajectory that built up over several turns, not a side effect of a single well-timed impression. Ahrefs' 2025 analysis found that AI search visitors generated a disproportionately large share of signups relative to their share of total traffic, a conversion edge consistent with the pre-qualification effect of a multi-turn dialogue.
The counterweight matters as much as the headline numbers. Independent agency tests have found B2B campaigns running against meaningful impression counts and producing zero conversions, a result consistent with ads placed against conversations whose intent trajectory never matched the offer. None of this amounts to proof that conversational advertising always outperforms its legacy counterpart. It is evidence, directional rather than definitive, that when context is matched well, the resulting signals carry real information, and when it isn't, performance collapses in ways legacy impression counts would never flag in advance.
How vertical context shapes which signals matter most
Signal weighting is not uniform across categories, and a framework that treats every vertical the same way will misread what it is looking at. The follow-up questions that mark high intent in a travel conversation differ in structure from the ones that mark high intent in healthcare or financial services, so collapsing them into one scoring model gives you a misleading performance read.
In travel, intent is typically inspirational before it becomes transactional. Consumers who visit a travel site after using an AI platform show substantially lower bounce rates than non-AI referrals, reinforcing that the early inspirational phase is doing real work even before a transaction shows up anywhere.
Healthcare and pharma compress that same trajectory into a shorter window.
Financial services shows a third pattern entirely, one built on deliberation. The category's share of ChatGPT ad impressions grew substantially between April and August 2026, suggesting that advertisers are finding the conversational format well suited to complex, consideration-heavy decisions where follow-up questions and topic deepening carry particular weight. The practical implication follows directly: a measurement framework built around one vertical's signal pattern will misread another's, and vertical-specific weighting is what separates an actionable read from a misleading average.
What the attribution gap means
Attribution in conversational AI advertising remains genuinely unsolved. The conversation gap, where a user engages in multi-turn dialogue, leaves, and converts somewhere else later, means last-click models structurally undercount the channel, and no platform has yet delivered a clean, cross-session solution to that problem.
The objection that follows naturally is that signals which cannot be cleanly attributed cannot be acted on. The conversion quality data described above suggests it is not. That reframes the comparison. The choice is not between perfect measurement and imperfect measurement; it is between imprecise measurement of a stronger signal and precise measurement of a weaker one, and the second option only looks safer because it comes with cleaner numbers attached to it.
OpenAI's "user objects" mechanism, introduced with its June 2026 product feed expansion, is an early attempt to close the conversation gap through better conversion matching. It points in the right architectural direction while remaining early in its development, and it should be read as a sign that the platforms themselves recognize the gap.
Transparency raises a related challenge alongside attribution: how ads are labeled. Academic framing around trustworthy commercial intervention points out that when ads are not clearly labeled or contextually appropriate, users attribute the unwanted recommendation to the AI itself, which erodes trust in both the platform and the advertiser. That makes transparent labeling a measurement requirement and not only an ethical one, because signal quality depends on users behaving authentically, and users stop behaving authentically the moment they suspect the conversation has been compromised.
Acting on signals now means leaning on proxy metrics that are already available: session continuation rates, follow-up question frequency, and the on-site behavior of AI-referred traffic. None of these close the attribution gap outright, but each one carries more information about intent quality than CTR alone, and none of them requires waiting for a platform to finish building a cross-session solution first.
Building a measurement practice around conversational signals today
Building a measurement practice for this channel starts with a shift in what gets measured first. Outcome-first measurement asks whether something clicked or converted. Trajectory-first measurement asks what the conversation revealed about where a given user was heading, and you need to instrument different moments and read different data than a legacy dashboard was built to surface to answer that.
Context hint construction is the first lever available to an advertiser, and it determines everything downstream. Prompt and intent mapping has to happen before a campaign launches: identifying the actual questions buyers put to AI tools at each stage of a decision, rather than the keywords they might have typed into a search box, surfaces paid visibility opportunities that keyword research alone was never built to find.
While platform-level attribution matures, three proxies are available right now. Session continuation rates after an ad-adjacent response tell you whether a placement felt native to the conversation. Frequency of follow-up questions within the same topical cluster measures depth of inquiry. Each of these runs meaningfully higher for AI-referred visitors in retail and travel, and each is measurable today without waiting on a platform update.
Holdout testing remains the most reliable incrementality method available: withholding the channel from a matched control group and measuring the resulting difference in conversion rate sidesteps the last-click problem entirely, because it measures channel-level lift rather than trying to trace an individual path through multiple sessions. Creative-level tagging, distinct UTM parameters applied to ad variants by conversational phase, vertical, and context hint, enables signal-level analysis even before platforms expose session-level engagement data natively.
All of this fits into a four-layer stack. Teams that build fluency in reading these signals now, rather than waiting for attribution to be solved for them, will be the ones positioned to act immediately as platforms like Microsoft and OpenAI continue to expose richer session-level data.


