Est.

Prompt-Level Signal Data in AI Ad Campaign Reports

Conversational AI ads need new reporting metrics built on session context, not keywords.

Staff Writer, Creative & Workflow · · 10 min read
Cover illustration for “Prompt-Level Signal Data in AI Ad Campaign Reports”
Measurement & Reporting · October 8, 2026 · 10 min read · 2,311 words

A keyword report tells you someone typed "college dorm essentials." A prompt-level signal tells you a person spent four turns explaining a move-in date, a modest budget, a roommate situation, and a preference for items that ship in under a week. It is a different kind of data, describing a different kind of interaction, and campaign reports now have to account for it. Prompt-level signals, the intent that surfaces turn by turn inside an AI conversation, mark a new reporting primitive because the unit being targeted has changed: it's a conversation, not a page, a keyword, or a stored user profile.

Legacy advertising reporting was built around surfaces that hold still long enough to be measured once: a search results page, a social feed, a landing page with a pixel on it. Large language model ad systems remove the page, the feed, the keyword, and the cookie at the same time, so none of the old measurement assumptions survive on their own, and a report built for one no longer applies to what's left. In their place sits conversational context: an assistant matches an ad to the meaning of a live exchange, not to a string a user typed into a box or a profile assembled from past browsing. Contextual targeting in ChatGPT ads works on conversation-level understanding rather than page-level content matching, so the targeting system has to read the current message against everything that came before it in that session. Microsoft Copilot uses the shape of an entire session, not one query, to decide which advertisers belong in the conversation.

The consequence for anyone reading a campaign report is direct. The number in front of you no longer answers "which keyword matched." It answers "what was this person trying to accomplish, and the report requires a different kind of attention to say so clearly.

The Conversational Context Hierarchy

Conversational signals build in layers, and a campaign report only makes sense once you can place a number inside the right layer. Three stages matter: the opening query, the progression that follows it, and the arc of the session as a whole.

The opening query sets the starting point. It establishes whether the person is browsing broadly or asking something specific, and whether the intent behind it is informational or closer to a purchase. A system reads this first message to judge first-impression relevance, the same way a search engine would read a typed query, except here it's one message inside a conversation that is expected to continue.

Progression is where the conversation actually moves. An exchange that opens with a restaurant recommendation can shift into trip planning, then into a budget conversation, then into questions about getting from the airport to the hotel. Each of those phases is a separate advertising opportunity, tied to a different stage of the same underlying decision, and a report that only logs the original query misses the turns where the person's actual need became clear.

Session arc ties the whole exchange together. Microsoft Copilot triggers ads based on the full session, not a single prompt, and a separate feature called "ad voice" introduces each sponsored block by explaining how it connects to the conversation that produced it, drawing on more than just the most recent message. That means an impression's value in a report has to be read against the direction the whole session was heading. A system can also pick up on a problem a user is describing at length even if the person never names the product category they eventually need, because the pattern in the conversation itself carries the signal. And because relevance comes from the live conversation rather than a stored browsing history, each session functions as its own self-contained targeting context, so reports describe what's happening in the moment, not an accumulated profile built over months.

A metric labeled "relevant impression" in a conversational AI platform is not the same thing as a "relevant impression" in a search report. One is a keyword match. The other is a semantic fit built across several turns of a live exchange, and treating the two labels as interchangeable will lead an advertiser to compare numbers that were never measuring the same thing.

What Each Live Platform Surfaces and Withholds in Its Reporting Today

How rich these conversational signals actually are and how much of that richness appears in a report differs a great deal from one platform to the next right now.

OpenAI's self-serve Ads Manager opened on May 5, 2026, and gives advertisers three objectives to choose from: Reach, billed on a CPM basis, Clicks, billed on CPC, and Conversions, billed on oCPC. The metrics available today are CPM and CPC, standard delivery numbers that describe how an ad was served and how often it was clicked, without describing the prompt that triggered it. Advertisers are already asking for more: which prompt types trigger a placement, how often those prompts occur, what time of day they cluster around, which regions they come from, how they shift seasonally, and how competitive a given query type is. Whether OpenAI builds that level of detail into reporting is still an open question. One constraint belongs in every report read from this platform: ads reach only logged-in adult users on the Free and Go tiers, while Plus, Pro, Business, Enterprise, and Education remain ad-free. Any impression share pulled from an OpenAI report has to be read against that ceiling, a detail that matters enormously for B2B advertisers whose buyers are disproportionately likely to sit on a paid tier.

Google's AI Overviews and AI Mode work differently still. Ads serve automatically from campaigns that are already eligible, meaning broad match or AI Max, Performance Max, or Shopping, with AI Mode placements specifically requiring Performance Max or AI Max, and smart bidding recommended though not required. There's no opt-in or opt-out control for any of it. More importantly for reporting, there's no way to separate what happened inside an AI surface from the rest of a campaign's performance. Impressions and clicks generated inside AI Overviews or AI Mode sit blended into standard reporting, with no segmentation available. A team can be running ads inside AI-generated answers right now with no visibility into which conversations produced those placements.

Microsoft Copilot currently offers the fullest picture of the three. All eligible campaign and ad types opt in automatically, and the available metrics run wide: impressions, clicks, CTR, conversions, conversion rate, CPC, and ROAS. Microsoft still blends these numbers with overall campaign performance rather than isolating Copilot specifically, but the metric set itself is the most complete of any live AI ad surface today. Because Copilot's "ad voice" targeting draws on the whole session rather than a single query, a conversion or CPA figure coming out of a Copilot report reflects a richer contextual match behind it than the same figure would carry from a standard search campaign.

Claude carries no advertising. Anthropic has stated that Claude will remain ad-free, with no sponsored links and no advertiser influence over responses, so it isn't part of this reporting conversation in any form.

Put the three live surfaces side by side: visibility into conversational signals varies enormously depending on where an ad ran, so a marketer reading a number needs to know which platform generated it before drawing any conclusion from it. Google's gap is the sharpest one: a brand can be generating impressions inside AI Overviews right now with no segmented way to see it.

Why the standard metrics, CPM, CPC, CPA, mean something different when the targeting unit is a conversation

Every metric in a campaign report carries an implicit story about what caused the number to look the way it does, a story tied to the targeting unit, whether conversation, keyword, or page. Reading a CPM or a CPA from an AI ad surface the same way you'd read one from a search campaign leads to optimization decisions built on a false comparison.

In keyword search, a CPC describes a user who typed a specific string, clicked a result, and landed on a page. The intent behind that click is narrow, and the click itself anchors attribution: everything measured downstream gets credited back to that single action. In conversational AI, a click, where one exists at all, happens after a multi-turn exchange in which the person has already revealed most of what they want. The click sits near the end of a decision process.

That compression changes what a cost-per-click comparison across channels actually tells you. Microsoft Copilot reports notably stronger CTR and higher conversion rates than traditional search ad placements, but setting that number directly against a search campaign's CTR ignores that a Copilot ad only serves after a session-level intent signal has already formed. The two numbers describe different points in a buying decision, even though they share a label.

Relevance also does work inside an AI ad auction that it doesn't do in keyword bidding: a more relevant ad can beat a higher bid. That means a low CPM inside a conversational AI report can reflect an unusually tight match between the ad and the conversation, not a weak auction with little competition, which is usually what a low CPM signals in search.

Impression share works differently too. A single ChatGPT response typically carries one sponsored card, though multiple cards from one or more advertisers can appear. An impression there means the ad was the only commercial unit sitting next to that answer, a kind of scarcity that has no real equivalent on a results page carrying several ad slots at once. The old instinct, that a click is a click regardless of channel, misses the point: what a click predicts about the conversion that follows depends on the context that preceded it, and a conversational exchange builds far more of that context before the click ever happens than a typed keyword does. Benchmarks built on search performance simply don't transfer.

Where attribution breaks down

Last-click attribution was built for a world where influence happened on pages a pixel could see. It cannot capture influence that happens entirely inside a conversation, and any team running AI ad campaigns without a replacement for it will undercount what the channel is actually doing.

The gap comes from a structural difference in how conversational signals work, not from a lack of better tagging. An assistant can research a category, compare options, and land on a recommendation without the user ever leaving the chat window, so the influence happens in a place no landing-page pixel ever fires. A meaningful share of what moves a purchase now happens inside that exchange, and a measurement system built around clicks has no way to register it.

The tension between paid and organic placement sharpens the problem further. An ad sits beside an AI-generated answer. A citation sits inside the answer itself. If a brand shows up both as a citation inside the answer and as a sponsored card below it, last-click attribution hands the ad credit for a decision that the citation may have already made. If the organic citation belongs to a competitor and the paid placement belongs to the brand, last-click attribution never registers that the competitor shaped the user's thinking before the brand's ad ever appeared, which makes the scenario run in reverse even worse.

Lift testing is emerging as the response. The method is straightforward in concept: hold out a comparable population that sees no conversational AI placement, expose a test population to the placement, and measure the gap in downstream conversion between the two groups. That delta gets attributed to the channel, capturing influence that never produced a click. Publisher-side data backs up why this matters: visitors arriving from AI platforms tend to convert at higher rates than visitors from traditional search, which suggests the conversation itself pre-qualifies intent before the visitor ever leaves it. That also means the click-through rate sitting in a standard report understates how much the channel is actually contributing.

None of this is settled practice yet. Lift testing is the right direction, but the methodology behind it, how long a holdout period should run, how to isolate conversational exposure cleanly, how to size a test population for a channel this new, is still being worked out across the industry. Teams building these measurement frameworks now are ahead of where the field currently stands, not following an established playbook. That gap will close as more advertisers run these tests, and the teams with a working baseline in place before the channel matures will be the ones able to tell, with real data, whether it was worth the investment.

How to read a prompt-level signal report in practice

Reading one of these reports starts with a different question than the one search marketers are used to asking. Instead of "which keyword triggered this," the question becomes: what was this person trying to accomplish, and how far along were they in that process when the ad appeared?

Answering that starts with locating the conversational stage behind the number. An impression triggered at the opening query reflects early, exploratory intent, someone still gathering information before narrowing toward a decision. An impression that appears after several follow-up turns reflects someone who has already refined their problem and moved closer to acting on it, which makes it the higher-value signal of the two, even if both get logged under the same metric label. A report that aggregates every impression into a single number without any indication of where in the conversation it landed flattens this distinction, so the first thing worth checking on any platform is whether early-session and late-session placements are broken out separately. Where that staging data exists, it tells you which numbers deserve the budget; where it doesn't exist yet, its absence tells you how much of this report still has to be read with judgment.

More in Measurement & Reporting