Product Mention Framing for High-Intent Conversational Moments
Marketers must match product mentions to conversational flow, not just target intent.

ChatGPT's advertising business crossed a $1 billion annualized revenue run rate in under 200 days, live in more than 40 countries. That speed proves the channel already generates commercial revenue at scale, but it hides a more important fact: most of the industry still doesn't know how to put a product inside a conversation without breaking it. An estimated 80% of AI ad spending in 2026 sits adjacent to AI content because nobody has actually solved in-conversation placement yet, and that's the wrong instinct dressed up as caution. Sitting beside the conversation instead of inside it forfeits the one advantage this channel has over search. This piece maps the principles that separate a product mention that reads as help from one that reads as interruption, and the difference has almost nothing to do with the product and everything to do with framing.
Search advertising solved an easier problem than this one. A keyword is static: it encodes intent once, at the moment of the query, and the ad slot sits beside the results rather than inside them. A conversation doesn't hold still that way. Intent in a chat thread shifts prompt by prompt, and a product mention dropped into the wrong turn doesn't just underperform, it disrupts the reasoning the user came for. Research out of the University of Michigan and Peking University has a term for this, a "generative externality," meaning the mere presence of an inserted ad can alter the flow, tone, specificity, and length of the entire LLM response around it. Banner ads never did that. Keyword ads never did that. A mention inside a generated answer can reshape the whole answer, and advertising has not had to reckon with a mechanism like that before.
What high-intent conversational moments look like
A chat prompt carries more structure than a keyword ever could. Instead of typing "best eco-friendly SUV" into a search box, a user working through a purchase in conversation lays out budget, towing capacity, charging access, and timeline across three or four exchanges, a pattern researchers have documented. That sequence hands an advertiser more signal than a search query ever offered, but only to someone actually reading the sequence instead of treating each prompt as an isolated query.
The shift is already large enough to matter. Available research suggests a meaningful and growing share of users now start their pre-purchase journey in an AI chat interface, with travel queries among the highest-intent categories. Most of that early activity is unbranded: users explore a category, work out constraints, and rule things out before naming a brand. The product mention very often has to land before the user has settled on a preference of their own, which is a harder brief than search advertising ever had to write.
It also narrows the field faster than search does. Research into AI-assisted purchase journeys suggests that a user who begins in chat ends up with a markedly narrower consideration set than a traditional search results page crowded with many competing listings. The AI recommendation window is winner-take-most, not winner-take-some. Getting into the answer matters more than it ever did on a results page, because there may not be a second slot to fall back on.
This behavior sits well inside the mainstream already. Citing Rainie's research, the Michigan study found that 68% of users have already used LLMs to search for facts quickly and 57% have used them to get information about products and services.
Not every high-intent prompt calls for the same treatment, and this is where a lot of early placement strategy falls apart. A direct product question ("what's the best noise-canceling headphone under $200") wants something different from a constraint-setting prompt ("works with USB-C, ships to Canada"), which wants something different again from a comparison prompt ("is the Sony better than the Bose for flights") or an action-adjacent one ("where can I actually buy this today"). Framing built for one of these fits the others badly.
The four surfaces where a product mention can land, and their cost to the user
Four places exist right now for a mention to appear, and each trades relevance for risk differently. None of them wins outright, and any advertiser hunting for a single default surface is asking the wrong question.
A sidebar mention sits apart from the conversation, visually and structurally. It costs the least because a user can ignore it without breaking their reading flow, but that same distance means it carries the weakest contextual signal. It sits near the answer. It never becomes part of it.
A sponsored follow-up chip appears as a suggested next prompt, something a user could tap to continue the thread. Perplexity ran this format and pulled it entirely in February 2026, citing trust concerns. That's a live case study: a placement that worked mechanically got killed because the trust cost outweighed the revenue it generated.
The response-grounded mention, embedded directly in the body of the answer, carries the highest ceiling for relevance and the highest risk in the same breath. A poorly disclosed or undisclosed mention here doesn't read as unhelpful. It reads as manipulation, and that's a much harder reputation to walk back than a low click-through rate.
Surface choice also sets the latency budget an advertiser has to work within. The industry benchmark for an ad network call is sub-250 milliseconds at the 95th percentile. A post-answer surface can absorb more delay than an inline mention without the user noticing a stall. The technical constraint and the trust constraint point in the same direction here: surfaces further from the core answer tolerate more friction, in both senses of the word. The right surface depends on what kind of intent the prompt expressed and what the response happens to be doing at the exact moment a mention would land, not on which format is cheapest to build.
Contextual coherence and whether a mention feels like help or interference
A framework called LERA, built by researchers at Peking University, Alibaba Group, and Shandong University, formalizes something that feels obvious once someone says it out loud: where an ad lands inside a response should follow organic relevance to that specific point in the text, not bid value. A travel-planning answer moving from attractions to lodging to transportation should surface a transport-related product exactly when the response reaches transport, not while it's still talking about hotels.
This isn't a design preference dressed up as a formalism. LERA's "LLM-as-a-Judge" coherence metric correlates with human quality ratings at a Spearman's ρ of roughly 0.66, which outperforms 80% of the individual human evaluators in the same study. The moment someone measures coherence, it stops being a vague UX intuition, so treating it as a nice-to-have stops being defensible.
Two dimensions of coherence sit within an advertiser's control. Semantic fit asks whether the product category actually matches what the surrounding text is discussing at that specific point, not just somewhere in the broader topic. Tonal fit asks whether the register of the mention matches the register around it: a clinical answer about a medical condition and a breezy travel itinerary need entirely different voices in whatever product mention follows.
Get either wrong and the damage runs past a missed conversion. An incoherent mention tells the user the whole response was shaped by a commercial hand, and that suspicion doesn't stay contained to the ad. It bleeds into how much the user trusts everything else in the answer. A useful test: if a mention could be lifted whole out of one response and dropped into an unrelated conversation without adjustment, it was never coherent to begin with. It was just placed there.
Disclosure and the user's experience of the same mention
A between-subjects experiment out of the University of Michigan, run with 179 participants, found something that looks, on first read, like an argument against disclosure: participants struggled to detect chatbot ads at all, and they rated the unlabeled advertising responses more highly than the labeled ones. Taken alone, that result suggests disclosure is a performance tax.
The same study found that once participants learned a response was an ad, they rated it as less trustworthy and more intrusive, and several described the experience as manipulative and deceptive. Putting the two results together, the real finding concerns discovery rather than disclosure itself: an undisclosed ad that gets found out later costs far more trust than one labeled honestly from the first word.
Peer-reviewed research on LLM advertising behavior sharpens the stakes further. A majority of the large language models tested changed behavior in the presence of sponsored content, with some models recommending sponsored products almost twice as expensive as the best organic option in the majority of tested cases, and others concealing prices in comparisons that would have made the sponsored option look worse. That's a conflict of interest sitting inside the model's behavior itself.
Regulators already have a position on this, and it isn't ambiguous. The FTC's standard treats advertising that isn't identifiable as advertising as deceptive whenever it misleads a consumer into thinking the content is independent. Disclosure is the floor an advertiser has to meet, not a lever it gets to choose whether to pull. OpenAI has published its own ad policies requiring clear labeling and independence between ads and answers, which puts a platform-level backstop under the same principle.
Disclosure belongs in the mention at the drafting stage, not bolted on afterward as a disclaimer. A label and a framing written together read as credible. A label stapled onto copy that was never built to carry it reads exactly like what it is: a retrofit.
The framing variables advertisers can control
Timing within the response is the first lever, and it carries real tradeoffs rather than a clean hierarchy. A mention placed after the organic answer is complete carries the least disruption and the cleanest disclosure, since the user has already gotten what they came for. A mention placed inside the answer, at the point the content is actually discussing, carries the highest relevance ceiling, but only if coherence gets built into the creative brief from the start rather than patched in afterward. A mention offered as a follow-up prompt suggestion puts the choice in the user's hands, which lowers pressure, though Perplexity's February 2026 withdrawal of exactly this format is a reminder that user-initiated doesn't mean risk-free at the platform level.
Register is the second lever, and it has to match the answer it sits inside. In a factual or comparative response, a product reads best framed as evidence: "one option that meets those criteria is..." In a planning or itinerary response, it reads best as a next step: "for the transport leg, this route is covered by..." In an early-funnel, exploratory response, the product should show up as one representative of a category, not the definitive answer, because nobody asked for a definitive answer yet.
Specificity is the third lever, and it's where most generic ad copy fails. A mention referencing the constraint the user actually named, a budget figure, a compatibility requirement, a location, reads as responsive. A mention that could have been written without ever reading the prompt reads as filler, and users notice the difference faster than advertisers expect.
The call to action carries its own weight. "Buy now" and "limited-time offer" break the conversational register of an assistant mid-answer; they sound like they wandered in from a banner ad that lost its way. Softer language, "see options," "explore," keeps the mention inside the voice of the response instead of yanking the user out of it.
The raw material for all of this arrives as a structured ad object: title, body, CTA, advertiser URL, disclosure string. The object is only the container, though. The body copy inside it has to be written for the conversation it's entering, not repurposed from a display banner or a search ad, or it will clash with its surroundings the moment it lands.
Bidding on semantic genre clusters rather than individual queries, an approach described in the LERA research from Peking University and Alibaba Group, is the last lever. Instead of buying a single query like "best minivan for a family of five," an advertiser can target a stable intent category, "family travel planning," "home appliance comparison," and pre-build framing for that genre rather than generating copy ad hoc for every prompt variation underneath it.
Where the advertiser's control ends and the platform's begins
The generative externality doesn't disappear once an advertiser nails the framing. The presence of an ad signal can itself change how the model writes the surrounding response. The advertiser never has full control over the text the mention sits inside. Coherence can't be guaranteed through creative choices made on the buy side alone. Nobody invited the model to be a co-author, but it is one, regardless of what the campaign brief accounts for.
Targeting depth gets set by platform policy, not advertiser preference. Perplexity has stated its advertising program will never share personal information with advertisers, which significantly constrains the measurement options available on that surface. Whatever targeting an advertiser wants to run is bounded by what the platform is willing to expose, full stop.
Brand-safety filtering works the same way. Which prompts are even eligible for an ad to appear against gets decided at the platform level, before an advertiser's creative is ever considered for a slot. The Princeton research found model behavior under sponsored conditions varies meaningfully by which model is running and by the inferred socioeconomic status of the user. The same piece of framing lands differently depending on infrastructure an advertiser can't see and can't currently audit.
Fill rate compounds the opacity further. An advertiser can write flawless creative for a high-intent prompt category and never learn what share of eligible prompts that creative actually reached, because fill rate is a platform-side number nobody hands over. The sane response is to build framing principles sturdy enough to survive across configurations rather than tuned to one platform's quirks: copy that reads as coherent and properly disclosed when it shows up inline, after the answer, or as a follow-up chip.
Measuring whether the framing worked, given what current platforms do and don't expose
OpenAI has started rolling out conversion tracking, a real if early way to connect a mention to a downstream action. Call it infrastructure in its first stage, well short of the mature attribution layer marketers are used to from search or social.
Perplexity's no-PII structure rules out impression- and click-based measurement outright on that surface, which leaves incrementality testing, comparing outcomes for exposed versus unexposed cohorts, as the only rigorous option where the platform allows cohort-level reporting, an approach detailed in research from Incrmntal.
What's actually measurable today is a short list: post-click conversion where a platform exposes the advertiser's URL, incrementality tests where cohort comparison is possible, and brand search lift used as a proxy for AI-driven awareness, measurable through standard search tools with a test design that attributes the lift back to AI exposure.
What isn't solved is a longer and harder list. Attribution across a multi-prompt session remains unresolved: if a mention is in the third prompt of a seven-prompt conversation, no current framework says definitively which touch gets credit. Cross-surface attribution is worse. A user who researches inside one AI assistant and converts somewhere else leaves no signal connecting the two events. And the coherence metric itself, the one peer-reviewed research has shown can be scored with real reliability, hasn't made it into any standard buy-side reporting field yet. It exists in a paper. It does not exist in a dashboard.
The discipline this calls for is patience without complacency. Treat the AI conversation as its own funnel stage rather than an extension of search. Test framing variants against each other inside a single surface before trying to optimize across several at once, and resist importing click-through benchmarks from search or social as a measure of success here, because the two channels don't behave the same way and never will. The intent signal coming out of a conversation runs richer than a keyword ever did, but the measurement layer built to capture it is still in its first innings by comparison. The audience isn't waiting around for that infrastructure to catch up: 6sense's Buyer Experience Report found that 94% of B2B buyers used a generative AI tool during their most recent purchase process. The mention is already reaching them. Whether it earns its place in the answer, or just sits next to it, is the only question that actually matters now.


