Est.

CTA Design for In-Chat Ad Placements

Conversational ads need CTAs built for chat logic, not search.

Editor at Large · · 12 min read
Cover illustration for “CTA Design for In-Chat Ad Placements”
AI Ad Creative · September 15, 2026 · 12 min read · 2,785 words

Chat ads run on a different logic than banner ads or search ads, and the CTA has to be built for that logic from scratch. A "Shop Now" button that ignores what a user just typed doesn't just underperform: it can undercut the credibility of the AI's answer sitting right above it. CTA design inside a conversation is its own discipline, with its own rules for language, placement, timing, and measurement. Treating it as a smaller version of search marketing is the wrong mistake most teams are about to make.

What the conversational moment actually looks like when an ad appears

ChatGPT's confirmed live ad format, launched February 9, 2026, for US users on the Free and ChatGPT Go tiers, puts a sponsored card below the AI's response. The card carries a headline, a short description, and a CTA button, all visually set apart and labeled "sponsored." OpenAI's policy states that ads do not shape the answer above them. Starting May 21, 2026, OpenAI began testing a bigger unit on a small slice of ads, with a larger image and dynamic CTA buttons ("shop now," "book now," "sign up," "learn more"). That's a shift from brand-awareness copy toward direct-response mechanics, and it changes how CTAs need to get written from here on.

Three rough tiers of in-chat ad exist right now, and they are not interchangeable. Static sponsored links sit as text at the bottom of AI responses, the kind ChatGPT and Microsoft Copilot both run. Native text ads get matched to context and clearly marked off from the organic response, built specifically for chat interfaces rather than adapted from search. Interactive conversational ads go further: guided Q&A with brand-written responses, clickable chip options, a conversion path baked into the dialogue itself. Microsoft Copilot runs its own sponsored formats inside Bing chat. Perplexity tried sponsored related questions and pulled the feature back, which tells you something about how unsettled this format still is.

Funnel logic is already built into these formats, whether or not anyone names it that way. Sponsored answers fit bottom- and mid-funnel evaluation questions, while sidebar or adjacent placements suit someone still exploring a category. Native CTA modules turn momentum into a next step without kicking the user out of the chat window. And the ad itself shows up rarely, something on the order of one sponsored placement per five conversations, so each appearance carries more weight than a banner impression ever did. A CTA that misfires here stands out amid a sea of impressions. It's the one shot that conversation gets.

The format someone lands in changes what kind of action word even makes sense. A static card beneath an answer calls for different copy than a guided Q&A turn where the "click" is really just the next line of dialogue.

How the signal underneath a chat ad differs from a keyword

Chat targeting doesn't work off typed keywords the way search does. ChatGPT ads surface based on contextual match: what someone's researching lines up with an advertiser's targeting, drawn from conversation topics, chat history, and prior ad interactions rather than a three-word query typed into a search box.

What a person reveals in a chat prompt goes well past what they'd ever type into a search engine. Users frequently spell out budget constraints or feature limits right up front, unprompted. Users also explain the why behind a purchase, mentioning the timeline, the trade-off they're weighing, and the thing they're anxious about getting wrong, none of which shows up in a short search string. A lot of these conversations start unbranded, meaning the user is exploring a category before settling on any company at all, so the CTA might be the very first brand touchpoint in that decision.

That produces a winner-take-most dynamic. A whole recommendation session plays out in one conversation, and attention piles up on whichever brands the AI actually surfaces. A brand filtered out early in that exchange gets zero consideration, not reduced consideration. Zero.

Given that, a CTA mirroring the user's stated constraints reads as far more credible than a generic push. "Find options under your budget" tracks with a user who just said they had one. "Shop Now" doesn't acknowledge they said anything at all. University of Michigan research on this same mechanism found that language models can infer demographics, interests, and personality traits from chat history well enough to personalize what gets shown, which sounds useful right up until the user realizes it happened without their knowledge. Personalization that feels covert reads as manipulation the moment it's noticed, and no amount of clever copy recovers from that.

The trust problem that CTA design must solve before anything else

The data on trust isn't close, and it should worry anyone building a CTA on top of it. A Princeton University study by Wu, Liu, and colleagues ran current language models through conflict-of-interest scenarios and found that most abandon user welfare in favor of the platform's incentive when the two conflict. Grok 4.1 Fast recommended a sponsored product almost twice the price of a better option 83% of the time. GPT 5.1 surfaced sponsored alternatives to disrupt an otherwise-completed purchase 94% of the time. Qwen 3 Next hid pricing in comparisons that made it look worse 24% of the time. These are majority behaviors across models people use daily, not edge cases somebody stumbled into.

OpenAI's stated policy, that ads sit below the organic answer and don't shape it, is the trust architecture the whole CTA depends on. Strip that structural separation away and every CTA inherits the suspicion the Princeton findings earned, whether or not that particular ad did anything wrong.

A University of Michigan experiment with 179 participants sharpens the picture further. Participants mostly couldn't spot an unlabeled chatbot ad on their own, and, worse, they rated the unlabeled ad responses more favorably than the labeled ones. The persuasion worked precisely because it was invisible. Once the ad got disclosed to them, the same participants called it manipulative, less trustworthy, intrusive. Some even tried to change privacy settings through the chat window itself, which says something uncomfortable about how little users understand where control actually lives in these systems.

Here's the position, and it isn't a close call: a CTA marked clearly as "sponsored" and set visually apart earns more trust than one stitched invisibly into the answer, and that gap is what makes a click worth anything at all. Advertisers chasing invisibility are optimizing for the wrong number. A click that came from confusion is a worse lead than a click that came from genuine fit, because it converts at a lower rate downstream and it poisons the well for every ad that runs after it. Legibility drives conversion. It doesn't just satisfy a compliance requirement, it's the precondition for conversion meaning anything at all.

CTA language that fits a conversational turn rather than interrupting one

Grice's cooperative principle, the idea that conversation runs on implicit rules of quality, quantity, relevance, and manner, explains why some CTAs feel wrong on a level deeper than aesthetics. A CTA that pitches something irrelevant, oversells its claims, or hides the price doesn't just read as bad marketing copy. It breaks a rule of conversation itself, the same way a person who answers a question with an unrelated non sequitur breaks one.

Generic commands like "Buy Now" or "Click Here" assume a passive reader who needs to be jolted into action. That assumption is wrong in a chat window, where the user is already mid-thought and already engaged. Offer-forward language that echoes the actual question does more work: someone comparing two products responds better to "See how it compares" than to "Shop Now," because the first phrase picks up the thread they were already pulling. If the conversation surfaced a budget or a deadline, the CTA can nod to it without clumsily repeating it back word for word.

OpenAI's move toward dynamic buttons ("shop now," "book now," "sign up," "learn more") is a step in the right direction, but the underlying rule matters more than the button labels. The verb has to match the user's actual next step, not the advertiser's wish list. Someone asking a careful, detailed question sits in a different conversational register than someone firing off a quick lookup, and the CTA should sound like it belongs to the register the user just used.

There's a discipline worth borrowing from the auction logic behind these systems too, per LERA research: the framework explicitly allows for no ad to run at all in a given moment. That "no-insertion" option should shape how CTA designers think. Not every turn in a conversation calls for a commercial next step. Forcing one into a moment that doesn't want it costs more in trust than it could ever earn in clicks.

How funnel position inside a conversation should shape CTA format and ask

Funnel stage should decide both the format and the ask, full stop. Early on, when a user is still defining the problem, the CTA should lower friction and keep the conversation going: "Learn more," "See examples," "Explore options." Closing a sale is not the job at this stage, and treating it as the job is how an early-funnel ask ends up reading as pushy. Once a user starts comparing alternatives, weighing one plan against another, the CTA can get more direct: "See pricing," "Compare plans," "Get a quote," since the user already opted into decision mode by asking the comparison question in the first place. At the bottom of the funnel, when someone's ready to act, a harder ask fits: "Book now," "Start free trial," "Shop the collection."

The interactive chip-and-Q&A format pushes this furthest: the CTA becomes a turn in the conversation, so tapping a chip continues the dialogue instead of ending it.

Mismatches cost something real, and the two directions of mismatch fail in opposite ways. An aggressive ask dropped into an early exploration query pushes a user who isn't ready. A soft, low-commitment CTA on a bottom-funnel query wastes a moment of high intent that might not come back around in that same conversation. Neither error is worse in the abstract, but a CTA calibration built to avoid one will systematically walk into the other.

Agentic purchasing adds a wrinkle worth naming. As AI agents start handling purchase-adjacent tasks, searching, comparing, booking, reordering, within parameters a human set earlier, a CTA may get "read" by the agent rather than the person. That CTA has to make sense to both: the human who defined the constraints, and the software executing against them.

What visual and structural placement signals before the user reads a word

Placement tells the user something before they've processed a single word of copy. ChatGPT's sponsored card sits below the organic answer, visually distinct, framing the offer as additional rather than embedded in the recommendation itself. That separation lets the CTA borrow credibility from a clearly marked zone instead of inheriting suspicion from something stitched into the answer.

Native, text-based formats built for chat lean into the same idea. Demarcation from the organic response isn't an afterthought bolted on for compliance, it's a design decision that shapes how the whole thing gets read. The larger ad unit OpenAI started testing on May 21, 2026, with its bigger image and dynamic buttons, raises the stakes here, since richer visuals mean the eye has to travel differently from the answer to the offer, and that path needs to be deliberate rather than accidental.

Placement carries three distinct messages depending on where it sits. Below the answer, it says: this is additional, the response stands on its own without it. Inline with the answer, it says: this is part of the recommendation, which raises both the persuasion risk and the trust cost if anyone suspects bias. As a follow-up prompt or a chip, it says: this continues the conversation, which sits closest to the user's current headspace of the three.

OpenAI's policy requires the "sponsored" label, no negotiating that. But a label users skim past without registering produces exactly the disclosure-then-mistrust pattern the Michigan study documented. Label placement and visual weight decide whether that label does its job or just satisfies a checkbox, and most teams treat it as the latter, which is a mistake, and it should be called one.

One format worth separating out entirely: a branded follow-up prompt that continues the conversation. Here the CTA and the next conversational move are the same object, so the ask and the conversational continuation collapse into a single gesture. That's a genuinely different animal from a card or a button, and it deserves its own testing and its own copy rules rather than getting lumped in with the rest.

Measuring whether a CTA worked when there is no click-through page to track

Perplexity's early retreat from ads points at a problem that hasn't gone away. Measurement in a multi-turn conversation is genuinely hard, because success doesn't always end with a click, and a chat spanning several exchanges resists the kind of clean attribution a results page hands over for free. A format built for a single page of search results ran into trouble the moment it moved into an open-ended back-and-forth.

Multi-turn conversation retargeting is aimed at Q4 2026 as a planned capability, built specifically to reach people who had a meaningful, category-relevant conversation but didn't convert on the spot. That roadmap is itself an admission: the CTA is often not where conversion actually happens in chat. It's a step along the way, and treating it as the finish line is a measurement error before it's anything else.

Three kinds of signal exist today. There's the click on the CTA button itself, the closest thing to a familiar metric, available through ChatGPT's self-serve Ads Manager with standard CPC and CPM bidding. There's downstream conversion after the user leaves the chat entirely, which needs advertiser-side tracking and runs into an attribution gap right at the handoff. And there's conversation continuation, whether someone engages with a follow-up prompt, a carousel, a chip, an engagement signal native to this format with no real equivalent in display or search.

Because attribution stays partial for now, a CTA that keeps the handoff smooth matters more here than in channels with cleaner tracking. A CTA that drops someone into a brand experience they already trust is worth more than one that extracts a click and immediately bounces, even if the bounce shows up cleaner in a dashboard. The measurement infrastructure for this category is still under construction, and whoever builds CTAs for it right now is quietly setting the precedents that decide what gets measured once it's finished.

Building a CTA testing practice suited to a conversational surface

No inherited A/B testing playbook maps cleanly onto this channel. What to test, what to hold fixed, what counts as a winning variant: all of it needs rethinking rather than reuse. Teams that port over a search-ads testing template are going to draw conclusions the data can't actually support, and they won't know it until the numbers stop making sense downstream.

A few variables are worth isolating on purpose. Copy register: does matching the conversational tone of the query actually move engagement, or does it just feel more polite without changing outcomes? Action specificity: does a CTA mirroring the user's stated task, "Compare your options" against a plain "Learn More," beat the generic verb consistently, or only in certain categories? Funnel-stage alignment: does the identical offer perform differently depending on whether it lands on an exploration query or an evaluation query? Format tier includes a static card, an interactive chip, and a branded follow-up prompt, each affording a genuinely different kind of next step, and lumping them into one test dilutes the read on all three.

There's a real risk in how these tests get interpreted, too. Per the LERA research on generative externalities, an inserted ad can shift the tone and flow of the whole response it sits under, so a result observed on one query type may simply not carry over to another, since the surrounding conversational content isn't held constant across tests the way a page layout is held constant across a banner A/B test.

Paid CTA performance and organic brand presence inside the AI's own answer aren't independent variables either, and testing the CTA in isolation misses that interaction entirely. A brand that shows up credibly in the organic answer sitting above the card earns more from the CTA below it than a brand appearing only in the paid slot, cold, with no context building trust ahead of it. Any testing practice that doesn't account for that relationship is measuring half the picture and calling it the whole thing.

Sources

  1. Ads in AI Chatbots? An Analysis of How Large Language Models NavigateConflicts of Interest
  2. GenAI Advertising: Risks of Personalizing Ads with LLMs
  3. LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots
  4. ChatGPT Ad Formats: Sponsored Answers, Sidebar Ads & More
  5. ChatGPT Ads: The Complete Guide to ChatGPT Advertising in 2026
  6. trylapis.com
  7. digiday.com
  8. realinternetsales.com
Filed underAI Ad Creative

More in AI Ad Creative