Campaign Launch Workflow for a First AI Ad Buy
Learn the funnel position where conversational AI ads work best and how to write for it.

ChatGPT ads went from zero to a billion-dollar annualized run rate in under 200 days, live in more than 40 countries. That speed is the whole argument for paying attention now, but it says nothing about how to actually spend money there. A marketer who opens a first AI ad campaign with a search or social mindset will pick the wrong surface, write copy that reads like an interruption, and end up staring at numbers that don't mean what they think they mean. The workflow below walks through the sequence that avoids that outcome, from mapping intent to pacing the budget.
The Different Starting Point for AI Ad Campaigns Versus Search or Social
Search built its entire discipline around the keyword. Someone types a phrase, the auction runs, four to seven sponsored results show up around the organic ones, and years of tooling exist to measure what happened next. Conversational AI surfaces don't work that way. LLM surfaces like Microsoft Copilot and Google's AI Mode typically run one or two ad placements per session. That's a formatting difference that means each placement is doing more work, carrying more weight per impression, and getting judged against a much less forgiving bar. It means each placement is doing more work, carrying more weight per impression, and getting judged against a much less forgiving bar.
The funnel position is different too. A search query usually shows up when someone has already narrowed things down and is close to deciding. A conversational prompt appears earlier, when someone is still working through the problem out loud, asking an assistant "what should I look for" or "is this worth it" before they've settled on a shortlist. That's a different moment in the buyer's head, and it calls for a different register of writing.
None of this makes AI ads harder than search or social. It makes them a different medium with its own grammar, and the marketers who do best here will be the ones who learn that grammar first instead of forcing a keyword-shaped peg into a conversation-shaped hole. What follows is the sequence for doing that: define intent, pick surfaces, build creative, set up ways to measure it, then size the test.
Defining conversational intent targets before touching a single ad platform
A keyword captures what someone typed. A conversational intent target captures what they're actually navigating: a question they're stuck on, two options they're weighing, a decision they haven't made yet. That's a meaningfully different unit to plan around, and it has to come first, before any surface gets picked or any budget gets set.
Start from the customer's decision journey rather than the product catalog. Consider the moments where someone would naturally turn to an AI assistant instead of a search bar: "what should I look for in X," "is this brand worth the price," "compare A against B."" Group those moments into three rough buckets, exploration, comparison, and near-purchase, because each one calls for a different creative register and, later, a different surface. Then write out actual sample prompts for each bucket. Those prompts are the targeting hypotheses that get tested once the campaign is live.
This groundwork matters because of how LLM ad targeting actually functions: relevance gets derived from the topic and progression of the live conversation, not from a persistent identifier following someone around the web. An advertiser who has already mapped intent buckets walks into the platform with a coherent structure to test against. One who hasn't is improvising ad groups on the fly. There's a genuine privacy upside buried in this too. Because targeting comes from what's being discussed right now rather than a cookie-based profile built over months, the whole mechanism sidesteps the cross-site tracking apparatus that search and social have leaned on for two decades.
Choosing which AI surfaces to buy and in what order
Two surfaces make sense for a first buy. Microsoft Copilot runs "Compare & Decide" ad formats along with shopping campaigns and a Copilot Checkout flow, sitting inside a Microsoft Advertising network that pulls in more than $20 billion a year. Google's AI Mode, the AI experience built into Search, reaches more than 75 million daily users and has already brought "Direct Offers" and in-AI-Mode checkout live with retail partners including Etsy and Wayfair. Ads inside the standalone Gemini app remain limited for now, which narrows where that budget can currently go.
For a first buy, start with the after-answer inline card. It's the highest-yield format with the lowest cost to the user's experience. It's the format OpenAI itself chose to build around, and it fits commercial-intent prompts, the near-purchase bucket from the intent map, better than any other slot available today.
Structuring creative for in-answer placements
Creative now drives something like 70% of how a campaign performs, which makes it the single biggest lever in this entire workflow, bigger than surface selection, bigger than targeting precision. Get the creative wrong and nothing downstream fixes it.
An in-answer ad is not a banner headline dropped into a chat window. The person reading it is mid-conversation with an assistant, and the ad has to read like a plausible next step in that answer, not like a pop-up that interrupted it. The inline card format itself is fairly fixed, made up of an advertiser name, favicon, a title, short copy, an image, and a link. All of it has to feel like a natural continuation of what the AI just said. Copy should match whichever intent bucket it's targeting. Exploration-stage prompts call for informational copy, something closer to "here's what to weigh," while near-purchase prompts can carry a sharper call to action.
A few rules follow directly from that. Write to the situation the user is in. Keep the copy short, because the AI's answer is the main event and the ad is a contextual add-on, not the headline. The "Sponsored" label is mandatory under OpenAI's ad policies, which require clear labeling and separation from the assistant's own organic response, so copy needs to hold up under that transparency rather than trying to blur the line between ad and answer. And avoid direct-response urgency, phrases like "limited time" or "act now," in exploration-stage placements specifically. That mismatch in register does more damage to trust than skipping the ad entirely would.
Going into a first test, prepare more than one execution: at least one visual variant and at least two distinct copy angles. Nobody has a proven formula for this format yet, so testing more than one direction from the start isn't optional; it's how the second wave gets smarter.
Setting up measurement proxies before the campaign goes live
Attribution here isn't solved, and pretending otherwise wastes a test. The channel is too new for the measurement infrastructure search and social spent years building, and that gap needs to be named before launch, not discovered afterward.
Standard metrics break down for a few specific reasons. LLM ads live inside a conversation, not on a page, so last-click attribution misses whatever role that conversation played in the decision. A lot of users move from an AI chat to a brand's site through paths that don't carry a referral at all, typing the URL directly or running a separate search, which severs the chain measurement tools normally rely on. And because ad load is so low compared to search, a first test will generate a small number of impressions; expecting search-sized statistical significance out of an early AI test sets the bar somewhere it can't be met.
A handful of proxies are still worth tracking, including branded search lift during and after the flight, UTM-tagged traffic from any su... Branded search lift during and after the flight is one: a rise in direct and branded search volume suggests the campaign built awareness even without a clean click path. UTM-tagged traffic from any surface that does pass a referral should be isolated and compared against other channels' conversion rates. Engagement depth on the landing page, time on page, pages viewed, matters too, since AI-sourced visitors are arriving from a different point in the funnel than a search click would. A lightweight post-campaign brand recall survey, even run against a small panel, rounds this out and is particularly useful for the awareness-stage bucket, where conversion was never the point.
Before launch, get alignment internally on what a "successful test" actually means. The channel can surface qualified intent, but that does not prove it pays back at scale, and the team needs to agree on that distinction before the numbers come in, not argue about it after. Agree on a learning period too. The platform's own context-matching gets better as it sees more data from the campaign running, so judging performance in week one punishes the channel for still being new.
Sizing and pacing a test budget for a first AI ad buy
Size the budget to generate enough impressions to compare creative across the intent buckets, not to chase search-scale reach. With only one or two placements per session, impression volume is structurally capped well below what a search campaign would produce, and that's fine: the value here is precision. Treat the first flight as a research budget rather than a performance budget. What it should produce is a validated intent map and a creative direction.
Don't front-load the spend at launch. Because the system's optimization improves as it accumulates data, early dollars are buying information as much as they're buying impressions. A longer flight at a modest daily rate beats a short burst, since conversational intent builds context over time in a way a one-week blitz can't capture. Hold back roughly a third of the total budget for a second wave, informed by which intent buckets actually performed in the first.
Getting this funded is its own conversation, and it's a leadership one. The IAB's 2026 Outlook found 66% of US ad buyers are increasing their focus on agentic AI for buying and campaign execution this year. This budget ask lands differently than a routine reallocation. Frame it as an innovation line rather than money pulled from search or social. Pulling dollars from a channel with a known ROAS to fund one that's still proving itself is a fight nobody needs to pick before the new channel has earned it.
The timing case is straightforward. eMarketer projects US AI ad spending will hit $68.25 billion by 2030, up sharply from standalone chatbot ad spending of just $0.96 billion in 2026. That gap is the window. CPMs and competition for inventory will only rise as more advertisers move in, so the intent signals available today, before the ad load thickens, are about as clean as they're going to get.
How the full workflow connects
The sequence runs in one direction: map intent, pick a surface, build creative for each intent bucket, put measurement proxies in place, size and pace the budget, then launch. Each step depends on the one before it, and the two most common failure points are skipping the intent map or skipping the measurement setup; both break the chain quietly enough that the damage becomes visible only when the results come in and nobody can explain them.
Once the first test wraps, audit which intent buckets drove the most engaged landing-page traffic. Those become the hypotheses the second wave expands on. Check branded search lift against the pre-campaign baseline; if it moved, that's real evidence of awareness, and it's worth quantifying carefully for whoever signed off on the budget. Note which creative angles pulled higher engagement and carry those forward. And write down what the platform's reporting actually showed, gaps included, because even an incomplete baseline becomes the reference point for every attribution conversation that follows.
Nobody has a finished playbook for this yet. Companies including Adobe, Ford, Target, Audible, and Mazda are among the early movers building theirs in real time, through the same kind of disciplined first tests described here. A marketer who runs one well ends up holding more practical knowledge about this channel than most of the market currently has, and given how fast CPMs tend to climb once a channel proves itself, that head start is worth more now than it will be in a year.


