Est.
FeaturesLong read

Budget Pacing Strategies for Conversational AI Ad Campaigns

ChatGPT ads mature faster than search, forcing marketers to rethink budget timing.

Contributing Editor · · 10 min read
Cover illustration for “Budget Pacing Strategies for Conversational AI Ad Campaigns”
Features · September 11, 2026 · 10 min read · 2,251 words

Conversational AI advertising sells something search and social never had to price: a single moment of high intent buried inside a conversation that could end at any turn. That changes what "budget pacing" even means. Instead of smoothing spend across a day of predictable slots, advertisers now have to guess when, inside an unfolding back-and-forth with a chatbot, a user's intent turns sharp enough to be worth an insertion, and that guess has to happen before the platform decides it for you.

How fast this inventory is scaling, and why that makes pacing decisions urgent now

Diagram: Chatbot Ad Revenue: From Zero to $1B Faster Than Any Channel Before It. Visualizes: Visualize a speed comparison showing how fast three platforms reached $1 billion in annualized ad revenue: ChatGPT reached it in under 200 days; Google…

ChatGPT's ad business hit a $1 billion annualized revenue run rate in under 200 days, with tens of thousands of advertisers buying across more than 40 countries, according to recent reporting. Google AdWords took four years to reach the same mark. That gap alone says something about how fast this inventory pool is filling in, and how little time advertisers have to build pacing habits before the channel matures around them.

The user base backing that ad growth is its own story. ChatGPT became the fastest mobile app ever to reach one billion monthly active users, doing it in three years, per Sensor Tower data from May 2026, faster than TikTok or Instagram managed the same climb. Users now submit 2.5 billion prompts a day. That's 2.5 billion discrete moments where intent might be forming, sharpening, or fading, and each one is a candidate for an ad insertion or nothing at all.

Only about 50 million of the 900 million weekly active users pay for a subscription. Advertising is the obvious lever for turning the other 850 million into revenue, which means the inventory pool keeps expanding whether or not advertisers are ready for it. eMarketer puts standalone chatbot ad spending up 1,641% in 2026, reaching $0.96 billion, while AI search-adjacent advertising (think AI Overviews and similar surfaces) grows 152% to $26.42 billion. Those are two separate inventory pools with two separate pacing problems, and more than 80% of 2026's AI ad dollars are landing in the search-adjacent pool rather than inside pure chatbot conversation, a distinction worth holding onto because the underlying auction mechanics differ.

By 2030, eMarketer projects a national market. AI ad spending reaches $68.25 billion. OpenAI, for its part, moved from a pilot phase requiring $200,000 minimum commitments (later cut to $50,000, then dropped entirely) to a fully self-serve platform with no spending floor as of May 2026. Budget size is no longer a barrier to entry. Pacing discipline, unfortunately, still is.

What the auction mechanics actually mean for when and how budget gets spent

Two research frameworks explain why this inventory behaves so differently from a keyword auction. LERA, developed by researchers at Peking University, Alibaba, and Shandong University, runs a two-stage retrieve-then-generate auction: embedding-based filtering narrows the candidate pool, then the LLM itself scores relevance, and that score combines with the bid to pick a winner. The payment rule is built to keep the auction truthful for advertisers trying to maximize utility, much like a second-price auction model keeps bidding honest.

LLM-OSDA, from a team at JD.com, goes a step further. It's a dynamic cost-per-click auction that folds in a Bellman optimal stopping problem, meaning the platform isn't just deciding who wins an ad slot, it's deciding when inside the conversation to show one at all. In testing, LLM-OSDA improved net revenue by 11% over the strongest fixed-timing baseline while holding user retention steady. That amounts to more than a marginal tweak. It means the timing decision now carries as much weight as the bid itself.

Here's the piece with no equivalent in search: LERA's authors call it a "generative externality." An inserted ad doesn't just occupy a slot next to the LLM's answer, it can reshape the tone, length, and specificity of the answer itself. A search results page never changes its own copy because an ad sits above the fold. A chatbot response might.

Practically, this means spend doesn't trickle in at a predictable per-query rate the way search CPC does. A single session might produce zero insertions or exactly one, and which of those happens depends on when, inside that specific conversation, intent crosses whatever threshold the LLM layer is scoring against. Bid level still matters, but it no longer controls spend velocity on its own, because the platform's timing decision sits between your bid and your actual spend. Even at a flat bid, daily totals can come in lumpy, simply because high-intent moments aren't evenly spread across sessions or across the day. One practitioner guide points to "conversation depth optimization," expected to roll out in some form by mid-2026, as a bidding option built around exactly this: targeting multi-turn dialogues that lead to conversions, which turns session length itself into a budget variable rather than just a targeting knob.

Why the standard pacing playbook breaks down in this environment

Conventional pacing math is simple: spend a linear share of the monthly budget each day (a thirtieth of it in a 30-day month, say), then adjust for known seasonality or day-of-week patterns. Improvado and most pacing tools built for search and social run on some version of that logic, and it works fine when impression supply is roughly stable and roughly predictable across the day.

That assumption doesn't hold here. Conversational inventory doesn't track a keyword demand curve, it tracks when people start and sustain multi-turn research sessions, and that can spike around a news event, a product launch, or some unpredictable cultural moment with no warning built into a day-of-week multiplier. Research on paid campaigns has shown week-over-week spend variance as high as 37% for unmonitored accounts, versus under 8% for accounts under active monitoring. That gap alone shows how fast pacing drifts without intervention, and conversational inventory is structurally more volatile than the search campaigns that produced those numbers in the first place.

Google's own May 2026 shift toward demand-led pacing, where AI shifts spend toward high-demand days and pulls back on slower ones without breaching a monthly cap, points at the right general direction. But that system was tuned on search demand signals. Conversational intent doesn't move the same way, and a recent change to Google's ad scheduling rules showed a 2 to 3 times higher overspend risk for advertisers relying on old scheduling logic, a useful reminder that pacing systems built for one kind of inventory can misfire badly when pointed at another.

Two failure modes show up specifically in conversational campaigns. Front-loading burn happens when budget runs out before the late-session moments that actually carry intent, since the strongest purchase signals tend to show up later in a conversation rather than at the start. Underspend waste happens in the opposite direction: a conservative daily cap treats a day with few early-turn insertions as slow and shuts off spend, missing the intent that was still building later in those same sessions.

Pacing principles built around when intent matures inside a conversation

Diagram: Where Intent Lives in a Conversation: The Case for Back-Weighted Spend. Visualizes: Show a single conversation arc divided into three stages — Early turns (exploratory: user setting context, broad questions), Mid turns (evaluative…

Budget deployment needs to track intent maturity, not clock time and not raw impression counts. Early turns in a conversation tend to be exploratory: the user is still setting context, asking broad questions, feeling out the topic. Mid-to-late turns tend to be evaluative: comparing options, asking for a recommendation, narrowing toward a decision. LLM-OSDA makes this explicit in its design. Inserting an ad too early wastes it on ambiguous intent, and the model's optimal stopping point only arrives once estimated click quality crosses a set threshold. Pacing logic should lean toward sessions that have reached that evaluative depth, not sessions that simply exist.

That points to a shift in the unit of account itself. Because LLM-OSDA's framework treats the session as the unit within which a native insertion opportunity is evaluated, the session, not the impression, is the more relevant thing to budget against. Daily targets should be framed as expected session volume at a given intent depth, not a raw impression number, which means advertisers need some read on the distribution of session lengths in their target inventory, a figure worth asking publishers or DSP partners for directly.

Time-of-day curves borrowed from search don't transfer cleanly either. Lowering bids overnight and raising them during a commute window assumes a demand pattern tied to when people are awake and searching. Conversational intent follows its own contextual rhythm: financial planning questions might cluster on weekday evenings, travel questions might spike on weekend mornings, but those patterns have to come from actual conversational data, not from a benchmark built for keyword search.

Early spend is best treated as tuition rather than performance. One practitioner guide suggests setting aside $2,000 to $5,000 purely as a learning budget in the first phase of a campaign. That money's job is to map the intent distribution of the actual audience: which turn depth produces the highest click quality, which conversation topics convert, whether an inline card or a branded follow-up prompt draws more engagement. Pacing during that phase should aim to cover as much session variety as possible, not to chase a low CPM or hit a ROAS number early.

Bid strategy and pacing are also more tangled together here than in search. Because LLM-OSDA treats timing as endogenous to the auction, a higher bid doesn't just improve odds of winning, it can also influence when during the conversation an insertion occurs. That relationship is non-linear, and it needs watching before anyone scales a bid change with confidence.

How to set budget controls that account for conversational inventory's lumpiness

Daily caps need real slack built in. A common alert formula from search pacing flags an account once cumulative spend passes target daily spend times days elapsed times 1.15. That's a fine floor to start from, but conversational inventory calls for a wider band given how uneven its structure is: a quiet day of few insertions can be followed by a high-intent day that burns through disproportionate budget if the cap assumed even distribution across the week.

Monthly caps make more sense here than rigid daily ones. A demand-led pacing model that spends more on peak days and less on slow ones inside a monthly ceiling is directionally right for conversational inventory too, even though existing systems weren't built with this channel in mind. The same logic applies: set the monthly number firmly, treat daily caps as soft guardrails with headroom (maybe double the average daily target on days that look intent-heavy), and check performance weekly rather than obsessing daily.

Early campaigns need tighter monitoring than mature ones. Four things generally decide how often to check pacing: budget size, bidding strategy, account maturity, and auction volatility. Conversational AI campaigns score high on volatility and low on maturity almost by definition, which argues for frequent, even daily, review. The earlier figure bears repeating here: unmonitored campaigns saw 37% week-over-week variance against under 8% for monitored ones. That's the argument for daily automated checks, even on a modest budget. Tools like Pace Ads, EDEE, and Optmyzr already recalculate daily caps against month-to-date actuals for other channels; applying that logic here means reconfiguring the adjustment rules around conversational inventory's lumpier curve rather than the smoother one search produces.

A "strong finish" pacing model is a reasonable default. Users who've been researching a topic across several sessions over multiple days are further along than someone on their first visit, so back-weighting spend toward the later part of a campaign period lines up with how intent actually builds. It also buys time to observe the intent distribution in the first half of a flight before committing the bulk of the budget.

Format mix adds another wrinkle. Different ad formats built for conversational surfaces, inline cards versus branded follow-up prompts, likely carry different session-depth profiles: a card might trigger early in a conversation, a follow-up prompt later. Running several formats at once blends those patterns together into something harder to read. Separating budgets by format during early testing isolates each one's spend pattern before combining them.

What measurement gaps mean for pacing decisions today

Pacing only works as well as the feedback loop underneath it, and that loop is still thin here. Agencies have been open about the measurement gaps and unproven outcomes in this channel, according to reporting from eMarketer in 2026, which means pacing built purely off click-through rate risks optimizing for the wrong signal entirely.

Click-through benchmarks for chatbot ads remain early-stage and not yet settled across the industry. For comparison, Some early data suggests organic traffic referred from AI assistants may convert at different rates than traditional search organic traffic, but that's a different path to a different outcome, and it can't be dropped in as a stand-in for how paid clicks in this channel should be paced.

Without conversion data feeding the loop, what gets called "metric-based" pacing, where the system learns from past outcomes to adjust budget mix, can't actually function as designed. Advertisers end up pacing toward impressions or clicks rather than toward anything resembling a real outcome, because the outcome data isn't there yet to pace against. OpenAI has started rolling out conversion tracking, per StackAdapt's 2026 reporting, which is a sign the infrastructure is coming. It isn't yet where search or social measurement sits.

Until it gets there, the workaround is proxy signals: site visits, upticks in branded search queries, spikes in direct traffic that line up with campaign activity. None of that replaces real attribution. But it beats pacing blind against a click-through number that lacks the reliability needed to carry the weight.

Sources

  1. LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots
  2. LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations
  3. What is LLM advertising? How LLM ads could reshape marketing
  4. Ads Inside AI: The Next Media Channel Marketers Can’t Ignore – Beet.TV
  5. enterprisedna.co
  6. intuitionlabs.ai
  7. beet.tv
  8. blog.google