Budget Throttling Tactics During LLM Inventory Spikes
When falling token prices hide rising total bills, pacing rules built on old assumptions collapse.

Two spike mechanisms fire at once inside LLM advertising, and each makes the other worse. No media channel before this one has produced that combination. On the infrastructure side, token costs can multiply within hours because usage dashboards from most providers update on a lag, so a spend spike is often over by the time anyone sees it on a screen. On the media-buying side, conversational inventory is scarce in a way display and search never were: a results page holds ten ads comfortably, but one conversational answer holds effectively one, which concentrates value into a single placement instead of spreading it across a page.
Those two forces do not take turns. A budget holder watching only the media dashboard sees rising CPMs and assumes a demand story. A budget holder watching only the infrastructure bill sees rising token spend and assumes an engineering story.
Token prices fell by roughly 80% between 2025 and 2026, and enterprise bills went up anyway because volume growth outran the price drop. Falling unit prices next to rising total bills need not be explained away as a contradiction. It is the defining shape of AI cost behavior in 2026, and any throttling system built on the old assumption, that falling prices mean falling risk, is already built on the wrong premise.
How the ChatGPT auction removes the controls advertisers are used to
Match types, negative lists, and search term reports assume a query string sits between user intent and ad delivery. ChatGPT's ad auction removes that object. Targeting runs on conversational context instead: advertisers bid on clusters of intent signals read from the arc of a conversation, submitted as a freeform natural-language hint rather than a keyword list. Ad eligibility follows the content and intent of what the user is actually saying, and OpenAI keeps final control over which eligible ad gets shown and inside which conversation.
That structure has a direct consequence for anyone trying to throttle spend the old way: there's no exclusion mechanism to lean on. The only way to keep an ad out of an unwanted conversation is to write the targeting hint narrow enough from the start that it never matches, which shifts the entire burden of precision onto campaign setup rather than ongoing optimization.
The auction mechanics compound that shift. It runs as a relevance-weighted second-price auction, so a tightly targeted ad can beat a much bigger bid if its relevance score is high enough. That rewards precision, but it also breaks the budget-modeling habits search advertisers bring with them, since bid amount alone stops predicting outcome.
The pricing model has already shifted once, from CPM pricing at launch to CPM rates eroding and OpenAI introducing cost-per-click bidding as an additional buying mechanism alongside CPM, a signal of ongoing instability for budget planning. A platform switching its core pricing model within months of launch is not a platform anyone should plan a three-year media strategy around without built-in flexibility.
Google AI Mode and Microsoft Copilot present a different problem, and arguably a more dangerous one for anyone not paying attention. Eligible campaigns on both surfaces serve automatically, with no opt-in and no opt-out control available to the advertiser. That means a search advertiser running a broad-match campaign on Google today may already be paying for conversational placements inside AI Mode without ever having made that decision. The first step for those advertisers is finding out they're already in the auction, not optimization.
Why conventional pacing rules fail when inventory can close overnight
Daily budget caps, frequency floors, dayparting: all of it assumes inventory is a stable pool that exists whenever the budget says to spend against it. Perplexity's exit from advertising is the clearest evidence that conversational inventory can disappear at scale without warning, and any budget architecture that treats LLM surfaces as stable inventory pools has already mispriced the risk.
Perplexity launched sponsored follow-up questions in November 2024 and built real scale, reaching 100 million users, before confirming in February 2026 that it was abandoning advertising entirely. The company cited user trust concerns, the phase-out had actually started earlier than the public announcement, and Perplexity pivoted toward a subscription-first, ad-free model with a substantial annualized revenue target. Perplexity's leadership told the Financial Times that sponsored placement risks making users suspicious of the entire answer they're reading. That is a trust argument that applies to every conversational AI surface carrying ads today. Perplexity's closure should read as a warning about structural volatility across the category rather than a one-off business decision.
Anthropic makes the same point from the opposite direction. At the 2026 Super Bowl, Anthropic drew a public line between Claude and OpenAI's ad launch, and Claude carries no ad inventory by design. That's a second major surface with zero available inventory because of a policy choice made in advance, not scale or user pushback. LLM advertising creates a cost environment where two separate spike mechanisms fire simultaneously and reinforce each other, which no prior media channel has produced.
Google's surfaces fail the pacing model in a quieter way. Google AI Overviews and AI Mode serve ads through existing campaigns automatically, but impressions are accumulating while clicks are falling. This means that inventory now functions more like awareness spend than direct-response spend, and budgets weighted toward performance KPIs are systematically misallocated here.
And the reversal risk runs both directions. OpenAI spent three years publicly stating ChatGPT would never carry advertising before launching ads on February 9, 2026. A platform that says no for three years and then says yes overnight can just as easily say yes today and no next quarter. Pacing rules assume continuity. This category has already demonstrated, twice, that continuity is not guaranteed.
The control architecture that replaces pacing rules: token-level signals, placement-value ceilings, and hard stops
Budget governance on LLM surfaces needs three controls, each operating at a different speed: token-level monitoring to catch the signal early, placement-value ceilings to stop overpaying during scarcity spikes, and hard stops that act without waiting on a dashboard to refresh.
Token-level monitoring comes first because it's the earliest available signal. Rate limiting, quotas, budgets, and spend caps are four separate controls, not variations on the same idea: rate limiting protects capacity within a time window, budgets protect money across a period, and treating them as interchangeable leaves exactly the gap where an unnoticed spike can run its full course. Because usage dashboards refresh on their own delayed schedule, token consumption data is available before any of that lag catches up, and organizations that route token-level alerts straight into bidding logic hold a real timing advantage over those still waiting on end-of-day reports.
Caching belongs in this same layer, and it's the most underused lever available. One analyzed production workload found a large share of LLM queries were semantically similar to prior ones, which is precisely the traffic pattern prompt caching is built to catch before it ever becomes a spike. Caching cuts costs meaningfully on cached input tokens with both Anthropic and OpenAI, but most teams running production workloads aren't using it well. Engineering teams could close that gap without touching the media-buying side.
Placement-value ceilings sit above the infrastructure layer and price the conversational match itself. Because the auction weights relevance rather than price alone, a precisely targeted hint can beat a larger raw bid, so the ceiling on what any placement is worth should track the quality of the conversational match rather than a CPM figure borrowed from another channel. An intent signal appears in a conversation that surfaces the user's budget, their stated constraint, their shortlist, and what they already ruled out, richer than a keyword ever carried, and that richness is what a ceiling ought to price. Intent forms layers within a single conversation too: informational intent suits awareness spend, comparative intent suits competitive positioning, transactional intent suits direct response, and a placement-value ceiling should differ across those tiers because the commercial value of the moment differs. Impression-based pricing behaves differently inside a conversational interface than inside a feed, a point OpenAI's own shift from CPM toward CPC as its primary buying mechanism confirms directly, so a ceiling built on CPM assumptions imported from display or social carries the wrong logic into this channel.
Hard stops are the enforcement layer, and they only work if they don't depend on the same lagging dashboards that miss the spike in the first place. A 24-to-48-hour reporting delay means a spend cap set inside a campaign manager only confirms what already happened; it doesn't prevent it. Hard stops need to sit at the API layer, before an inference call ever fires, rather than at the reporting layer after the spend is already committed: a governor that limits an engine's speed works differently than a speedometer that only tells you how fast it already went. Caching helps here too, since fewer inference calls reaching the system means fewer calls ever approaching a hard-stop threshold, so it functions as both a cost control and a spike throttle at once. On surfaces where opt-out isn't available at all, Google AI Mode and Microsoft Copilot among them, hard stops on bid values and creative eligibility are the only levers an advertiser has, and auditing which campaigns currently qualify for AI placement is the necessary first step before setting any ceiling.
Budget allocation across multiple AI surfaces when inventory can disappear
Because any single conversational AI surface can close, pause, or reprice faster than a planning cycle, cross-surface budget allocation should be structured around surface type and intent tier rather than channel-level CPM comparisons.
The surfaces carrying real, open inventory in 2026 differ sharply in how they can be bought and who they reach. ChatGPT opened to advertisers on February 9, 2026, buyable self-serve through ads.openai.com or programmatically through Criteo, a named partner since March 2, 2026, and StackAdapt, added as of May 5, 2026. Ads there run only on the Free and Go tiers; Plus, Pro, Business, Enterprise, and Education subscribers see no ads at all. Self-serve access expanded to Europe, India, the Middle East, and North Africa on August 31, 2026, and OpenAI reports a $1 billion annualized run rate reached in well under a year.
Google AI Overviews and AI Mode work on a different model entirely. AI Overviews has carried ads since May 2024 and AI Mode since May or June 2025, both served automatically through campaigns advertisers already run, with no opt-in and no opt-out.
Microsoft Copilot has carried ads since 2023, inherited from Bing Chat, making it the most established conversational ad product in the category by tenure. Recent additions include Compare and Decide ad units, shopping campaigns, and Copilot Checkout, an in-conversation purchase flow that launched in January 2026.
The rest of the landscape is thinner. Google's Gemini app carries no ads as of August 2026, though Google's AI Overviews and AI Mode inside Search still carry live inventory through Performance Max and AI Max. Meta AI's inventory status is unconfirmed, with no announced buy path, which argues for planning around eventual availability rather than budgeting against it now. Perplexity has abandoned advertising entirely as of February 2026 and moved to a subscription-first model, leaving organic and answer-engine optimization as the only routes to visibility there. Anthropic's Claude carries no inventory by policy, with none planned.
The entire ad-eligible ChatGPT audience is on the Free and Go tiers, concentrating the risk.
Demand for entry into this channel is already high regardless. One agency executive's canvass of 45 enterprise clients found all 45 willing to shift Bing budgets into ChatGPT's ad pilot, graded against existing search performance benchmarks. That comparison frame matters on its own, since it sets the performance bar the new channel will be judged against from day one, often unfairly given how different the auction mechanics are. eMarketer's forecast for US AI search ad spend has it rising from a modest base in 2025 toward a meaningfully larger share of all search ad spend by 2029. The exact timing carries some uncertainty, but the direction is settled, and that direction is the part planners actually need for budgeting purposes.
Put practically: budgets on Google's AI surfaces should lean toward awareness outcomes, since impressions there keep climbing while clicks fall, and ChatGPT or Copilot budgets carrying direct-response expectations should run on the placement-value ceiling logic described above rather than CPM benchmarks pulled from display advertising. Spreading budget across more than one AI surface is a structural necessity for planners, not a cautious hedge for the risk-averse. It's a structural necessity, given that Perplexity's ad inventory closed at real scale with no advance signal to any advertiser running campaigns there.
Brand safety and measurement gaps that throttling alone cannot solve
Every control described above governs cost and placement frequency. None of it governs what happens to a brand's reputation once an ad appears next to an answer the model generated on its own, in language no creative team wrote and no one previewed before it went live. A spend cap can stop a budget from overrunning. It cannot stop a sponsored placement from appearing beside a response that turns out to be wrong, tone-deaf, or adjacent to a topic the advertiser would never have chosen to sit next to.
Perplexity's own stated reason for leaving advertising was a trust concern: sponsored placement risks making users suspicious of the whole answer, not just the ad within it. That is a measurement gap no throttling architecture touches, because throttling controls how much gets spent and how fast, not whether the placement itself damages the credibility of the surface it sits inside. A hard stop set at the API layer will halt an inference call before it fires. It has no mechanism for judging whether the ten calls that did fire put a brand next to content that undermines the message the ad was trying to send.
Measurement carries its own separate hole. Google's AI Overviews and AI Mode show rising impressions alongside falling clicks, and that pattern gets read here as evidence the inventory behaves like awareness spend. But awareness spend has historically been measured through tools built for a page: viewability standards, brand lift studies, panel-based recall surveys, none designed for a single-turn conversational answer that no two users ever see rendered quite the same way. Attribution built for keyword-driven search or feed-based social doesn't map cleanly onto an interface where the ad appears once, inside a conversation shaped by everything the user said before it, and then is gone.
None of that argues against building the control architecture laid out above. Token-level monitoring, placement-value ceilings, and hard stops remain the right response to a cost structure that can spike without warning and inventory that can vanish overnight. Cost control and brand safety are different problems solved by different tools, and an advertiser who treats a well-built throttling system as a complete answer to conversational advertising has solved only the part of the problem that shows up on an invoice.
Sources
- LLM Ads Explained: How AI Advertising Works in 2026 | guptadeepak.com Guides
- Weekly LLM Advertising Brief - by robbie caploe
- How to Build an LLM Advertising Stack: Tools, Workflow, and Budget (2026) | Lapis
- LLM Advertising: Every Platform, Format and Cost in 2026 - AI Ad Placements
- Rate Limiting in AI Gateway : The Ultimate Guide
- Media Buying


