Bid Shading in Real-Time Auctions for AI Ad Inventory
Bid shading struggles when inventory never repeats, like in live conversational AI ads.

Bid shading is a predictive algorithm that lowers a submitted bid to just above the level needed to win an impression, cutting overspend without giving up volume. The logic only exists because of how first-price auctions settle: the winner pays exactly what they bid, so anyone bidding their true valuation is leaving money on the table the moment they win. Three bidders compete for the same impression, the highest bid takes it at full price, and there's no natural ceiling stopping the winner from paying far more than the second-highest bidder would have required to lose. That gap between what you pay and what you needed to pay is pure waste, repeating across every auction you win.
Second-price, or Vickrey, auctions were designed to close that gap. The winner still takes the impression, but pays the second-highest bid rather than their own, which makes truthful bidding the mathematically dominant strategy. There's no incentive to shade because shading buys you nothing: bid your real valuation, and the mechanism handles the discount for you.
Most major ad exchanges and supply-side platforms abandoned that protection starting around 2019, moving to first-price auctions across the board. That single shift is why bid shading stopped being a nice-to-have and became infrastructure. Any demand-side platform that didn't build shading models into its core bidding logic was, by definition, overpaying relative to competitors who had.
Bid shading reduces publisher revenue and forces advertisers to maintain shading models, a layer of computational overhead that produces no benefit for the end consumer. Transparency rules have started to catch up with this reality. MRC standards adopted in 2026 now require platforms to disclose which auction model they're running, and that disclosure requirement has surfaced just how much variance still exists across networks, even years into the first-price transition.
The clearest evidence of what auction model choice actually does to market behavior comes from a mid-market marketplace that reversed course, moving from first-price back to second-price in June 2022. Advertiser spend held steady, but clicks rose substantially and CPCs fell sharply, about as clean a demonstration as exists of how much auction mechanics shape outcomes independent of demand. The rest of this piece tests whether machinery built for a signal environment of page URL, user profile, and keyword still functions when the signal is a live conversation.
How bid shading models work: landscape modeling, surplus optimization, and the two-stage pipeline
Standard bid shading runs on a two-stage pipeline. Stage one models the bid landscape, the distribution of competing bids likely to show up for a given impression, and stage two takes that distribution and optimizes surplus using techniques borrowed from operations research. Nearly everything downstream depends on getting that first stage right, since the bid landscape is, by definition, the distribution of every live, eligible bid on one inventory unit at the instant the auction fires. Surplus, in this context, is what an advertiser privately values an impression at minus what they end up paying for it. Shading tries to maximize that gap without losing the auction. To make its call, the model needs historical win rates at various price points, floor price signals, and impression-level features like page context, user profile, time of day, and device type.
The two-stage approach has known weaknesses. Surplus optimization routines often assume the surplus curve is unimodal, and that assumption breaks down when the real curve is non-convex, producing predictions that are confidently wrong. Errors also cascade: whatever mistake gets baked into the bid landscape model in stage one flows straight into the surplus optimization in stage two, since the two stages run as a decoupled sequential workflow with no mechanism for the second stage to correct the first. Add a sample selection problem that's structural rather than incidental: only winning bids get observed. Losing bids leave no trace, so the data used to model the competitive landscape is censored in a way that systematically understates how competitive that landscape actually is.
The research lineage tackling these problems shows a field iterating rather than standing still. The MEOW nonparametric algorithm and Deep Distribution Networks for first-price auctions both came out of KDD in 2021, followed by a multi-slot, end-to-end bid shading approach called MEBS at CIKM in 2023. The current state of the art is a generative bid shading framework from researchers at Meituan and Chongqing University, presented at SIGIR in Melbourne in July 2026. Instead of the two-stage pipeline, this model generates shading ratios autoregressively, step by step, with no predefined priors constraining the shape of the output. It pairs that generative core with a reward alignment system, a channel-aware hierarchical dynamic network acting as the reward model, tuned through group relative policy optimization, letting the system optimize short-term and long-term surplus at once rather than trading one off against the other. At inference time it outputs the shading ratio directly, skipping the two-stage handoff entirely and cutting latency. It's already running in production on Meituan's demand-side platform, serving billions of bid requests a day.
A paper published in Information Systems Research tackles nonstationary Bayesian multiarmed bandit methods for first-price real-time bidding. Taken together with the generative shading work, it signals that researchers are reworking the foundational assumptions of how shading should work at all, not just tuning parameters on an established method.
All of this machinery, the landscape models, the surplus curves, the bandit algorithms, assumes the underlying signal is stable and dense. Historical win rates need many auctions on genuinely comparable inventory to mean anything. Floor prices need a defined clearing history to be set sensibly. What happens to all of it when the inventory being sold is a conversation that has never happened before and will never happen again in quite the same form?
What makes conversational AI inventory structurally different from the inventory bid shading was built for
Bid shading sits in the demand and auction layer, but its inputs come from the context and targeting layer, which is categorically different from web inventory.
The signal itself has changed shape. Web advertising targets a page URL or a user profile built from browsing history. Conversational advertising targets intent expressed in language, in real time, with no persistent record required. A user who never types "project management software" into a search box might spend several minutes describing a team that keeps missing deadlines, coordination breaking down across time zones, and updates getting lost between tools. A targeting system built for this environment reads that as a high-intent signal every bit as strong as a branded search query, maybe stronger, because it captures the problem rather than the solution keyword. That's a shift from keyword match types and demographic buckets to semantic intent clusters and problem narratives, and because the relevance signal comes from the conversation happening right now, persistent identifiers and cross-site tracking aren't doing the targeting work here at all.
That absence of comparable inventory units is the deeper structural problem. A homepage leaderboard on a known domain has a pricing history stretching back years. A bid request for "a turn in a conversation about team coordination friction" has no historical twin to price against, because it may be the first time that exact combination of context has ever come up. Bid landscape modeling depends on volume and repetition, and this inventory doesn't offer much of either yet. There's no established click-through or conversion rate baseline for LLM advertising at all; the channel is simply too new to have produced the longitudinal data other formats take for granted. Floor pricing runs into the identical wall. Standard floor-setting logic works backward from historical clearing prices, and AI publishers in 2026 are setting floors without anything close to the depth of auction data retail media or the open web has built up over more than a decade.
The physical shape of the inventory adds another layer of complexity. There are at least four distinct ad surfaces in play: an inline sponsored card that appears after the answer, a sidebar placement, a sponsored follow-up suggestion chip, and a response-grounded brand mention woven into the answer itself. Each carries its own latency budget, its own UX cost, and its own revenue ceiling, so a bid request has to encode which surface is actually being auctioned, because winning a sidebar placement and winning a brand mention inside the answer are not the same prize. Latency itself becomes a constraint web advertising never had to reckon with in the same way. Serving an ad inside a conversational response can require a second call to the language model, adding real latency to what the user experiences as a single, continuous reply. A shading model that's computationally expensive to run can't simply be lifted out of standard RTB and dropped into this environment; the cost of running it competes directly with the cost of keeping the conversation responsive.
None of this would matter quite as much if the auction rules were settled, but they aren't. ChatGPT runs a relevance-weighted second-price auction, other platforms are running different designs or haven't announced anything at all, and open-web RTB's convergence on first-price as the near-universal standard simply hasn't happened here. Bid shading strategy depends entirely on knowing which auction model you're actually bidding into, and right now that answer changes by platform.
The auction models AI platforms have chosen and their implications for bidders
OpenAI moved fastest and most publicly. Plans to test ads in ChatGPT were announced in January 2026, testing began February 9, 2026, and ChatGPT Ads became widely available in May 2026. OpenAI's default max bid for CPM campaigns has been described in the $60 CPM range, though observed clearing rates have run as low as $25 CPM, telling you the gap between ceiling and floor is still wide while the market finds its footing. Criteo came on as the first technology partner, and Adobe followed with a partnership announced the same day testing began, expanding into performance marketing integration by early May.
The theoretical implication is interesting on its face: second-price auction mechanics should theoretically reduce the incentive to shade at all, since truthful bidding is rational under this logic, though relevance weighting introduces a quality score dimension that interacts with bid in ways similar to Google's Ad Rank, creating a different optimization surface. The effective bid becomes a function of price and predicted relevance together, behaving a lot like Google's Ad Rank formula, where a lower bid paired with high relevance can beat a higher bid paired with weak relevance. That's a genuinely different optimization surface than a pure second-price auction, and it means bidders can't just relax into truthful bidding and call it done.
Google took the opposite path, extending rather than reinventing. Ads inside Gemini-powered AI Overviews launched in October 2024, expanded to desktop in May 2025, and reached eleven additional countries by December 2025. Eligible formats are limited to text and shopping ads pulled from existing Search, Shopping, and Performance Max campaigns, labeled "Sponsored," available in English only, and excluded from sensitive verticals including adult content, alcohol, gambling, finance, healthcare, and politics. Because this format inherits Google's existing Search auction mechanics wholesale, it's the least foreign environment for any bidder already running Search campaigns; shading strategies built for Search apply here with little modification.
Microsoft's approach centers on Copilot, where ads were announced in March 2025 and have taken the form of "Showroom," an interactive format that surfaces rich sponsored content and images at the bottom of a Copilot answer when the system detects buying intent. Microsoft has also floated brand agents that users could engage with directly, but the auction mechanics for that format haven't been announced.
Perplexity moved early and then reversed. Sponsored follow-up questions launched on a CPM model in late 2024, but the company pulled back from advertising entirely in February 2026, citing trust concerns, and currently has no ad inventory running. Its audience, concentrated in technology, finance, law, and healthcare, remains relevant background for whenever inventory returns, even if it isn't buyable today.
Anthropic has staked out the opposite end of the spectrum entirely, positioning Claude as explicitly ad-free and reinforcing that stance with a Super Bowl advertisement that ran shortly after OpenAI announced its ChatGPT ad plans. There's no inventory to bid on there, by design. Standalone Gemini sits in a stranger spot: Google told advertisers ads were coming in 2026, then publicly denied it, leaving that inventory unconfirmed for now.
Putting these together shows a demand-side platform buying across AI surfaces is juggling at least three fundamentally different auction designs at once: relevance-weighted second-price on ChatGPT, search-inherited first-price mechanics on AI Overviews, and an emerging, unannounced format on Copilot. A single shading model applied uniformly across all three would be wrong for at least two of them. Shading has to be surface-aware from the start.
How bid shading logic must adapt when the signal is conversational intent
Comparability is the core problem. Standard shading estimates win probability at a given price using historical data from comparable inventory, and "comparable" has always meant matching on page-level or user-level features that simply don't exist in the same form for a live, unrepeated conversation.
The most workable substitute is semantic clustering: grouping past conversations by intent rather than by URL, so every past auction tied to, say, "team coordination friction" becomes the historical cohort used to price the next one that appears in that cluster. Floors can follow the same logic, set at the level of the intent cluster rather than at the level of a fixed placement. Recency shapes pricing more heavily here than on the open web, too. A page's context is stable for as long as the page exists; a conversation's context can shift the moment the user changes the subject, and when it does, the relevant pricing cohort has to shift with it in real time.
Relevance weighting adds a second axis bidders have to manage simultaneously. Where an auction weighs relevance alongside price, as ChatGPT's does, shading the bid down without also improving predicted relevance can cost you an impression you would have won at a lower effective bid had relevance been stronger. The dynamic echoes Google's Quality Score system from Search, except the relevance input here is conversational context rather than landing page and keyword match.
Genuinely novel topics create a cold-start problem that has no clean solution yet. When a conversation touches a new product category or a breaking news event, there's no historical auction data to draw on at all, and shading models are forced back onto prior distributions or bandit-style exploration. Accepting higher variance in outcomes is the price of learning. The nonstationary Bayesian bandit work published in Information Systems Research in June 2026 speaks directly to this scenario, since it's built for a first-price environment where the underlying distribution keeps moving.
Latency turns from a nice-to-have optimization into a hard constraint. The generative bid shading framework (Huang et al., SIGIR '26) reduces inference time by outputting the shading ratio directly rather than running a two-stage pipeline. The ad has to clear within the response time budget of a live conversation, so a shading model that adds meaningful delay isn't just a performance regression, it's a UX failure the user notices immediately.
Non-stationarity becomes structural rather than an occasional nuisance, because conversational context shifts far faster than standard signals do. Standard RTB landscapes shift gradually, tracking seasonality or competitive pressure over weeks. In a conversational interface, intent context can shift within a single session, so a shading model must treat each conversation turn as a potentially new pricing environment rather than assuming the auction dynamics that held at the start of the session still hold later.
That's precisely why the autoregressive, generative approach looks better suited to this inventory than the older two-stage pipeline. It carries no predefined priors about the shape of the surplus curve, so it can capture the non-convex, non-stationary patterns that conversational auctions actually produce. That it's already running at scale on Meituan's platform, processing billions of requests a day, is real evidence the architecture holds up under volume, even though that deployment isn't drawing on AI chat data specifically. The absence of pricing history cuts the other way for publishers rather than bidders: set floors too conservatively and you leave revenue on the table during a scarcity period, set them too aggressively and you choke fill rate before the bid landscape has enough depth to support the price. Soft floors paired with feedback loops that tighten as auction data actually accumulates are the sensible starting point.
What the absence of established performance benchmarks means for valuation
Start with the gap rather than a number, because there isn't a reliable number yet. No established click-through or conversion rate baseline exists for LLM advertising; the channel hasn't been running long enough to produce the longitudinal data that would make one meaningful. Any shading model that optimizes toward expected conversion probability is, right now, optimizing against a genuinely thin prior, and that estimate shouldn't be dressed up as more solid than it is.
That thinness cuts straight to the heart of what shading does. The whole exercise depends on calibrating the shade amount against the expected value of the impression. If conversion rate is unknown, expected value is unknown too, so the optimal shade amount can't be derived from first principles the way it can on inventory with years of clearing data behind it. Google Search, by contrast, has click-through rates that have historically averaged above 3% across industries. That figure doesn't transfer cleanly to conversational surfaces, where the ad format, the surrounding context, and the user's relationship to the interface are all different enough that borrowing Search's numbers as a stand-in would be a guess dressed up as an estimate.
The honest position for anyone valuing this inventory right now is to treat every input as provisional and build in feedback loops that let pricing tighten as real auction data accumulates, rather than importing benchmarks from web advertising and hoping they hold. They probably won't, not exactly, and the platforms still finding their footing on auction design make that clear enough on their own.
Sources
- Retail Media Auction Mechanics: Bids, Floors, MRC 2026
- An Efficient Deep Distribution Network for Bid Shading in First-Price Auctions | Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining
- Generative Bid Shading in Real-Time Bidding Advertising
- searchlab.nl
- dl.acm.org
- pacvue.com


