BROUGHT TO YOU BY
iROAS is one of retail media's most important metrics is also its most misunderstood. Research from Albertsons Media Collective and Ovative Group found that how you measure incrementality can swing results as much as the campaign itself.
Methodology is the Missing Variable in Retail Media Measurement
Across the 42 campaigns, the gap between the highest and lowest iROAS generated by different methodologies averaged 6.5x. What’s more, 83% of campaigns flipped from positive to negative incremental return based solely on how the math was done. The ads, shoppers, and sales data stayed the same. Only the methodology changed, and that alone was enough to meaningfully shift the result.
The IAB defines incrementality as measuring marketing’s causal impact: the additional business outcomes a campaign drove compared with what would have happened without it. One group sees the ad, one group doesn’t, and the difference between the two is measured. That’s the basic idea.
In practice, though, it gets a lot messier. Every retail media network (RMN) defines and measures incrementality differently, which means results can vary dramatically even when the campaign itself stays the same.
Same campaign, different numbers
iROAS carries a lot of weight in retail media measurement. Brands rely on it to evaluate campaigns, decide where budgets go, and choose partners. The problem is that all of those decisions assume the metric means the same thing everywhere; but often, it doesn’t.
A joint analysis from Albertsons Media Collective (The Collective) and Ovative Group, conducted with professors from Northwestern University’s Kellogg School of Management, looked at 42 digital ad campaigns to see how different methodological choices can affect iROAS results, regardless of actual campaign performance. The conclusion: If you change the methodology, you can change the outcome.
Four levers that change everything
Matching approach
This analysis compared propensity score matching (PSM) and clustering.
PSM produced match quality roughly 12 times better than clustering, based on the statistical similarity between test and control groups.
It also produced lower iROAS estimates: $0.73 for PSM (1:many) versus $1.80 for clustering.
Clustering is faster and more scalable; it is also more likely to overstate incremental revenue. PSM is more stringent and tends toward more conservative results.
Feature selection
Within PSM, the features used to match customers can completely change the outcome.
One of the biggest examples was historical brand sales. When that variable was removed from the model, average iROAS shifted from $1.23 to -$0.14, while making the same change barely moved clustering results.
That’s because PSM is highly sensitive to the inputs used to calculate the propensity score, which estimates how likely someone was to see an ad in the first place. Remove a strong predictive feature, and the entire matching model changes. In The Collective’s analysis, past brand purchase behavior was the strongest predictor available, so leaving it out was enough to flip the result from positive to negative
Revenue calculation
Once matched groups are built, analysts still need to estimate how much revenue was actually incremental. The study looked at two common approaches:
The first was observed sales performance, which simply compares actual sales between exposed and control groups during the campaign period.
The second was Bayesian Structural Time Series (BSTS) modeling, which uses historical sales patterns to estimate what sales would likely have looked like without advertising. That model explicitly accounts for things like trends and seasonality.
Across The Collective’s campaigns, BSTS produced estimates that were, on average, 90% lower than observed sales comparisons, with iROAS averaging $0.97 versus $1.56. BSTS is more precise, but it's also more complex to build, harder to explain, and not every RMN has the infrastructure to run it consistently.
There’s no single “right” methodology for measuring incrementality. Every approach comes with tradeoffs between rigor, scale, speed, and practicality. The bigger problem is that advertisers often have little visibility into which tradeoffs their RMN made in the first place.
That’s what needs to change. Before taking an iROAS result at face value, advertisers should understand a few core things:
What advertisers need to ask
Retail media built its reputation on measurement. Closed-loop data and shopper insights gave brands something most other channels couldn’t: a clearer connection between ad exposure and actual purchases. iROAS was supposed to take that a step further by isolating the sales lift advertising truly drove, instead of just measuring what happened near an ad.
The problem is that the methodology behind most iROAS numbers is still largely a black box. Advertisers struggle to interpret the results, while RMNs face growing skepticism about whether reported performance is actually comparable or even accurate. Over time, that risks weakening the measurement story that made retail media so valuable in the first place.
The answer isn’t forcing the industry into one universal methodology. It’s setting a baseline expectation around transparency. Every iROAS result should come with clear documentation explaining how it was measured.
Once that becomes standard practice, advertisers can make more informed decisions about what their retail media investments are actually delivering. Until then, many are comparing numbers that may not be measuring the same thing at all.
The bigger picture
Albertsons Media Collective is a unique tier-one growth partner with a local legacy of trust you can’t buy—you can only build. We unite 22 iconic grocery banners serving millions of households in the country’s most coveted markets, turning reach, loyalty and hometown relevance into real results for brands hungry to grow.
By bringing merchandising and media to the same table, we collaborate with you at every step. Fueled by real shopper behavior and rich insights from across our banners, we help you show up in the moments that matter most—wherever, whenever and however your customers choose to shop.
This immersive experience was produced for Albertsons Media Collective by EMARKETER Studio, an in-house creative studio within research company EMARKETER. EMARKETER is the leading provider of research, data, and insights for marketing, advertising, and commerce professionals. Our data-driven forecasts and rigorous analysis empower revenue-driving teams to make strategic decisions with confidence. Through expert context from our analysts, carefully vetted data sources, and a proprietary research methodology, EMARKETER delivers forecasts, reports, and benchmarks that help companies anticipate tomorrow’s market trends today. EMARKETER is a division of Axel Springer S.E.
Design
Miri Kramer - Creative Director, Content Studio, EMARKETER
Anthony Wuillaume - Art Director, Content Studio, EMARKETER
* Copyright © 2026 EMARKETER Inc. All Rights Reserved.
Privacy Policy
Terms of Service
Sitemap
Your Privacy Choices
The industry generally relies on three main approaches, each landing at a different point on the spectrum between methodological rigor and what’s actually practical to execute.
Three measurement approaches, different tradeoffs
Cluster matching.
Customers with similar characteristics, like demographics or purchase history, are grouped into clusters. Customers who saw an ad are then matched with similar customers in the same cluster who did not see the ad, creating comparable test and control groups.
In retail media, there are two common forms of matching:
Propensity score matching (PSM).
Customers are matched based on how likely they were to purchase or be exposed to an ad. That likelihood, called a propensity score, is calculated using factors like demographics or past purchase behavior. Customers with similar scores are then paired to compare outcomes between exposed and unexposed groups. PSM can be done as 1:1 matching, where each exposed customer is paired with one control customer, or 1:many matching, where one exposed customer is matched with multiple control customers.
Randomized controlled trials (RCTs) are considered the gold standard. Audiences are randomly split before a campaign runs, which creates the clearest picture of causal impact. Ecommerce retailers with mostly online businesses can run these tests relatively easily and at scale. Omnichannel retailers have a tougher challenge because customers might see an ad online and then make a purchase in-store, making it much harder to maintain clean test and control groups. At large scale, RCTs often become difficult to execute.
Matching methods fill that gap. Instead of randomizing audiences upfront, matching builds a control group after the campaign by pairing exposed customers with unexposed customers who share similar traits, like purchase history or demographics. It’s scalable, can be done retrospectively, and has become the most common approach in retail media. But it’s also where a lot of methodological variation shows up.
Synthetic controls take a different approach by creating a weighted blend of unexposed markets or users whose pre-campaign behavior closely matches the exposed group. That blended group is then projected forward as the “what would have happened otherwise” scenario. The method requires a longer history of data and is less common in retail media, but it can work well for market-level testing.
Pre/post analysis, which simply compares performance before and after a campaign, sometimes gets treated like an incrementality method, but it really isn’t one. It can provide useful context, but it can’t isolate what the advertising actually caused.
Method Overview
Operational Complexity vs. Causal Inference Rigor
Method 1 RCT Definition
DESIGN NOTE:
(Animated version of EXHIBIT 2, going step by step)
Method 2
Matching Method Definition
DESIGN NOTE:
(Animated version of EXHIBIT 3, going step by step)
Method 3
Synthetic Control Definition
DESIGN NOTE:
(Animated version of EXHIBIT 4, scrolling through the stages)
There’s no single “right” methodology for measuring incrementality. Every approach comes with tradeoffs between rigor, scale, speed, and practicality. The bigger problem is that advertisers often have little visibility into which tradeoffs their RMN made in the first place.
That’s what needs to change. Before taking an iROAS result at face value, advertisers should understand a few core things:
What methodology was used to measure incrementality?
How were the test and control groups built?
What assumptions or modeling decisions influenced the outcome?
How was incremental revenue actually calculated?
Those questions can help separate a result driven by genuine incremental impact from one driven mostly by methodological choices.
Consistency matters too. If an RMN applies different methodologies across campaigns, performance comparisons become unreliable And even when two RMNs use approaches that sound similar on paper, differences in attribution logic or measurement scope can still make apples-to-apples comparisons difficult.
Those questions can help separate a result driven by genuine incremental impact from one driven mostly by methodological choices.
Consistency matters too. If an RMN applies different methodologies across campaigns, performance comparisons become unreliable And even when two RMNs use approaches that sound similar on paper, differences in attribution logic or measurement scope can still make apples-to-apples comparisons difficult.
What methodology was used to measure incrementality?
How were the test and control groups built?
How was incremental revenue actually calculated?
What assumptions or modeling decisions influenced the outcome?
This article is based on “Retail Media iROAS Demystified,” a March 2026 report from Albertsons Media Collective and Ovative Group, developed in partnership with professors from Northwestern University’s Kellogg School of Management. To access the full report, visit albertsonsmediacollective.com.
' If the retailer leverages dynamic pricing, RCTs should be structured to ensure consistent pricing across customers / markets to eliminate contamination.
Audience filtering
Before matching, analysts decide who goes into the test and control groups.
Filtering to past brand buyers reduced average sample size by 83% and dropped average iROAS from $2.27 to $0.22.
A category buyer filter was more moderate, with iROAS falling from $2.27 to $2.20.
Tighter filters mean fewer sales get counted as incremental, while ad spend stays the same. So iROAS drops, not because the ads suddenly worked less well, but because the measurement got narrower.
The Collective analyzed 42 campaigns across 54 different methodology combinations, testing how four key methodological choices influenced the final results. Here’s where the numbers started to shift.
Revenue calculation
Once matched groups are built, analysts still need to estimate how much revenue was actually incremental. The study looked at two common approaches:
The first was observed sales performance, which simply compares actual sales between exposed and control groups during the campaign period.
The second was Bayesian Structural Time Series (BSTS) modeling, which uses historical sales patterns to estimate what sales would likely have looked like without advertising. That model explicitly accounts for things like trends and seasonality.
Across The Collective’s campaigns, BSTS produced estimates that were, on average, 90% lower than observed sales comparisons, with iROAS averaging $0.97 versus $1.56. BSTS is more precise, but it's also more complex to build, harder to explain, and not every RMN has the infrastructure to run it consistently.
Feature selection
Within PSM, the features used to match customers can completely change the outcome.
One of the biggest examples was historical brand sales. When that variable was removed from the model, average iROAS shifted from $1.23 to -$0.14, while making the same change barely moved clustering results.
That’s because PSM is highly sensitive to the inputs used to calculate the propensity score, which estimates how likely someone was to see an ad in the first place. Remove a strong predictive feature, and the entire matching model changes. In The Collective’s analysis, past brand purchase behavior was the strongest predictor available, so leaving it out was enough to flip the result from positive to negative.
Matching approach
This analysis compared propensity score matching (PSM) and clustering.
PSM produced match quality roughly 12 times better than clustering, based on the statistical similarity between test and control groups.
It also produced lower iROAS estimates: $0.73 for PSM (1:many) versus $1.80 for clustering.
Clustering is faster and more scalable; it is also more likely to overstate incremental revenue. PSM is more stringent and tends toward more conservative results.
