Back to blog
Insights

Why Incremental Revenue Matters in Ecommerce (More Than Conversion Rate Ever Will)

Learn what incremental revenue really means, why attribution often overestimates marketing performance, and how holdout testing reveals true ecommerce growth.

Every ecommerce dashboard tells a story. Facebook says it drove $80,000 last month. Google Ads claims $65,000. Klaviyo reports $40,000 from email flows. Add them up and the numbers rarely match the store's actual revenue growth — usually by a wide margin. Most teams have learned to shrug this off as "attribution overlap" and move on. That shrug is one of the most expensive habits in modern ecommerce.


The problem isn't that these platforms are lying. Each one is measuring something real: a click, a view, a session that ended in a purchase within a set window. But none of them are measuring the question that actually matters to a business: what happened because of this spend, versus what would have happened anyway?


That distinction — between revenue that is attributed to a channel and revenue that is caused by it — is the difference between running a company on fiction and running one on evidence. This article is about why that distinction deserves far more attention than it currently gets, and what changes once a team starts asking the right question.


The Metric That Quietly Misleads Everyone


Attribution models exist because measurement had to start somewhere. Last-click, first-click, linear, time-decay — each one assigns credit for a sale to one or more touchpoints in a customer's journey. They are useful for understanding sequence. They are far less useful for understanding cause.


Here's a scenario that plays out in nearly every mid-market ecommerce company. A brand runs retargeting ads to people who already added a product to their cart. Conversion rate on that campaign looks excellent — often 8 to 12 times higher than prospecting campaigns. The finance team sees the ROAS and pushes more budget into it. But a large share of those shoppers were already going to buy. The ad didn't create the sale; it took credit for a sale that attribution software would have counted regardless.


This isn't a hypothetical edge case. It's the default behavior of retargeting, branded search, and any channel that reaches people deep in the purchase funnel. Attribution rewards channels for showing up right before a conversion, not for causing one. Over time, budgets drift toward the channels that are best at claiming credit rather than the channels that are best at generating demand.


Correlation Isn't Causation: The Coffee Shop Problem


A useful way to see the flaw is to leave ecommerce entirely for a moment. Suppose a coffee shop notices that on days when it plays jazz music, average purchase size is higher. Someone on the team proposes playing jazz every day to lift revenue.


But what if jazz happens to play mostly in the late afternoon, when the shop's highest-spending regulars — professionals stopping in after work — are already more likely to order a pastry with their coffee? The music didn't cause the bigger order. Time of day did, and jazz was just a marker correlated with that time slot. Play jazz at 8am instead, and nothing changes, because the actual driver was never the music at all.


Ecommerce attribution makes the same mistake at scale, every single day. A channel that reaches people who were already going to buy will always look like it's driving revenue, the same way jazz music looks like it's driving order size. The only way to know if the channel is the cause, rather than a correlated marker, is to remove it and see what happens.


What Medicine Already Figured Out


Pharmaceutical research solved this exact problem more than a century ago, and the solution is worth borrowing directly. When a new drug is tested, researchers don't just give it to sick patients and measure how many get better. Many people recover on their own. Some improve simply because they believe they're being treated — the placebo effect is well documented and can account for a meaningful share of perceived improvement in a trial.


So trials split participants into two groups. One receives the drug. One receives a placebo. Both groups are tracked identically. The difference in outcomes between the two groups — not the raw recovery rate of the treated group — is the actual effect of the drug. Everything else is noise, bias, or the body's own tendency to heal.


Marketing has an equivalent structure, and it's called a holdout group. A portion of the audience is deliberately withheld from a campaign, a channel, or a feature. Their behavior becomes the placebo arm — the baseline of what would have happened with no intervention at all. Compare that baseline against the group that received the marketing, and the gap is the true, causal lift. Everything an attribution model reports, by contrast, is closer to counting every patient who got better and crediting the drug for all of it.


Incrementality, Defined Precisely


Incremental revenue is the revenue that would not have occurred without a specific action — a campaign, a discount, a product recommendation, a channel. It is calculated as the difference between a treatment group's outcome and a comparable control group's outcome, not as a sum of touchpoints along a purchase path.


The formula is simple to write and hard to execute well:


Incremental Revenue = Revenue (Exposed Group) − Revenue (Holdout Group)


The complexity lives in the word "comparable." If the two groups differ in any systematic way — one skews toward existing customers, one gets exposed to a promotion at a different time of month, one is geographically different — the comparison breaks. Rigorous incrementality testing spends most of its effort on making sure the groups are genuinely equivalent before spend even begins, through random assignment rather than convenient segmentation.


Control Groups and Holdout Testing in Practice


A holdout test for a marketing channel typically works like this. Before a campaign launches, the audience is randomly split — say, 90% exposed, 10% held out. The holdout group is deliberately excluded from seeing the campaign, whether that means suppressing ads to them, excluding them from an email send, or withholding a promotional offer.
At the end of the test window, both groups' purchase behavior is compared. If the exposed group generated meaningfully more revenue per person than the holdout group, that gap is incremental. If the two groups performed almost identically, the channel likely wasn't driving new revenue — it was capturing demand that already existed.


Why the Split Ratio Matters

A 50/50 split gives the cleanest statistical read but sacrifices half the campaign's reach while the test runs. A 95/5 split protects revenue but takes longer to reach statistical confidence, especially for lower-traffic stores. Most mature testing programs settle somewhere between 80/20 and 90/10, adjusting based on baseline traffic volume and how quickly a confident answer is needed.


Why Sample Size Isn't Optional

A holdout test on a store doing 200 orders a month will struggle to detect anything short of an enormous effect. Statistical power depends on sample size, baseline conversion rate, and the size of the effect being measured. A store running a holdout test without first estimating the sample size it needs is likely to end the test with a result that looks meaningful but is actually noise.


Why Meta and Amazon Don't Trust Their Own Dashboards


It's worth noticing that the companies with the most sophisticated attribution technology in the world are also the most vocal about its limits. Meta has published extensively on conversion lift studies, which are essentially large-scale holdout tests run on top of ad delivery, specifically because platform-reported attribution tends to overstate the causal impact of ads. Amazon is known for running large numbers of randomized controlled experiments across its business — a practice that has been described publicly by Amazon leadership as core to how the company makes decisions, precisely because intuition and correlated metrics are unreliable guides at that scale.


The lesson isn't "use the same tools Amazon uses." Few merchants have that infrastructure. The lesson is that organizations with the deepest data access and the most resources to build better attribution have concluded that attribution alone isn't good enough, and have invested instead in experimentation as the source of truth. That's a strong signal about where the ceiling of dashboard-based measurement actually sits.


Why Conversion Rate Alone Is a Dangerous Metric


Conversion rate is one of the most watched numbers in ecommerce, and one of the easiest to improve in ways that quietly hurt the business. Add a 15% discount banner site-wide, and conversion rate climbs. Simplify checkout to one click, and conversion rate climbs. Remove a shipping cost by baking it into the product price, and conversion rate climbs.
None of these actions tell you whether more people are buying who wouldn't have otherwise, or whether people who were already going to buy are simply paying less, or converting slightly faster, while margin quietly erodes underneath a metric that keeps going up.


A conversion rate lift of 20% sounds unambiguously good in a slide deck. It says nothing about whether the increase came from incremental demand or from discounting away revenue that was already coming. Those are opposite outcomes for the business, and conversion rate cannot tell them apart on its own.


Profit Versus Revenue: Where Incrementality Actually Bites


Revenue growth is easy to celebrate. Profit is what survives. A campaign that generates $50,000 in attributed revenue at a 20% discount rate, where 70% of buyers would have purchased anyway at full price, has quietly destroyed margin on every non-incremental order while claiming full credit for the sale.


Run the arithmetic on a simplified version of this. Suppose a store runs a 20% off promotion to 10,000 customers. Attribution shows 800 orders at an average order value of $60 — $48,000 in "driven" revenue. A holdout test reveals that the control group, which received no discount, still converted at a rate producing 550 equivalent orders. Only 250 orders were truly incremental. The other 550 would have happened anyway, just now at 20% less margin.


That reframes the campaign entirely. Instead of $48,000 in new revenue, the honest picture is roughly 250 incremental orders, plus a permanent 20% margin discount handed to 550 customers who needed no incentive to buy. The campaign might still be worth running — but that's a very different decision than the one the attribution dashboard implied.


Discount Addiction and What It Trains Customers to Expect


Discounting has a compounding cost that rarely shows up in a single campaign's numbers: it changes customer behavior over time. Shoppers who are repeatedly offered discounts learn to delay purchases until the next one arrives. A brand that promotes 15% off every few weeks eventually finds that full-price conversion quietly declines, because its own customer base has been trained to wait.


This is a second-order effect that attribution can't see, because it only ever looks backward at a single campaign's performance rather than forward at how the campaign reshapes future demand. Incrementality testing, run consistently over time, is one of the few tools that can actually detect this — by tracking whether the incremental lift from discount campaigns is shrinking release over release, even as attributed revenue looks stable or grows.


Long-Term Customer Behavior Beyond the First Sale


A promotion's real cost or benefit often doesn't reveal itself until months later. A discount that pulls forward a purchase a customer was going to make anyway isn't creating new revenue — it's borrowing revenue from the future and often lowering the price at which it arrives. Measuring incrementality over a longer window, looking at repeat purchase rate and customer lifetime value rather than just the initial order, tends to tell a more honest story than a 7-day attribution window ever could.


Statistical Significance and the Cost of False Positives


Not every difference between two numbers is meaningful. If a holdout group converts at 3.1% and an exposed group converts at 3.3%, that gap might reflect a genuine lift, or it might be ordinary random variation that would disappear if the test were repeated. Statistical significance testing exists to separate the two, typically by calculating whether the observed difference is larger than what chance alone would plausibly produce given the sample size.


Teams that skip this step and act on any positive-looking gap expose themselves to false positives — concluding a campaign worked when it didn't, then scaling a strategy that has no real effect. This is a well-known failure mode in A/B testing broadly, not just in marketing: small sample sizes and short test windows produce results that look decisive and often aren't. The fix isn't complicated in principle — larger samples, longer windows, and pre-registered thresholds for what counts as significant — but it requires discipline that a lot of "just check the dashboard" cultures don't have.


Survivorship Bias and Selection Bias in Ecommerce Data


Two related biases quietly distort ecommerce analysis. Survivorship bias shows up when a team studies its best customers to figure out what marketing "worked," without noticing that those customers might have converted regardless of marketing, while the campaign's actual effect on the much larger group of people who never became best customers goes unmeasured.


Selection bias shows up when the exposed and unexposed groups being compared were never actually comparable to begin with. A classic version: a brand notices that customers who open its emails have a much higher repeat purchase rate than customers who don't, and concludes email marketing is driving loyalty. But people who open emails are, on average, already more engaged with the brand — email engagement may be a symptom of loyalty rather than a cause of it. Without random assignment to an exposed and holdout group, it's very difficult to tell which direction the causation runs.


How Bad Measurement Quietly Destroys Marketing Budgets


Put these problems together and a predictable pattern emerges across ecommerce companies of almost every size. Budget flows toward channels that are best at claiming attribution credit — retargeting, branded search, email to existing customers — because they show the highest ROAS on a dashboard. Budget flows away from channels that build new demand but convert less immediately, like upper-funnel prospecting or brand awareness, because their attributed ROAS looks weaker even when their incremental contribution is larger.


Over a year or two, this reallocation can leave a company running an increasingly efficient-looking marketing engine that is, in incremental terms, barely growing the business at all. The dashboards keep reporting green. The actual growth rate tells a different story. This gap is one of the most common reasons a company's reported marketing efficiency and its actual revenue growth start to diverge, and it's rarely obvious until someone runs the holdout test that finally exposes it.


How Incremental Thinking Changes Product Decisions


Incrementality isn't only a marketing measurement question — the same logic applies to product and feature decisions. A recommendation widget that shows "customers also bought" items might be credited with a large share of attached-item revenue in a dashboard, but the honest question is whether those customers would have added a second item anyway, simply through normal browsing. Running the widget as an A/B test with a true control group, rather than just measuring attach rate among people who saw it, answers that question directly instead of assuming it.

The same applies to loyalty programs, free shipping thresholds, and personalization features. Each one can be evaluated by attributed engagement, which almost always looks favorable, or by incremental lift against a holdout, which tells you whether the feature is actually changing behavior or just riding along with behavior that was already happening.


How Incremental Revenue Changes Executive Decisions


At the executive level, the shift from attribution to incrementality changes what gets asked in a budget meeting. Instead of "which channel has the best ROAS," the question becomes "which channel, if we cut it entirely, would actually hurt revenue." Those two questions frequently point at different channels, and the second one is the one that actually protects the business from cutting something valuable or funding something inert.


It also changes how growth targets get set. A team that has only ever looked at attributed revenue tends to overestimate how much of its growth is actually driven by marketing spend, versus organic demand, word of mouth, or product quality. Incrementality testing puts a number on that split, which tends to be humbling the first time a company runs it seriously.


Why Investors Are Paying More Attention to Incrementality


Investors evaluating ecommerce and DTC businesses have grown more skeptical of growth stories built entirely on attributed marketing performance, partly because the last decade produced enough cautionary examples of companies that scaled spend on channels that looked efficient on paper and collapsed in growth the moment spend was cut. A company that can show incremental lift — that its marketing genuinely creates demand rather than just capturing demand that existed anyway — is making a fundamentally stronger case about the durability of its growth. This is one of the reasons due diligence processes increasingly ask not just for ROAS by channel, but for evidence of holdout testing and true incremental contribution.


Privacy Changes and the Slow Death of Attribution


The attribution infrastructure most ecommerce brands rely on was built in an era of persistent third-party cookies and largely unrestricted cross-site tracking. That era has been ending for several years, through browser-level tracking restrictions, platform-level privacy changes, and regulatory pressure across multiple markets. Each of these developments makes last-click and multi-touch attribution progressively less reliable, because the underlying tracking signal they depend on is thinner than it used to be.


Incrementality testing is largely immune to this shift, because it doesn't depend on tracking an individual user's journey across touchpoints at all. It only requires knowing which group a customer was randomly assigned to and what they purchased. As tracking signal continues to degrade, that structural advantage becomes more, not less, important.


Where Ecommerce Measurement Is Heading


The direction of travel is fairly clear across the industry. Platforms and tools are beginning to move past dashboards that simply report attributed conversions, toward built-in experimentation — holdout groups, causal analysis, and lift measurement designed into the product rather than bolted on as a separate research project. This shift mirrors what happened in product analytics a decade ago, when A/B testing moved from a specialized discipline into a standard, expected capability.


Peloran is one of the platforms working in this direction, building incrementality measurement — holdout-based testing and causal lift analysis — directly into its ecommerce intelligence layer, rather than treating it as a separate audit a merchant has to commission externally. The goal is the same one this article has been making the case for: replacing "which channel gets the credit" with "which channel actually creates the revenue."


Frequently Asked Questions


Is incremental revenue the same as attributed revenue?

No. Attributed revenue assigns credit for a sale to a touchpoint based on a rules-based model, such as last click. Incremental revenue measures the actual causal lift a campaign or channel produced, calculated by comparing an exposed group against a genuinely comparable holdout group.


Why does attributed revenue usually overstate marketing impact?

Because it counts sales that would have happened anyway as if the marketing caused them, particularly for channels like retargeting and branded search that tend to reach people already close to purchasing.


How long does a holdout test need to run?

It depends on baseline traffic, conversion rate, and the size of the effect being measured. Low-traffic stores generally need longer test windows to reach statistical confidence than high-traffic stores testing the same effect size.


Can small ecommerce brands run incrementality testing, or is it only for large companies?

Smaller brands can run it, but they need larger effect sizes or longer windows to detect a reliable result, since statistical power depends heavily on sample size. It's more constrained than for a high-traffic retailer, but not out of reach.

Does incrementality testing replace attribution entirely?

Not necessarily. Attribution still has value for understanding sequence and touchpoint behavior. The shift is in trusting attribution for budget and strategic decisions less, and trusting incrementality testing for those decisions more.


The Bottom Line


Attribution answers a narrow, mechanical question: which touchpoint was closest to a sale. Incrementality answers the question that actually determines whether a business is growing: what would have happened without this spend, this discount, or this feature. Those are not the same question, and treating them as interchangeable is how marketing budgets quietly drift toward channels that are good at claiming credit rather than channels that are good at creating demand.


None of this requires abandoning dashboards or attribution software. It requires holding them at arm's length, testing their claims against a real control group often enough to know where they're reliable and where they aren't, and building decisions — budget, product, growth targets — around the answer to the harder, more honest question: what actually moved because we acted, and what would have happened either way?

Ready to transform your e-commerce business?

Join hundreds of Shopify stores using AI to grow smarter.

Get Started Free