A/B Testing Your Return Policy
Most return policies are not decided; they are inherited. A merchant copies the window a competitor publishes, or lands on whatever the legal minimum allows, writes it once, and never touches it again. That is strange, because the return policy is one of the highest-leverage numbers in the business — it sits directly on the conversion path at checkout and directly on the cost line in the warehouse. A team that would never ship a checkout redesign without an experiment will change a 30-day window to 60 days on a hunch, or leave a paid-return policy in place for years without ever measuring what free returns would do. The return policy is a growth lever, and like any growth lever, it should be tested, not assumed.
The return policy is a two-sided lever
What makes return-policy testing harder than testing a button color is that almost every change moves two numbers in opposite directions. Make returns free and conversion rises, because the perceived risk of buying drops — and return volume rises too, because the friction that used to suppress marginal returns is gone. Extend the window from 30 to 60 days and you lift conversion a little while lengthening the liability tail and nudging up the return rate. Add a restocking fee and you claw back per-return cost while dampening the very purchases the generous policy was winning. You cannot judge either side in isolation. A test that only watches conversion will tell you free returns are a triumph; a test that only watches return rate will tell you they are a disaster. Both are half-readings of the same coin.
The way out is to commit, before the test starts, to a single metric that nets the two sides against each other. The right north-star is net contribution: the incremental gross profit from the conversion lift, minus the incremental return-and-fulfillment cost the change creates, extended where you can measure it by the effect on repeat purchase and lifetime value. A policy change that lifts conversion two points but raises return cost enough to erase the added margin is a losing change, and only a netted metric will tell you so. A change that lifts conversion less but converts refunds into exchanges can win on net contribution even though the raw return rate barely moves.
| Lever | Test variants | Expected conversion effect | Expected return-cost effect |
|---|---|---|---|
| Return window length | 14 vs 30 vs 60 days | Longer lifts conversion modestly | Longer raises return rate and liability tail |
| Return shipping cost | Free vs customer-paid | Free clearly lifts conversion | Free raises return volume and cost |
| Default outcome | Refund-first vs exchange-first | Neutral to slight lift | Exchange-first cuts net refund cost |
| Restocking fee | None vs small fee | Fee dampens conversion | Fee recovers per-return cost, may cut volume |
What to test, and what to expect
Start with the levers that move the most money, and form a directional hypothesis for each before you run it so you are confirming a prediction rather than fishing in the results. Window length is the natural first test, and the counterintuitive finding merchants keep rediscovering is that a longer window often lowers the return rate rather than raising it, because urgency fades and the item becomes part of the customer's life — a dynamic worth reading up on before you test, in the work on optimal return window length. Free versus customer-paid return shipping is the highest-variance test on the list, moving both conversion and volume hard, so it deserves the most careful measurement. The default outcome test — presenting exchange and store credit ahead of a cash refund — is often the best risk-adjusted bet, because it tends to protect conversion while cutting net refund cost. And a restocking fee test almost always dampens conversion, so it should be judged strictly on net contribution, never on the return-rate improvement alone, which will always look flattering in isolation.
If a policy test only reports the return rate, it is designed to be passed. Judge it on net contribution or do not run it.
The north-star metric
Net contribution per exposed visitor is the number to instrument. For each variant, take the gross profit from orders placed, subtract the fully loaded cost of the returns those orders generate — shipping both ways, handling, inspection, markdown on returned units — and, where your data allows, add the difference in repeat-purchase value between the cohorts. The discipline of controlled experiments is not new, and general treatments of experimentation from sources such as Harvard Business Review make the same core point that applies here: an experiment is only as trustworthy as the metric it optimizes and the rigor of its design. If you optimize return rate, you will systematically choose worse policies, because the cheapest way to cut returns is to cut sales. If you optimize net contribution, the trade-off is priced in automatically. This is also why policy testing should sit alongside the operational work of reducing your return rate rather than substituting for it — a good policy and a low return cause are complementary, not interchangeable.
Measurement pitfalls that will fool you
Return-policy tests are unusually easy to misread, because the two sides of the lever resolve on different clocks. Conversion moves the moment a visitor sees the policy; the return cost lands weeks later, when the item actually comes back. Read the test at the point where conversion looks great and returns have not matured, and every generous policy looks like a winner. Three pitfalls in particular catch teams out.
- Return lag: the conversion lift is immediate but the return-cost delta arrives across the following weeks. Do not call a test until the exposed cohort's returns have substantially matured, or you will bank a win that erodes.
- Seasonality and contamination: a window-length or holiday-adjacent test spans periods with very different return behavior, and returns generated during the test may be governed by whichever policy the customer saw at purchase — keep cohorts cleanly separated by exposure, not by return date.
- Small samples and long feedback loops: because you must wait for returns to mature, policy tests need more traffic and more patience than a typical UI test. Underpowered tests read noise as signal, especially on the return side where events are rarer than conversions.
This is where returns intelligence stops being a reporting nicety and becomes the instrument the whole test depends on. ResReturn's analytics are built to attribute returns back to the order and the cohort that produced them, track net contribution per variant as returns mature rather than at the flattering early moment, and hold structured return reasons so you can see not just that a policy changed the return rate but why. Pair that with the platform's ability to run different outcome rules — exchange-first, store-credit, window length — as configurable variants, and the return policy becomes something you can actually experiment on with the same rigor you bring to the rest of the funnel, instead of a static paragraph nobody is allowed to touch.
- Treat the return policy as a testable growth lever, not a fixed legal footnote copied from a competitor.
- Pick net contribution — conversion-driven margin minus fully loaded return cost, plus LTV where measurable — as the single north-star before the test starts.
- Form a directional hypothesis per lever: window length, free versus paid shipping, exchange-first default, restocking fee.
- Wait for returns to mature before reading results, because the conversion win and the return-cost bill land on different clocks.
- Separate cohorts by policy exposure at purchase, power the test for the rarer return events, and never let return rate alone decide a winner.
What metric should I use to judge a return-policy test?
Net contribution per exposed visitor: the incremental gross profit from the conversion effect, minus the fully loaded cost of the returns that orders generate, plus the difference in repeat-purchase value where you can measure it. Optimizing return rate alone systematically selects worse policies, because the cheapest way to cut returns is to cut sales. A netted metric prices the trade-off in automatically.
Which parts of a return policy are worth A/B testing?
The levers that move the most money: return window length, free versus customer-paid return shipping, the default outcome (refund-first versus exchange-first), and whether to charge a restocking fee. Each moves conversion and return cost in opposite directions, so each needs a netted metric. The exchange-first default is often the best risk-adjusted test because it tends to protect conversion while cutting net refund cost.
Why do return-policy tests take longer than normal A/B tests?
Because the two effects resolve on different clocks. Conversion moves the instant a visitor sees the policy, but the return cost only lands weeks later when items actually come back. If you read the test before the exposed cohort's returns have matured, every generous policy looks like a winner. Policy tests therefore need more traffic and more patience than a typical UI experiment.
Can a longer return window actually reduce returns?
Often, yes. A longer window removes urgency, the item becomes part of the customer's routine, and the impulse to send it back fades — an endowment effect that can lower the return rate even as it lifts conversion. It is one of the clearer cases where the intuitive prediction is wrong, which is exactly why the lever is worth testing rather than assuming.
See it on your own returns.
Start freeKeep reading
From Apology to Advocacy After a Return
A great return recovery creates advocates. Learn the service-recovery moves that turn a disappointed returner into a repeat buyer and a referral, not a churn.
Building a Branded Returns Portal Customers Trust
A branded returns portal keeps shoppers on-brand through the refund moment. See how logo, domain, and tone in your returns portal build repeat trust.
