Return Fraud: The Signals, and How to Score Them
Return fraud is the line item nobody puts on the P&L, because it hides inside a number everyone already accepts as normal. Fashion and apparel run 30 to 40 percent return rates, and inside that band sits a smaller, more expensive slice that is not dissatisfaction at all. It is a customer who wore the dress to the wedding, a box that came back a pound too light, an order of five sizes designed to keep one and expense the rest to your reverse-logistics budget. You paid for the flexibility your honest buyers love, and a small malicious minority is cashing it in. This is a field guide to recognizing that minority by its signals, and scoring it out of the flow, without laying a finger on the customers who make up the overwhelming majority of your volume.
One principle sits above everything that follows: a fraud-prevention system that punishes everyone costs more than the fraud it catches. Demand a photo from every returner, manually approve every parcel, tax every refund with a fee, and you will save a few hundred dollars in fraud while shedding thousands in conversion and repeat purchase. So test every control against a single question. Is this targeting risk, or is it spreading friction across people who never did anything wrong? If the answer is the second one, it is not a control. It is a leak.
The six shapes fraud actually takes
You cannot score what you cannot name. Each of these has an operational signature, a specific thing it does to your data that a rule can see even when a warehouse clerk cannot.
Wardrobing
The most common and the hardest to catch. Buy the item, use it once, the suit for the interview, the jersey for the one match, the dress for the event, then send it back as did not like it. It may look new and still be dead stock: sweat, perfume, makeup transfer, a stretched seam, a faint wash line. The tell is timing. Wardrobing clusters around events and the very end of your return window, because the customer needed the item until the last legal minute. Short hold time on a high-value garment returned days before the window closes is not a coincidence.
Empty-box and wrong-item returns
The parcel comes back with nothing in it, or with a brick, an old shoe, a competitor's product standing in for yours. This one is brutal in any operation that refunds on scan instead of on inspection, because the money is gone before anyone opens the box. The signal is physical: the return weight does not match the outbound weight. Weigh on intake and this fraud has nowhere to hide. Refund before intake and you will never even know it happened.
Worn or damaged returns
The item comes back used up, and the customer insists it arrived that way or was defective out of the box. The hard part is that genuine defects exist and deserve fast, no-questions refunds. What separates the two is not any single return, it is the rate. A customer whose defective reason fires far above your baseline is telling you something a one-off claim never could.
Bracketing abuse
Ordering the same item in several sizes and keeping one is normal shopping, not fraud, and any control that treats it as fraud will torch your best customers. It only crosses the line when it is systematic and excessive, four or five sizes every single order, every time, keeping one and returning the rest as a matter of routine. The control has to be aimed at the pattern, not the behavior, because the behavior itself is something you want to accommodate.
Receipt and label fraud
Returning goods bought elsewhere at your full price, altered receipts, a copied return label riding on someone else's order. An order-based return portal largely closes this on its own, because a return can only exist against a real order record. If the request cannot bind to an order you actually shipped, it never enters the flow.
Serial returners
A small group generates a wildly disproportionate share of your returns, usually by combining several of the behaviors above. Any single return from this profile looks innocent. The pattern only appears once you accumulate order and return history per customer over time. This is the one type that is invisible at the parcel and obvious in the data, which is exactly why scoring beats inspection here.
Fraud type, signal, control
Map each type to the signal that exposes it and the control that answers it. Read across a row and you see which lever actually works against which risk, and why buying one blanket control never covers the board.
| Fraud type | Warning signal | Control that answers it |
|---|---|---|
| Wardrobing | Post-event return, signs of use, short hold time near window close | Warehouse inspection and grading, tag integrity, tighter window in high-risk categories |
| Empty box or wrong item | Return weight mismatch, item mismatch on intake | Weigh and photo on intake, no refund before receipt |
| Worn or damaged return | Defective reason firing above baseline, repeat damage claims | Standard grading, photo requirement on damage claims |
| Bracketing abuse | Many sizes of one item, high return share, systematic repeat | Size recommendation, exchange-first flow, threshold tracking |
| Receipt or label fraud | Return that binds to no real order record | Order-based portal verification |
| Serial returners | Abnormal return rate per customer over time | Rule-based risk scoring, tiered controls |
The controls, in the order you should build them
None of these is a magic switch. They work as layers that feed each other, and they work in this sequence because each one makes the next one cheaper.
1. Write a policy you will actually enforce
Policy is the first line, and an unenforced policy is not a deterrent, it is a suggestion. Set the window to the category and shorten it where wardrobing risk runs hot. Require sellable, tagged, unused condition in plain language. Name the items that cannot come back at all, hygiene, personalized, final sale. Then enforce it the same way every time, because inconsistency is what abusers probe for.
2. Inspect and grade at the warehouse
Most fraud only becomes visible when the product lands. Run the same checklist on every parcel: does the weight match, is it the right item, is the tag intact, are there signs of use. Grade the outcome consistently, resellable, discounted, non-returnable, so the decision is a category and not a mood. Crucially, do not refund before this step on anything your scoring has flagged.
3. Track return rate per customer
A single return means nothing. The accumulated ratio means everything. Roll up each customer's order-to-return rate, their reason mix, and how often the defective reason fires. Without this history you literally cannot tell a serial returner from someone who had two bad weeks, and you will end up guessing, which means you will end up wrong on both sides.
4. Score every return with rules you can read
This is the core of ResReturn. Every return request gets a risk score built from three families of signal. Velocity, how fast and how often this customer returns. Amount, the value at stake on this specific request. Pattern, the fingerprints of the behaviors above, size spread, defective-reason frequency, hold time, address and payment anomalies. The scoring is rule-based on purpose, so you can read exactly why any given return landed where it did, and defend the decision to a customer or a chargeback team without shrugging at a black box.
The score resolves to one of three outcomes. Auto: the return clears frictionlessly, refund or exchange, no human touch, which is where the honest majority lives. Instant-credit-off: the return proceeds but instant credit is withheld until the parcel is received and graded, closing the empty-box and wrong-item window without refusing anyone outright. Flagged: the return is routed to manual review, and a merchant can require a photo, hold the refund, apply a fee where law allows, or refuse. Thresholds are yours to set, because your tolerance is not our tolerance.
The point of scoring is not to catch everyone. It is to let the ninety-plus percent who are honest never notice a control exists, while the small minority that signals risk meets exactly as much friction as its score earns.
5. Ask for reasons and photos, but only where warranted
A structured reason, chosen from a list rather than typed as free text, buys you both analytics and deterrence in one field. Requiring a product photo on a damage claim visibly deflates false defect reports. But attach that photo requirement to the flagged flow, not to everyone. The moment you ask a clean customer to photograph a jacket that fit fine, you have converted a fraud control into a conversion problem.
6. Use fees and exchanges as the deterrent, not the punishment
Restocking fees are a legitimate lever against wardrobing and bracketing in the categories and jurisdictions that allow them, provided you comply with consumer law and disclose the fee up front. Hidden fees cost more trust than they recover margin. For most brands an exchange-first flow with a store-credit incentive does the same job more gently: it turns the refund reflex into a swap for the right size, which also attacks the root cause of bracketing, uncertainty about fit, before the return ever exists.
Protecting the fit-graph from poisoned returns
There is a second, quieter reason flagged returns matter, and it is one most fraud discussions miss entirely. ResReturn learns fit from returns. When a customer returns an item as too small or too tight, that reason feeds the fit-graph that recommends sizes to the next shopper. That loop is only as trustworthy as the returns feeding it, which means a fraudulent return is not just a lost margin event, it is a data-poisoning event.
A wardrober who returns a perfectly fitting dress as too small is lying to your recommendation engine, not just your refund desk. A bracketing abuser who sends back four sizes muddies the signal about which size actually fits their body. Left in the dataset, these events drag the fit-graph toward noise and eventually steer honest shoppers toward the wrong size, which generates more genuine returns. Fraud that corrupts fit data compounds.
So flagged returns are auto-excluded from the fit-graph. The moment scoring flags a return, its fit reason stops counting toward recommendations. The signal that trains the model is drawn only from returns that cleared as genuine. You get two protections from one score: the margin defense on the flagged return itself, and the integrity defense on every future recommendation that would otherwise have been trained on a lie.
Tiering: friction proportional to risk
Everything above collapses into one operating rule. Do not apply controls uniformly. Apply them in tiers, matched to the score.
- Low-risk majority: instant approval, easy exchange, fast refund, zero added steps. Most of your customers should never learn the scoring exists.
- Medium-risk: structured reason required, instant credit held until receipt, an exchange or credit incentive offered before a refund.
- High-risk minority: flagged for manual review, photo required, refund held pending grading, fee or refusal where justified and legal.
Spread friction evenly and you punish the many to inconvenience the few. Route it by score and the honest customer experiences a security layer they never even feel, while the abuser meets a wall calibrated precisely to what their behavior earned. That asymmetry, invisible to the good and unavoidable for the bad, is the entire goal.
What actually counts as return fraud?
Any abuse of the return process for illegitimate gain. The common shapes are wardrobing, using an item once then returning it as new, empty-box and wrong-item returns, passing off worn or self-damaged goods as defective, systematic bracketing built purely to expense reverse logistics, receipt and label fraud, and serial returners who chain several of these together at high volume.
Which type hits fashion merchants hardest?
Wardrobing and bracketing abuse. Wearing a garment once and returning it, and ordering many sizes to keep one as a routine, are the behaviors apparel brands see most often and the ones that erode margin fastest, because both target exactly the flexible policies that drive your legitimate sales.
How do I stop fraud without punishing honest customers?
Score risk instead of applying controls uniformly. Rule-based scoring over velocity, amount, and pattern resolves each return to auto, instant-credit-off, or flagged. The honest majority clears automatically and never notices a control. Only the small high-score minority meets photos, holds, or review. You target risk, you do not tax everyone.
Why exclude flagged returns from the fit-graph?
Because ResReturn learns fit from return reasons, and a fraudulent return carries a false reason. A wardrober returning a fitting dress as too small trains the recommendation engine on a lie, which steers future shoppers toward the wrong size and manufactures more genuine returns. Auto-excluding flagged returns keeps the fit signal drawn only from returns that cleared as real.
Which controls give the most protection per unit of friction?
In order: an enforceable order-based return policy, consistent warehouse inspection and grading, per-customer return-rate tracking, rule-based risk scoring with tiered outcomes, reasons and photos on flagged flows only, and fees or exchange incentives where they fit your law and category. They compound as layers and underperform in isolation.
See it on your own returns.
Start freeKeep reading
From Apology to Advocacy After a Return
A great return recovery creates advocates. Learn the service-recovery moves that turn a disappointed returner into a repeat buyer and a referral, not a churn.
Building a Branded Returns Portal Customers Trust
A branded returns portal keeps shoppers on-brand through the refund moment. See how logo, domain, and tone in your returns portal build repeat trust.
