All articles
IntelligenceAug 18, 2026 · 7 min

How to Grade Returned Inventory Accurately

DA
Defne Aksoy
Reverse Logistics Manager

Two warehouse graders inspect the same returned jacket. One marks it 'Grade A — resell at full price.' The other marks it 'Grade B — discount 30%.' Multiply that inconsistency across thousands of SKUs a month and the gap between what a store could recover and what it actually recovers widens into real money left on the table. Grading is the single decision point in reverse logistics that determines whether a returned item becomes a full-price resale, a discounted outlet sale, a refurbishment candidate, or scrap — and most merchants still leave it to gut feel.

Inconsistent grading is one of the most under-discussed causes of margin leakage in returns operations. A recommerce operations study cited by industry analysts found that subjective, ungoverned grading destroys recoverable resale value at a scale most finance teams never trace back to the returns desk — because the loss shows up as 'lower than expected recovery rate' rather than as a line item anyone audits. Fixing it doesn't require new headcount. It requires a rubric two different people can apply and reach the same answer.

Why grading consistency is worth solving first

Grading sits upstream of every downstream decision: pricing, channel routing, and even whether an item is worth touching at all. If grading is inconsistent, the returns disposition rules that route items to resale, liquidation, refurbishment, or recycling are working off bad inputs — so no disposition engine, however smart, can fully compensate. Get grading right and everything downstream gets more accurate for free.

  • Overgraded items get listed at a price the condition can't support, driving return-of-return cycles and refunds.
  • Undergraded items get routed to liquidation or outlet channels at a steep discount when they could have sold at near-full price.
  • Grader-to-grader variance makes recovery-rate reporting unreliable, so leadership can't tell if a channel or a policy change actually worked.
  • Inconsistent grades erode trust between operations and finance, since neither side can explain swings in recovered value.

A five-tier grading rubric that travels

The rubric below is deliberately simple: five tiers, each with an objective test rather than a subjective impression. The goal is that a new hire on day three and a five-year veteran assign the same grade to the same item, every time.

GradeDefinitionObjective testTypical disposition
A — New/UnwornNo signs of wear, tags attached, original packaging intactTag verification + zero visible marks under standard lightFull-price resale
B — Like NewTried on, no wear marks, packaging may be openedNo stains, no odor, no thread pulls under close inspectionResale at 10-15% discount
C — Light WearMinor cosmetic flaw: light crease, small mark, missing hangtagFlaw is under 2cm and does not affect functionOutlet channel or bundle resale
D — Damaged/FunctionalVisible damage but item is usable: stain, small tear, broken zipper pullDamage documented with photo, function test passedRefurbishment candidate
E — Non-RecoverableStructural damage, contamination, or safety issueFails function or hygiene testRecycle or scrap

Each tier maps to a default disposition, but the mapping should stay a default, not a hard rule — a Grade C electronics item might still be worth refurbishment if the part cost is low and resale value is high, while a Grade C garment might go straight to outlet. The rubric standardizes the grade; the disposition engine still optimizes the channel.

The cheapest returns loss to fix is the one caused by two people looking at the same item and writing down two different answers.

The photo standard that makes grading auditable

A grade without a photo is an opinion with no paper trail. Every graded item should generate a consistent photo set so disputes, audits, and — increasingly — computer-vision return grading automation have something reliable to work from. Standardizing the capture is what makes automated grading models trainable in the first place; inconsistent lighting, angles, and backgrounds are the number one reason vision models underperform in pilot programs.

  1. 1Full front shot on a neutral, consistent background with standardized lighting.
  2. 2Close-up of any flaw, damage, or wear mark with a ruler or reference object for scale.
  3. 3Tag/label shot showing SKU, size, and original packaging condition.
  4. 4Interior or underside shot for footwear, bags, and electronics where wear concentrates.
  5. 5Timestamp and grader ID embedded in the file metadata, not just the filename.

Building grader agreement into the process

Rubrics fail in practice when they aren't reinforced. Retailers that get grading consistency right typically run three things in parallel: a shared reference photo library showing what each grade looks like across categories, periodic calibration sessions where graders re-score the same sample set and compare notes, and a spot-audit process where a supervisor re-grades a random 5% of items each week. None of this is expensive. It's discipline, not tooling — though tooling makes the discipline easier to sustain at volume.

The financial case for this discipline is not abstract. Retail returns represent a substantial share of e-commerce revenue, and NRF's annual returns research consistently shows that a meaningful percentage of returned merchandise loses resale value before it ever reaches a resale channel — much of it due to handling and grading delays rather than actual product damage. Every day an item sits waiting for a confident grading decision is a day closer to it aging out of full-price eligibility.

Where automation fits — and where it doesn't yet

Computer vision models can now flag obvious flaws, measure stains and tears against reference scales, and pre-sort items into likely grade bands faster than a human can pick the item up. What they can't yet do reliably is judge borderline cases — a B versus C call on a garment with a barely visible pull, for instance — without a human final check. The most effective setups today treat vision models as a first-pass triage layer that speeds up the obvious 80% of decisions, freeing graders to spend their attention on the 20% that actually require judgment. That's a very different design goal from full automation, and it's realistic for most mid-size retailers within a single budget cycle.

Putting it into an operating rhythm

A grading rubric only pays off if it's embedded in the daily workflow, not laminated and pinned to a wall. Practical rollout steps that work across store sizes:

  • Print the five-tier rubric and photo checklist at every grading station, not just in a shared drive.
  • Require the photo set before a grade can be submitted in the system — make it a hard gate, not a suggestion.
  • Route Grade D items automatically to a refurbishment queue rather than letting them sit in a general holding area.
  • Review grade-to-recovery-rate correlation monthly so pricing and disposition rules can be recalibrated against real outcomes.
  • Recalibrate the rubric itself once or twice a year as product mix and channel economics shift.

Grading is the least glamorous step in reverse logistics and the one with the most leverage over recovered value. Get it standardized, photographed, and lightly automated, and the disposition, pricing, and refurbishment decisions that follow all get sharper — without touching headcount.

How many grading tiers should a mid-size retailer use?

Five tiers (A through E) is enough granularity for most categories without slowing graders down. Fewer than four tiers tends to blur pricing decisions; more than six creates disagreement at the boundaries and slows throughput without adding useful precision.

Should grading rubrics differ by product category?

The tier structure should stay consistent across categories so reporting stays comparable, but the objective tests within each tier should be category-specific — a scuff test for footwear looks different from a function test for electronics or a stain test for apparel.

Can computer vision fully replace manual grading?

Not yet for most retailers. Vision models are strong at flagging obvious damage and speeding up triage, but borderline grade calls still benefit from human judgment. Treat automation as a first-pass filter, not a full replacement, until accuracy on your specific catalog is proven over several months.

What's the fastest way to reduce grader disagreement?

Run a calibration session where multiple graders independently score the same sample set of 20-30 items, then compare and discuss any grade that differs by more than one tier. Repeating this quarterly keeps standards from drifting as staff turns over.

See it on your own returns.

Start free