This website uses cookies

Read our Privacy policy and Terms of use for more information.

🧩 The answer

More products can make a store more useful. A shopper is likelier to find the right size, style or price; a marketplace can attract people it could not serve before. But the benefit changes once someone has to choose one item from a crowded set. Ten nearly identical products can be harder to buy from than fifty products with differences you can spot in seconds.

That is why “How many SKUs should we stock?” is usually the wrong first question. The shopper sees a particular set, for a particular task, with a particular level of confidence. The studies below show where more choice starts to impose a cost — and how to test a smaller view without discarding the value of range.

Short answer

There is no universal number where choice becomes too much. The danger rises when products are hard to compare, the shopper is unsure what matters, or they want to finish quickly. In those situations, a focused starting set may help; for someone who knows exactly what they want, the same cut can remove a valuable option. One hotel-platform experiment found precisely that split: reducing the displayed range helped less familiar shoppers but hurt more familiar ones. So keep breadth where it creates discovery and access, then test how much each shopper must evaluate at once. Judge the result by completed purchases and repeat orders, not clicks or a single average conversion rate.

🗺️ Jump to a recommendation

🧭 01. Find the hard choices

👉 Do this

Before cutting products, find the moments when the shopper has the hardest decision. Start with a category where many options look alike, important attributes are difficult to judge, or people often leave after comparing several items. Ask what the shopper is trying to do: browse for ideas, find a known product, or make a purchase now. These are different tasks and should not receive one blanket assortment rule.

Treat a large range as a possible source of friction only when the decision context points that way. If shoppers can spot a clear winner or already know their requirements, removing options may destroy useful variety. Make the test about the set they have to evaluate at once, with the full range still reachable while you learn.

💡 Why

A meta-analysis of 99 observations from earlier studies found that four conditions made overload more likely: a more complex choice set, a more difficult task, uncertain preferences, and a goal of minimizing effort. The work explains why a big assortment can be attractive when people are browsing yet costly when they must commit to one product. It does not identify a magic number of SKUs; the same count can feel simple or exhausting depending on the decision.

🛠️ How

Choose one category and record its current visible product count, key differences and decision path. Segment sessions by intent signals you can actually observe, such as a precise search versus broad browsing. Compare the current view with a focused starting view while keeping search and ‘see all’ available. Read completed purchases, choice deferral and subsequent returns together; a faster click is not enough if the purchase does not follow.

🖼️ Good / Bad illustration

One side: a shopper facing near-identical products with unclear differences. Other side: the same stocked range, but an initial set organized around the few attributes that matter for this task. The image will label the tested decision context, not suggest a universal SKU limit.

📖 Evidence

Chernev and coauthors synthesized 53 studies across 21 articles (7,202 participants). Their four moderators were statistically significant in their model. The evidence mostly comes from controlled choice tasks, so the proposed ecommerce test is an application to validate locally. Chernev, Böckenholt & Goodman, 2015.

👋 02. Start smaller for unfamiliar shoppers

👉 Do this

Give a shopper who knows little about the category or destination a smaller, relevant first set to inspect. The aim is to lower the work needed to start choosing, not to make the rest of the range disappear. Leave an obvious route to search, filters and the full catalog. Avoid a global cut that also constrains experienced shoppers who may have come for a specific niche option.

Define ‘unfamiliar’ with a signal connected to the actual decision. Someone can be a loyal customer and still be new to a category; a first-time visitor can already know exactly which model they want. Validate that your signal predicts familiarity before using it to change the view.

💡 Why

On a hotel-booking platform, a randomized 40–60% reduction in the displayed assortment raised completed bookings by 0.4 percentage points for less familiar shoppers, while lowering them by 0.6 points for more familiar shoppers. The average alone would have hidden two opposite outcomes. The study’s familiarity measure was tied to distance from the destination, so it cannot be replaced automatically with a generic new-versus-returning customer flag.

🛠️ How

Pick a search result or category with a meaningful range. Randomize the focused initial set against the current view and keep total inventory accessible in both arms. Predefine the familiarity signal and analyze completed orders by that segment, not only in aggregate. Check whether shoppers use ‘see all’ or search to recover hidden options, and stop the rollout if the knowledgeable segment loses purchases.

🖼️ Good / Bad illustration

One side: a first-time destination shopper receives an unfiltered wall of hotels while a frequent destination shopper loses niche hotels after a universal cut. Other side: a focused first view for the unfamiliar shopper and full access for the familiar shopper.

📖 Evidence

Aparicio, Prelec and Zhu ran the Despegar experiment over 47 days with 322,800 users. The opposite booking effects were segment outcomes from that platform and familiarity proxy; they are not a general 40–60% reduction recipe. Aparicio, Prelec & Zhu, 2025.

🔍 03. Make differences easy to judge

👉 Do this

When products differ on attributes shoppers cannot evaluate confidently, make the differences legible before deciding that the whole assortment is too large. Show a small number of decision-relevant attributes in the same order and units across products. Explain what an unfamiliar specification means in use, and give a simple way to compare a few candidates side by side. Let shoppers narrow the view by the job they need done.

This is a design hypothesis to test, not a claim that one particular comparison widget has already been proven. Start where low evaluability is obvious — for example, technical materials, compatibility or performance claims — and avoid adding extra labels that create yet another thing to decode.

💡 Why

Across four experiments, Wang and coauthors found that larger assortments made choice harder for products that participants found difficult to evaluate. When the products were easier to evaluate, the difference between large and small sets was not significant in their reported choice tasks. That points to a useful distinction: the same number of products can carry a different decision cost depending on how easy it is to understand their attributes.

🛠️ How

Select a category with complex specifications. Keep the actual products and prices fixed, then test a clearer comparison view against the current presentation. Track completed purchases and time to decision, with returns or support questions as follow-up signals where available. If clarity improves outcomes, test whether a larger visible set can be retained; if it does not, try a smaller initial set rather than assuming more labels solved the problem.

🖼️ Good / Bad illustration

One side: product cards with incomparable technical claims and no plain-language meaning. Other side: the same products aligned on three useful attributes, with a short explanation of what each means for the buyer. This is an illustration for a future test, not a studied interface.

📖 Evidence

Wang and coauthors report four experiments with fictional and real products, comparing larger and smaller sets under different levels of product evaluability. Their reported outcome concerns experimental choice, not a live comparison-layout conversion test. Wang, Wang, Han & Wang, 2025.

📊 04. Measure new and repeat buyers separately

👉 Do this

When you add products, keep two scorecards. First, ask whether the wider range attracts buyers you could not serve before: new customers, first orders and the revenue they bring. Second, ask what happens to people who already buy: order frequency, retention and revenue per existing customer. A range expansion can win on the first scorecard and lose on the second, so a single blended conversion number can mislead the team.

Look closely at added options that are substitutes for products existing customers already know. They may add selection work without adding a new reason to visit. Preserve the options that bring genuine new demand, then test a more focused path for returning buyers rather than deleting products solely because the average looks flat.

💡 Why

In a food-delivery platform study, Natan found that assortment expansion increased acquisition of new consumers but reduced purchase frequency among consumers who stayed on the platform. The paper models costly attention as one explanation and uses a model counterfactual to assess targeted reductions. It does not show that every retail category has the same tradeoff, or that a particular first-screen design will reverse it.

🛠️ How

Before an expansion, define cohorts and a time window long enough to observe repeat behaviour. Compare exposed and unexposed markets, categories or randomized views where feasible. Report new-customer orders separately from existing-customer order frequency, retention and contribution after costs. If acquisition rises while repeat use falls, test a returning-customer starting view that surfaces likely fits while keeping discovery of the wider range possible.

🖼️ Good / Bad illustration

One side: a single green ‘conversion’ line hides weaker repeat purchasing. Other side: new-buyer acquisition and existing-buyer order frequency appear as separate outcomes beside the same expanded catalog. No illustrative number will be presented as study data.

📖 Evidence

Natan analyzes changes in restaurant availability over consumer lifetimes on one food-delivery platform. The direction of acquisition and repeat-frequency results is stated in the published abstract; targeted-reduction revenue is a model counterfactual, not a live test. Natan, 2024/2025.

📚 Sources