A personalized list can help a shopper find a suitable product. It can also move attention toward items the retailer wants to sell. A higher click rate alone does not tell you which of those happened.
The decision is how to evaluate personalized order. Compare it with a useful baseline, follow the path from discovery to purchase and inspect both the products chosen and the distribution of sales across customers.
🔎 What research says
Compare personalization with a credible alternative and separate product discovery from purchases. Check whether the chosen products offer better fit or value, while monitoring store-level sales concentration. Retailer gains and shopper gains need their own evidence.
🗺️ In this guide
1. Compare personalization with a useful baseline

📈 Recommendation
Test personalized order against a credible alternative, such as your current bestseller ranking. Define the customer and retailer outcomes before comparing them.
Keep the available products and commercial conditions constant; assign visitors to the competing orders.
Measure purchases and product discovery, then check whether the selected products fit the shopper’s task.
🎓 Findings
Donnelly, Kanodia and Morozov (2024; online 2023) studied a large retailer’s randomized introduction of personalized rankings. Compared with a uniform bestseller order, personalization increased search and purchases.
Their search model also estimated a gain in consumer surplus even though the algorithm gave positive weight to retailer profitability. The surplus result is a model estimate, not a directly observed cash saving for each customer.
🧠 Why it works
A bestseller list is a practical benchmark that already helps some shoppers. Beating it is more informative than beating an arbitrary order. More search can mean useful exploration, so a faster visit is not the only sign of success.
✋ Limitations
The result comes from one retailer and one implemented ranking system. It does not establish a general uplift or identify the best profitability weight for your store.
2. Separate discovery from completed purchases

📈 Recommendation
Track what happens after a recommendation is seen. Distinguish product views, purchases after a view and the final purchase rate.
Define the denominator for every measure before testing; report how many visitors actually reach each stage.
Compare relevant product types and review profiles, rather than assuming one average effect fits the range.
🎓 Findings
Lee and Hosanagar (2021; online 2020) ran a randomized retailer experiment with 184,375 users and analyzed 37,125 products. Their purchase-based recommender improved views, conversion after views and final conversion.
The effects differed by product attributes and review ratings. Products that gained more attention were not necessarily the ones with the largest improvement in conversion after a view.
🧠 Why it works
A recommendation can help someone discover a product without helping them choose it. Separating those steps shows where the system adds value and where a more attractive card merely draws attention.
✋ Limitations
The result concerns the tested recommendation algorithm and retailer. These stages use different denominators; they must not be added or treated as one sales forecast. Detailed uncertainty still needs full-paper extraction.
3. Check whether the suggested product fits better

📈 Recommendation
Evaluate the products a ranking helps people find. A useful suggestion should improve suitability, price or another relevant part of the choice.
For a concrete task, compare suggested products with the alternatives the shopper would otherwise inspect.
Watch use of search and filters alongside recommendations; a change in channel use may explain the result.
🎓 Findings
Wan, Kumar and Li (2024; online 2023) ran a randomized retailer experiment with an item-based recommender. Using algorithmic affinity scores, they developed measures of product value and taste fit.
Their analysis linked the purchase benefit mainly to finding lower-priced products, better-fitting products or both, rather than easier navigation alone. Shoppers also substituted recommendations for other search tools.
🧠 Why it works
A smoother route helps only if it leads to something worth choosing. Reviewing product fit connects ranking quality to the shopper’s task. Our proposed task-based check is a store application of the finding.
✋ Limitations
Taste fit and net value were inferred partly from algorithmic measures. They are not direct reports of every shopper’s needs. Effects varied with price dispersion and differences in tastes across categories.
4. Watch sales concentration across the whole store

📈 Recommendation
Check the range of products sold across shoppers as well as within one person’s basket. A varied-looking recommendation row can still concentrate sales on the same few products.
Track product and supplier sales shares alongside total sales and individual basket variety.
Separate niche products’ absolute sales from their share of all sales, and compare categories separately.
🎓 Findings
Lee and Hosanagar (2019) studied 1,138,238 users and 82,290 products in a randomized retail experiment. Collaborative-filtering recommendations could increase diversity for an individual while concentrating aggregate sales.
Niche products could gain sales in absolute terms while losing share. A wider personal basket and a broader store-level distribution are different outcomes.
🧠 Why it works
Several shoppers can each buy a slightly wider range yet converge on the same popular products. A store needs both views to judge whether discovery is broadening or repeatedly pointing everyone to the same winners.
✋ Limitations
This does not show that every personalized ranking concentrates sales. Algorithm, category and the diversity measure matter. Shared authors and retailer settings also require a dataset-overlap check before treating these papers as independent replications.
💡 The takeaway
Judge personalization by the choices it helps people make.
Test against the existing useful order, report the purchase stages with clear denominators and examine product fit alongside revenue. Track the range sold across shoppers as well as within individual baskets before expanding the ranking.
📚 Sources and study types
Robert Donnelly, Ayush Kanodia, Ilya Morozov (2024). Welfare Effects of Personalized Rankings. Marketing Science, 43(1), 92–113. Large randomized retailer experiment plus structural search/ranking model; One retailer; surplus is model-based; no universal conversion forecast.
Dokyun Lee, Kartik Hosanagar (2021). How Do Product Attributes and Reviews Moderate the Impact of Recommender Systems Through Purchase Stages?. Management Science, 67(1), 524–546. Randomized retail field experiment; Exact percentage denominators and uncertainty not yet extracted; no guarantee across algorithms/categories.
Xiang (Shawn) Wan, Anuj Kumar, Xitong Li (2024). How Do Product Recommendations Help Consumers Search? Evidence from a Field Experiment. Management Science, 70(9), 5776–5794. Randomized retail field experiment plus affinity-based mechanism measures; Fit/net value partly inferred from algorithmic measures; exact local test proposed separately.
Dokyun Lee, Kartik Hosanagar (2019). How Do Recommender Systems Affect Sales Diversity? A Cross-Category Investigation via Randomized Field Experiment. Information Systems Research, 30(1), 239–259. Randomized retail field experiment; Possible dataset overlap with same-author retail paper not resolved; no independent-replication count.

