The Recommendation System Needed More Than Good Metrics, it Needed to Actually Know Its Customers.
How layering category based personalization on BPR finally made the system feel like it understood who it was talking to.
The Recommendation System Needed More Than Good Metrics, it Needed to Actually Know Its Customers. (Part 2).
How layering category based personalization on BPR finally made the system feel like it understood who it was talking to.
Photo by Startaê Team on Unsplash
Brief recap of Part 1: I have created a cross-selling recommendation system for one of the FMCG companies, switched from User-Based Collaborative Filtering to Bayesian Personalized Ranking and got stuck at the point when no matter what customer I was working on, the results suggested recommending globally popular products.
All the metrics were looking good. But not the recommendations themselves. It was not a metrics issue to get a baby store recommended dishwashing liquid. It was a personalization issue.
The Fix: Hybrid Scoring with Product Categories
The solution I came to is incorporating a personalization component based on customer categories onto the BPR scoring.
Basic principle: before presenting recommendation options to a particular client, apply another rank score to those candidates, where such score will show how well that product fits into a category the customer usually purchases. To do this, two steps were necessary. One, developing a category profile of the customer. Two, devising a score formula that would balance the collaborative part of the ranking and the similarity of product categories.
The product list included hierarchy of category IDs — PH_DESC2 and PH_DESC3. PH_DESC2 was a broader category, such as “Hair Care”, “Baby Diapers”, “Noodles”. PH_DESC3 was a narrower one, e.g., “Baby Happy Pants”.
As per each customer, we determined how many percent of unique products he/she bought belong to each PH_DESC2 category. Let’s say that out of 10 unique products that person bought 6 belonged to the category “Baby”. So his/her category weight would be 0.6. Similar process was followed for PH_DESC3.
That resulted in a clear and precise vector of the category identity for each customer.
The score formula The final score for each candidate product from the recommendation list was calculated based on the following formula:

Here, the PH2 Score refers to the historical weight of the customer’s products within the PH_DESC2 category. Explanation: BPR serves as a source of collaborative filtering (similarity between customers’ purchases) and the category score serves as a personalizing factor (is the product relevant to this customer’s category of goods?). The ratio of 70% BPR and 30% PH2 is intentional. It means that we still rely mainly on BPR as our source, but at the same time, the category score could be a significant influencing factor in the final decision making process.
First Results: Better, But Still Not Right
There were slight improvements in the numbers of evaluation.

Cross Sell Hit Rate went up from 6.2% to 6.31%. Cross Sell NDCG from 7.23% to 8.5%. Restock Hit Rate increased from 6.13% to 7.5%. Restock NDCG from 18.72% to 21.66%.
Not huge leaps, but the right kind of movement.The baby store had more product suggestions that were baby adjacent. The algorithm was finally thinking of the category fit. However, when I went over specific customers’ cases, it became apparent that PH2 by itself did not have enough granularity.
That one level of granularity was the difference between a recommendation that makes sense and one that would make a sales give wrong recommendation. The fix was straightforward: add PH3.
Going One Level Deeper: PH2 + PH3 Personalization
The updated scoring formula:

PH3 now carried the most weight at 50%. BPR dropped to 30%. PH2 stayed at 20%.
BPR is still in the formula because it provides the signal about which specific products within a category are worth recommending. PH3 tells the system which subcategories are relevant. BPR tells it which products in those subcategories to surface.
What I Would Do Differently
Reflecting on this project overall, there are a few takeaways worth mentioning. It was wise to start with qualitative output analysis rather than aggregate Hit Rate metrics. As most of Month 2 was dedicated to UBCF optimization and aggregate Hit Rate metrics analysis, the category alignment issue could have been found weeks earlier if I had examined individual customer recommendation cases. Metrics are great until you get distracted from the output.
As we discussed the PH2 to PH2+PH3 iteration case, the PH2-only recommendation results still being okay is a good illustration of how many applied ML projects go: solving one issue, discovering another one under the hood and solving that as well.
There is nothing wrong with that; however, curiosity should be kept high to see the problems coming.
If you have worked on recommendation systems for B2B retail specifically, I would be genuinely curious how you handled the category relevance problem. It seems like a common issue but I have not seen it discussed much compared to the usual B2C recommendation literature.
메타데이터
- post_id
- 6a38d9a7a3f1
- slug
- the-recommendation-system-needed-more-than-good-metrics-6a38d9a7a3f1
- url
- https://medium.com/@nando196/the-recommendation-system-needed-more-than-good-metrics-6a38d9a7a3f1
- canonical_url
- https://medium.com/@nando196/the-recommendation-system-needed-more-than-good-metrics-6a38d9a7a3f1
- author_url
- https://medium.com/@nando196
- status
- ok
- fetched_at
- 2026-06-09 14:34:10