Diagnosing Cross-Encoder Rerankers in Cold-Start Recommendation: Coverage, Score Signals, and Exposure Effects
Диагностика реранкеров на базе cross-encoder в условиях «холодного старта» рекомендаций: покрытие, сигналы оценки и эффекты экспозиции
2026-02-25
SCID: 54.1/keuy9nru
Discuss with AI
Cross-Encoder rerankersSerendipity-2018cold-start recommendationexposure-aware post-processingrecall@K
Figures from the paper
Abstract (AI)
Reranking with cross-encoders (including LLM-derived rerankers) is a popular design for retrieval-based recom-menders, but strict cold-start exposes a key limitation: rerankers can only reorder retrieved candidates. We present an empirical diagnosis of a CrossEncoder reranker in a cold-start movie recommendation pipeline on Serendipity-2018. In an updated evaluation over 500 cold-start users (3 seeds), a popularity baseline strongly outperforms reranking (HR@10: 0.268 vs. 0.008; nDCG@10: 0.224 vs. 0.005). Diagnostics show that (i) ground-truth items frequently sit deep in the candidate pool (median rank 6717), (ii) hybrid candidate generation yields low recall@K relative to ANN baselines, and (iii) reranker scores barely correlate with relevance (Spearman r ≈ 0), while exposure concentrates on a handful of items. To reflect the algorithm tweak that produced the new results, we also include a compact comparison to an earlier pilot run (30 users/seed) to illustrate how scaling and retrieval/candidate logic shifts observed quality. We conclude with practical mitigation steps: strengthen retrieval first, tune candidate pool size, calibrate scores, and apply exposure-aware post-processing.
Key Findings
1
Ground-truth relevant items are often ranked very low in the candidate pool (median ground-truth rank 6717), limiting reranker potential.
2
Hybrid candidate generation achieves low recall@K compared to ANN baselines, reducing the chances rerankers can surface relevant items.
3
In cold-start movie recommendation on Serendipity-2018, a popularity baseline strongly outperforms a Cross-Encoder reranker (HR@10: 0.268 vs. 0.008; nDCG@10: 0.224 vs. 0.005).
4
Practical mitigations include: strengthen retrieval, increase/tune candidate pool size, calibrate reranker scores, and apply exposure-aware post-processing.
5
Reranker scores show almost no correlation with relevance (Spearman r ≈ 0), and model exposure concentrates on a small subset of items.
Research Object
Cross-encoder reranker within a cold-start movie recommendation pipeline (on Serendipity-2018)
Research Subject
Diagnostic evaluation of reranker effectiveness in cold-start recommendation, including candidate coverage/recall, reranker score–relevance correlation, exposure concentration, and impact on top-K ranking performance
Publication Details
Publication Date
2026-02-25
Journal
Publisher
ISSN
Cited by
0
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest