Diagnosing Cross-Encoder Rerankers in Cold-Start Recommendation: Coverage, Score Signals, and Exposure Effects

Диагностика реранкеров на базе cross-encoder в условиях «холодного старта» рекомендаций: покрытие, сигналы оценки и эффекты экспозиции
Ekaterina Lemdiasova, Nikita Zmanovskii
2026-02-25

Cross-Encoder rerankersSerendipity-2018cold-start recommendationexposure-aware post-processingrecall@K
Reranking with cross-encoders (including LLM-derived rerankers) is a popular design for retrieval-based recom-menders, but strict cold-start exposes a key limitation: rerankers can only reorder retrieved candidates. We present an empirical diagnosis of a CrossEncoder reranker in a cold-start movie recommendation pipeline on Serendipity-2018. In an updated evaluation over 500 cold-start users (3 seeds), a popularity baseline strongly outperforms reranking (HR@10: 0.268 vs. 0.008; nDCG@10: 0.224 vs. 0.005). Diagnostics show that (i) ground-truth items frequently sit deep in the candidate pool (median rank 6717), (ii) hybrid candidate generation yields low recall@K relative to ANN baselines, and (iii) reranker scores barely correlate with relevance (Spearman r ≈ 0), while exposure concentrates on a handful of items. To reflect the algorithm tweak that produced the new results, we also include a compact comparison to an earlier pilot run (30 users/seed) to illustrate how scaling and retrieval/candidate logic shifts observed quality. We conclude with practical mitigation steps: strengthen retrieval first, tune candidate pool size, calibrate scores, and apply exposure-aware post-processing.
1
Ground-truth relevant items are often ranked very low in the candidate pool (median ground-truth rank 6717), limiting reranker potential.
2
Hybrid candidate generation achieves low recall@K compared to ANN baselines, reducing the chances rerankers can surface relevant items.
3
In cold-start movie recommendation on Serendipity-2018, a popularity baseline strongly outperforms a Cross-Encoder reranker (HR@10: 0.268 vs. 0.008; nDCG@10: 0.224 vs. 0.005).
4
Practical mitigations include: strengthen retrieval, increase/tune candidate pool size, calibrate reranker scores, and apply exposure-aware post-processing.
5
Reranker scores show almost no correlation with relevance (Spearman r ≈ 0), and model exposure concentrates on a small subset of items.

Cross-encoder reranker within a cold-start movie recommendation pipeline (on Serendipity-2018)

Diagnostic evaluation of reranker effectiveness in cold-start recommendation, including candidate coverage/recall, reranker score–relevance correlation, exposure concentration, and impact on top-K ranking performance

Publication Details
Publication Date
2026-02-25
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Ekaterina Lemdiasova
Nikita Zmanovskii
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%