Résumé
According to the European Commission, the number of scientific publications produced annually worldwide has more than tripled between 2000 and 2022. This rapid growth has made it increasingly difficult for researchers to identify and keep up with relevant work. Recommender systems (RSs) have emerged as a promising solution to this problem by helping users navigate the expanding scientific literature. In this paper, we introduce ScientiaRec, a novel scientific article RS that combines user-item interactions and content-based features, including descriptive tags and keywords automatically extracted using a large language model (LLM). At the core of our approach is a serendipity-aware matrix factorization model, designed to recommend relevant items while actively promoting serendipity. The goal is to help users discover novel and potentially insightful papers that go beyond their immediate research interests. We evaluate the performance of ScientiaRec against several baseline models using the publicly available CiteULike dataset, employing a comprehensive set of both accuracy-oriented and beyond-accuracy evaluation metrics. In addition, we conduct an LLM-based study to assess the serendipity of ScientiaRec Top-N recommendations. The experimental results demonstrate that ScientiaRec achieves a strong balance between relevance and beyond-accuracy objectives. Moreover, the inclusion of LLM-derived keywords significantly enhances the RS's overall performance.