Fully Algorithmic Librarian: Large-Scale Citation Experiments
Full text
Fully Algorithmic Librarian: Large-Scale Citation Experiments Tomasz Stompor, Janina Zittel, Thorsten Koch and Beate Rusch Cooperation: KOBV, Kompetenznetzwerk Bibliometrie Fully Algorithmic Librarian: Large-Scale Citation Experiments Tomasz Stompor, Janina Zittel, Thorsten Koch and Beate Rusch Cooperation: KOBV, Kompetenznetzwerk Bibliometrie Project Outline The Fully Algorithmic Librarian (FAN) is an interdisciplinary project at the Zuse Institute Berlin, combining mathematics and library & information science to explore how libraries can use algorithmic methods to better support academic research. We analyze large-scale citation networks with up to 119M records and 1,4B citations using data from Web of Science [1] and OpenAlex[2], building acitation graph as a foundation for future AI-powered services in libraries. In this process, we develop mathematical methods that can be applied to interdisciplinary citation graphs. The development of the knowledge graph is accompanied by exploiting abstract vectorization using embedding models to compare how they align. We aim to enable libraries to offer scalable,data-driven, and open bibliometric services, empowering researchers and institutions to better understand and shape the academic landscape. Our mission is to develop an open prototype that can be reimplemented by the scientific community. Evaluating Scientific Prestige Using Citation Graphs Evaluating scientific impact requires precise measurement of individual article influence, traditionally assessed through citation-based metrics. Recently, approaches have shifted toward leveraging citation graph structures rather than relying solely on raw citation counts — for example, by employing the PageRank method to assess scientific prestige [3]. Beyond ranking influence, PageRank also serves as a valuable tool for comparing bibliometric databases, revealing that citation-based prestige inherently depends on the completeness and accuracy of the chosen dataset. The PageRank computation for Web of Science (2000–2021) and OpenAlex (1950–2020) highlights differences in both temporal coverage and citation network structure. Notably, no PageRank is calculated for the most recent 10 years in either dataset, as the metric requires a 10year citation window. Beyond this temporal aspect, the results also reveal structural variations between the two citation networks. Most prominently, the WoS is more strongly connected reflected by a higher average PageRank. The structural differences of bibliometric datasets illustrated by a PageRank metric following Chen et al. (2023) with 10 years citation span and a damping factor of 0.5 on WoS and OpenAlex. Visualizing Publication Networks Visualization for a subgraph using citation data from the IPCC Assessment Report 6 (2021) in the field of climate research. The IPCC data was matched with the data from OpenAlex. Funding Clustering Techniques Analyzing large citation graphs requires automated classification. Comparison of algorithmic clustering labels with WoS labels for a subgraph of WoS on Mathematics and Operation Research & Management Science. Disciplines often overlap, so we use a multi-label clustering approach based on article similarity derived from citations [4]. The clustering task is to find a soft cluster assignment X∈[0,1]C×Nthat minimizes the discrepancy between a similarity S∈RN×Nobserved from citations and the similarity from a predicted distribution: f(X) = N ∑ i=1 N ∑ j=1 si j − C ∑ k xkixk j!2 . The optimization problem is non-convex and large-scale, with number of articles N>107and number of clusters C∼100–500. We use GPU-accelerated gradient descent (CUDA) to scale to massive graphs: 700k-node subgraph: clustering in 30 seconds Full OpenAlex (60M nodes, 1.2B links): in progress Article 1 Article 2 ... Article N Cluster 1 x11 x12 ... x1N Cluster 2 x21 x22 ... x2N . . .. . .. . .. . . Cluster C xC1xC2... xCN Structure of the cluster assignment matrix X Use-Case Scenarios Unlike the evaluative application of bibliometrics through statistical tools, our approach focuses on exploratory methods that could be used in the following scenarios: automatic clustering and classification detection of emerging research topics visualization of interdisciplinary networks thematic search assistant for potentially missing citations reviewer and collaborator recommendation detection of blind spots: what is not represented in the data? Publications [1] We acknowledge the use of WoS through the Kompetenznetzwerk Bibliometrie. Supported via the German Competence Network for Bibliometrics funded by the Federal Ministry of Education and Research (Grant: 16WIK2101A). [2] J. Priem, H. Piwowar, R. Orr. OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. ArXiv.https://arxiv.org/abs/2205.01833, 2022 [3] Y. Chen, T. Koch, N. Zakiyeva, K. Liu, Z. Xu, C-h. Chen, J. Nakano, K. Honda: Article’s scientific prestige: Measuring the impact of individual articles in the web of science. Journal of Informetrics, 17, 101379, 2023 [4] T. Nepusz, A. Petróczi, L. Négyessy, F. Bazsó: Fuzzy communities and the concept of bridgeness in complex networks, Phys. Rev. E 77, 016107, 2008. [5] V.T. Huong, I. Litzel, T. Koch: Similarity-based fuzzy clustering scientific articles: potentials and challenges from mathematical and computational perspectives, arxiv:2506.04045, 2025.