scieee AI-readable full text Open interactive document viewer

Why keep data science local? Case Studies from Two Universities Building Scholarly Indices with Open Data to Improve Institutional Data Literacy

Teykl, Katharine; Hemingway, Jordan; Clark, Jason; Hutchens, Chad

Abstract

Two mid-sized libraries created indices of institutional scholarship using open code and the OpenAlex dataset. This panel explores the implementation of these applications, examines the open data sources used, and offers a candid discussion on what is gained and lost when choosing to outsource data science research and development.

Full text

WHY KEEP DATA SCIENCE LOCAL? Case studies from two universities building scholarly indices with open data to improve institutional data literacy INTRODUCTIONS Jason A. Clark Montana State Jordan Hemingway University of Wyoming Katharine Teykl University of Wyoming Chad Hutchens University of Wyoming Institutional Context Wyoming •Rural, Carnegie R1, Land Grant, only 4 year college in state •Student FTE: 9255 •Dept of 8 people, data, digitization, digital scholarship, scholarly communications Montana State •Rural, Carnegie R1, Land Grant •Student FTE: 17,165 •Dept of 6 people, data, machine learning, scholarly communications Bozeman Laramie Data Source: OpenAlex Overview •https://openalex.org •Free & Open SKG •Open & documented data sources •Alternative to Scopus, WoS, Dimensions, etc. OpenAlex Data Comparisons "ScholarNodes recommends collaborations, identifies local and external researcher networks, and visualizes an interdisciplinarity graph for individual researchers." Md Asaduzzaman Noor, Jason A. Clark, and John W. Sheppard. 2024. ScholarNodes: Applying Content-based Filtering to Recommend Interdisciplinary Communities within Scholarly Social Networks. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '24). Association for Computing Machinery, New York, NY, USA, 2791–2795. https://doi.org/10.1145/3626772.3657668 Machine-ready reports for searching the literature Recent works and authors used to identify MSU acknowledgements and attribution. www.lib.montana.edu/msu-researcher-harvest Discussion What is gained by taking a DIY approach? What is lost? How do you convince library decision-makers to support a DIY approach? Why invest in expanding library expertise? Why work with open data? What factors should you consider? What was it like to work with collaborators with more data science experience? So what? Next steps? Takeaways •Don't give your labor away •Don't miss your chance to create technical capabilities •Build transparent & interoperable metadata •Collaborate on data science projects to communicate & extend libraries' existing expertise •Embrace our data teaching (and learning!) moment(s) What would help your users? •Find your starting point •Use open infrastructures •Explore OpenAlex, CrossRef, DataCite •Download your dataset •Use GUIs as an on-ramp •Plug in & build community Resources for getting started: •OpenAlex GUI, Data Sources & Documentation •Library Carpentry & Software Carpentry •DLF Digital Library Pedagogy Group (#DLFteach) & Code4Lib •Persistent Identifier (PID) cookbook •UCLA's Open Science for Librarians •The Turing Way …want resources tailored to you? Please reach out!