scieee AI-readable full text Open interactive document viewer

Once upon a Data Steward - A PSDI Webinar

Pearman-Kanza, Samantha

Abstract

The Physical Sciences Data Infrastructure (PSDI) aims to accelerate research in the physical sciences by providing a data infrastructure that brings together and builds upon the various data systems researchers currently use. Our webinar series provide updates on our exploratory pathfinder work, and on relevant tools and technologies that have been developed by members of our community. This record contains a more detailed version of the presentation that was presented at the "Once upon a Data Steward" webinar presented by Dr Samantha Pearman-Kanza on 10th July 2025. The recording of this webinar is available on YouTube: https://youtu.be/e0T4XEsMFFA Abstract: This webinar explored the growing importance of data availability and management in modern research. The session introduced the emerging role of the Data Steward, a new type of digital Research Technical Professional (dRTP) who ensures data quality and usability throughout its lifecycle. The webinar also unveiled CaSDaR (Careers and Skills for data-driven Research), a new four-year initiative designed to support and empower Data Stewards as central figures in the research lifecycle.

Full text

Once upon a Data Steward Dr Samantha Pearman-Kanza University of Southampton PSDI Webinar 10th July 2025 About Me: The Storyteller Introducing our characters: The FAIR Data Witch, Datacat and The Data Steward Sorceresses The origin story: The complex journey of FAIR data The twist in the tale: Introducing Data Stewards Sequel Setup: Introducing the CaSDaR Network Presentation Outline CC BY-ND 4.0 Errant Science https://errantscience.com/ Senior Enterprise Fellow at University of Southampton Principal Investigator for CaSDaR (Careers and Skills for Data-driven Research) Pathfinder Lead on Process Recording for PSDI Research Interests: Semantic Web Technologies, IoT, Research Data Management, Digitisation, Lab of the Future, Paperless Labs, Re-use of Technology About Me Once upon a time , in the vast and ever - expanding Research Kingdom , a magical force called Data began to grow at an exponential rate . It sparkled with immense potential , promising to unlock secrets of the universe , underpin magical spells , and build previously unimagined wonders . Once upon a time... DATA Although the data multiplied faster than an enchanted forest, it was scattered across the Research Kingdom, hiding in unreachable towers, forgotten caves and within dusty scrolls buried deep in the magical archives. Brave scholars and curious minds searched far and wide to try and find these vital datasets, but sadly much of it could not be found, and even when they thought they’d found it, they inevitably discovered, like Prince Mario, that what they were looking for actually lay in another castle. Even when a scroll was uncovered, its language was often cryptic—written in unintelligable symbols and codes only a few could decipher. The knowledge of the ancients, instead of guiding the future, lay dormant and remained misunderstood. And thus, the kingdom suffered. Brilliant magical discoveries faded into the past, and the same spells were cast again and again, never building upon the magical lore that came before. But all was not well in the Research Kingdom SORRY MARIO BUT YOUR DATA IS IN ANOTHER CASTLE https://en.wikipedia.org/wiki/Our_princess_is_in_another_castle! Having heard of the Research Kingdoms plight, the Queen of the distant and orderly land of Metadataia sent one of their wisest and most organised consorts, the FAIR Data Witch to visit the Kingdom. Filled with a determined purpose to spread the wisdom of Metadatia’s vision of kingdom wide Findability, Accessibility, Interoperability, and Reusability, she declared: But one day, a spark of hope emerged FAIR No more shall knowledge be lost in the shadows! Let every piece of data be FINDABLE, every scroll ACCESSIBLE, every language INTEROPERABLE, and every discovery RE-USABLE! The Witches Plan The FAIR Data Witch introduced a set of principles to ensure that all of the data in the Research Kingdom is: F – Findable A – Accessible I – Interoperable R – Reusable Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016). https://doi.org/10.1038/sdata.2016.18 The FAIR Data Witch was delighted with the opportunity to share her FAIR principles with the Research Kingdom, and was on the verge of mounting her broomstick back to the land of Metadataia, when she came across a clever and curious creature named Datacat, who stopped her in her tracks. Datacat was beloved by scholars and scribes alike, known for curling up beside glowing scrolls and purring softly when the data was clean, and believe me they had learned to give her a wide berth if their data was sub par. Datacat had wandered the halls of the Research kingdom for years, watching, listening, and learning, and she could see what the FAIR Data Witch, so blinded by her current achievements, could not. However, Datacat was not conviced The Magic of FAIR is not so easily cast All throughout the Research kingdom, scholars and scribes were struggling to understand the FAIR Witches complex incantations. The Metadatia scrolls of guidance were long and written in a complex tongue that many did not understand. Numerous scholars and scribes found themselves lost in the labyrinth of standards, unsure how to begin, and many had to carve out new pathways just to find a solution to their particular questions. Others lacked the tools or time to follow her path, and so the Witches vision, while admired, was only partially fulfilled. Datacat had heard about the Witches plan, and unconvinced, had undertaken her own investigations. Hmm...there must be some solution... This isn’t just about Data, you need to consider Tools, Code, Methods, Context. How could/would your work be reused, replicated, reproduced or repurposed across the Kingdom? Re-use: Re-use the data (or run the software) in the same manner Replicate: Repeat entire research from scratch including data collection and analysis Reproduce: Reanalyse the existingdata in the same manner Repurpose: Use existing data R1. meta(data) are richly described with a plurality of accurate and relevant attributes R1.1. (meta)data are released with a clear and accessible data usage license R1.2. (meta)data are associated with detailed provenance R1.3. (meta)data meet domain-relevant community standards CC BY-ND 3.0 Dataedo - https://dataedo.com/cartoon/ R is for Reusable It is important to ensure both READMe and Metadata files have appropriate information Metadata should enable DISCOVERY of the data, and provide an UNDERSTANDING of its contents/provenance etc READMEs should explain HOW to make use of the data R is for Reusable At first, it seemed the Witches magic was finally taking hold. The scholars and scribes, invigorated with a new understanding of the complexities of FAIR were more determined than ever to put them into practice. But as the moons passed, Datacat began to notice something strange. Despite all of these efforts, many scrolls still vanished into obscurity. Some were shared too late, others were missing key details, and many were still written in ancient languages, citing spells that no one else could understand. Even those who embraced FAIR often struggled to apply it fully. Datacat sat crestfallen beneath a starry sky one evening, too despondent to even eat her tuna, and sighed. But issues still remained Why does it still feel like we ’ re chasing shadows? Datacat ventured out across the moonlit corridors, knowing many of her fellow scholars and scribes would be burning the midnight oil. She began to observe the scholars, working feverishly on creating their new data. She watched as scholars scribbled notes on scraps of parchment, titled scrolls with names such as “Final_Final_REALLYFINAL_V2_SPKComments” and witnessed first hand as they almost immediately forgot the spells of yesterday that they had used to conjure their results. And slowly, Datacat formed a new realisation. To truly enact the FAIR vision, we must start not at the end, but where every fairy story starts, at the begnning, where the data is first created. We must teach our scribes and our scholars to weave FAIRness into their work from the first stroke of the quill. Let them not wait until the scroll is sealed, but guide them as they write, label, and store their knowledge, to make data that is FAIR from its very inception. And then suddenly, the penny dropped It’s not just the scrolls, it’s the scribes. We’ve been trying to polish the leaves, but the roots are tangled. If we want FAIR to truly flourish, we must go back to the origins of our data stories. Datacat understood that Data creation requires: a. Proper planning (DMPs) b. How are you going to collect your data? c. How are you going to standardise data entry? d. How are you going to mitigate against missing or incomplete data records? e. What are the ethical considerations? Are you collecting personal data? f. How are you going to organise and manage your data throughout the project? Data Creation Many of the issues with data re-use start at the creation or collection stage FAIR Requirements Datacat also understood that Implementing FAIR fully throughout the entire research data lifecycle requires knowledge of: a. Data Management Planning b. Data Collection & Curation c. Metadata Generation and appropriate schemas for generic/domain specific metadata d. Ethics & Data Governance e. Publication & Data Sharing f. Repositories CC BY-ND 4.0 Errant Science https://errantscience.com/ However, unfortunately, this still did not solve this deep rooted and complex problem. Whilst scholars and scribes welcomed this new information, and indeed some disparate scrolls started to gleamed with clearer labels and shared knowledge, this was not the case across the entire Research Kingdom. Scholars were tired, they were trying and failing to juggle their research with the growing demands of trying to implement all of the aspects of FAIR, some of which they had very limited knowledge on. Some gave up completely, others quietly slipped back into old habits, and the shadows began to return. The realisation struck Datacat deeply, that, even with clearer guidance, FAIR could not thrive if it relied solely on the shoulders of individual scholars. Something more was needed—support, structure, and perhaps… a whole new kind of magic. Datacat thought she had finally found the answer We’ve asked too much. To expect every researcher to be a metadata witch, a keeper of standards, to understand the complex spells of ethics and responsibility, and be guardian of reusability… it’s just too much! Datacat sought advice far and wide If anyone says “ throw it at an AI ” I ’ ll be very displeased … ! Datacat invited representatives from neighbouring lands to discuss this deeply concerning matter. Magical officials came far and wide including from the Mystical Realm of Libraries, the ever bountiful Kingdom of Funding, and the futuristic land of AI. The wise officials gave Datacat their consult, and despite the insistance from the Grand Vizier of the Land of AI that their people could magically solve this problem, Datacat realised that this was not an issue that could be solved simply by spells and scrolls alone — no amount of technical magic could replace the need for human care and guidance. FAIRness, Datacat realized, required more than just good intentions; it needed skilled hands and thoughtful minds. However, expecting every researcher to master the full range of FAIR practices on top of their already complex research responsibilities was simply too much. That ’ s when it struck Datacat, what the Kingdom needed, were data stewards. Spoiler alert, the answer isn ’ t just “ AI ” Data Steward Sorceresses Create valuable data that is worth sharing Avoid repetition due to lack of reusability We shall manage the data of the Research Kingdom to ensure its quality, accessibility and usability The role of Data Stewards The need for Data Stewards across the research data lifecycle Ensure that data is suitably documented, described and licensed for re-use. Possess skills required to facilitate data re-use. Given appropriate recognition for their role in the research process. Aid with the preparation of data for publication, including identifying appropriate licenses. In depth knowledge of identifying and using repositories for long-term data preservation. Advise on packaging up datasets with relevant (machine-readable) metadata and documentation. Understand data and research requirements for proposed work. Support creation of viable iterative Data Management Plans. Influence collection and capture processes for data. Tailor these to the nature of the research and data. Understand the principles of responsible research and integrity, and when ethics is required. Advise on best practices and tools for collaboration, e.g. secure storage processes, data transform mechanisms. This image is based on Jisc's Research Data Lifecycle taken from - https://www.jisc.ac.uk/guides/research-data-management-toolkit Data Steward Fellowship Support Data Steward positions across different domains & UK research institutions Data Steward Internships Internship opportunities linked with projects that require different types of data stewardship Open Funding Calls Proposals for data stewardship projects CaSDaR Funding Calls These will be reviewed by an expert panel with interdisciplinary experiences in data stewardship Join us at our launch on the 18 September to find out more! th CaSDaR’s Core Values & Commitments 1. Data Collection 2. Performance Metrics 3. Analysis Tools 4. Reporting Embed EDI principles across all activities Provide inclusive, flexible support for all activities Accessible Venues for Meetings & Hybrid Events Bursary fund for event attendance Foster a sense of belonging for all data stewards Promote teamwork and knowledge sharing within the data stewardship community Promote and upskill researchers in responsible and ethical data stewardship practices Recognise and credit all contributions to research data Targeted support for early career researchers for progression and acknowledgement Ensure long-term sustainability of data stewardship Support long-term preservation and accessibility of research data 5. Optimization Inclusivity Collaboration Integrity Recognition Sustainability Sonraí (The Irish Data Stewardship Network) Website: https://datastewards.ie/ Aims to promote and support the development of data stewardship skills across Ireland's research landscape, ensuring the curation, preservation, and dissemination of national data assets in alignment with FAIR principles Data Steward Interest Group (The Netherlands) Website: https://tdcc.nl/dsig/ This group is aiming to establish a community hub for data stewardship that enables informal and inclusive knowledge and experience exchange. NFDI (The German National Data Infrastructure) Website: https://www.nfdi.de/?lang=en The NFDI plays a crucial role in data stewardship by developing and implementing standards, tools, and best practices for managing research data. Many groups within the NFDI support data stewardship including NFDI4Chem https://knowledgebase.nfdi4chem.de/knowledge_base/docs/data_steward/ and NFDI4BioImage https://nfdi4bioimage.de/about-us/data-stewardship-team/ World Data System Website: https://worlddatasystem.org/ The World Data System (WDS) enhances data stewardship by promoting the long-term stewardship and accessibility of quality-assured scientific data and run the prestigious Data Stewardship Awards. Research Data Alliance – Professionalising Data Stewardship Interest Group Website: https://www.rd-alliance.org/groups/professionalising-data-stewardship-ig/members/all-members/ This is an RDA interest group on professionalising data stewardship Digital Data Curation (DCC) Website: https://www.dcc.ac.uk/ DCC is dedicated to improving data stewardship in the research community and runs regular events to support this Physical Sciences Data Infrastructure Website: www.psdi.ac.uk PSDI is working hard to establish a community of data stewards across the physical sciences STEP-UP Website: https://step-up.ac.uk/ Imperial College London / University College London / Kings College London / University of Westminster – This project is developing communities, training and career opportunities to support “digital Research Technical Professionals” (dRTPs) National & International Data Steward Networks & Supporting Initiatives Save the Date for our Townhall Launch! 18th September 2025 Hybrid Townhall Launch Birmingham Library & Online Lunch & Refreshments Provided Opportunities to present Lightning Talks & Posters Hear inspirational stories from Data Stewards about their journeys Gain insights from our expert project partners about the vital role of Data Stewardship Hear about related Data Stewardship initiatives Network with likeminded people and meet useful contacts across the breadth of data stewardship Discuss key issues and areas of data stewardship including training, career progression and the role of data stewardship Find out all about our upcoming funding call! CaSDaR Details www.casdar.ac.uk CaSDaRNetwork @CaSDaRNetwork @casdarnetwork.bsky.social @[email protected] @CaSDaRNetwork https://www.jiscmail.ac.uk/CASDAR