CLARIN Resource Families: Future Improvements and Collaborative Work
Lenardič, Jakob; König, Alexander; Van Uytvanck, Dieter; Pahor de Maiti Tekavčič, Kristina
- Publisher
- Zenodo
- Language
- en
Abstract
This poster, presented at the Bazaar of the CLARIN Annual Conference 2025 (https://www.clarin.eu/content/bazaar-clarin2025), gives several suggestions for future work on the CLARIN Resource Families (CRF). We first propose potential improvements to the CRF overviews with regard to classification and nomenclature while defining precise inclusion criteria (i.e., ‘What is a CLARIN corpus anyway?’). We also propose a way to streamline curation between the CRF metadata on the one hand and the original repository records on the other while touching on key differences and overlaps between the two. Finally, we present the use of the new GitHub-based workflow for generating CRF overviews and discuss both ongoing and new opportunities for external involvement.
Full text
•User-friendly overviews of resources and tools in the CLARIN infrastructure. •Organised into three data types for easier discovery and comparison. •Manually curated, highlighting essential metadata. •Highly visible in browser searches. •Promoting re-use of high-quality content. NEW in 2026: Collaborative Curation •GitHub-based workflow ensuring transparency and easy issue management. •Supporting higher consistency in descriptions across families. •Seamless integration into website overviews for always up-todate overviews. •Leveraging local expertise of CLARIN centres to ensure optimal curation. •Promoting community building and knowledge exchange across CLARIN centres. •Different levels of involvement to encourage collaboration. SHARE YOUR VISION for Redesigning the CRF Taxonomy and Presentation •Do you find the current high-level division helpful enough or you would prefer higher granularity? •Do you miss a particular resource family? •Do you see any resource family as redundant or unhelpful? •Should families be exclusive or is overlap acceptable, e.g., how should we handle units like a corpus of historical press or a tool that performs both PoS and NER? •Should modality be treated as a separate resource family or only as a metadata label? •What metadata for an entry matters most for your research needs? •What subcategories would be helpful in a resource family you are most closely linked to (mono/multilingual resource, online/local tool instalation)? •How should we improve the presentation on the website: faceted search, collapsable entries …? •What is the added value for you of CRF overview compared to VLO-based discovery?