Tour de CLARIN Volume 5
Full text
Tour de CLARIN VOLUME FIVE Edited by Kristina Pahor de Maiti Tekavčič, Jakob Lenardič and Karina Berger CLARIN 2025 Tour-de-CLARIN-5-cover.indd 8 19/9/25 07:41:52
1 Tour de CLARIN volume FIVE edited by Kristina Pahor de Maiti Tekavčič, Jakob Lenardič and Karina Berger
2 CoNSoRTIA K-CeNTReS / B-CeNTReS 3 Tour de CLARIN Foreword ---------------------------------------------------------------------------------------- 4 PART 1 | CoNSoRTIA -------------------------------------------------------------------------- 6 Austria Introduction | CLARIAH-AT -------------------------------------------------------------------- 8 Tool | Tool Chain for Sentiment Analysis --------------------------------------------------- 11 Resource | Austrian Media Corpus ---------------------------------------------------------- 13 event | #digitalDHaustria Twitter showcase 2021 ---------------------------------------- 15 Interview | Amelie Dorn ----------------------------------------------------------------------- 16 Iceland Introduction | CLARIN-IS --------------------------------------------------------------------- 24 Tool | AMI’s Corpus Web --------------------------------------------------------------------- 26 Resource | Icelandic Gigaword Corpus ----------------------------------------------------- 27 event | What can CLARIN do for you? ------------------------------------------------------ 29 Interview | Sigríður Ólafsdóttir -------------------------------------------------------------- 31 Lithuania Introduction | CLARIN-LT ---------------------------------------------------------------------- 38 Tool | Lithuanian Spelling Checker V.1.0.45 for macOS ---------------------------------- 42 Resource | English-Lithuanian Language Resources for the Cybersecurity Domain --------44 event | Annual CLARIN-LT Seminar 2024 ---------------------------------------------------47 Interview | Erika Rimkutė and Virginijus Dadurkevičius ---------------------------------- 49 Poland Introduction | CLARIN-PL ---------------------------------------------------------------------- 56 Tool | CompCorp ------------------------------------------------------------------------------ 58 Service | CLARIN-PL Helpdesk -------------------------------------------------------------- 60 event | CLARIN in Research Practice -------------------------------------------------------- 62 Interview | Agnieszka Hess ------------------------------------------------------------------- 64 Switzerland Introduction | CLARIN-CH -------------------------------------------------------------------- 70 Tool | LCP -------------------------------------------------------------------------------------- 76 Service | CLARIN-CH FAIRification Pipeline ------------------------------------------------ 78 event | CLARIN-CH Day 2024 ----------------------------------------------------------------- 80 Interview | Jannis Vamvas -------------------------------------------------------------------- 82 PART 2 | KNowledge (K) ANd SeRvICe PRovIdINg (B) CeNTReS ----------------- 88 K-centre: DANSK Introduction ----------------------------------------------------------------------------- 90 Interview | Sidsel Boldsen ------------------------------------------------------------- 93 K-centre: Phonogrammarchiv Introduction ----------------------------------------------------------------------------- 98 Interview | Fayrouz Kaddal ----------------------------------------------------------- 101 K-centre: CLARIN-SMS Introduction --------------------------------------------------------------------------- 108 Interview | Andrea Fried and Arne Jönsson ---------------------------------------- 113 B-centre: CLARIN-LV Introduction --------------------------------------------------------------------------- 120 Interview | Kristīna Korneliusa ------------------------------------------------------ 123 Colophon ------------------------------------------------------------------------------- 131 KNowledge (K) ANd SeRvICe PRovIdINg (B) CeNTReS
5 4 Foreword This fifth volume of Tour de CLARIN comes after a four-year hiatus in the print editions, which were originally published each year between 2018 and 2021 The Tour de CLARIN initiative, aimed at increasing the visibility of CLARIN’s members and their diverse activities, continued beyond the last print publication in 2021, but the stories were made available only in an online format. While the online presence certainly allows for the broadest outreach in real time, the sheer volume of information offered in the digital space increases the risk of some of the content being overlooked or underappreciated. The Tour de CLARIN stories are not only intended as periodic updates on centres’ activities, but also as a venue to showcase flagship initiatives that have proven valuable for the community and may serve as inspiration for other consortium members. For this reason, even in the digital era, we remain committed to publishing in print—in the hope of highlighting the significance of these contributions and recognising the effort invested by each centre in the consortium. This volume follows the format established in 2021 and features the work carried out by CLARIN national nodes, Knowledge Centres (K-centres), and Service Providing Centres (B-centres). The stories span the years 2022 to 2024, which means that the volume represents a kind of time capsule—offering insight into how centres presented themselves and engaged with their communities during this period—but it should not be taken as the most up-to-date source of information on their current activities. For the latest developments, we encourage readers to visit the centres’ official webpages. The volume consists of two parts. In Part 1, we visit Austria, Iceland, Lithuania, Poland, and Switzerland. Each national consortium is introduced through five chapters: a presentation of its structure and members, a key resource, a flagship tool or service, a successful user involvement event, and an interview with a researcher who has fruitfully engaged with the consortium’s infrastructure. Part 2 turns the spotlight on three K-centres and one B-centre: the Danish K-centre DANSK, the Austrian Phonogrammarchiv, the Swedish CLARIN-SMS centre, and the Latvian B-centre CLARIN-LV. Each centre is presented with a dual focus: a description of its mission, services, challenges and user base, followed by an interview with a researcher or expert who has collaborated closely with the centre and benefited from its offer. This volume would not have been possible without the dedication and efforts of the national coordinators and centre staff. We also warmly thank the researchers who shared their experiences and insights in the interviews—their contributions are not only informative but also a source of inspiration for the wider CLARIN community. As Tour de CLARIN continues its journey, we remain committed to amplifying the voices of the communities that make CLARIN a vibrant, user-driven research infrastructure. We hope this volume provides valuable insights into how CLARIN continues to support cutting-edge research and international collaboration in the study of language and society—bringing together both longstanding contributors and newly established centres in a shared vision for the future. Kristina Pahor de Maiti Tekavčič, Jakob Lenardič and Karina Berger September 2025
7 Consortia featured in this volume: Austria Iceland lithuania Poland Switzerland PART 1 CONSORTIA 6
8 9 • AuSTRIA Introduction | ClARIAH-AT written by Karlheinz Mörth and Walter Scholger Thanks to the support of the (former) Federal Ministry of Science, Research and Economy, Austria was able to become a founding member of the two research infrastructure consortia CLARIN and DARIAH in 2012 and 2014 respectively. From the very beginning, the two infrastructures have worked together as a single entity, and in recent years, many Austrian Digital Humanities (DH) activities have been coordinated and driven by CLARIAH-AT. In 2014, the Austrian Academy of Sciences was entrusted by the then Federal Ministry of Science and Research with the coordination of the Austrian DH activities. Karlheinz Mörth (Austrian Academy of Sciences) was appointed as the national coordinator of both infrastructures. In 2024, Tanja Wissik (Austrian Academy of Sciences) was appointed national coordinator for CLARIN, and Walter Scholger (University of Graz) was appointed national coordinator for DARIAH. The national consortium has been managed by its speaker, Walter Scholger (University of Graz), since 2019. And since 2024, Tanja Wissik has been deputy speaker. The consortium brings together Austrian institutions with relevant expertise in research, development, and teaching, aiming to establish and operate sustainable research (data) infrastructures. The involved institutions also play an active role in the development of the DH landscape in Austria and the establishment and expansion of technical and social infrastructures in this area. In 2025, the members of the CLARIAH-AT Consortium are the Austrian Academy of Sciences (ÖAW), the Austrian National Library (ÖNB), the University for Continuing Education Krems (UWK), the University of Graz, the University of Innsbruck, the University of Klagenfurt, the University of Salzburg, and the University of Vienna. Furthermore, the University of Music and Performing Arts Vienna (mdv), the Central European University (CEU), and the Natural History Museum (nhm) are currently observers in CLARIAH-AT. Representatives of the CLARIAH AT Consortium and Working Group leads from left to right in 2023: Andreas Baumann (university of vienna), max Kaiser (Austrian National library), Anja grebe (university for Continuing education Krems), günter mühlberger (university of Innsbruck), Karlheinz mörth (Austrian Academy of Sciences), walter Scholger (university of graz), and vera maria Charvat (Austrian Academy of Sciences). The consortium has recently drafted a strategy paper and work programme, the ‘Digital Humanities Austria Strategy 2021+’,1 based on efforts by the consortium and a consultation of the professional Austrian DH community. It defines four guidelines for the development of Digital Humanities in Austria, and serves as a work programme for working groups of the same name: 1 http://gams.uni-graz.at/o:clariah.dha-strategie-2021-en
10 11 AustriA • Research infrastructures and networks, • Research data and repositories, • Methods, services and tools, • Education, training, and knowledge transfer. Austria currently operates two CLARIN B-centres: the Austrian Centre for Digital Humanities and Cultural Heritage,2 and the Centre for Information Modelling at the University of Graz.3 Both B-centres run Core Trust Seal certified repositories: ARCHE4 and GAMS.5 Additionally, there are two Austrian K-Centres; one is operated by the Phonogrammarchiv at the Austrian Academy of Sciences,6 while the CLARIN K-Centre for Terminology Resources and Translation Corpora is based at the Centre for Translation Studies at the University of Vienna. CLARIAH-AT members participate in CLARIN committees, such as the User Involvement Committee, the CLARIN Legal and Ethical Issues Committee, the CLARIN Committee for Standards and Interoperability, and the Standing Committee of CLARIN Technical Centres. Austria also hosts several language resources, language corpora, lexical resources, and pertinent tools. Activities by the national consortium include an annual Summer School, dedicated project funding calls7 in the context of ensuring interoperability and reusability of national digital humanities data and resources and funding opportunities for early career researchers.8 With close ties to Austrian cultural heritage stakeholders, CLARIAH-AT also stands at the forefront of the currently ongoing national initiatives regarding the digitisation of cultural heritage and the subsequent re-use, processing, and making available of these data. 2 https://www.oeaw.ac.at/acdh/acdh-home 3 https://digital-humanities.uni-graz.at/en 4 https://arche.acdh.oeaw.ac.at/browser 5 https://gams.uni-graz.at/context:gams?mode=&locale=en 6 https://www.clarin.eu/blog/tour-de-clarin-phonogrammarchiv-austrian-academysciences-clarin-knowledge-centre 7 https://clariah.at/en/project-funding 8 https://clariah.at/en/funding-opportunities-for-junior-researcher Tool | Tool Chain for Sentiment Analysis written by Martina Scholger Thanks to CLARIAH-AT project funding, dedicated sentiment dictionaries and a freely available tool chain for conducting sentiment analysis on Italian, French, and Spanish Spectators periodicals from the digital scholarly edition ‘The Spectators in the International Context’9 were developed. The Spectator press is a journalistic-literary genre of the 18th century, popularising enlightened ideas among a non-academic audience in an entertaining way. The transmission of values, social norms, as well as positive and negative behavioural patterns and (character) traits, was communicated through emotional engagement with the audience. This makes the study of the material by means of sentiment analysis particularly interesting. Sentiment words highlighting. Initial experiments in the project ‘Distant Spectators. Distant Reading for Periodicals of the Enlightenment’10 have shown that freely available sentiment dictionaries and tools are especially tailored to modern languages, mostly English, and widely used in the context of social media analysis. Therefore, they are not suitable for 18th-century literary texts that show, for example, orthographic variances and shifts in meaning. Hence, the project created its own sentiment dictionaries. For this purpose, a certain number of seed words were selected from the corpus and then manually annotated by experts with respect to their polarities (positive, negative, neutral). Based on these seed words, word embeddings were trained, and a machine learning classifier was used to transfer the sentiment score to 9 http://gams.uni-graz.at/spectators 10 http://gams.uni-graz.at/context:dispecs
12 13 AustriA other words occurring in a similar context. In this way, the list of seed words was expanded using computational methods and the time-consuming manual annotation process was shortened. In addition to the introduction of reusable sentiment dictionaries, a freely and publicly available tool chain based on Jupyter Notebooks was developed, enabling researchers to apply 1) the dictionary creation process, and 2) the actual sentiment analysis methods (i.e., import of dictionaries, computation of sentiment, data preparation, and various visualisations) to their own material. The notebook contains executable code as well as tutorial-style introductions to concepts such as word embeddings, k-nearest neighbour classification and dictionary-based sentiment analysis. Thus, the notebooks can be used in teaching and training, but also as a basic framework for researchers in the field, who can replace certain components with more sophisticated methods as needed. overview of the dictionary Creation Pipeline. The project was supplemented by a three-day online workshop on ‘Sentiment Analysis for Literary Studies’, which introduced the methods and the developed tool chain. The workshop materials, including slides, exercises, and videos of the evening lectures, are available on the DiSpecs website.11 The dictionaries and the tool chain were published in the new publication format of a code experiment12 in Melusina Press, as part of the virtual DHd 2021 conference. 11 https://gams.uni-graz.at/o:dispecs.blg.sa.3#materials 12 https://doi.org/10.26298/ezpg-wk34 Resource | Austrian Media Corpus written by Hannes Pirker A public-private cooperation between the Austrian Academy of Sciences (ÖAW) and the Austrian Press Agency (APA)13 has made it possible to provide the scientific community with the Austrian Media Corpus (amc),14 a unique corpus that almost fully covers the whole country’s print media production of the past 30 years. The Austria Press Agency (APA) initiated the collection of the original data, beginning with their press releases in 1986. From 1992 onwards, print media, such as newspapers and weekly or monthly magazines, were added to the collection, as well as selected transcriptions of interviews and news stories from several television channels, reaching an almost complete coverage of Austria’s print media production from 1998 onwards. The amc is a plain text corpus comprising born-digital data only. As of 2022, the corpus contained 47 million articles from 58 different media outlets, constituting more than 11 billion tokens. Thus, the amc ranks among the largest collections of its kind. It is annually updated with new data provided by APA, further increasing the amc by approximately 500 million tokens a year. 13 https://apa.at/ 14 https://amc.acdh.oeaw.ac.at
14 15 AustriA Annotations, Legal Aspects and Applications The texts in the corpus are provided with basic metadata, i.e., name of the news media, date of publication, geographical region of origin, and a rough classification of texts based on the newspaper section they were published in. In terms of linguistic annotation, the texts have undergone some de-duplication heuristics and have been tokenised, lemmatised, and annotated with different part-of-speech taggers, a dependency parser, and a named entity recogniser. The conditions of use are specified by APA, the collector of the original data and holder of the rights to the collection. Access to the amc is exclusively granted for the purpose of linguistic research. The majority of users, i.e., academics and students, can access the corpus free of charge. Fees apply when the amc is used within a funded project. The majority of projects which make use of the amc are concerned with the analysis of linguistic variation in general and lexical variation in particular. The amc is also used professionally as a source of information by the Digitales Wörterbuch der Deutschen Sprache (DWDS) project and by the Council for German Orthography for monitoring the application of orthographic rules in everyday life. Event | #digitalDHaustria Twitter Showcase 2021 written by Sabrina Melcher From 14–16 April 2021, CLARIAH-AT hosted a Twitter showcase event, which introduced more than 100 Austrian Digital Humanities projects from all across the country through a series of tweets. The thematic focus ranged from 3D reconstructions of Arabic inscription impressions, through a number of digital editions of resources from various time periods and humanities disciplines, to the application of distant reading methods to journals of the Enlightenment, to name just a few of the diverse and groundbreaking projects involved. Each project was promoted by up to three tweets, which were often accompanied by additional Twitter and other dissemination activities by the researchers involved in the projects. The audience was directed to individual project pages on the Digital Humanities Austria Projects website,15 containing further information and resources for each individual project. The three-day event featured 214 tweets, creating more than 1800 retweets and more than 2600 likes. In addition, it led to almost 1000 visits to the digital-humanities.at/clariah.at website,16 ample evidence that the event drew a lot of international attention! For up-to-date information, follow us on LinkedIn at CLARIAH-AT.17 15 https://clariah.at/en/projects 16 http://digital-humanities.at/ or http://clariah.at 17 https://www.linkedin.com/company/clariah-at Delete paragraph break, move to the previous line.
28 29 icelAnd • IGC-Adjud: Adjudications (CC BY license), • IGC-Books: Published books (restricted licence), • IGC-Laws: Law, bills and resolutions (CC BY licence), • IGC-Journals: Scientific/academic journals (CC BY licence), • IGC-News1: News (CC BY licence), • IGC-News2: News (restricted licence), • IGC-Parla: Parliamentary speeches (CC BY licence), • IGC-Social: Forums, blogs and tweets (CC BY licence), • IGC-Wiki: Texts from the Icelandic Wikipedia (CC BY licence). The corpus was first published in 2018, and for the first five years, a new version was published every year, containing more texts and using newer technology for PoStagging and lemmatising the text. The latest version of the corpus is the 2024 version, which is an extension of the 2022 version, mainly containing text from the years 2022 and 2023. Each subcorpus is published in two versions, one with annotated text and another with unannotated text. In both cases, the TEI-conformant XML format is used. They are also published in JSONL format, which is suitable for LLM training, and made available on HuggingFace.29 As already mentioned, the corpus can be queried using the KORP tool, but in addition to this, two websites have been created for further investigation of the texts, one with an N-gram viewer30 and the other with information about word frequency.31 Since the publication of the first version of IGC in 2018, the corpus has shown itself to be a valuable resource, both for building LT tools of Icelandic as well as for linguistics research. 29 https://huggingface.co/datasets/arnastofnun/IGC-2022-1 30 https://n.arnastofnun.is 31 https://ordtidni.arnastofnun.is With the n-gram viewer, one can find and compare the frequency of individual words or word combinations from the Icelandic gigaword Corpus in a historical context. Event | What Can CLARIN Do for You? written by Starkaður Barkarson The humanities conference is held every year by the University of Iceland. In 2024 AMI, CLARIN-IS’s B-centre, held a seminar titled ‘What can CLARIN do for you?’. The seminar started with Starkaður Barkarson, the national coordinator of CLARIN-IS, introducing CLARIN and explaining how the infrastructure could be used to support research in the fields of the humanities, social sciences, and language technology. The main objectives of the infrastructure were discussed, and the services provided by both CLARIN ERIC and CLARIN-IS were presented. Auður Pálsdóttir, a teacher at the School of Education, University of Iceland, presented the Corpus of Icelandic Academic Vocabulary (MÍNO)32 and the List of Icelandic Academic Vocabulary Level 2,33 both related to academic vocabulary. The project received assistance from the CLARIN B-centre in both preparation and presentation, and its outputs are hosted by the CLARIN repository. 32 http://hdl.handle.net/20.500.12537/299 33 http://hdl.handle.net/20.500.12537/307
30 31 Finally, Einar Freyr Sigurðsson and Steinþór Steingrímsson, both researchers at AMI, presented research conducted using corpora. They demonstrated how to use AMI’s corpus search engine34 in relation to biases and various linguistic variables, such as the incorporation of the prefix endur- (‘re-’). Additionally, they also discussed the historical development of the reflexive passive in Icelandic, and the n-gram viewer concerning political discourse, used to explore interesting examples and study the language. For example, Einar Sigurðsson demonstrated, among other things, how discussion about two political parties had evolved in recent years by using the n-gram viewer on the Icelandic Gigaword Corpus. mentions of two political parties in the Icelandic gigaword Corpus. 34 https://malheildir.arnastofnun.is Interview | Sigríður Ólafsdóttir The conversation was led by Karina Berger Sigríður Ólafsdóttir is an associate professor at the School of Education, University of Iceland. She has benefited from CLARIN-IS tools in creating MÍNO, the first corpus specifically targeted for compulsory school learners, and the derived Icelandic Academic Word List, LÍNO-2. Using data from the CLARIN-IS repository, Sigríður Ólafsdóttir and Auður Pálsdóttir, associate professors at the University of Iceland, have created and published a list of Icelandic academic words – the first of its kind in Iceland. The researchers first developed a new corpus, MÍNO,35 from which a word frequency list of Icelandic vocabulary36 was created. The texts in the MÍNO corpus were obtained from two corpora already available in the CLARIN-IS repository, the Icelandic Gigaword Corpus37 and the Tagged Icelandic Corpus.38 The list of Icelandic academic words (LÍNO-2)39 was systematically selected from the frequency list of MÍNO. Starkaður Barkarson, the National Coordinator for CLARIN-IS, assisted the researchers in creating the new corpus. All three products, the MÍNO corpus, the MÍNÓ frequency list, and the LÍNO-2 word list, have been uploaded to the CLARIN-IS repository. 35 http://hdl.handle.net/20.500.12537/299 36 http://hdl.handle.net/20.500.12537/306 37 https://clarin.is/en/resources/gigaword 38 https://clarin.is/en/resources/mim 39 http://hdl.handle.net/20.500.12537/307 icelAnd
32 33 Sigríður Ólafsdóttir, what is your research area and academic background? < My initial research focused on teaching Icelandic to children of foreign origin in Icelandic schools. My current study area is mainly the following: active use of the Icelandic language by children and adolescents, vocabulary acquisition, reading comprehension, writing proficiency, the knowledge of Icelandic academic vocabulary, and the use of Icelandic academic vocabulary in school activities. I currently work as an associate professor at the School of Education at the University of Iceland. > Tell me about the Academic Word List project.40 You began by creating a new corpus. < Yes, first we compiled a new Icelandic language corpus: MÍNO (Ice. Málheild fyrir íslenskan námsorðaforða). This is the first corpus to be targeted at Icelandic students, and it is designed to enhance their studies in all areas. MÍNO includes selected student and academic texts, as well as texts from societal discussions, e.g., parliamentary debates, from two existing corpora: the Tagged Icelandic Corpus (MÍM) and the Icelandic Gigaword Corpus (IGC). All the texts are from this century. The new MÍNO corpus includes a total of 31,680,235 running words. Words (lemmas) that appear 100 times or more in the corpus are listed in the order of their frequency. The total frequency list includes 10,314 words (Tier 1, 2, and 3). > What can MÍNO be used for? < MÍNO is a valuable contribution to the education of Icelandic learners, with two main target groups. The first one is students whose first language is Icelandic. They benefit from it because it reinforces their command of important Icelandic vocabulary. Now that students speak so much English outside of school, and the Icelandic youth is practically bilingual, it is crucial to enhance their understanding of Icelandic 40 The project was funded by the university of Iceland Research Fund (Ice. Rannsóknasjóður Háskóla Íslands), the Icelandic language Technology Fund (Ice. markáætlun um tungu og tækni) and the Icelandic language Fund (Ice. Íslenskusjóðurinn). It received the Science and Innovation prize at the university of Iceland in the society category in spring 2023. vocabulary and their proficiency in using Icelandic words effectively. Unfortunately, we have seen a rapid decrease in performance in the PISA41 survey from 2000 to 2022, particularly in reading literacy and science literacy, so we hope that LÍNO-2 can improve literacy skills in particular. The second target group is students who have Icelandic as their second language, i.e., firstgeneration and second-generation immigrant students. These groups generally have poor skills in the Icelandic language, even after having spent 3 to 4 years in kindergarten and 10 years in compulsory school. This is the result of a variety of poor policy decisions, which have affected school activities at all levels. These include a low rate of qualified teachers in kindergarten, and the fact that teachers have generally not been provided with enough knowledge and skills to teach bilingual and multilingual students from immigrant backgrounds. This, in turn, makes it difficult to effectively prepare them for further studies in Icelandic schools. Proficiency in the Icelandic language, particularly in academic vocabulary, is pivotal for their academic progress. > How did you develop the Icelandic Academic Word List? < The Icelandic Academic Word List, LÍNO-2, was developed by selecting Tier 2 words from the frequency list of MÍNO, that is, words that are beyond the most frequent words and words that are used across different subject areas. Tier 2 words play a fundamental role in learners’ active participation in school, for the development of their reading comprehension, as well as discussion and writing skills. Thus, these words lie at the heart of academic success at all school levels and academic fields. A rule of thumb is that the more frequent the words are, the greater their importance. This means that the youngest learners and those who are in their first stages of Icelandic studies should be taught the most frequent words, and then gradually words of lower frequency. Both objective and subjective approaches were used42 to categorise the most frequent 5,000 words on the MÍM frequency list, i.e., to distinguish between words belonging to Tier 1 and Tier 2, as well as words belonging to Tier 2 and Tier 3. The Icelandic Academic Word List, LÍNO-2, contains 2,294 Tier 2 words (lemmas). > 41 https://www.oecd.org/en/publications/pisa-2022-results-volume-i-and-ii-country-notes_ ed6fbcc5-en/iceland_4e941265-en.html 42 https://netla.hi.is/greinar/2023/alm/09.pdf icelAnd
34 35 You then developed an Icelandic academic vocabulary test. < The Icelandic Academic Vocabulary Test aims to investigate and compare the comprehension of LÍNO-2 words by Icelandic learners of different ages. The main questions are whether older learners know more words than younger ones, and if there is a relation between the word frequency and learners’ knowledge of the words. The words on LÍNO-2 were divided into five bands depending on their frequency in MÍNO. The first frequency band includes the most frequent 1,000 words, the second band includes words from the next 1,000 words, and so on. Words from each frequency band were selected and included in the new Icelandic academic vocabulary test. The first administration of the test was a multiple-choice questionnaire, where participants were asked to select the sentence that correctly used the stimulus word among three possibilities, based on its usage in a sample sentence. The test was given to 851 learners in grades 4, 7, 9 and 10, in spring 2023. The same ten words from all frequency bands were included in all tests; additionally, the youngest learners received more words from the most frequent bands, the 7th graders from the middle bands, and the 9th and 10th graders from the lowest frequency bands. Findings demonstrated poor inter-validity for the 4th graders, indicating that their answers were, to a large extent, based on guessing. However, this was not the case among the older groups, as the 9th and 10th graders outperformed the 7th graders, whereas there was no difference between the two oldest age groups. No relation was detected between the frequency of the words and the extent to which the learners understood the words. The second version of the test builds on findings from the first phase and was collected in May 2024 with approximately 2,000 participants in grades 7, 8, 9 and 10, in 18 compulsory schools around the country. The test was divided into two sections: • A multiple-choice questionnaire, with the same format as in the first version; • A self-evaluation questionnaire, in which participants were asked to report their degree of understanding of the word (on a 4-point scale) and the frequency with which they use the word (on a 3-point scale). The third version, a multiple-choice questionnaire based on the findings of the first two versions, was administered to 350 learners in grades 7, 8, 9 and 10 in January 2025. The final version of the test, based on the findings from the three pre-tests, will then be used in an intervention study in autumn 2025. > You have also developed quality texts, illustrating the use of the words from the list. Why? < As part of the project, master’s students in creative writing at the University of Iceland composed short stories and expository texts with selected words from LÍNO-2, especially those that fell within the most frequent bands. There is a need for Icelandic texts that address interesting topics for young learners, while also including selected important Icelandic words based on their frequency in published texts from the current century. In this way, the acquisition of the Icelandic language, whether as a first or second language, becomes more effective. The overall aim is to strengthen students’ vocabulary skills, which is foundational for their academic achievement in Icelandic schools, at all levels and fields of study. The Icelandic Academic Word List and the quality texts are now being translated into six languages spoken by Icelandic immigrants: English, Filipino, Polish, Spanish, Thai, and Ukrainian. We plan to compose more texts with words from LÍNO-2 further down the frequency list. > What were some of the main outputs of the project? < The next step in the research project is an academic vocabulary intervention study, conducted in autumn 2025. The aim is to measure the impact of an explicit instruction of LÍNO-2 words, with learners in grades 7, 8, 9 and 10. Teachers in the intervention schools will take a two-day course by the end of August, and will be provided with teaching material that includes: • Written short stories and expository texts containing LÍNO-2 words; • Exercises to facilitate the learning: definition of words, morphological awareness, Frayermodel, filling-in blanks, as well as ideas for discussions and writings in which the learners are encouraged to use the words taught. icelAnd
36 37 To assess the effectiveness of the intervention, preand post-tests will be administered to both the research group and a comparison group with learners of the same age: • The Icelandic Academic Vocabulary Test; • Writing test with a list of the taught LÍNO-2 words and academic words not taught, during which learners are asked to use as many of them as they can; • Reading comprehension test available from the Centre for Education and Schools Services (Ice. Miðstöð menntunar og skólaþjónustu). From the intervention and interviews with participating teachers, teaching guidelines will be developed to accompany LÍNO-2 and the related quality texts, encouraging discussions and writing activities in school settings. The aim is that learners will master the words, both in terms of comprehension as well as active usage. > What are your plans for the future? < We will apply for funding to create a new Icelandic Academic Word list with Tier 2 words that are of lower frequency than the most frequent 5,000 words of MÍNO, meaning that the new list will proceed from LÍNO-2. This must be followed by a new Icelandic academic vocabulary test, intended as a continuation of the other one, which can be administered to older learners or those with more developed Icelandic language skills. Icelandic Tier 2 words lie at the core of the Icelandic language, making it essential that they are widely understood and used. These words enable discussions of complex issues in Icelandic, e.g., philosophy, anthropology, economics, and medicine at a university level, as well as geography, literature, social studies, and mathematics in compulsory schools. Knowing these words, both in terms of comprehension and active use, is therefore fundamental to all study fields taking place and being conducted in Icelandic. > icelAnd
38 39 • lIThuANIA Introduction | CLARIN-LT written by Jurgita Vaičenonienė Welcome to CLARIN-LT! In 2024, we celebrated our 10th anniversary: Lithuania became a full member of CLARIN ERIC on 25 October 2014, and a year later, a consortium of three partners was established. At present, the consortium includes six full partners: Vytautas Magnus University (coordinating institution), Kaunas Technology University, Vilnius University, and the newer consortium members Mykolas Romeris University, Baltic Institute of Advanced Technologies, and the Institute of Baltic Region History and Archaeology. Although the composition of our team members has changed over time, we have always been an interdisciplinary and international consortium. We are happy that the consortium has been recognised by the Research Council of Lithuania as an exemplary infrastructure with Landmark status, which secures CLARIN-LT stable funding until 2030. CLARIN-LT has also been included in the new Roadmap. CLARIN-LT team at the CLARIN Annual conference 2024. Our main services offered to academia, as well as the public, industrial, and cultural sectors, include a repository with various Lithuanian datasets and analysis tools. This is especially important in view of Lithuanian being among the lesser-resourced languages in terms of language resources and tools. Our priority goals are: • Preservation, accessibility, and visibility of the Lithuanian language in the digital environment, • Ensuring that the data stored matches the principles of FAIRness, • Encouraging the reuse of existing resources, • Increasing the international participation and collaboration of Lithuanian researchers, • Building a community around the general CLARIN ERIC goals, • Sharing our expertise and knowledge with all interested parties. Since the presentation in the Tour de CLARIN in 2018,43 the CLARIN-LT consortium, apart from expanding to six members, has undergone other major developments. One of them was becoming a part of the virtual CLARIN Knowledge Centre for Systems and Frameworks for Morphologically 43 https://www.clarin.eu/sites/default/files/Tour-de-CLARIN-Lithuania.pdf
40 41 Rich Languages, SAFMORIL.44 Apart from other tasks, CLARIN-LT in SAFMORIL offers a helpdesk for Lithuanian-related user questions and support to users in applying corpus linguistics and natural language processing methods, as well as consulting the users on the Lithuanian morphology, syntax and semantics. This cooperation led to a vivid information exchange on user requests and resulted in joint presentations at the conferences, for example, at the CLARIN Bazaar 2023. In 2022, the CLARIN-LT centre became a constituent part of the Institute of Digital Resources and Interdisciplinary Research45 (SITTI) at Vytautas Magnus University, thus expanding the collaborative potential with researchers from three research strands: (1) language use research, resources and technologies, (2) machine learning for language technologies, and (3) applied psychoand sociolinguistic research. This collaboration has already resulted in mono or multilingual resources and tools for Lithuanian, but also some other languages, such as Scottish Gaelic. The visibility of CLARIN-LT has also been enhanced by the University library, which promotes our services on their website, and all the latest news is posted on both our Facebook page and CLARIN-LT website, the latter being bilingual (Lithuanian and English). user involvement event at the vytautas magnus university that hosts ClARIN-lT. CLARIN-LT members are active in various CLARIN ERIC initiatives and committees, for example, the ‘Trainers’ Network and Knowledge Infrastructure Committee’ (KIC). One of the most recent outputs of such cooperation was a joint presentation given by KIC committee members at the 2024 Annual DHNB conference in Reykjavik, ‘Transnational 44 https://www.kielipankki.fi/safmoril 45 https://sitti.vdu.lt/en Research Infrastructure: A Journey Through CLARIN Knowledge Centres’, by Jurgita Vaičenonienė, Michal Kren, Vesna Lušicky, and Vincent Vandeghinste. The information about the concept of the K-centre and its value for researchers was new for the audience of participating Digital Humanities researchers. The listeners found it useful to hear a presentation on the operation of the network of K-centres, as well as an overview of existing centres and their areas of expertise. All interested were invited to contact us either as potential users or future K-centres. More broadly, CLARIN-LT is proud to have an active collaboration both in academia and the industry sector. The most popular tools stored in the CLARIN-LT repository that appeal to the general public and industry are Lithuanian Spelling Checker for Macintosh computers, Lithuanian Spelling Checker V.1.0.42 for LibreOffice and OpenOffice, and Lithuanian speechto-text transcriber produced in the project SEMIANTIKA-2. The speech-to-text transcriber is an exceptional case as, since its release, it has already been implemented at the Lithuanian Parliament, National Radio and Television, Police Department of Lithuania, universities and some companies. In 2020, the tool received a national award of ‘science-based business service of the year, 2020’ from the Lithuanian Business Confederation (Petrauskaitė et al. 2022: 523). We are pleased that Lithuanian researchers are increasingly becoming aware of our services and choose to store their research outputs in our repository. It is interesting to observe how the most recently uploaded corpora and datasets reflect the key thematic project trends and mirror current geopolitical tensions and realia, such as cybersecurity,46 the Russo-Ukrainian war,47 and pandemics.48 Currently, one of the most recent projects by SITTI and CLARIN-LT researchers, ‘Morphologically and Syntactically Annotated Corpora Models for Training (Gold Standards)’ (2024–2026) has been launched, and thus, more resources are going to be added to the CLARIN-LT repository. One of the most recent highlights is the Morfuoklis tool, a morphological analyser for Lithuanian developed by the CLAIRN-LT team. Reference: Petrauskaitė, R., Amilevičius, D., Dadurkevičius, D., Krilavičius, T., Raškinis, G., Utka, A. and Vaičenonienė, J. 2002. CLARIN-LT: Home for Lithuanian Language Resources. Eds., Fišer, D., and Witt, A. CLARIN: The Infrastructure for Language Resources, Berlin, Boston: De Gruyter. https://doi.org/10.1515/9783110767377 46 https://clarin.vdu.lt/xmlui/handle/20.500.11821/59 47 http://hdl.handle.net/20.500.11821/57 48 http://hdl.handle.net/20.500.11821/54 litHuAniA
42 43 Tool | Lithuanian Spelling Checker V.1.0.45 for macOS written by Virginijus Dadurkevičius The Lithuanian CLARIN-LT repository49 hosts several related resources consistently ranked among the most accessed items over recent years. The common denominator is the Lithuanian morphology rules and the corresponding dictionaries implemented on the Hunspell platform. These most popular items are the ‘Lithuanian Hunspell dictionary’,50 ‘Lithuanian Spelling Checker V.1.0.45 for LibreOffice and OpenOffice’,51 ‘Lithuanian Spelling Checker V.1.0.45 for Linux’,52 and ‘Lithuanian Spelling Checker V.1.0.45 for macOS’.53 The latter, developed specifically to enable system-wide Lithuanian spell checking on Apple computers, has proven especially popular. Since 2019, the tool has been downloaded 2,385 times (i.e., 35 times per month), and is the overall leader of all CLARIN-LT resources. The tool’s popularity stems from the limited native support for the Lithuanian language on Apple computers. Installing this CLARIN-LT tool automatically allows you to spellcheck almost all programs used on Apple computers, and this explains the popularity of the tool. The morphological rules and the dictionary have been compiled through extensive academic research, regular cooperation with the State Commission of the Lithuanian Language, and years of monitoring of Lithuanian corpora. Every effort has been made to include in the dictionary all words in actual use (i.e., ‘real information circulation’), provided they conform to language norms. Deprecated loanwords or extremely rare, exotic, obsolete, jargon, or insulting forms were discarded from the list. The resulting dictionary consists of more than 181,000 lemmas: 48,000 common nouns, 75,000 proper nouns, 15,000 adjectives, 54 pronouns, 158 numerals, 38,000 verbs, 4,000 adverbs, and 2,000 other items (prepositions, conjunctions, particles, onomatopoeias, interjections, acronyms and abbreviations). 49 https://clarin.vdu.lt/xmlui/?locale-attribute=lt 50 http://hdl.handle.net/20.500.11821/64 51 http://hdl.handle.net/20.500.11821/24 52 http://hdl.handle.net/20.500.11821/23 53 http://hdl.handle.net/20.500.11821/22 Although Hunspell is primarily a spelling checker, it also supports morphological analysis and synthesis. Moreover, the ability to efficiently perform lemmatisation (stemming) makes this platform the best option for text search engines (e.g., Solr/Lucene) and information retrieval. Taggers, grammar checkers and other basic natural language processing tools can also be created using properly built Hunspell language resources. The platform also has a dedicated Python module for using Hunspell dictionaries, further facilitating integration into language technology pipelines. Every Hunspell language resource consists of two files: a dictionary and affixes (which may be empty). The dictionary (.dic filename extension) contains main forms (i.e., lemmas), whereas the affixes contain the morphological rules to generate all possible forms. The second component, the so-called ‘affix file’ (.aff filename extension), contains information on metadata, preferable suggestions for spelling correction, grouping of rules, explicit tags for flexing and non-flexing properties, and rules for suffix and affix alteration. In order to make the Hunspell resources suitable for creating basic language tools, the following principles were kept in mind: • Every flexion paradigm (consisting of one or more rules) should be thoroughly generated from one single lemma in the dictionary file (especially for irregular verbs); • Every individual alteration should have its own morphological tag, e.g., ‘Masc_Sg_Il’ for masculine+singular+illative; • Every dictionary item should have references for part of speech and other non-flexing information; avoid prefixation via rules, use the dictionary instead – affixed forms may have completely different meanings and using them under a single lemma may cause problems for text search engines; • Calling depth can be no more than 1. The macOS operating system provides native support for Hunspell-based resources. Once the dictionary and affix files for a given language are placed in the appropriate system directories, spell checking functionality is activated automatically. ‘The Lithuanian Spelling Checker V.1.0.45 for macOS’ as well as the *.dic and *.aff files for the Lithuanian language were developed during the Semantika 254 project by Virginijus Dadurkevičius, Danielius Algirdas Ralys, Arūnas Samuilis, Jonas Vaičiulis, Franciška Ralienė and Linas Valiukas. The tool is distributed under the open licence ‘CLARINLT PUBLIC END-USER LICENCE (PUB)’. 54 https://github.com/Semantika2 litHuAniA
44 45 lithuanian spelling checker in action – downloaded from ClARIN-lT repository, installed on macoS, automatically launched in word processor Pages and suggesting the right correction (gražiausias) for the misspelt Lithuanian word (gražeusias). Resource | English-Lithuanian Language Resources for the Cybersecurity Domain written by Sigita Rackevičienė55 Between 2022 and 2024, five English-Lithuanian language resources were deposited in the CLARIN-LT repository, with one more being currently prepared for deposit. All of these resources have been developed by the same team of researchers from Mykolas Romeris University and Vytautas Magnus University, starting with the joint research project DVITAS56 (Bilingual Automatic Terminology Extraction) and continuing through post-project initiatives (see Rackevičienė et al., 2021).57 The cybersecurity domain was chosen for its relevance in today’s digitalised world, characterised by increasing global connectivity, cloud services and the challenges of securing sensitive data, all of which require cyber awareness, cyber hygiene, and knowledge of relevant terminology from all internet users. 55 Sigita Rackevičienė is affiliated with the Mykolas Romeris University, Lithuania. 56 https://sitti.vdu.lt/dvitas/en 57 https://doi.org/10.5755/j01.sal.1.39.29156 The first resources, the ‘English-Lithuanian Parallel Cybersecurity Corpus - DVITAS’ (2022)58 and the ‘English-Lithuanian Comparable Cybersecurity Corpus - DVITAS’ (2022),59 were compiled to develop a deep learning-based bilingual terminology extraction methodology. The combination of these two types of corpora ensured the collection of terminology used not only in translated texts but also in original language texts, allowing for the inclusion of a wider range of discourses. The parallel corpus primarily consists of EU documents on cybersecurity issues in English and their Lithuanian translations, while the comparable corpus includes a broader variety of texts from legislative, administrative, informative, academic, and media discourses. In 2024, an updated version of the parallel corpus (‘English-Lithuanian Parallel Cybersecurity Corpus - DVITAS v.2.0’)60 was deposited in the CLARIN-LT repository, expanding the number of EU documents and metadata. Additionally, a new parallel corpus of Lithuanian cybersecurity documents issued by the institutions of the Republic of Lithuania, along with their English translations, is being prepared for deposit. The corpora have been the primary sources for developing the Lithuanian-English Cybersecurity Termbase,61 the TBX version (‘Lithuanian-English Cybersecurity Termbase v.0.1,’ 2023),62 of which was uploaded to the CLARIN-LT repository. The TBX provides data structured according to the ISO 30042:2019 standard for representing and exchanging terminological information exported from terminology databases. The TBX contains 233 concept records, each featuring Lithuanian cybersecurity terms and their English equivalents, definitions in both languages, their sources, and contextual usage examples extracted from the aforementioned corpora. Finally, the wide variety of Lithuanian terminology inspired a survey on the preferences of synonymous cybersecurity terms among different user groups. A total of 593 respondents participated in the survey, sharing their insights on the most suitable Lithuanian terms for 10 cybersecurity concepts. This led to the development of another sociolinguistic resource – the terminological survey dataset, which was deposited in the CLARIN-LT repository in 2024 (‘Survey Data on Preferences of Lithuanian Cybersecurity Terminology’).63 The developed resources – corpora, termbase, and survey dataset – have enabled research into the cybersecurity domain from various perspectives: examining source availability and data curation issues relevant to the compilation of parallel and comparable corpora (see Utka et al., 58 http://hdl.handle.net/20.500.11821/46 59 http://hdl.handle.net/20.500.11821/47 60 http://hdl.handle.net/20.500.11821/63 61 https://www.terminologue.org/csterms/ 62 http://hdl.handle.net/20.500.11821/55 63 http://hdl.handle.net/20.500.11821/59 litHuAniA
46 47 2022),64 conducting experiments on automatic term extraction using deep learning systems (see Rokas et al., 2020),65 analysing the density and diversity of cybersecurity terminology across various text genres (see Rackevičienė et al., 2022),66 developing data collection and structuring principles for the compilation of a bilingual cybersecurity termbase (see Rackevičienė et al., 2023),67 and analysing users’ choices of the most suitable Lithuanian cybersecurity terms and the motivations behind their selections (see Rackevičienė & Utka, 2024).68 The compiled resources are expected to be reusable in future studies as well. english-lithuanian Cybersecurity Corpora System used for Bilingual Terminology extraction and Termbase Compilation. Scheme developed by Andrius utka (vytautas magnus university). 64 https://doi.org/10.3384/ecp18912 65 https://doi.org/10.3233/FAIA200600 66 https://doi.org/10.15388/RESPECTUS.2022.41.46.105 67 https://doi.org/10.31724/rihjj.49.2.12 68 https://doi.org/10.5755/j01.sal.1.44.36235 Event | Annual CLARIN-LT Seminar 2024 written by Jurgita Vaičenonienė One of the long-established CLARIN-LT traditions is an annual seminar, usually organised at the end of the year. The seminar aims to gather all CLARIN-LT community members and all those interested in our activities, to share the main highlights of the year. The 2024 seminar was quite exceptional, as CLARIN-LT celebrated its 10th anniversary. During this hybrid event, we were particularly happy to welcome more than 40 registered participants from six Lithuanian educational institutions, such as Vytautas Magnus University, Vilnius University, Mykolas Romeris University, Kaunas University of Technology, Institute of the Lithuanian Language, Kauno kolegija Higher Education Institution, guests from the Research Infrastructure Lithuanian Data Archive for Social Sciences and Humanities (LiDA), and, of course, members of different CLARIN-LT consortium partner institutions. Although we were concerned about the hybrid mode of the event, the follow-up feedback of the participants showed us that this is a preferred mode and that it should be maintained in the future. To commemorate the 10th anniversary of CLARIN-LT, we showcased the newest projects and open access resources in Lithuanian that are being developed to promote more effective use of data in academic and educational activities, hoping that the presented information will be of interest to researchers, teachers, and students in the social sciences and humanities. For those participants who were unable to attend, all presentations were made available on our website. First, the CLARIN-LT national coordinator, Jurgita Vaičenonienė, gave a welcome speech thanking everyone for their contribution and interest in the infrastructure over the years. Second, the CLARIN-LT project leader, Andrius Utka, presented the newest project ‘Membership in the international scientific research infrastructure CLARIN ERIC (Common Language Resources and Technology Infrastructure) Plan’ (joint no. VS-15). The project is funded by the Lithuanian Research Council for the period 2024-2029, with three members of the national CLARIN-LT consortium, Vytautas Magnus University (coordinator), Vilnius University, and Mykolas Romeris University, being mostly involved in the project activities. Furthermore, the major developments of the infrastructure since its establishment in 2014 were examined. The seminar highlighted projects and initiatives led by different consortium partners. For example, Erika Rimkutė (Vytautas Magnus University) presented challenges in the development of the new litHuAniA
60 61 POlAnd Service | CLARIN-PL Helpdesk written by Krzysztof Hwaszcz The ClARIN-Pl helpdesk team: Krzysztof Hwaszcz and Julia Klyus. CLARIN-PL has established a dedicated and active helpdesk to support researchers working in the humanities and social sciences. This initiative is designed to facilitate access to language technologies and digital tools that are essential in the field of DHSS (digital humanities and social sciences). Our helpdesk has become a central point for sharing knowledge, resources, and practical guidance on language processing in Poland by providing customised support and promoting collaboration. An important goal of the CLARIN-PL helpdesk is to promote the skills and expertise necessary for researchers to engage effectively with our wide range of language processing tools and services. To achieve this goal, we have created a dynamic system that is responsive to the evolving needs of the scientific community. Our helpdesk continuously adapts to infrastructure developments and user feedback, ensuring that the support provided remains current and relevant. The helpdesk is characterised by several important features: • Fast responsiveness: We prioritise timely assistance and guarantee responses to all queries and requests within 48 hours. • Multiple contact channels: Users can reach out directly to one of our two dedicated helpdesk employees or submit their inquiries via an easy-to-use ticketing system. • Comprehensive information services: Our staff provide detailed information about CLARINPL resources, including access to documentation, terms and conditions of use, legal frameworks, and training materials. • Issue management and technical support: Any reported defects or issues with tools and services are promptly forwarded to the relevant technical teams for resolution. • Advanced consultation options: For more complex inquiries that require specialised knowledge, the helpdesk organises individual or group consultations; these can sometimes take the form of mini-workshops, offering in-depth support tailored to the specific needs of the researcher or institution. • Workshop development: User inquiries often serve as the inspiration for broader educational initiatives; based on recurring requests, the helpdesk helps coordinate workshops, either onsite at the researcher’s institution or at the CLARIN-PL headquarters at Wrocław University of Science and Technology (WUST). The organisation of these events is managed by PolLinguaTech, the CLARIN Knowledge Centre for linguistic resources and technologies. • Research support and training development: PolLinguaTech also addresses requests for new training materials and provides support in preparing research funding applications that involve language technology components. • Data-driven improvements: All interactions, whether initiated via email or the ticketing system, are archived and used to generate statistical data; these metrics inform future improvements and support reporting obligations. Overall, the CLARIN-PL helpdesk offers comprehensive, personalised support for researchers at every stage of the scientific process, from planning and data collection to the analysis and interpretation of linguistic data. We want to help researchers get the most out of digital tools and infrastructure, so they can drive progress in the field of DHSS in Poland. For any questions or to request assistance, please contact us at [email protected].
62 63 POlAnd Event | CLARIN in Research Practice written by Krzysztof Hwaszcz 14th edition of the ‘CLARIN in Research Practice’ Workshop in Poznań. On 23 and 24 September 2024, the Faculty of Modern Languages and Literatures at Adam Mickiewicz University in Poznań hosted the 14th edition of the workshop ‘CLARIN in Research Practice’. The event was organised by the CLARIN-PL Language Technology Centre at Wrocław University of Science and Technology, in collaboration with the Faculty of Modern Languages and Literatures and the Faculty of Polish and Classical Philology at AMU, as well as other CLARIN-PL consortium partners. The workshop brought together researchers, developers, and infrastructure users from across Poland and beyond. As a flagship initiative of CLARIN-PL, this recurring event promotes connections between the development of language technologies and their application in the humanities and social sciences. Each edition provides a unique setting for knowledge exchange, training, and discussion, attracting a diverse academic audience. This year’s program upheld that tradition, offering a rich lineup of lectures, presentations, and consultation sessions. Participants had the opportunity to explore both the tools and services developed by CLARIN-PL and the research projects utilising its infrastructure. This blend of perspectives highlighted the broad potential of language technologies for a range of academic goals, from digital philology and corpus linguistics to sociolinguistics and historical linguistics. A particularly valuable aspect of the workshop was the opportunity for one-on-one consultations. Attendees could discuss their projects with experts and receive tailored advice on integrating CLARIN-PL resources into their workflows. These sessions addressed technical and methodological challenges and helped participants identify the most suitable tools and datasets for their needs. The workshop’s practical and interactive format is consistently praised as one of its greatest strengths. The success of the event was made possible by the contributions of many individuals and organisations. The organisers extend their sincere thanks to all who made the workshop a success: invited guests, participants, moderators, and administrative staff. Events such as ‘CLARIN in Research Practice’ are central to CLARIN-PL’s mission of promoting the use of language technologies in academic research and ensuring that these resources are accessible and useful to a broad scholarly audience. As interest in digital tools in the humanities and social sciences continues to grow, such initiatives are becoming increasingly vital. The CLARIN-PL team looks forward to the next edition and to continuing the dialogue between infrastructure developers and the academic community in future events.
64 65 POlAnd Interview | Agnieszka Hess The conversation was led by Karina Berger Agnieszka Hess is Professor of Social Communication and Media Science at Jagiellonian University, where she also directs the Institute for Journalism, Media and Social Communication. Her research explores political communication, civil dialogue, and the role of NGOs in democratic processes. Through collaboration with CLARINPL, she applies corpus linguistics tools to study parliamentary and media discourse, notably within the Civil Dialogue Observatory project. Please describe your academic background and current position. < I studied political science at the Jagiellonian University in Krakow, obtained my PhD degree in political science at the University of Vienna, and then returned to my home university, where I have been working for more than 25 years. I am Professor of Social Communication and Media Science at Jagiellonian University, chair of the Discipline Council for Social Communication and Media Sciences, and Director of the Institute for Journalism, Media and Social Communication at the same university. My research interests include political communication, the mediatisation of social reality, civil dialogue, and populism. My focus is on the analysis of the relations between institutions functioning within a country’s democratic system. I am particularly interested in the communicative aspect of these relations and the role that the media play. I am the author of more than 100 scientific publications in social communication and media sciences. I am also an expert, executor, and advisor in international and national research projects > Can you tell us a bit about how your research benefits from the CLARIN infrastructure? < I have been analysing the functioning and the role of non-governmental organisations in my country in the context of the democratisation process since around 2004, when Poland joined the structures of the European Union. I am interested in political, social, and media conditions for the development of the third sector and inter-sectoral cooperation in Poland. This is the most important research that I am carrying out using various quantitative and qualitative research methods. As part of the project, I conducted a multi-faceted study on non-governmental organisations as participants in political discourse. I analysed their social representations, the way they are portrayed in the media, and their communication strategies. I also studied the relationship (including communication) between the NGO sector and public administration, as well as between representatives of NGOs and journalists. The results of this study are presented in my monograph, ‘Social Media Participants of Political Discourse in Poland: The Publicity and Communication Strategies of Non-Governmental Organisations’ (in Polish). I am currently conducting research on parliamentary discourse, which is a continuation of the aforementioned project. I am interested in the institutionalisation of the functioning of NGOs and civil dialogue in Poland after 1989. These are long-term studies, covering a period of more than 30 years, that require working with large text resources, including the corpus of parliamentary discourse that consists of transcripts from the work of the Sejm and the Senate (plenary sessions and committee meetings of both chambers of the Polish parliament). This study would not be possible without the cooperation of the CLARIN-PL team and the possibility of using the computational linguistics tools provided by this consortium. The selected tools have been adapted to the needs of my research. Based on the analyses carried out with the help of these tools, I was able to characterise the political conditions for the development of NGOs and the features of the discourse of political decision-makers, concerning NGOs and cross-sectoral cooperation, in particular terms of office of the Polish parliament. Photo credit Janusz Wilkoński
66 67 POlAnd I also collaborated with the CLARIN-PL team as part of the Civil Dialogue Observatory73 (ODO), which I run. It is a scientific and didactic project implemented in 2015 by Jagiellonian University in consultation with the municipality of Kraków. In 2021, the subject of research of the ODO team was multiculturalism, or the changing structure of the inhabitants of Kraków towards a multicultural community. This is a topic that has become extremely important in the face of the war in Ukraine, since Kraków became a place where thousands of people fleeing from the war sought shelter. As part of the ODO project, we analysed both the Kraków City Council discourse during the period when city authorities introduced a programme for intercultural integration, as well as the media discourse at the time. The ability to use the infrastructure provided by the CLARIN consortium and their technical support allowed us to study entire corpora, including texts documenting the work of the Kraków City Council and media materials from that period. This resulted in a publication (in Polish) from December 2022, documenting the results of these analyses, including recommendations for city authorities regarding the examined issues. The monograph also contains a detailed description of the research methodology, the tools used, and the cooperation process with the CLARIN-PL team. > How did you hear about CLARIN, and how did you get involved? < I participated in some CLARIN-PL workshops at my home institute in 2018. It was an initiative of a friend who had already worked with this team. During the workshops, various research tools that are available as part of the CLARIN-PL infrastructure were presented. A few of them caught my special attention. I started to think about the possibility of using computational linguistics tools in my research, which enables working with large text resources. After the workshop, I contacted the people responsible for cooperation with users. After sending the initial research concept, I first received the information I needed, and then help, for instance, on the choice of tools, and how to adapt them to the needs of my project. This is how my scientific adventure with corpus research began. > 73 https://dialogobywatelski.org/ Which CLARIN tools and resources have you used, and how did you integrate them into your research? < My cooperation with CLARIN began with research on parliamentary discourse. With the support of the CLARIN team, I used the TOPIC tool for topic modelling, thanks to which it was possible to define and illustrate the structure of topics related to the issues of the analysed corpus. The TermoPL tool, which was used to analyse the language of politicians, for instance, topics, contexts, terminology, and vocabulary, also turned out to be very useful. In this study, I also tried to use sentiment analysis tools (SENTEMO, MULTIEMO, and WYDŹWIĘK) to determine the emotional overtones of a given statement. It did not work for this study, due to the specific nature of formalised parliamentary discourse. However, these tools worked well in the initial analysis of the media discourse, which was one part of the research on multiculturalism carried out by the Civil Dialogue Observatory. > What are the methodological and technical challenges that you face in your particular field? < The analysis of media discourses often requires working with vast resources of various types of text and other forms of communication, e.g., visual. Significant technical problems involve the collection, accumulation, and storage of research material. Deciding on the criteria for selecting the sample in a situation where it is impossible to cover the entire research material is often a methodological challenge. In addition, media experts have to face the specific character of their research subject. On the one hand, an analysis of the media and changes occurring in the media and under the influence of the media, requires the use and combination of various research methods, which extends the research process. On the other hand, in order to capture the dynamics of these changes, short-term, fragmentary research must be designed, which cannot cover the complexity of the analysed problem. >
68 69 What do you think needs to be developed to enrich CLARIN and make it better known within your research community? < I have been watching the growing popularity of the tools available as part of the CLARIN infrastructure among media and communication scholars in Poland. Researchers who are using the infrastructure and collaborating with the consortium and who present their research results in presentations and publications, are important ambassadors for CLARIN in the scientific community. I think that an even greater participation of representatives of the consortium in scientific events would help to boost CLARIN’s recognition among social science scholars. Conference programmes often lack practical presentations of research tools that are used in projects in the field of social sciences. > What is in store for your future collaboration with CLARIN? < I hope to continue corpus research in cooperation with CLARIN-PL, and I plan to use some of its tools, above all TOPIC and TERMO PL. In 2025, I was invited by professor Lucía Caro Castaño of the Universidad de Cádiz to do comparative research on how populist radical right parties are narratively constructing crises through political discourse. This is an extremely important and timely issue in view of the growth of the radical right in Europe (both in terms of the number of voters and the diversity of parties). Thus, we first want to examine the speeches of leaders of populist right-wing parties in parliamentary debates in Spain and Poland. These countries are currently experiencing a situation in which radical right-wing parties (Spain’s VOX and Poland’s Konfederacja) have unprecedented public support. The analysis will cover extensive material from 2019-2025, so I hope to work with CLARIN on this project. First of all, we will study the resources of the Polish Corpus of Parliamentary Discourse,74 which are provided by CLARIN-PL. In the first stage, we plan to conduct quantitative research using the TOPIC tool.75 In the second stage, we will conduct qualitative research on a smaller sample, generated purposely according to the categories determined by interpreting the results of thematic modelling. > 74 http://hdl.handle.net/11321/467 75 https://services.clarin-pl.eu/tools/topic POlAnd
70 71 SwITzeRlANd Introduction | CLARIN-CH written by Cristina Grisot and Clemens Lutz After being an observer of CLARIN since January 2023, Switzerland joined as a member on 1 May 2025. In Switzerland, CLARIN and its Swiss node were evaluated by the Swiss National Science Foundation to be of strategic importance for the Swiss scientific community working with language data in the context of Open Science. On the basis of this evaluation,CLARIN-CH was included on the 2023 Swiss roadmap for research infrastructures. CLARIN-CH76 represents an ecosystem of national infrastructure and networks which work in close partnership: • The CLARIN-CH Consortium, • The CLARIN-CH Coordination Office, • The Linguistic Research Infrastructure LiRI of the University of Zurich, • The Digital Discourse Lab of the Zurich University of Applied Sciences, • The Language Repository of Switzerland LaRS. 76 https://clarin-ch.ch The CLARIN-CH Consortium, the kernel of CLARIN-CH, was founded in 2020. To date, the Swiss Academy for Humanities and Social Sciences and nine Swiss higher education institutions are members of the consortium, and they co-fund CLARIN-CH activities by paying affiliation fees. These are: • University of Zurich (hosting institution), • University of Basel, • University of Bern, • University of Geneva, • University of Fribourg, • University of Lausanne, • University of Neuchâtel, • Università della Svizzera italiana, • Zurich University of Applied Sciences. The Coordination Office77 is hosted by the University of Zurich and is in charge of the operation of CLARIN-CH. It is led by Cristina Grisot, national coordinator, and consists of CLARIN-CH staff: technical officer, content officer, and training officer. The Linguistic Research Infrastructure LiRI,78 hosted by the University of Zurich, is a technology platform that acts as a technical data and service providing centre. It offers a comprehensive suite of services that covers every stage of a research project’s life cycle, from the initial planning and experiment design, to the acquisition and processing of data, as well as language technology services, and statistical consulting. 77 https://clarin-ch.ch/governance-and-coordination 78 https://www.liri.uzh.ch/en.html switzerlAnd
72 73 switzerlAnd The Digital Discourse Lab,79 hosted by the Zurich University of Applied Sciences, runs the Swiss-AL multilingual corpus platform and provides solutions for the linguistic analysis of the digitised public sphere and allows partners from the realms of research, business, politics and society to access discourse research in a targeted manner. The results include data-based situation reports, innovative analysis practices and the moderation of strategy dialogues. The Digital Discourse Lab has been certified as a CLARIN K-centre in April 2025. The Language Repository of Switzerland LaRS80 is a national platform for the publication of linguistic research data, and it uses the repository system of the national repository SWISSUbase. LaRS provides Swiss research institutions with a reliable data infrastructure, and facilitates access to research data and projects across the linguistic domain. LaRS has received the CoreTrustSeal Certification in March 2025. LiRI and LaRS are in the process of certification as a distributed CLARIN B-centre. Building the Community Since the creation of CLARIN-CH in 2020, we have worked to build a strong partnership between the four entities mentioned above. The current configuration of CLARIN-CH enables the participation of Switzerland in CLARIN and the existence of a national ecosystem of research infrastructure and network to better support Swiss scholars in their research and in managing their language data in the spirit of Open Science and FAIR principles. One important initiative in which CLARIN-CH actively took part is the ‘RIs for the SSH in Switzerland’ initiative. In spring 2022, a coordination group consisting of the directors of national research infrastructures (RIs) in the SSH, the national coordinators of international SSH RIs with Swiss participation (CESSDA, CLARIN, DARIAH, ESS, SHARE, GGP), and the representatives of the Swiss Academy for the SSH was created. This working group started a reflection to gather the SSH community and to defend its interests with respect to RIs. This is being done through a series of events and a position paper. The work resulted in the foundation of the SSHOC-CH81 cluster in April 2024, which is accompanied by a white paper that describes its mission and objectives. 79 https://www.zhaw.ch/en/linguistics/business-services/digital-discourse-lab 80 https://www.lars.uzh.ch/en.html 81 https://sshoc.ch User-Driven Focus Shortly after CLARIN-CH started, we realised that CLARIN-CH’s offerings should be user-driven and adapted to the needs of the Swiss researchers working with language data. For this, we carried out a survey to get to know the target user community and learn about their needs. 92 researchers from all Swiss higher education institutions participated in the survey. The answers to one of the questions regarding their needs are shown in the word cloud below. Needs formulated by the CLARIN-CH scientific community (early 2023). To answer these needs, CLARIN-CH offers several solutions based on identified users’ needs: • National working groups addressing topics such as the management of sensitive and personal data, ethical and legal issues for linguistic data,82 and increasing the FAIRness of (Swiss) learner corpora and second language acquisition.83 82 https://clarin-ch.ch/working-groups/sensitive-personal-data 83 https://clarin-ch.ch/working-groups/swiss-learner-corpora
74 75 switzerlAnd • A Documentation Platform, hosted on the CLARIN-CH website, which provides useful information relevant to the different steps of the data life cycle, as well as information about copyright, licences, data protection, data access and security, metadata and data standards, and data stewardship services in Switzerland. The platform offers best practices and resources to support researchers to engage in FAIR-compliant data management, and also offers relevant webinars, for example, on the topic of legal aspects for language data in the Swiss context, and an FAQ section. • Options to increase the degree of FAIR-ness of language resources,84 by type: for corpora, for tools, and for lexical resources. For instance, corpora may be: • Published and archived with LaRS@SWISSUbase, • Included in the LiRI Corpus Platform (LCP), • Added to the SSH Open Marketplace, • Added to the CLARIN Resource Families. Creating and Preserving Knowledge With the foundation of the CLARIN-CH consortium in 2020, Swiss higher education institutions started to work together to build a FAIR-compliant, sustainable, and expandable CLARIN-CH ecosystem of federated infrastructure to answer the needs of researchers and professionals using language data in Switzerland and beyond. This ecosystem will be interoperable at the national and European levels. The concrete work towards achieving this goal started in 2023 with a series of projects funded through the Swiss national programme for Open Research Data (ORD) and will continue in the years to come. In the long term, these projects will significantly contribute to developing a strong foundation for a sustainable ORD strategy for linguistic data in Switzerland. 84 https://clarin-ch.ch/documentation-platform/data-sharing Among the infrastructure components addressed in these upgrading ORD projects are: • The LiRI Corpus Platform LCP,85 which is a software system for handling and querying multimodal corpora. Expected during summer 2024, the LCP has three interfaces: CatchPhrase (for text), SoundScript (for audio), and VideoScope (for video). • Swiss-AL86 consists of (1) a family of multilingual Swiss public communication corpora, (2) a corpus and computational linguistics pipeline for compiling and processing these corpora, and (3) a browser-based workbench for analysing these corpora. The workbench allows for quantitative linguistics analysis, including recent machine-learning-based methods. It contains Swiss journalistic media, media releases, news reports and blog posts from players in the worlds of politics and administration, industry, academia, and civil society. To analyse discourses, a flexible processing pipeline makes it possible to model tailor-made sub-corpora. • The Swissdox@LiRI87 is the largest database of Swiss media texts. It allows the distribution and the use for academic purposes of journalistic content despite its regular restrictive copyright conditions. The database contains a daily growing collection of around 24 million media articles from more than 250 media titles. The Swissdox@LiRI can currently be accessed via a query interface. Due to copyright restrictions, data can only be downloaded with a subscription fee, and only derivative outputs can be shared. • LaRS@SWISSUbase88 will be harvested by the CLARIN VLO starting in autumn 2024, and it can be accessed via API. LaRS@SWISSUbase also provides a series of guides about preparing and submitting data, understanding metadata and principles of data curation for linguistics. Finally, a frontend for multimodal interaction data called VIAN-COSSIO is under construction, and it will allow the visualisation of annotation layers, transcriptions in various formats and data analysis. 85 https://www.liri.uzh.ch/en/services/liRI-Corpus-Platform-lCP.html 86 https://www.zhaw.ch/en/linguistics/research/swiss-al 87 https://www.liri.uzh.ch/en/services/swissdox.html 88 https://info.swissubase.ch/resources/?sr=929
76 77 switzerlAnd Tool | LCP written by Seraina Nadig lCP’s main page. The LiRI Corpus Platform (LCP)89 was developed by the Linguistic Research Infrastructure (LiRI), which serves as the technical centre of CLARIN-CH, to facilitate the handling and querying of diverse corpora, including text, audio, and audiovisual data. LCP is designed to support linguistic research by providing tools for corpus creation, annotation, management, and complex querying. It addresses the need for accessible, richly annotated corpus data through a user-friendly interface that allows for sophisticated linguistic queries. lCP’s three interfaces: Catchphrase for text corpora, Soundscript for audio and text corpora, and videoscope for text, audio and video corpora. 89 https://lcp.linguistik.uzh.ch LCP offers three specialised web applications, each optimised for different data modalities: Catchphrase:90 Tailored for text corpora, enabling analysis of monoor multilingual texts of any size. Soundscript:91 Designed for audio corpora, facilitating the analysis of speech recordings along with their transcriptions and annotations. Videoscope:92 Focused on audiovisual data, allowing users to view and query video corpora with associated annotations. The platform’s backend comprises an asynchronous Python web server, a PostgreSQL database, and a Redis instance for caching. This infrastructure supports efficient querying, even for complex searches involving nested logical operations and quantifiers. LCP employs a dedicated query language known as DQD,93 which allows users to perform intricate searches across different corpora. The platform provides example queries and guidance to assist users in crafting their searches. Users can access the LCP through their web browsers and are required to log in using SWITCH edu-ID or institutional credentials for full functionality. The platform supports the import of usergenerated corpora via a command-line interface, facilitating personalised research projects. LCP hosts several publicly accessible corpora, such as the British National Corpus (BNC), the ArchiMob corpus, representing German linguistic varieties spoken within Switzerland, or OFROM, an oral corpus with various recordings from the French-speaking part of Switzerland. LCP is actively developed and maintained by the LiRI team, with ongoing efforts to enhance its features based on user feedback. The platform is currently in beta, and the team encourages users to participate in the hands-on training workshops offered by LiRI to learn how to use the platform and to be able to contribute to its evolution. 90 https://catchphrase.linguistik.uzh.ch 91 https://soundscript.linguistik.uzh.ch 92 https://videoscope.linguistik.uzh.ch 93 https://lcp.linguistik.uzh.ch/manual/dqd.html
78 79 switzerlAnd Service | CLARIN-CH FAIRification Pipeline written by Alexandru Craevschi At CLARIN-CH, the Swiss national node of the European CLARIN research infrastructure, we are actively engaged in promoting the FAIR principles, making linguistic and languagebased research data Findable, Accessible, Interoperable, and Reusable. Our FAIR-ification pipeline is a hands-on, researcher-centric approach that assists scholars in enhancing the long-term usability and visibility of their language data. The ClARIN-CH FAIRification pipeline: a researcher-centered process supporting outreach, consultation, data transformation, and archiving to ensure FAIR-compliant language resources. Our engagement begins with personalised outreach. Researchers are contacted directly and invited to consider four options for making their data FAIR-compliant. These options are not mutually exclusive and include: • Publishing and archiving datasets on SWISSUbase,94 a national, FAIR-compliant repository harvested by CLARIN’s Virtual Language Observatory (VLO), ideal for finalised datasets requiring long-term preservation in compliance with SNSF requirements and FAIR principles. 94 https://www.swissubase.ch/en • Contributing to the Linguistic Corpus Platform (LCP), a corpus analysis environment hosted by the Linguistic Research Infrastructure (LiRI), particularly suitable for corpora that would benefit from web-based querying and corpus-linguistic analysis. This makes a resource immediately available for queries of various complexity and is especially useful when the raw data cannot be shared for legal or other reasons. • Showcasing resources via the SSH Open Marketplace, recommended for tools, services, and workflows that support reuse and method visibility across the social sciences and humanities. • Featuring datasets on the CLARIN Resource Families webpage, best suited for mature resources that fit within established categories like treebanks, speech corpora, or multilingual lexicons, and where increased international visibility is desired. Once a researcher expresses interest, we schedule a consultation to review their data and determine the best dissemination pathway. We evaluate the state of the dataset, including format, completeness of metadata, and potential need for preprocessing. For example, preparing a dataset for the LCP may require transforming files into a specific format of a set of relational CSV files, a task that can be challenging for researchers unfamiliar with tools such as Python or R. Where needed, CLARIN-CH provides direct assistance in formatting, metadata enrichment, and the technical steps required for depositing. In some cases, especially for concluded projects where researchers lack the time or capacity, we manage the entire FAIR-ification process on their behalf to avoid the loss of a resource or tool. Beyond one-on-one collaboration, CLARIN-CH also supports the broader research community through thematic working groups, training activities, and extensive documentation. We facilitate interdisciplinary working groups on sensitive data management, learner corpora, and legalethical challenges, acting as intermediaries to support collaboration and project development. Our training activities cover topics from corpus querying and statistical methods to multimodal data processing. In addition, we maintain a comprehensive documentation platform that provides best practices across the entire data lifecycle, helping researchers see the full picture of sustainable and reusable data management.
92 93 The CLARIN-DK infrastructure also offers online tools, such as: • Annotation tools, namely tokeniser, PoS tagger, lemmatiser, named entity tagger, parsers, TEI annotation, • Corpus search and visualisation, • Workflow manager, Text Tensorium:107 for annotating texts with many types of linguistic information. The workflow offers tools for texts in several languages and in different formats. Text Tonsorium is also included in the Text Normalisers CLARIN Resource Family. CLARIN-DK Text Tonsorium workflow. CST also supports researchers who want to produce, share or use FAIR resources distributed in the CLARIN-DK repository by providing guidance, as well as curation of resources and/or participating as members in research projects. 107 https://clarin.dk/clarindk/tools-texton.jsp Interview | Sidsel Boldsen The conversation was led by Karina Berger At the time of the interview, in 2022, Sidsel Boldsen was a PhD Student in Natural Language Processing (NLP) and digital humanities, with a special interest in historical languages and linguistic knowledge representation. She was part of the interdisciplinary research project ‘Script and Text in Time and Place’108 at the University of Copenhagen. Please describe your academic background. < My background is in historical linguistics and comparative linguistics, which I did for my BA. But then I moved on to language technology and did a Master’s in IT and Cognition at the University of Copenhagen. In my PhD, I focused on language technology and I had a special interest in language change, which comes from my background in comparative linguistics. I became interested in language technology because I thought that programming and scripting could offer interesting research avenues for linguistic studies involving digital corpora. > You were part of the research project ‘Script and Text in Time and Place’ at the Department of Nordic Research at the University of Copenhagen. Could you describe the project? < The project was very interdisciplinary and is a qualitative and quantitative study of about 300 medieval Danish charters from the thirteenth to the sixteenth centuries. The goal was to study the script and language of medieval Denmark through these resources. These charters are very interesting from both a historical and a linguistic point of view because they have not been edited in any way, so they are direct sources of language and history. They are also dated and geographically localised based on where they were produced, allowing for a very nuanced picture. 108 https://nors.ku.dk/english/research/projects/script-and-text-in-space-and-time dAnsK
94 95 Often when you work with historical texts, it is an edition of an edition, and it can be difficult to say what the ‘real’ language actually is, and what the later additions are. So philologists were working on the project, as well as historians looking into the monastic history. The project ended in May 2022. The main output was an open-source, digital scholarly edition of the charters. This scholarly edition enables scholars within philology and history to search these charters and to see the texts in different layers, where we have annotated the different features of the script, and other levels, too. For instance, we lemmatised the texts so that one can search for word forms, and we also annotated the people and places that occur in the text. > What was your role in the project? < One focus of the project was to develop tools for automatic linguistic analysis of texts, as well as automatic dating, localising, and identifying scribal schools. Such customised digital tools should improve the quantitative analysis of historical sources. I was involved in that part. There were not many tools available at that time to date or localise Danish script, and little systematic analysis has been conducted on Danish medieval texts. We were working towards new methods for automated dating, localisation and grouping of texts based on machine learning techniques (MLT), which improved our understanding of the relevant factors for establishing the date, place, and scribe of primary sources from the Middle Ages. The benefit of using MLT in the project was that it allowed us to take advantage of the dated charter material for building a reference and training corpus. What I did was to look at how language changed and whether it was possible to develop tools to date texts without a date. These corpora are all dated, so one can use these dated corpora to develop tools to automatically date texts that do not have a date. The tool development was the starting point for my thesis. But then my work became more theoretical, and I began to study how language change is captured in language models, and what kind of features the models recognise or are sensitive to. I have tried to look at different layers. Of course, there is topical change, with different places being named or different expressions being mentioned differently through time, a sort of topic model. But then I also looked at sound change. In that period, we know that some sound changes were supposed to have happened, but when we look in the corpus, can we actually identify those? So I was also interested in change on a phonological level. > Can you say a little more about the tool you developed? < The tool109 uses support vector machines (SVMs) to automatically assign the manuscripts to a specific time period, or bin. It represents a text in a vector space, either the words it contains or smaller segments such as character n-grams, and then it projects these into space and tries to create learning boundaries between different classes. In my case, the classes are the specific time periods: it could be centuries, or it could be spans of 50 years. And then the tool tries to learn how to divide those that are projected within that given space. When you receive a new document, you map that new document into that space, and then you can evaluate how well the space was constructed with respect to how well documents can be divided into those periods. We received pretty good accuracy. We reached around 75%, which means we were able to date almost 75% of the charters with a 25-year error margin, which is used by philologists as a standard of the precision with which medieval texts can be dated manually. But it is a bit complex also as to how these so-called bins are constructed. And we did not find a way to address that, as explained in our paper110 on the topic. If I were to develop the tool further, I think the work should focus on what type of features it actually recognises. For it to be useful, it would have to work for another corpus, which had been annotated using different schemas, for example. So, the big question would be: how well is the tool able to generalise across corpora? We have another corpus of charters or medieval documents called ‘Diplomatarium Danicum’,111 and it would be interesting to test the tool on this resource because it is a much bigger corpus spanning a broad period. It would be interesting to see how well the tool that has been trained on a very specific corpus would transfer to a bigger one. In principle, I think the tool could be useful for other corpora as well, at least within the same domain. > 109 https://github.com/syssel/twec 110 https://ceur-ws.org/Vol-2364/5_paper.pdf 111 https://diplomatarium.dk/english dAnsK
96 97 You were applying machine learning techniques to the analysis of medieval Danish texts with the cooperation of CLARIN-DK. How did you start collaborating with them? Have you used any specific CLARIN tools as part of your research? < One of our project members, Bart Jongejan, is in the CLARIN-DK team. So when we needed tools, we used one that was in the CLARIN-DK repository. Although the charters had already been transcribed, they were annotated in a CSV-like format, in which each row represents one token, and they needed to be converted to XML format for the actual edition. We used the workflow manager for NLP called Text Tonsorium to automatically convert the format. We used the same tool for automatic part-of-speech (PoS) tagging of the Latin charters. For the Danish charters, we wanted to annotate PoS and lemma manually and, in this case, we used the tools offered in Text Tonsorium as a starting point for the annotation. This was very useful as it dramatically reduced the workload of the manual annotation. One of the great features of the Text Tonsorium is that it offers many different pipelines and workflows, so you can quickly test out different parsers or lemmatisers. Otherwise, you would have to set up all the different tools and try them out, but this is one common format where you can try them all out in one go. For my research area, CLARIN-DK provides everything I need. They have trained both an old Danish lemmatiser and a Latin one, and also offer PoS-tagging. > Why is it important to take a computational approach, such as natural language processing, in the humanities? < I think the contribution from computational methods is two-fold: the most important, in my view, is to make corpora more accessible and searchable, so that qualitative researchers can use these resources in a more focused way. I think it is more about how language technology can assist certain research questions that qualitative researchers work with. Therefore, a way to assist, not to replace different methods. You can work with larger resources and filter them, for example. And the other reason is to actually use those tools in order to carry out humanities research, which can also be very interesting. For example, if you develop a tool to date text manually, you could maybe also learn from those tools: what are the predictive features of language change? In that way you use those tools not only to be able to annotate, but also to learn from those models and use them to study language and text. But I think the first contribution is more important. > dAnsK
99 K-centre: Phonogrammarchiv Introduction written by Kerstin Klenke In April 2024, the Phonogrammarchiv112 of the Austrian Academy of Sciences (OeAW) celebrated its 125th anniversary. Official celebrations were postponed, however, as the archive is already busy preparing for a long-awaited move to new premises after almost a century in its current location. The move to the former home of the OeAW’s Stefan Meyer Institute for Subatomic Physics in Vienna’s 3rd district will come with greatly improved facilities: not only will the Phonogrammarchiv have enough room to keep all its collections in its own storage space, it will also have more studios for the digitisation of historical recordings. Moreover, a chemical lab and restoration workshop will provide a better basis for projects researching the materiality of sound and video carriers. In general, at its new premises, the Phonogrammarchiv will have a more user-friendly layout, more room for hosting guest researchers, and more space for offering training. Since our last contribution to Tour de CLARIN in 2020, several new funding projects have enabled the Phonogrammarchiv to greatly improve its technical infrastructure by adding state-of-the-art cylinder and wire transfer machines to its equipment park. Among these projects is an infrastructure and capacity-building project with the Academy of Sciences of the Republic of Uzbekistan and two German partners, funded by the Volkswagen Foundation: ‘The Fonoteka of the Uzbek Academy of Sciences’ Institute for Art Studies – a Trilateral Infrastructure and Capacity Building Project (D-UZ-A)’, which involves training and supporting the Uzbek partners in digitising and cataloguing a unique archive of historical sound carriers with research recordings from the 1930s to 1990s in Tashkent. 112 https://www.oeaw.ac.at/en/phonogrammarchiv A technician working on a reel-to-reel tape recorder used in audio preservation and archival work. Similarly devoted to infrastructure and training is the project ‘Applied / Experimental Sound Research Laboratory (ÆSR Lab)’,113 a three-year cooperation between the Phonogrammarchiv, the University of Applied Arts Vienna and the University of Music and Performing Arts Vienna. Funded by the Austrian Ministry for Education, Science, and Research, the project cooperates in the establishment of easily accessible technical infrastructures for various forms of research into sound. According to its special expertise, the Phonogrammarchiv contributes to the tripartite lab structure of this project with a ‘Field Recording & Digitisation/Restoration Lab’. In the sphere of heritage science, the Phonogrammarchiv is one of three partners in a cooperation project that explores the materiality as well as the cultural meaning of analogue sound carriers used for audio letters in the 20th century: ‘Sonic Memories: Audio Letters in Times of Migration and Mobility (SONIME)’.114 Among other equipment, the project has allowed the Phonogrammarchiv to acquire a Fourier-transform infrared spectroscopy (FTIR) device, which greatly improves the analysis facilities the archive can offer in the realm of material science. Besides these more technically oriented projects, the Phonogrammarchiv continues to be involved in various projects and initiatives that aim at the critical contextualisation and (re-) circulation of its collections. Here, the Phonogrammarchiv has supported efforts by communities to offer local or online access to audio resources linked to their history, such as the Virtualni arhiv 113 https://aesr-lab.uni-ak.ac.at 114 https://sonime.at 98
101 Stinjaki (Virtual Archive of Stinatz)115 or the Burgenlandi Magyar Kultúrlevéltár (Burgenlandian-Hungarian Cultural Archive).116 This is also the sphere of activity which the interview following this text focuses on: Fayrouz Kaddal speaks about her project on Anna Hohenwart-Gerlachstein’s recordings from historical Nubia. Some of the Phonogrammarchiv’s own recent initiatives in the sphere of critical contextualisation and (re-)circulation have been a project on the archival traces of ‘Rudolf Pöch’s Papua New Guinea recordings (1904–1906)’ and the colonial context of their production, as well as the decision to go online with the edition series ‘The Complete Historical Collections 1899–1950’ after 25 years on CD format. The first edition in new form will have a focus on linguistics with ‘Adolf Dirr’s Recordings from the Caucasus (1909–1910)’ scheduled to be published in 2025 in English and Georgian. Another project with a focus on language and linguistics, which has been one of the Phonogrammarchiv’s main collection fields since its founding in 1899, is a cooperation with the Austrian Centre for Digital Humanities and Cultural Heritage (ACDH-CH)117 at the OeAW, which engages with Austrian dialect recordings and is devoted to improving its metadata and accessibility. In the field of academic events, in 2024, the Phonogrammarchiv launched its new series of ‘Advanced Seminars’, which will regularly bring together a group of experts on topics relevant to the various fields of its archival practice. The inaugural ‘Advanced Seminar’ in 2023 was devoted to ‘(Re)Contextualising and Recirculating Historical Sound Recordings: Experiences and Approaches’, with Fayrouz Kaddal, interviewee in the following section, among the invited participants. The topic of its 2025 edition will be ‘Dealing with Affective Resonances and Dissonances when Engaging with Archival Sound’. In addition, 2025 will see the Phonogrammarchiv’s move to new premises, followed by a belated international anniversary symposium, 125+2, in 2026. 115 https://arhivstinjaki.at 116 https://bukv.at 117 https://www.oeaw.ac.at/acdh/acdh-home Interview | Fayrouz Kaddal The conversation was led by Jakob Lenardič Fayrouz Kaddal is a PhD researcher in Cultural Anthropology at Duke University, specialising in Nubian music, displacement, and sound archives. Specifically, her auto-ethnographic work explores Nubian musical practices shaped by migration. She works with the Phonogrammarchiv’s archival recordings to reconnect Nubian communities with their musical heritage. Please introduce yourself – your background, academic and otherwise? < My name is Fayrouz Kaddal. I am an Egyptian Nubian born and raised in Alexandria. I am also a flautist. My musical training was in Western classical music, but my musical inclination is toward Nubian pentatonic melodies and polyrhythms. I was the flautist of a Nubian band called High Dam Band for several years, which paved the way for my interest in anthropology and ethnomusicology. My musical and academic interests emerged from my music practice and my family’s Nubian heritage. The scope of my academic work focuses on themes of displacement, migration, Nubia, music, sound, salvage anthropology, and the circulation and repatriation of sound archives. > Can you present your auto-ethnographic work and what motivated you to pursue it? < An auto-ethnographic scholarship is a qualitative research method that focuses on the positionality of the researcher to shape the ethnography produced. This means the personal story of the ethnographer, their relationship with the wider community and the research question of inquiry all shape the knowledge produced. My MA dissertation, ‘On Displacement and Music: Embodiments of Contemporary Nubian Musical Practices in the Nubian Resettlements’ (2021), is an example of auto-ethnographic work. In this, PHOnOgrAmmArcHiv 100 Photo credit: daniel merrill
103 I explore the relationship between migration, displacement and contemporary musical practices in Nubia, Egypt, based on my personal account as a Nubian musician. The ethnographic fieldwork for this project began from the house of my great uncle in the Nubian village Toushka ‘El Tahgeer’ (Toushka the displaced), to which he and his family, alongside 50,000 Egyptian Nubians, were displaced to in 1964 due to the construction of the Aswan High Dam. The construction of the dam was the signature project of the former Egyptian president Gamal Abdel Nasser, and was intended for the regulation of the Nile River flood, the expansion of agricultural lands, and the introduction of hydroelectric power. Therefore, On Displacement and Music: Embodiments of Contemporary Nubian Musical Practices in the Nubian Resettlements (2021) is an ethnography of the relationship between music, displacement, and Nubia based on my positionality. As a Nubian musician, I was able to go back to the village of my ancestors, listen to, and join musicians in their practices. I have attended Nubian weddings, listened to the music played, and danced with women during these events. My ethnography was also shaped by the intimate conversations I had with family relatives and other Nubians about the Nubian displacement, old patterns of migration, and music. > An image of Toushka el Tahgeer (Toushka the displaced). Taken by Fayrouz Kaddal during fieldwork in 2019. How did you get involved with the Phonogrammarchiv? < In 2011, a book titled Nubian Encounters: The Story of the Nubian Ethnological Survey 1961–1964, written by Nicholas Hopkins and Soheir Mehanna, was published. I remember how I immediately rushed to buy a copy. I was surprised to learn of the work of Anna Hohenwart-Gerlachstein, an Austrian anthropologist (1909–2008). In the very first chapter, it is mentioned that she had a tape recorder during her fieldwork and that the recordings are kept at the Phonogrammarchiv of the Austrian Academy of Sciences. Anna Hohenwart-Gerlachstein was interested in the Nubian language and salvage anthropology, a practice in anthropological studies that focuses on researching and preserving the culture of endangered communities. She joined UNESCO’s Save Nubia campaign to preserve and document life in Nubia before the latter’s displacement in 1964. As a Nubian musician, I was not aware of the existence of these recordings, or of any recordings of Nubian music made before the displacement in 1964. I can confidently assume that the majority of Nubians are not aware of the presence of Nubia’s endangered sound archives collected in the early 1960s. As a Nubian musician, it was very heartwarming and important to be able to listen to recordings of the ancestral voices and music. I recall that the first recordings I requested access to from the Phonogrammarchiv were made in the village of my ancestors, Toushka. Listening to recordings of Old Nubia meant a possible conversation between what I have learned of Old Nubia from both my family and from actual fragments of Nubia from those who lived in Old Nubia. For my generation, learning Nubian music is generally based on popular Nubian music and oral history. Hohenwart-Gerlachstein’s collection at the Phonogrammarchiv offers a new tool for us to learn the ancestral music and to engage with it. Studies on Nubia have rarely made use of field recordings, especially before 1964. That is why knowing that there might be field recordings resulting from UNESCO’s famous Save Nubia campaign in the 1960s came as an important surprise. In 2018, I contacted the Phonogrammarchiv for the first time, and I was very happy with the prompt reply and attention that Kerstin Klenke (Head of Phonogrammarchiv), Clemens Gütl (Curator and Researcher), and Gebhard Fartacek (Curator and Researcher) gave me, helping me access some of HohenwartGerlachstein’s recordings and learn more about Anna Hohenwart-Gerlachstein. After getting to know the collection, I started exploring ways to circulate Anna Hohenwart-Gerlachstein’s sound collection with Nubian communities. > PHOnOgrAmmArcHiv 102
105 Fayrouz Kaddal, Stephanie Wiesbauer and Clemens Gütl (left to right) at the vienna Phonogrammarchiv. Photo credit: Christian liebl. How has the Phonogrammarchiv supported your fieldwork? < The Phonogrammarchiv have been supporting my work since 2019 in different ways. For my MA thesis, I was allowed access to a selection of the recordings. They have also shared information on the protocols of the recordings, as well as information about Anna Hohenwart-Gerlachstein in terms of her career and publications. In 2022, when I was working on a paper about the repatriation and circulation of sound archives, I wanted to learn more about the training that Anna Hohenwart-Gerlachstein received as an anthropologist and how these recordings were made. For this reason, I decided to visit the Phonogrammarchiv and spend a week in Vienna. During that week, Clemens Gütl and Gebhard Fartacek took me on a guided tour around the Phonogrammarchiv, where I met the different specialists and learned how they maintain very old recordings in excellent conditions. I also learned about the process of digitising recordings and the length of time it can take, something that I was completely unaware of. I was introduced to different writings by and about Anna Hohenwart-Gerlachstein, and I learned more about the history of the Phonogrammarchiv. I was fascinated by the fact that part of Anna Hohenwart-Gerlachstein’s training as an anthropologist included learning how to make field recordings and keep meticulous logs of them. Clemens Gütl helped me get in touch with Stephanie Wiesbauer, Anna Hohenwart-Gerlachstein’s niece. We met at the Phonogrammarchiv, where I interviewed her for my research. Mrs Wiesbauer explained the profound and strong connection that Anna Hohenwart-Gerlachstein had with Nubia and the Nubians, which extended to Mrs. Wiesbauer even after the passing away of Anna Hohenwart-Gerlachstein. I learned more about Anna Hohenwart-Gerlachstein’s role during the Nubia salvage campaign led by UNESCO (1961–1964) before the drowning of Nubia under Lake Nasser, which was created by the construction of the Aswan High Dam. I also learned that Stephanie Wiesbauer had handed over 60 kilos of valuable photographs, field notes and sound recordings to Professor Mohamed Riad, who deposited the materials at CultNat in Egypt, in order to make them available to anyone interested in Anna Hohenwart-Gerlachstein’s work. This information (collections/recordings) is valuable to my research. In addition, Clemens Gütl introduced me to Tobias Mörike, Curator of the Collection on North Africa, West and Central Asia, and Siberia at the Welt Museum in Vienna. Tobias Mörike gave Clemens Gütl, Mr and Mrs Wiesbauer, and me a guided tour in the Welt museum’s holdings of physical objects that Anna Hohenwart-Gerlachstein had collected from Old Nubia. The collection includes Nubian pottery, instruments, and jewellery. During this same visit in 2022, I had the chance to discuss with Kerstin Klenke, the head of the Phonogrammarchiv, the different potentialities of working closely with the collection. Kerstin Klenke encouraged me to think of continuing my studies focusing on Anna HohenwartGerlachstein’s sound archives. She explained to me the different potential support that the Phonogrammarchiv can offer in the future. I learned of the possibility of a PhD placement in Vienna to work closely on the archives and of a seed grant. After this trip, and seeing the level of interest and the collaboration the Phonogrammachiv has shown me, I realised that there might be a strong potential to work on circulating Anna Hohenwart-Gerlachstein’s sound archives. Thanks to Dr Klenke’s support for my PhD applications, as well as the support of my MA thesis committee Dr Hanan Sabea, Dr Reem Saad, Dr Tom Western and Dr Manuel Schwab, I am currently enrolled as a doctoral student at Duke University’s Cultural Anthropology department in the US. My current project examines the making and the content of Anna Hohenwart-Gerlachstein’s sound archives, as well as the circulation of the archives amongst Nubians in Egypt. The latest of the Phonogrammarchiv’s support was in 2023, when I took part in a one-and-a-halfday workshop organised by the Phonogrammarchiv as part of an advanced seminar titled ‘ (Re)Contextualising and Recirculating Historical Sound Recordings: Experiences and Approaches’. The advanced seminar was attended by the most experienced experts on the circulation of sound archives. The workshops allowed me to learn more about the most recent debates and work on PHOnOgrAmmArcHiv 104
the topic. It was also an opportunity for me to think about and share my PhD proposal with great scholars working on the circulation of sound archives. > Could you present the collection you have worked with that is hosted by Phonogrammarchiv? < The collection of recordings available at the Phonogrammarchiv is the outcome of Hohenwart-Gerlachstein’s trips and field work in Nubia between 1962 and 1964. She was invited by UNESCO to save Nubia’s heritage before it was submerged under Lake Nasser. As an anthropologist, her work was part of the international efforts to save Nubia, as well as part of an older history of salvage anthropology. Anna HohenwartGerlachstein believed in her duty as an anthropologist to preserve elements of endangered cultures that are under the threat of vanishing. Therefore, the content of Nubia’s sound archive is a reflection of her role as a salvage anthropologist. Anna Hohenwart-Gerlachstein spent most of her time in Al Derr village in Nubia, Egypt. Yet, the recordings were collected from various Nubian villages, such as Toushka, Adendan, and El Malki. The archive includes interviews with people in Nubia, speeches by Nubians welcoming Anna Hohenwart-Gerlachstein to Nubia, language lessons, recordings of wedding ceremonies and musicians playing music. She was interested in documenting and learning the Nubian language. The recordings of Kenzi language, Nubian songs, and wedding rituals follow a longer Western academic practice of preserving endangered languages around the world. > What is the benefit of archiving ethnographic materials in the Phonogrammarchiv? < The Phonogrammarchiv did an amazing job by working on preserving and digitising the material they had. Digitisation of sound on historical carriers is a process that takes time and requires highly knowledgeable technical skills to ensure a high-quality outcome. They also have special expertise and storage conditions for the long-term preservation of analogue carriers. The Phonogrammarchiv, as an interdisciplinary institution with 125 years of experience, provides all of that, in addition to the expertise of the curators working on the recordings’ content and contexts. As for Anna Hohenwart-Gerlachstein’s collection of Nubia’s heritage in particular, I believe that, thanks to the efforts of the Phonogrammarchiv, young Nubians will have a base to learn the music of their ancestors. Circulating Anna Hohenwart-Gerlachstein’s sound archive will enable the oral history of Nubia to be in conversation with the archives. Hegemonic understanding of Nubia could be challenged. On a personal level, Anna Hohenwart-Gerlachstein’s collection encouraged me to think of how my personal archive, collected from fieldwork and research, can be of benefit to others, although my collection is more than 50 years younger than Anna HohenwartGerlachstein’s and my research questions are very different from hers. However, I started to think of the responsibility we have as anthropologists in depositing our field recordings in archival institutions and the wider interest in the data we collect during our fieldwork. That said, we must be very considerate and careful about the ethical implications of granting access to archives. These processes require the expertise of archival institutions, which are actively working on developing various licensing protocols. > How can institutions like the Phonogrammarchiv further support you in the future? < Archival institutions have a lot to offer to scholars and the public, as they can facilitate access to the material they hold, always with careful attention to necessary ethical considerations. Archived voices could mean a lot more to the communities of origin than we think. If we take the case of Nubia as an example, Anna Hohenwart-Gerlachstein’s collection can be heard as the sounds or voices of Nubians. However, Nubians would hear the recordings as the voices of their relatives. That is why archival institutions should be able to reach out to communities of origin, and/or those that might have an interest in the holdings of the archives. I believe that the Phonogrammarchiv has been working on bridging the gap between the institution and communities of research by responding to the different calls for repatriation and circulation of sound archives. Regarding my work, they continue to help me think of the methods and implications of circulating Anna Hohenwart-Gerlachstein’s sound archive. > References: Kaddal, F. (2021).On Displacement and Music: Embodiments of Contemporary Nubian Music in the Nubian Resettlements [Master’s Thesis, the American University in Cairo]. AUC Knowledge Fountain. https:// fount.aucegypt.edu/etds/1591 Mehanna, S. & Hopkins, N. S. (2011). Nubian Encounters: The Story of the Nubian Ethnological Survey 1961– 1964. American University in Cairo Press. PHOnOgrAmmArcHiv 106 107
109 K-centre: ClARIN-SMS Introduction written by Arne Jönsson The CLARIN Knowledge Centre for Swedish in a Multilingual Setting (CLARIN-SMS) is primarily directed at researchers in the social sciences and humanities and beyond with a need for analysis, annotation or data mining of Swedish or multilingual texts, and Swedish Sign Language. CLARIN-SMS makes resources, in the form of tools for linguistic processing and corpora, available for research in the humanities and social sciences. The resources include monolingual (mainly Swedish) and multilingual corpora across several domains, and tools for the basic processing of text, including tokenisation, morphological analysis, part-of-speech tagging, syntactic parsing, and named entity recognition. Main Areas of Expertise CLARIN-SMS offers special expertise: • For researchers interested in exploring Swedish texts, by providing support for the creation and processing of Swedish texts with a variety of computational methods, such as linguistic annotation at different levels, or sentiment analysis. • For researchers interested in comparative analyses, by providing support for the creation and processing of parallel and comparable corpora, including alignment and machine translation, as well as cross-linguistically consistent annotation within the framework of Universal Dependencies, which allows for easy comparative analyses. • For researchers interested in education and content accessibility, by providing support for the computation and evaluation of measures of text complexity. • For researchers and users of Swedish Sign Language (SSL), by providing support for the creation of lexicons and corpora for SSL, and annotation of SSL (including glosses, part-ofspeech tagging and syntactic structure). The support is provided by several partners participating in the CLARIN-SMS distributed Knowledge Centre: • Linköping University, Department of Computer and Information Science, • Stockholm University, Department of Linguistics, • Uppsala University, Department of Linguistics and Philology. Current Challenges Each CLARIN-SMS node works as a separate unit, and promotes its services and resources in various ways, including outreach tours at universities, web pages presenting projects and resources, and presentations at CLARIN-related events. One challenge that we face is promoting the K-centre as a common resource. The K-centre has its own web page, but we do not believe that anyone has reached out to us via the website. For this reason, we work towards a targeted promotion of the K-centre, although the focus remains on promoting the resources that we offer through our partners. This is far from saying that CLARIN-SMS is not a vibrant community. Following CLARIN’s general mission of creating and promoting language resources, a variety of activities have been carried out at the respective nodes, including tool and resource development for language analysis, both multilingual and Swedish only. An Active Research Hub Several activities are focused especially on promoting the use of language technology in the social sciences and humanities. For instance, one of the projects includes analysing the development of the concept of handicapped from a Swedish parliamentary perspective. In this project, we help researchers process and analyse the Swedish Government’s official reports from the early 1900s to the present day with a variety of SweClarin resources and language technology tools, such as the SPARV pipeline. 108
111 Another example is the analysis of the protocols of the Swedish National Bank (Sw. Riksbanken), where we compare protocols from the period when they were anonymous to protocols from the period when they were not. One of the goals of this study is to see if we can identify individual speakers from the period of anonymous protocols. Another goal is to provide the National Bank with information about potential differences and similarities in argumentation between the two types of protocols. To this end, we use a variety of SweClarin resources, such as the sensaldo-v02 sentiment lexicon or the SPARV pipeline for parsing, in combination with, for instance, topic and sentiment analysis models. Another example is a project that is led in cooperation with management researchers, in which we are analysing Swedish companies’ adherence and adoption of the information security standard ISO 27001. The project aims to examine the communicative constitution of preventive innovation in organisations. For this project, we helped create a corpus and analyse it from multiple interdisciplinary perspectives using SweClarin tools and resources, such as the sensaldo-v02 sentiment lexicon, or the SPARV pipeline for parsing, as well as other language technology tools, including word clouds. example of two word clouds from the ISo adoption analysis type called Stewards. The left word cloud is from companies’ web pages with indirect economic benefits resulting from preventive communication adoption, and the second is from direct economic benefits. Some Flagship CLARIN-SMS Tools and Resources Tools and models: • SWEGRAM118 provides a tool for text analysis in Swedish and English. The user can upload one or several texts and annotate them at different linguistic levels with morphological and syntactic information. The annotated texts can then be used to extract statistics about the text properties concerning text length, number of words, readability measures, part-ofspeech, and much more. The SWEGRAM annotation workflow. Created by: Beáta Megyesi. • Sapis - StilLett API Service119 is a web service (REST API) including tools for measuring text complexity and text simplification. • Gold standard alignments120 for 1164 English-Swedish sentence pairs are used in word alignment software tasks. Source data is from Europarl v.2. • Universal Dependencies represent a framework for consistent annotation of grammar (parts of speech, morphological features, and syntactic dependencies) across different human languages. UD is an open community effort with more than 300 contributors producing nearly 200 treebanks in over 100 languages. CLARIN-SMS provides the application of both Swedish-specific as well as UD-based annotations.121 The resource is useful for studies of NLP applications, such as multilingual parsing and language typology. Moreover, parallel UD treebanks can be used for studies of human and machine translation. 118 https://www.su.se/english/research/research-projects/swegram-a-tool-for-text-analysis-forswedish-and-english 119 https://www.ida.liu.se/projects/stillett/Publications/SAPIS_user_manual.pdf 120 https://www.ida.liu.se/divisions/hcs/nlplab/resources/ges 121 https://www.ida.liu.se/~larah03/transmap/Corpus clArin-sms 110
125 124 After completing my MA studies at the University of Latvia (UL), I applied for the PhD programme there, because my adviser Zigrīda Vinčela encouraged me to continue working on the topic of my thesis. It was also my dream to teach at UL one day. I expanded the boundaries of my MA thesis and started exploring not only sentiment, but also subjectivity in language. I was not satisfied with the tools and resources I had used before, especially the sentiment lexicons, as they only list words in isolation and do not take context into account. The tagging of the lexicons is also limited to the three-way distinction of positive, negative, and neutral. I felt that this distinction was too general for the exploration of subjectivity in language. I also realised that I needed to learn the theoretical background so that I could identify which linguistic features I am interested in, and which tools to use to extract them. Therefore, I turned to theoretical sources and started looking for definitions of subjectivity across different areas of expertise, and in linguistics in particular. This is how I came across Douglas Biber’s work on the dimensions of the English language. There are two dimensions related to subjectivity that I became interested in: overt expression of persuasion on the one hand, and informational versus involved on the other. Informational expression refers to written documents that contain factual details on a particular topic, while involved expression refers to interactive, spontaneous, spoken texts (in the majority of cases). > How did you get to know CLARIN? Are you involved with the Latvian CLARIN consortium in any way? < During the first semester of my PhD studies, I signed up for a MOOC course called Corpus Linguistics: Method, Analysis, Interpretation, which is offered every year by Lancaster University. There, I was introduced to two CLARIN-UK tools, the #LancsBox software for corpus analysis and the CQPweb concordancer. In addition to these two tools, I now also use the semantic tagger WMatrix (also a CLARIN-UK tool), which my adviser, prof. Zigrīda Vinčela had introduced me to. I later on mastered these tools, along with CLAWS and Voyant, at a PhD course taught at UL by my adviser, together with Ilze Auziņa, a member of CLARIN’s User Involvement Committee. It was Ilze who introduced me and my coursemates to the wider CLARIN infrastructure. As one of the few students regularly using LancsBox, Wmatrix, and Voyant in my research and daily work, I was approached in spring 2023 by Ilze regarding the participation in the PhD section of the CLARIN conference, and in the end, I was selected to participate. I would also like to thank our national coordinator, Inguna Skadiņa, for her support and mentorship in this. From September 2023 to January 2024, I participated in the project Language Technology Initiative, which also included CLARIN Latvia. I was a part of the working group of educators developing study courses involving digital technologies. My adviser, Zigrīda Vinčela, and I were developing study materials for the master’s study programme course Corpus Linguistics. The tasks we developed involve CLARIN tools such as #LancsBox, CLAWS, and Wmatrix. Since September 2024, both of us have been part of the project Latvian Diaspora Identity Transformation: Text, Language, Digital Environment (LATDIT),145 working on the corpus compilation and analysis of Latvian diaspora authors writing in English. Our primary choice for detecting linguistic markers of national identity elements is Wmatrix. > Please present the Latvian and American Political and Sports News corpus. How did you build it – what were the corpus creation tools; what were the sources; was anyone else involved in the preparation of the corpus? < Latvian and American Political and Sports News (LAPIS) is a micro-corpus consisting of 24 texts and around 12,000 tokens. I compiled this corpus on my own simply by creating txt files and uploading them into #LancsBox and Wmatrix. The main sources include the English version of the lsm.lv portal (Public Broadcasting of Latvia), from which I manually extracted six samples of political news, based on their date of publication (I selected the most recent ones at the time of compilation) and six samples of sports news focusing specifically on the bronze medal that the Latvian hockey team had won at the World Championships in 2023. This is an event that was also covered internationally. The American political news was extracted from the Financial Times; the sports news related to our achievement in hockey was extracted from various sources, including Sports Illustrated, the Washington Times, and NBC Sports. > 145 https://latdit.lu.lv/en clArin-lv
127 126 How concretely did you avail yourself of CLARIN tools for the preparatory extraction? < I used CLARIN UK’s #LancsBox tool to extract features relevant to Biber’s aforementioned informational vs. involved dimensions. The features were extracted from my political and sports news corpus, LAPIS. Apart from basic syntactic categories, such as prepositions, nouns, and adjectives, the tool allowed me to extract finegrained morphosyntactic features, such as first and second person pronouns, the use of the auxiliary ‘do’ in elided verbal phrases (e.g., he did that too), and contracted forms. These features are among those that are more predictive of the involved texts, which are more interpersonal and perhaps less formal than the informational ones. The tool is also easy to use – for instance, for extracting second person pronouns, I simply specified the regular expression ‘you|your|yourself|yours|thou|thee|thine|th y|thyself’, or ‘I|my|me|myself|we|our|ours|ourselves+mine P.* + us P.*’ for first person pronouns. There are some things that I was not able to do, however. One pertains to cases where the #LancsBox tagset does not define a category that would otherwise be relevant for the informational vs. involved dimension, such as discourse particles. Another pertains to the lack of syntactic annotation, so it was impossible to accurately extract prepositions that are separated from their nominal complements (who … to?), which is less formal than the variant where the preposition and the who word are directly adjacent (to whom?). Lastly, it was difficult to extract constructions with ‘deleted’ features (e.g., the optional deletion of the subordinator ‘that’ in dependent statements, such as ‘John said (that)’, which is, by assumption, also more characteristic of less formal language). > How concretely did you use CLARIN tools for the semantic tagging? < My first semantic tagging experience was with the CLAWS demo tool, but it did not allow me to upload large volumes of text. Wmatrix, which largely uses the same syntax as CLAWS, does not have this problem and is the best semantic tagging tool I have discovered so far. With Wmatrix, I compared the use of emotional expressions and expressions denoting psychological processes in four subcorpora: Latvian politics (texts from the LSM portal), Latvian sports (again, from the LSM portal), US politics (the Financial Times), and US sports (NBC Sports, Sports Illustrated, and the Washington Times). What I was able to show, for instance, is that the US politics subcorpus contains a greater number of expressions of emotions and psychological processes than the Latvian political subcorpus. For instance, the average number of emotional expressions in the Latvian subcorpus was 26.31 tokens, whereas it was 104.67 in the US subcorpus. This presumably has to do with the fact that the articles in the Financial Times are aimed at providing an analysis of events, and do not simply report them (which means that they constitute ‘involved’ discourse rather than ‘informational’ discourse in Biber’s terms). By contrast, the Latvian sports subcorpus contains more expressions of emotions than the US sports subcorpus, presumably because the event described in the corpus, i.e., the victory of the Latvian hockey team at the 2023 World Championships, was such an important and joyful event for our country. > Do you see any room for improvement for the tools that you have used? < The tools that I have used, i.e., #LancsBox and Wmatrix, are not designed to perform statistical analysis, so you cannot use them to calculate the weight of each linguistic feature, which would show how generalisable the feature is from the selected group of texts. To do the calculations, you have to build a correlation matrix of all the features, and there is software specifically designed for that. The applications I have tried are not a part of the CLARIN infrastructure, and they also have a serious drawback: they are designed specifically for Biber’s dimensions. They do not allow me to add my own linguistic features, for instance, my dimension of subjectivity that I am working on, and which does not wholly correspond to any of the dimensions in Biber’s original proposal. I would therefore be happy to find a multidimensional analysis tool in the CLARIN infrastructure that would allow me to add new dimensions and linguistic features, and would help me to do the necessary statistical calculations. Something that I would like to see added to the Wmatrix tool specifically are tags related to sentiment (for instance, the difference in connotation between the near synonyms unpleasant and horrible), as well as the differences between comparative and superlative forms of adjectives, and a more fine-grained classification of modals, such as the distinction between epistemic and non-epistemic uses. > clArin-lv
129 128 What are your plans for future work, especially regarding the use of CLARIN tools and resources? < I am creating other subcorpora, and plan to create additional ones which I expect to analyse in my research on subjectivity. Currently, I am working not only on my dissertation corpus, but also the LATDIT corpus mentioned previously. I will continue exploring semantic tagging in Wmatrix despite the fact that I miss some features, which I mentioned above; the tool is great for grouping the words belonging to the same semantic fields. The tool proves to be handy not only for analysing subjectivity, but also identity. I am also eager to start exploring other tools for multidimensional analysis and sentiment analysis provided by CLARIN. So far, I have encountered several sentiment analysis tools in the context of the CLARIN Resource Families, but they are all for languages other than English, while the Etuma Customer Feedback Analysis tool, which does include English, is limited to analysing customer feedback specifically. I am nevertheless curious to see how it compares to tools such as Wmatrix for researching subjectivity. > clArin-lv
131 130 COLOPHON Coordinated and edited by Kristina Pahor de Maiti Tekavčič (Institute of Contemporary History, Ljubljana; Faculty of Arts University of Ljubljana), Jakob Lenardič (Institute of Contemporary History, Ljubljana) and Karina Berger (CLARIN ERIC) Proofread by Laura Gusan (CLARIN ERIC) designed by Tanja Radež Cover image National and University Library of Iceland (row 3, image 3) Österreichische Nationalbibliothek: Tagebücher Andreas Okopenko (row 2, image 1) online version www.clarin.eu/Tour-de-CLARIN/Publication Publication number CLARIN-CE-2025-2598 September 2025 ISBN/eAN 9789082990942 This work is licensed under the Creative Commons Attribution-Share Alike 4.0 International licence. Contact CLARIN ERIC c/o Utrecht University Drift 10, 3512 BS Utrecht The Netherlands www.clarin.eu
132
Tour de CLARIN VOLUME FIVE Edited by Kristina Pahor de Maiti Tekavčič, Jakob Lenardič and Karina Berger CLARIN 2025 Tour-de-CLARIN-5-cover.indd 8 19/9/25 07:41:52