Vademecum and Survey outcomes - D4.1.2 Whitepaper on the results using DaMSym on Nicene-Constantinopolitan Creed
Abstract
Vademecum for Latin and Ancient Greek and outcomes of the surveys made for D4.1.2 Whitepaper on the results using DaMSym on Nicene-Constantinopolitan Creed. Surveys are derived from Google Forms submitted to experts in the linguistic field and finally turned into excel sheets with one participant for each column.
Full text
Vademecum Test Introduction ● We expect the test to last 20-25'. ● The tests will be collected via email but when we use them for our analysis the references will be anonymously. ● The test consists of two main parts and an overall assessment. ○ The first part consists of a hypothetical workflow that a scholar might perform using the tool we are working on. ○ The second offers a space for “free play” of which we ask you to annotate the steps Our analysis will reference these steps, so please remember to annotate them critically and clearly. 1. What you are testing You are looking at a user interface that allows you to query a corpus of sources, specifically, the Corpus Corporum. Both your query and the corpus are transformed into multidimensional vectors that numerically condense various explicit and implicit information within the sentences. In this way you can perform searches not strictly bound to verbatim rules or n-grams but expand the result queue by including correlations that are also relevant to semantics. 2. Guided test The first part of the test offers you to follow this guided case study so as to help you become familiar with the tool. 2.1. Corpus Corporum A good way to evaluate the functionality of our tool might be to compare the results obtained during the test with those offered by the online resources of the Corpus Corporum website at this link. 2.2 Test case In the development of the test, consider that, in the final version, the selection of authors, periods, etc. will be included. In this version, we still need to implement such content filtering. In Tertullian’s doctrinal and ecclesiological thought, the apostles and the disciples are those who founded the churches and transmitted through them the true faith. How would you retrieve the passages in which Tertullian presents his idea using traditional computing tools? Example of workflow with DamSym: enter in the "enter query" field:
“Ecclesiae antiquissimae ab apostolis et discipulis Christi institutae sunt.” Select “Tertullianus” among the authors. Enter as additional phrase: “Apostoli ecclesias condiderunt.” How would you conduct a search for expressions of the same literary motif using the tools already available to scholars for Latin? This is an open space for your research. We suggest you copy and paste the sentence so that it is embedded and its meaning can be searched within the corpus. We ask you to check the first 20 results, approximately. DamSym is an extremely flexible tool. It is possible to use meaningful phrases in synthetic Latin to retrieve concepts that we are interested in. You can check if the results change by entering contextual information in Latin in the second query box, adjusting the value to assign to this context by moving the slider below. We recommend starting by setting the slider halfway. 3. Free play You now have a window of time at your discretion in which we invite you to conduct some free-form testing, possibly, if you can think of it, drawing on previous work experience in which you have needed to conduct semantic retrieval. We ask you to operate and perform a judicious check on at least 5 searches (if you wish to conduct more, that is perfectly fine). If you needed ideas, here are a couple more guided tests Option A ● Word in context One of the possibilities offered by the tool is to sum inputs. A wide-ranging search on a single term, such as “water” in a baptismal context, would not return too precise information. The model needs to be helped by “summing” meanings. You might therefore try to see what happens by performing these queries, or even combining the elements contained in here. aqua aqua + baptismus aqua + sacramentum aqua + Credo in unum Dóminum Iesum Christum, Fílium Dei Unigenitum + … aqua + sacramentum + credo … Option B ● Syntactical structure It may be interesting to try to collect expressions that exceed the search for n-grams. You can search for them individually, or try to collect x-by-x formulas by summing them in your query. Can you find additional x-da-x formulas than those included in your query? Deum verum de Deo vero lumen de lumine Deum de deo
Deum de Deo, lumen de lumine, virtutem de virtute, integrum de integro, perfectum de perfecto
Vademecum Test Introduction ● We expect the test to last 20-25'. ● The tests will be collected via email but when we use them for our analysis the references will be anonymously. ● The test consists of two main parts and an overall assessment. ○ The first part consists of a hypothetical workflow that a scholar might perform using the tool we are working on. ○ The second offers a space for “free play” of which we ask you to annotate the steps Our analysis will reference these steps, so please remember to annotate them critically and clearly. 1. What you are testing You are looking at a user interface that allows you to query a corpus of sources, specifically, the greek Corpus TLG. Both your query and the corpus are transformed into multidimensional vectors that numerically condense various explicit and implicit information within the sentences. In this way you can perform searches not strictly bound to verbatim rules or n-grams but expand the result queue by including correlations that are also relevant to semantics. 1. Guided test The first part of the test offers you to follow this guided case study so as to help you become familiar with the tool. 2.1. TLG - opzionale A good way to evaluate the functionality of our tool might be to compare the results obtained with those offered by the Thesaurus Linguae Graecae (TLG). If you already have credentials of your own, you can access the online software directly and skip this section; if you need them, however, here are Elia Scapini's private credentials provided through the University of Vienna (with the implicit agreement to forget them once you have completed this testing path). ● At thislink: here ● click on Thesaurus Linguae Graecae ● Log-in: username “scapinie94” password “Vienna1994!” ● On the TLG username “EliaScapini” e password “Vienna1994!”
2.2 Test case In conducting the test, consider that in the final version, there will be a more structured selection of authors, titles, periods, etc. In this version, we have yet to implement such content filtering. Let's imagine that, while reading Eustatius' fragments, you come across the following passage ● Phrase condensing meaning: Those whom the gentiles call gods are actually demons τοὺς δαίμονας, οὓς αὐτοὶ οἱ Ἕλληνες νομίζουσιν εἶναι θεούς, τούτους οἱ χριστιανοὶ ἐλέγχουσιν, οὐ μόνον μὴ εἶναι θεούς, ἀλλὰ καὶ πατοῦσι καὶ διώκουσιν, ὡς πλάνους καὶ φθορέας τῶν ἀνθρώπων τυγχάνοντας G.J.M. Bartelink, Athanase d'Alexandrie, Vie d'Antoine [Sources chrétiennes 400. Paris: Éditions du Cerf, 2004]: 124-376. Retrieved from: http://stephanus-tlg-uci-edu.uaccess.univie.ac.at/Iris/Cite?2035:128:147329 ● How would you operate a search for expressions of the same literary motif on the TLG? This is a free space for your search. ● We propose that you copy and paste the phrase for it to be embeddable and its meaning to be searched within the corpus. We ask you, indicatively, to check the first 20 results ● OPTIONAL: You can check whether the results change by entering Greek language contextual information in the second query box, then adjusting the value to be given to this context by sliding the bar down. We recommend starting by putting the bottom bar in the middle 2. Free play You now have a window of time at your discretion in which we invite you to conduct some free-form testing, possibly, if you can think of it, drawing on previous work experience in which you have needed to conduct semantic retrieval. We ask you to operate and perform a judicious check on at least 5 searches (if you wish to conduct more, that is perfectly fine). If you needed ideas, here are a couple more guided tests Option A ● Word in context One of the possibilities offered by the tool is to sum inputs. A wide-ranging search on a single term, such as “water” in a baptismal context, would not return too precise information. The model needs to be helped by “summing” meanings. You might therefore try to see what happens by performing these queries, or even combining the elements contained in here. ὕδος + βάπτισμα ὕδος + λύτρον + βάπτισμα ὕδος + Πιστεύομεν εἰς ἕνα Κύριον Ἰησοῦν Χριστόν, τὸν Υἱὸν τοῦ Θεοῦ τὸν μονογενῆ [+...] …
Option B ● Syntactical structure It may be interesting to try to collect expressions that exceed the search for n-grams. You can search them individually, or try to collect x-by-x formulas by summing them in your query. φῶς ἐκ φωτός, Θεὸν ἀληθινὸν ἐκ Θεοῦ ἀληθινοῦ, θεὸν εκ θεοῦ εκ θεοῦ θεοῦ θεὸν ἐκ θεοῦ, ὅλον ἐξ ὅλου, μόνον ἐκ μόνου, τέλειον ἐκ τελείου, βασιλέα ἐκ βασιλέως, κύριον ἀπὸ κυρίου Grazie per il tuo prezioso contributo!
Latin survey Informazioni cronologiche 19/03/2025 16.11.20 21/03/2025 13.57.37 24/03/2025 13.12.56 28/03/2025 16.58.29 16/04/2025 18.15.40 What tools do you typically use to analyze the Latin sources you work on? Pdf, Excel, Word, Online databases on manuscripts I used the tool "search" on Digital Library of the Catholic Reformation (but I do not have access to it anymore). Fortunatly, most of the sources I use are available on Google Books, and its OCR is quite good, hence I can use the tool "search" quite well. DLCR allowed to look for specific words or expressions within the entire corpus of sources Thesaurus Linguae Latinae (TLL) / Library of Latin Texts (LLT) / Perseus Digital Library/ PHI Latin Texts/Lemlat I don't have a specific type of tool I usually work with Latin sources through direct reading of manuscripts, often digitized by the Institutions where they are archived. For Latin sources that have editions, I make use of the relevant printed editions or, when possible, digital editions. For many works, I take advantage of access through online repositories, including, for example: Corpus Corporum; The Roman Law Library; dMGH; Clavis Canonum; Library of Latin Texts; Brepolis databases; Mansi online (http://mansi.fscire.it/); etc. Latin survey - 1
Latin survey What features do you most look for when using software or online services to work on Latin sources? word and concept search and recognition "Download" (I like to build my own 'digital library' of the sources I used), and "search" reliability and authoritative sources, preferring academic databases such as the Thesaurus Linguae Latinae or the Library of Latin Texts (LLT) to ensure accurate and well-curated texts. Advanced morphological and syntactic analysis, with tools that help recognize declensions and conjugations, facilitating text interpretation. Semantic search engines, allowing searches for words in specific contexts or through synonyms to identify particular lexical usages. Support for textual variants and comparison between editions, which is essential for studying differences between manuscripts or critical editions. Nothing specific A software or online tool should have a simple and intuitive interface, providing access to features that speed up research that would otherwise rely on traditional tools. The operating instructions should be clearly specified through concise and straightforward guidelines. Search results should be presented just as clearly and with relative certainty, so as to exclude, as far as possible, the need for further searches using similar queries. In working with digital tools for Latinsource analysis, what improvements or additions do you find necessary? manuscript ocr improvement An ideal tool should provide the results within the entire paragraph in which the expression I look for is present; the possibility to combine two queries; the possibility to download the source; the reference to the critical edition, if any; the reference to the relevant literature on the source (maybe with an automatic search in Google Scholar or Jstor) semantic search capabilities, such as contextual understanding of phrases and idiomatic expressions. refinement to handle abbreviations, ligatures. No improvements or additions to suggest A digital tool for the analysis of Latin sources should: 1) allow for broad-range semantic search; 2) fundamentally ensure access to the widest possible number of sources; 3) with regard to texts not included in the original database, allow for the upload of additional texts in PDF (or other formats) in order to apply its features to new works as well. Latin survey - 2
Latin survey Describe in your own words what the tested tool does recognizes phrases and concepts, in association or not, within various Latin sources available in the software's database It searches within the entire corpus of sources a specific set of words and their variants, provided in the "query" and in the "additional phrase". It should also search for words semantically related (e.g.: praedestinatio > praescientia) This is a Semantic retrieval tool for Latin It provides a list of sources that use part of the words written in the quaery Describe your general impressions Is a good starting tool for doing initial source survey Overall, the tool is well designed, but it did not work quite well with the query I used. he idea can be extremely useful, but the corpus of authors should be expanded, and there should be a way to search for multiple authors at the same time without necessarily selecting 'all' of them. Due to a lack of information, the search on the site is not very precise, while it would be interesting to search for semantically similar phrases across two or three authors at a time. I have good impressions The tests conducted on the tool have provided very satisfactory results in relation to expectations. Most importantly, the tool offers an analytical capability that is currently unavailable on large repositories of Latin sources. Rate how intuitive the tested software is. 1 = very intuitive 4 = not intuitive 32112 Latin survey - 3
Greek survey What tools do you typically use to analyze the Greek sources you work on? TLG TLG TLG, Scaife, own scripts TLG, BibleWorks TLG What features do you most look for when using software or online services to work on Greek sources? La presenza di alcune parole o sintagmi negli scritti di alcuni autori/epoche/aree geografiche Lately, just text browser or simple lexical search good search possibility, download of results, retrieval of context Lemmatisation, dictionaries, English translation, intertextual search, parallel view of texts, user-friendly interface a well cured corpus to search across In working with digital tools for Greek source analysis, what improvements or additions do you find necessary? Sarebbe bello se si potesse fare ricerche per argomento, ossia individuare testi che trattano di un determinato argomento, a partire dalla presenza in essi di almeno un certo numero di vocaboli che appartengono a quella sfera di contenuto e impiego // easy use, more sophisticated search, open source More critical editions of the same ancient work, critical apparatus, less rigid research a more content related retrieval Describe in your own words what the tested tool does Cerca di recuperare testi con un certo tipo di significato seeks parallel passages. It does not limit itself to a lexical search but grasps the semantics of the text entered in the query and looks for passages of similar meaning. The tool provides an interface for semantic search The tool searches for similarities in topic or literary motifs it retrieves sentences accross a corpus based on semantic similarity to the query Greek survey - 1
Greek survey Describe your general impressions Non ho praticato molto il programma, ma non riesco bene a districarmi The tool works and can be useful for an initial phase of the search for parallel texts. The results are interesting, but lack a bit of context. I could not in each case understand why a result was retrieved. If I look for a rather simple and overused theme, such as the creation of heaven and earth, the results are good. Literal quotations are included if I enter a Greek phrase to search for in the additional query. I am surprised for how well the tool can collect similar sentences Rate how intuitive the tested software is 3 2 3 2 2 Suggestions for making the tool more practical and intuitive Inserire una finestra con uno o più esempi di ricerca; inserire un link a un tutorial No clue More explanations (on the input and selection of values); but also especially on the sorting of the results (why are they sorted this way?) and on what basis they were retrieved; at the moment it looks a bit like a black box. It is practical, but I would give suggestions on how best to set up the query to obtain good results. adjust authors and works, with more selection What is your level of familiarity with technologies based on artificial intelligence models? Sono poco familiare 0 low to intermediate Average familiarity low-medium Are you familiar with the basic mechanisms governing the operation of models such as BERT (transformerbased)? Beginner Beginner Intermediate Intermediate Intermediate Greek survey - 2
Greek survey ott-20 15/20 15/20 15/20 15/20 Do you have any impressions to report about the particular type of answers provided by the model? Often they sticked too much to the original text (lexical rather semantical); when I entered famous passages I had many direct quotations, and that wasn't what I was searching for Some results are relevant, but mixed with others far removed from the query the additional query is very influencing the output Greek survey - 3
Greek survey Describe the queries you decided to prompt La presenza in un autore di alcuni sintagmi I searched for a biblical passage, Sirach 24:7-8, in which the idea expressed is that of God sending his Wisdom to dwell in the world and in particular among the people of Israel. I was hoping that the search engine would find other passages associated with this idea, but instead it mostly found a more lexical type of correlation (passages containing the words ‘tent’, ‘land’, ‘inheritance’, etc.). The corpus is exclusively patristic. Having edited the entry for διαχωρίζω in the HTLS, I became interested in studying the cosmogonic use of this verb, and in general the idea of ‘division’ associated with ‘generation’, like in mitosis. This type of cosmogony is present in Aristophanes, Thesm. 14: αἰθὴρ γὰρ ὅτε τὰ πρῶτα διεχωρίζετο. So I looked up this sentence. The parallels all have to do with meteorological issues, but I don't think any of them are related to the cosmogonic dimension of the passage. The same verb is present in a passage from the Timaeus in which the changing nature of reality is questioned: After having presented some principles, Plato, Timaeus 58a poses a question: how is it that individual bodies (ἕκαστα) are not divided into four types (some made only of earth, others made only of water, etc.? ), I've took some passages from Ps.Chrys., in epiph (one bible verse, a phrase, the two words οἱκονομία βάπτισμα) φῶς ἐκ φωτός in Clemens Alexandrinus; ἔξοδος + βίος i followed the indications of the vademecum Greek survey - 4
Greek survey and do they not cease from their continuous moving and passing through one another (ἕκαστα πέπαυται τῆς δι᾽ ἀλλήλων κινήσεως καὶ φορᾶς)? So I searched this part of the Timaeus for parallels in the corpus of this idea of the becoming of reality as a continuous merging and not remaining divided of the elements. Predictably, he found parallels almost exclusively in the Platonic corpus. It would be necessary to include the possibility of excluding some texts from the search (the alternative seems to me to be just between ‘All’ and ‘single author/text’).Following the same line of reasoning in relation to the same verb but in a very different context, I searched for Gen 1:4 ’ καὶ εἶδεν ὁ θεὸς τὸ φῶς ὅτι καλόν καὶ διεχώρισεν ὁ θεὸς ἀνὰ μέσον τοῦ φωτὸς καὶ ἀνὰ μέσον τοῦ σκότους’. For the most part, the search engine found me commentaries that reported the passage from Genesis verbatim, and not passages whose meaning could be linked to the cosmogonic use that the LXX makes of the verb. Compared to a normal search engine, however, I see that the search engine also finds passages where the Genesis passage is not quoted literally but only alluded to. Greek survey - 5
Greek survey mar-20 mag-20 mar-20 ott-20 Greek survey - 6
Greek survey Do you have any impressions to report about the particular type of answers provided by the model? As a last step I searched for Jewish War II.154, in which Josephus reports the belief of the Essenes about the immortality of the soul: Καὶ γὰρ ἔρρωται παρ' αὐτοῖς ἥδε ἡ δόξα, φθαρτὰ μὲν εἶναι τὰ σώματα καὶ τὴν . αὐτοῖς ἥδε ἡ δόξα, φθαρτὰ μὲν εἶναι τὰ σώματα καὶ τὴν ὕλην οὐ μόνιμον αὐτῶν, τὰς δὲ ψυχὰς ἀθανάτους ἀεὶ διαμένειν, ecc. I wanted to see other texts that, even with different words, expressed the same meaning. The first three are direct quotes of the passage (you should find a way to exclude them from the search, if one is not interested in this kind of search, more useful for the philological edition). If I only enter two words or concepts, I struggle to grasp the connection no Greek survey - 7
Greek survey Did you find any errors or inconsistencies in the sentences suggested by the model? If so, can you describe them? Avendo formulato una query che conteneva anche un termine generico e dall'ampio spettro, ciò ha fatto si che i risultati non avessero attinenza con il contenuto che avevo in mente I found very good and clever that among the passages he included Pseudo-Galenus Med., which not only talks about the immortality of the soul but does so in exactly the same way and context as the passage by Josephus Flavius: by reporting the opinions of different philosophical schools. In several passages the tool has actually picked up this nuance of ‘indirect discourse’ (e.g. ‘they say that’, ‘among them the belief has become established that’, ‘they believe that’...) not limiting itself to simply reporting passages in which the immortality of the soul is spoken of. I expected that for the bibleverse the NT passage would be a result; the search with οἱκονομία and βάπτισμα mostly resulted in sentences with some form of οἱκονομία If I only enter two words or concepts, I struggle to grasp the connection it is hard to understand how the model derives similarities How well do you think the model captures the contextual meaning of the sentences in Ancient Greek? Abbastanza bene 08-ott difficult to answer Sometimes I don't think so pretty good Did you find difficulties in interpreting the results produced by the model? 3 3 3 3 3 Greek survey - 8
Greek survey What improvements would you propose to improve the generalization of the model so to scale it to your colleagues' needs? spiegare meglio come collegare la query con la domanda opzionale The possibility to exclude some authors from the corpus for the search; the possibility to exclude verbatim quotations. is the model to general? (built from the whole TLG from Homer to Byzantium? balance in view of greek as a changing language over time?) improving the corpus and making all the answers clearer. Add more indications on the work and passage In what ways do you think this model can enrich your research work? Per ora ho poca esperienza con questo modello, e non posso fornire una risposta adeguata I'm not doing research on this field There are not many tools that really include semantics ir really helps in finding similar content Do you find that this model offers hitherto unseen functionality? Sì, può aprire nuove prospettive No clue; I'm not interested on this kind of development of the research method. I hope so yes Do you think specific training should be organized to be associated with the distribution of the tool to the public? 1 3 2 2 2 Free comments It would be good to be able to change also the parameter for similarity; to be able to select more than one author for comparison Greek survey - 9
Church Slavonic Survey What features do you look for most when using software or online tools to work with Old Church Slavonic and Church Slavonic sources? manuscript repositories, aligned sources (with Greek original), annotated corpora the ability to identify different grammatical forms linked to the dictionary entry; the identification of biblical quotations Cyrillic and Glagolitic script support, proper rendering of fonts Proper visualization; accurate reference to grammatical features. When working with digital tools for analyzing Old Church Slavonic and Church Slavonic sources, what improvements or additions do you consider necessary? alignment, segmentation (word and sentence), tagging A standardisation of orthography, notifications and grammatical forms More accurate OCR for Glagolitic and early Cyrillic, Better HTR models trained on Slavonic manuscripts Visualizations – fonts; access to original manuscripts. Describe your general impressions of the tool you tested. the converter didn't work for my input text It seems very useful for slavic philology's scholars These tools are goog in supporting historical scripts and handle basic Cyrillic and Glagolitic input It would be beneficial to add additional fonts. Please provide any suggestions for improving the tool. mapping of all existing characters and considering existing converters (e.g. https://zenodo.org/records/7821 709) It is necessary the integration with OCR/HTR modules Extending the mapping to additional texts How accurate were the results you obtained? 5 7 6 6 Describe your general impressions of the tool you tested. the lemmatizer worked very well when providing a clean text. When providing raw HTRoutput and even GT-data it fails . It certainly useful, but needs implementations and improvements. They are good in supporting historical scripts and handle basic Cyrillic and Glagolitic input. The tool is useful, but it should be improved by expanding the range of scripts and words. Church Slavonic Survey - 1
Arabic survey Do you have any impressions to report about the particular type of answers provided by the model? Se la stringa in input è composta da pochi termini le troppe similarità semantiche e la polisemia tipica delle radici arabe non permette un facile riscontro. Lo stesso vale per stringhe in input molto coprose e che portano il sistema ad approssimare troppi termini. In alcuni casi un match perfetto si trova in stringhe contenenti testo che è sicuramente presente in alcuni testi (e.g., ميحرلا نمحرلا ﷲ مسب) ma questo non può riflettere l'effettivo funzionamento overall dello strumento. No, it didn’t no coherent and contextually relevant response. Only few keywords The tool answered me by quoting a prophetic tradition (hadith)in the first place, So it had vaguely identified the religious topic, but semantically the text was far off. Describe the queries you decided to prompt I searched for the following Arabic text taken from a tafsīr: يتوق فعض وكشأ كيلإ مهللا سانلا ىلع يناوهو يتليح ةلقو محرأ تنأ نيمحارلا محرأ اي نيفعضتسملا بر تنأو نيمحارلا Query composte da una singola parola, fino a stringhe consistenti in un paragrafo intero A Qur’anic verse and its commentaries by some of the most important classical exegetes I inputted specific passages from Sahih Al Bukhari texts into the search or analysis tool to evaluate how effectively the tool can handle and analyze classical Arabic texts I submitted for research the first verse of the Qur'an form the Sura of the Elephant Arabic survey - 6
Arabic survey Arabic survey - 7
Arabic survey Do you have any impressions to report about the particular type of answers provided by the model? no relevant results for the submitted query. In some cases it proposes a short section of the text of the identified works, in others a very long section of text: I would standardize in a defined number of lines. Come anticipato, il modello gestisce bene strighe brevi e porta al risultato, gli output limitati però al singolo secolo pre-impostato (difficile anche permettere una selezione mutliclasse temporale dal momento che il corpus OpenITI è così indicizzato), non garantiscono un accesso diretto alla porzione di testo in cui l'output e contenuto e serve comunuqe effettuare una ricerca anche nel corpus OpenITI che, essendo in txt su github, non sempre molto fluido. No, because it isn’t precise in answering I tested with the simple and commonly used word "Allah," but the outputs did not include this word It seems to me that it does not recognise the passage I am proposing (i.e. whether it is Qur'aninc verse or passages from classical works) and, consequently, is unable to respond adequately. Did you find any errors or inconsistencies in the sentences suggested by the model? If so, can you describe them? I found no major errors in the proposed texts Ci sono errori dovuti alla capacità del modello di generare match semantici in determinati casi most of the sentences were not connected to the inputs I gave to the model Must be improved Arabic survey - 8
Arabic survey How well do you think the model captures the contextual meaning of the sentences in Arabic? The model does not capture meaning because the results are not relevant in terms of text/contextual reuse In alcuni casi funziona correttamente riportando adeguatamente il contesto semantico, in altri sembra non riuscire a cogliere correttamente la semantica provando comunque a far corrispondere stringhe di testo che non sembrano essere effettivamente in relazione semantica. The model captures the genre of the sources, but not the real meaning and the context of the sources Not really well. The model struggles to accurately capture the contextual meaning of Arabic sentences As I said before, I think the tool must distinguishes the texts it reads Did you find difficulties in interpreting the results produced by the model? 1 = Easy to interpret 4= Hard to understand 4313 Arabic survey - 9
Arabic survey What improvements would you propose to improve the generalization of the model so to scale it to your colleagues' needs? cf. the answers to these questions: Suggestions for making the tool more practical and intuitive Do you have any impressions to report about the particular type of answers provided by the model? Maggiore addestramento non solo esclusivamente su un singolo corpus, nonostante non sia qualcosa di semplice data la scarsità di dataset per questa tipologia di testi per numerose applicazioni computazionali. I think a model like this is most useful for finding links between sources, but it absolutely should be optimized in the meaning of the sources it is asked to process I would suggest incorporating a wider variety of sources, including texts from different genres, time periods, and both with and without diacritical marks It must improve semantic reconnaissance and it needs to learn to recognise the sources it uses placing them by subject, so as to focus on a more specific topis (qur'anic exegesis, historial material, philosophical or literary texts) In what ways do you think this model can enrich your research work? In finding intertextual references between texts Connettere significati semantici in testi diversi portando a risultati nuovi per il ricercatore It would greatly speed up the search for sources related to the ones I am studying It has the potential to assist with text comparison, keyword searches, and tracking recurring themes across various Arabic sources, which can support deeper analysis of Islamic text like Sahih al Bukhari collection Arabic survey - 10
Arabic survey Do you find that this model offers hitherto unseen functionality? not entirely, but if applied to a specific literature with better performing results than general tools, yes Non allo stato attuale, ma potrebbe offrirle in futuro anche con la possibilità di combinare più fasce temporali e con la possibilità di raggiungere la zona di testo contenente la stringa di output, nonostante non sia completamente nuova come funzionalità (un lavoro interessante era stato fatto con Qawl software sviluppato a Lovanio da Sebastien Moreau nel 2014-2015, su un corpus preso dal sito alWarraq, una volta open ora non più. Questo permetteva la ricerca testuale in un corpus di circa 2000 testi nonostante non adottasse BERT o altri transformer e quindi senza generalizzazioni semantiche. Yes, the idea is new For now, the model does not introduce any significantly new features or capabilities." Do you think specific training should be organized to be associated with the distribution of the tool to the public? 12212 Arabic survey - 11
Arabic survey Free comments Arabic survey - 12
Sankrit survey What tools do you typically use to analyze the Sanskrit sources you work on? Dictionaries, electronic texts, online texts Thesaurus Indogermanischer Textund Sprachmaterialien Monier-Williams dictionary online version, sometimes comparing other Sanskrit-English dictionaries Dictionaries, electronic dictionaries, electronic texts, online texts What features do you most look for when using software or online services to work on Sanskrit sources? Occurrences of words, isolated or in groups. Formation of compound words. Division of the texts without sandhi For me, it is crucial to identify where specific words appear, in which text, and in what context. I also want to analyze which other words they are associated with to determine whether there is an increase in their usage and whether clusters of contexts emerge where these words are commonly found. Precision and completeness Identification of word occurrences, whether individually or in clusters. Analysis of compound word formation. Segmentation of texts without applying sandhi rules In working with digital tools for Sanskrit source analysis, what improvements or additions do you find necessary? Having the complete texts that have been published recently. Having the possibility to undo the sandhi. Metrical indications of poetic texts. A complete and uniform list of abbreviations for the texts. To explore semantic aspects, identify synonyms, and analyze clusters or pairs of words that frequently appear together. Widening the corpus of digitalised Sanskrit texts Access to the complete collection of recently published texts. The ability to reverse sandhi. Metrical annotations for poetic texts. A standardized and comprehensive list of text abbreviations. Expanding the corpus of digitized Sanskrit texts to explore semantic nuances, identify synonyms, and analyze clusters or word pairs that frequently co-occur. Sanskrit survey - 1
Sankrit survey Describe in your own words what the tested tool does The tested tool searches for occurrences of single words or groups of words, indicating also the position of the word(s) in the verse or line. The tool is useful because it identifies words within texts and specifies exactly in which texts they appear and their positions. It becomes even more valuable as it allows searching for two or three words together, pinpointing the texts where these words co-occur. Additionally, it shows their relative positions to one another, enabling me to easily determine whether their connection is coincidental or if there is a semantic link worth exploring. The tested tool allows to map the occurrences of a Sanskrit term in a wide range of textual sources, highlighting its position in the sentence. The tested tool effectively searches for single words or groups of words in Sanskrit texts, pinpointing their exact position within verses or lines. It allows users to map the occurrences of terms across a wide range of textual sources and highlights their placement within sentences. A key advantage is the ability to search for multiple words simultaneously, identifying cooccurrences and showing their relative positions. This functionality helps in determining whether the connection between words is coincidental or holds a deeper semantic significance Describe your general impressions A very useful and powerful tool The tool is useful. Regarding queries about single words, it is not particularly original, as other websites offer similar services with different types of texts. However, this tool stands out by allowing queries across multiple texts simultaneously, which is a significant advantage. Another highly useful feature is the ability to perform queries involving multiple words together. This functionality is unique, and I don’t recall encountering a similar service elsewhere. It appears to be an efficacious tool to work on the enormous Sanskrit textual corpus. It allows the researcher to quickly identify the contexts in which a term is found. This tool is effective for working with the vast Sanskrit textual corpus, enabling researchers to swiftly identify the contexts in which a term appears. While similar platforms exist for singleword queries, this tool allows searches across multiple texts simultaneously. A particularly unique feature is the ability to search for multiple words together and analyze their relative positions, which helps in identifying potential semantic connections. This functionality is not commonly found in other tools. Sanskrit survey - 2
Sankrit survey Rate how correct the tested tool performs. 1 = correct. 4 = not correct 4 1 4 3 Explain your previous answer It is in accordance with what I expected, it's very fast and precise I tested the tool with very specific queries, and the responses were both precise and satisfactory. The results of my queries were relevant and correct. The tool met my expectations, demonstrating both speed and accuracy. The search results were relevant and precise, even when handling highly specific queries. Overall, the tool provided satisfactory and reliable responses. Do you have any suggestions for making the tool more practical and intuitive Not really no particular suggestion, because the tool is practical and intuitive. Maybe, if possible, it could be useful to organise the textual corpora appearing in the results in estimated chronological order: first Vedas, then epics, and so on. I found the tool practical and intuitive, with no major suggestions for improvement. However, I believe that organizing the textual corpora in the search results by estimated chronological order could be beneficial. What is your level of familiarity with technologies based on artificial intelligence models? I'm not very familiar with technologies I use some tools powered by AI systems, but I wouldn't consider myself an expert. Very low familiarity. I have limited familiarity with technology and I not consider myself an expert. While I use some AI-powered tools, my experience remains minimal Are you familiar with the basic mechanisms governing the operation of models such as LSTM? Beginner Beginner Beginner Beginner Sanskrit survey - 3