„Knygos nobažnystės“ (1653) morfologijos tyrimo galimybės naudojant reliacinę duomenų bazę
Full text
ACTA LINGUISTICA LITHUANICA LI (2004), 1-14 Klaipėdos universitetas The article presents the relational database created for the investigation of the DALIA JAKULYTĖ morphology of old Lithuanian writings. A logical data model was developed and implemented within the Microsoft Access database management system (DBMS). All the words of two parts of the Knyga Nobažnystės were entered into the database. An engine was created for morphological (inflectional) data selection, so that the words can now be selected by any assigned property or set of properties. Such a database serves as a tool for data sorting, selection and correction, as well as Jor text and index media storage. It allows the researcher to deal with various linguistic problems, such as noun stem alternation, authorship of texts etc. in an efficient, guick and accurate manner. 0. INTRODUCTION One of the biggest problems facing linguists studying ancient writings is the selection of material. The main sources of material are the texts themselves (e.g. photocopies of books), dictionaries, word registers. The data is transcribed (by hand or by computer), paper files are created, they are sorted and arranged in a certain order. The selection of material from such a library is a long and rather complicated task: the entire library is reviewed, the necessary sheets are selected, they are divided, etc. ; at the end of the work, the selected sheets must be returned to their places in the common file, otherwise it will not be suitable for use. Such problems are solved by computer technology. For lexicographical research, computer texts, databases of ancient writings, etc. are created. One of the essential characteristics of computer textbooks is the ability to perform automatic text analysis, for example, to automatically make text concordances, i.e. to present the searched word in a minimal context. For automatic language analysis, texts must be marked so that the concordance program searches for language units not only by formal, but also by grammatical characteristics. (Marquardt 1997). However, it is only for the normative language that automatic recognition of grammar forms or similar programs can be developed and applied, the grammatical features of the Oo old scripts must be marked manually. For example, the 100 million word textbook of Lithuanian language (donelaitis.vdu.lt/ tekstynas) prepared by the Computer Linguistics Centre of the Faculty of Humanities of Vytautas Magnus University consists of publications from the independence period, selected in such a way that as much as possible
2 | DALIA JAKULYTĖ riau reflect the current written Lithuanian language. The textbook search engine allows you to search for a word, find out statistical information about the selected word, search for the selected word in a specific grammatical form (e.g. independence), or all forms of the word using the symbol "*" (independence*). The Institute of Mathematics and Informatics of the University of Latvia has also made available on the Internet the texts of ancient 16th-18th century writing monuments with similar search options (you can search for a specific word or part of it with the symbol “96”) and archived dictionary and index files of the frequencies of word forms (www.ailab.lv/senie). Similar work is carried out at the Lithuanian Language Institute, where the texts of ancient writings are collected or scanned by computer and indexes of word forms are created. They are created semi-manually, i.e. the computer only arranges the word forms in alphabetical order according to a special program, indicating their place in the text and providing context. Textological information as well as the grammatical and lexical meaning of a particular word form are written down by hand (Ambrazas 1998). The morphology of ancient writings is traditionally studied using the same library as for dictionaries, but with a different mechanism of their use, for example, they are sorted not only according to lexems, but also according to other characteristics (parts of the language, morphemes, verbs, form, etc.). Therefore, the structure of the computer database and the material selection options must be adapted accordingly. The idea of creating a database of the morphology of ancient scripts arose as a continuation of the research of Knygos nobažnystė (1653) by Antanas Jakulis. Until 2004, when the facsimile edition of the Book of Nobažnystė (Pociūtė, ed., 2004) appeared, Lithuanian linguists had access only to a microfilm made in 1968 (preserved in the Central Library of the Lithuanian Academy of Sciences). In the same year, the Baltic Studies Centre of Klaipėda University had photocopies of the Book of Nobažnystė (Suma Gospels and Krikscionisky Prayers) made by A. Jakulis, and the Library of the Faculty of Arts had photocopies of the hymnal. From these, A. Jakulis compiled the register of the Book of Nobažnystė lexicon, of which only one part has been published — the Summa Gospels lexicon (Jakulis 1995). For this lexical register, a paper record of the entire lexicon was compiled, i.e. words were written on separate sheets, indicating their basic forms, meaning, grammatical form (Figure 1), these sheets were grouped according to the basic forms and arranged in alphabetical order. However, the compiler, further investigating the language of the Book of the Church, himself used the card library and rearranged the sheets in the order known only to him. Basically it had to be rebuilt, but we wanted to make it suitable for many uses. Therefore, it was decided to create an electronic repository — a database in which various morphological, lexical, semantic and other data would be collected, sorted or filtered and which would serve as a tool for the researcher of the morphology of ancient writings. The creation of such a database consists of two main stages — design and implementation. During the design process, it is decided what kind of data is needed by the researcher of the morphology of ancient writings, the specificity, characteristics and interrelationships of those data are defined, and the questions that the database must answer are foreseen. The result of this step is a logical data structure. It is described in this article Section 2.1. For more information about the Book of the Church and its history, see In the works of Dainora Pociūte (2004), Inga Lukšaite (2001), Zigmas Zinkevičius (1988: 213-223), Juozas Tumelis (1967, 1968) and others.
BOOK OF THE CHURCH DATABASE | 3 sis stage —realisation of the created data structure in one of the database management systems —described Section 2.2. The created and implemented morphology database of ancient scripts can already be used for research, i.e., after uploading the data and marking the necessary characteristics, the data can be selected, processed and sorted in the desired order. The second section of the article describes the application of the created database of morphology of ancient writings to the morphological research of a specific source — the Book of the Church. 1. MORPHOLOGY DATABASE OF OLD WRITTINGS DEVELOPMENT 1.1. Logical data structure The main elements of the lexical register are the forms of the words used in the text (hereinafter, each word form, i.e., each case of use of any lexicon, will be referred to as a “word”). One card — one word and its characteristics (attributes) (see Figure 1). 1 PAV. BOOK OF THE CHURCH LEXICON CARD basic word kara form source, page, line part of the Chiefs language id sas, (E50 kt „someone mimatytas, mistatytas, skutas metas, | termmas““ — meaning Pagižink tq čiefą oF atlankima faw/o duolds jam ii kolap i éigfas ird dffiraft —| Bais Eshis Each word is unique. Even if the same word is used several times in the same form and meaning, each time it will be in a different place, in a different context; if the cards were numbered, the same word forms would have different numbers. Therefore, the main element of the database must be a word with a unique number. Its attributes — form, source, basic forms, meaning, etc. Some attributes (e.g., part of language, paradigm) are common to the entire lexicon. To define the attributes of a lexem, an element “Lexeme” is created, linked to the element “Word” in a “one-to-- many” relationship, i.e. one “Lexeme” record can correspond to several “Word” records. A summary of the logical structure of the morphological data of ancient scripts is shown in Fig. 2.
2 PAV. LOGIC DATA STRUCTURE? DALIA JAKULYTĖ Word I Zisman Lexicon of the Word ID Tumikabis mmeris Leksemos ID Main f. N Socket Value 4 Pagr. f, ormos Function Part of speech Changes imekat ; links. ; asmen. Kilme Family, L ed. ; neither. ; General Structure Number {ns dgs., i.e. Paradigm Degree į of liar. ; high. word Definition /identify ; rt ormantas Form of action Sbendr. asm dal; f sly ... : Contribution(s) Reflectivity ~~ /sanex; Don. Species fvelk ; reveik, Determine ties, Lip; tar. ; Time to fes. ; be. ; e.g. ; will be. Laiko/nuos.forma | comes. sud atlikt.sud sākt.,... Person 11;2;3 The la i trunk goes ow... Valdym as You have a father. ; su kilm. .... Context / text string sentence ar The text į of the Gospel; sermon prayer... Author Section Page Row More about the specifics of morphological data, the elements, their attributes and interrelationships is described in the article on the creation of a database of morphology of the Book of Mormon (Jakulytė 2001). 1.2. Morphology database implementation The electronic equivalent of the archive is a database. For processing, viewing, selection, printing and entering new data, specialized software is used — database management system (DBMS). Dabartiniu metu plačiausiai taikomos reliacinės duomenų bazių valdymo sistemos (pvz., Visual 2 Duomenų bazėje saugomos informacijos elementas —esybė —loginėje duomenų struktūroje žymimas stačiakampiu. At the top is the name of the entity, at the bottom are the properties of the entity. The one-to-many relationship is represented by a branching web.
BOOKS CHURCHES DATA BASE | 5 FoxPro, dBase, Paradox, etc.). MS Access, a database management system included in the Microsoft Office suite, has all the characteristics of a relational database management system and a good integration with other programs in this suite. It stores data in tables. Each row in the primary table a separate record — (main element ir jo attributes- ). Lists of duplicate items (attributes) are stored in auxiliary tables. Elements are written to the main table, and attributes are selected from the auxiliary tables. By linking tables, you can make various queries that allow you to select information from one or more tables (i.e. select the required cards from the card library and arrange them in the desired order). Based on the logical data structure, the morphology database of ancient scripts was realized using Microsoft Access 2003 DBMS. Two main tables are created here: "Words" (for word forms) ir ‘Pagrindinės formos’ (leksemoms)’. In the auxiliary tables — lists of attributes. All auxiliary tables with a parent linked relationship "one su among many". Tables of the morphology database of ancient scripts and the relationships between them are shown in Figure 3. 3 PAV. DATA BASE TABLES IR COMMUNICATIONS TARP JŲ DBVS MS ACCESS SCHEMA Author Section Page Row Self-sufficiency Basic forms Definition aldymas Verb for Linksniavino, asme Daryba Reference word Action formant Base root1 Base root 2 Priesaga (hours) Prefix (ii) operational basis Lako/nuosakos for Person Notes Jaknavičius Sirvydas Daukša 3 Here are the original table names for now. They, well, ko changeable will not change the — name quite complicated, the user usually does o not see. jo
6 | DALIA JAKULYTĖ Some of the auxiliary tables are already completed, others are being adapted for each writing monument. The tables of grammatical categories (shown in Figure 3 on the left) are filled in. Each of them consists of several graphs: identifier, attribute name(s), and notes. Only the identifier is entered into the main table from the auxiliary, but for the gum some attributes are defined by duomenų atrankoje galima naudoti bet kurios grafos duomenis, todėl vartotojų patoboth Lithuanian and international terms (see the list of verb categories in Table 1). The Context, Section, Text, and Author tables apply to each monument of writing. The “Lexemy” and “Value” tables are empty at the beginning, but data can be loaded from the database created for another script monument. TABLE 1. LIST OF LINKS CATEGORY ATTRIBUTE Links category ID Links Latin Notes 0 i 1 ward. N. Nomen 2 kilm. G. Kilmininkas 3 assists D. Uudnieks d 4 gal. A. Rear end 5 innag. I. Įnagininkas 6 ines. In. Inesyvas 7 iliat. It. Iliad 8 ades. Ad. Adesyvas 9 alliat. Al. Aliative 10 tablespoons. 21st Century Fox. Mr. Possess. Characteristic causative agent
BOOKS OF THE NEW CHURCH DATABASE | 7 2. APPLICATION OF THE DATABASE TO BOOK COLLECTIONS FOR MORPHOLOGY STUDIES 2.1. Data entry Before entering the morphological data of a particular script monument, the database is adapted. For example, when studying the KN language, the text structure and authorship must be taken into account (see Jakulis 1982, 1984). Different parts of the book were prepared or written by different authors, so the table “Text” lists the types of text (gospel, sermon, prayer, hymn...), the table “Authors” lists the known authors of the KN. The ‘Chapter’ table is also filled in, listing all CN headings (see Figure 6). Data entry starts with typing text. The collected text is loaded into a table in rows, with identification numbers assigned to the rows. The text is then broken down into words, assigned identification numbers, marked with line IDs, text IDs, section IDs, and original page and line numbers. All these actions are performed automatically by the author's created macros of the text editor Microsoft Word. Since the word boundaries in this text often do not coincide with the spaces between words, the table is reviewed and corrected. The viewed tables are loaded into the corresponding tables of the morphology database “Context” (it has been renamed “KNtext”) and “Words”. Other attribute IDs (genus, number, verb, time, etc.) are recorded manually (manually) in the database table “Words”. For ease of entry, IDs are numbers, the sequence numbers of the corresponding category, so for example, the word Kryftaus form (vyr. vns. kilm.) simply enter “1”, “1”, “2” in the appropriate columns (see Figure 4). 4 PAV. FRAGMENT OF THE TABLE "WORDS". FORM INDICATION BSR | Word I Context cision im ea oy i Via ro | | 20030701 dteyfiancia 200307 2 I 301 3 7 O etti(at) ateiti 2 1 2 1.12 3 0 | | 20030702 Mrs 200307 2 I S01 3 7 Mr Mr la 1 1 1 2 1 0 P| 20030703 Kryftaus 0007 2 1 Sot 3 7 O Kristus 102/1 1 10 0 || 20030704 ant 200307 2 1 s01 3 7 0 ant 2804/1 0 g-.0 '0 0 0 |_| 20030705 luda 200307 2 1 S01 3 7 0 sūdasi sūdas 1 1 || 2 1 0 Before specifying the basic forms, the corresponding entry is created in the table "Basic forms", and it is entered or selected by expanding the list in the column "Basic forms" of the table "Words" (Figure 5).
8 | DALIA JAKULYTĖ 5 PAV. FRAGMENT OF THE TABLE "WORDS". | | i [K m: | a Sa Fb lg Zi Meaning HOE my ar ira 2 1 3 7 0 go(at) come 2 1 1 0 | | 20030702 Pond Ar 2 1 ui 3 7 0 Mr. Mr. la 1 to 1; to 0 |-7| 20030703 Kryftaus 200307 2 1 sol 3 7 0 o 0 | | 20030704 ant 200307 2 1 S01 3 7 0 1 0 0 | | 20030705 fuda 200307 2 1 S01 3 7 0 Knstu Knstusas, Christo 1 1 0 | | 20030802 Kayp 200308 2 1 S01 3 8 0 Crossing crosses crosses, crosses 1 0 0 |_| 20030803 tatay 200308 2 1 S01 3 8 O Įkrosyti krosyt krosyti* 3 0 0 | | 20030804 ira 200308 2 1 S01 3 8 0 Įkrūpauti krūpai krūpauti $ 1 0 Currently, the SE, Pas, MKr and K words are included in the database, and those of their features that do not require additional research or for which research has already been carried out are marked. Therefore, for example, according to the hypothesis of Antanas Jakulis (Jakulis 1982), the authorship of SE parts is marked, and in the table “Main forms” only the identifier, the nest, the implied main form and the language part ID are entered. The main forms may be adjusted and other features are indicated later, after reviewing all forms of the word (see section 3.2). Not all signs are listed for several reasons: —the characteristics (location, form and main forms) specified for the selection of material for morphological studies are sufficient, many features are reserved for other research, such as lexicology or semantics; —some characteristics are needed only by one other researcher (e.g. construction, base word, etc. will concern only word construction specialists), some should be specified by the relevant specialist (e.g. meanings — lexicologist), etc. ; —many characteristics are controversial or doubtful, some depend on the starting point of the research: for example, the word kalnas should be considered a root or an adverb, a derivative of the adverb, i.e., is it associated in the speaker- ’s consciousness with rise, rise or not? Or delnas? How to indicate the root — as it is (kal-) or with the basic degree of vowel change (kel-) — also depends on the needs of the researcher, for example, whether he will want to select words with a in the root or all derivatives of a certain root." —it is often difficult to define the type of mockery (it is indicated in the table “Basic forms”), because in the Book of the Church there are many different forms, e.g.: ugly / ugly, honour / honour, etc. In each such case, it is necessary to first carry out detailed research, to determine what this variation depends on (maybe it is determined by semantics, maybe different forms are used by different authors, etc.); It seems, how much less problems should arise when indicating the specific form of the stem (in the table "Words"- ), but here too we have to doubt: if the ears —i stem, ears —io stem, ears —adhesive, then the ear —? The investigation is therefore left to the investigator himself. For this purpose, it is convenient to use different modes of viewing the database management system material. 4 It is true that two rows are provided for this purpose: “Root” and “Footnote root”.
BOOKS OF THE NEW CHURCH DATABASE | 9 2.2. Material review, sorting and adjustment The material can be viewed directly with the main one in the associated auxiliary tables. In the “Chapters” table we can see the full text of each chapter (Figure 6). 6 PAV. FRAGMENT OF THE TABLE "SECTION". CHAPTER 13 OVERVIEW | | Chapter ID | Book | Chapter Title | Gospel | Beginning p: | +1S12 1 On the fifth week after the three days Matthew 13, 24:30 37 =!5813 1 Antennas of the old usage Matthew 20, 1:16 4C | ID | Rows | Text + 204005 ON WEEK 8 = 204006 OLD CREDIT. g | + 204007 Gospel of Matt. 20. 8 mf 204008 PRiliginta ir karalifte dangaus žmoguj kutlay 1 E 204009 iBeio tšbay anklti/famdit darbinikus ing 1 EF 204010 winničią [awą o kad [udereia darbinikus iž gra-1 pip 204011 fia ant dienos/nuliunte iuos winnicion lawa. | + MAN? TeiRaias dna alina tvačia imide Ihe Rosse 1 By looking through the text, we can see and edit all the words and you characters in each line (Figure 7). 7 PAV. REVIEW OF THE DATA IN THE TABLE ‘CN TEXT* | ID | Eilutés | Tekstad Puslapis | Eituté | + 207005 uzutaykis/no taste [death on the age/kal-1 = 207006 beia then him židay. Now pažind iuog welnia 1, D | Word [Tekst[Autor] Chapter [Puslag| Reilut{Savar|Pagrindines fo] Meaning |Kaity| Gimin] Skaid Link{Kamien|Laips] Ap |-| 20700601 then 1 1 S20 70 6 O then 0 0 {0 lo io 0 |-| 20700602 iam 1 1 S00 7 6 0 iis 0 1 2 11 3 o 0 |-| 20700603 židay 1 1 S20 70 6 0 žydas 0 1 i 13 id 10 |_| 20700604 Now 1 1 50 70 ls O now now 1/o 9 io 10 o 10 |» | 20700605 paind 1 1 S10 70 6 0 = zinti(pa) 0 2 | 0—13 io gv o | | 20700606 iuog 1 1 S20 70 6 0 jog 0 0:0 {0 0 o 0 | | 20700607 welnia 1 1 S20 70 6 0 = velnias 0 1 1 1 4 2 0 |%) 0 0 0 0 0 0 0 10 ip 0 gig + 207007 holds. Abrahomés numire and prénaBay/o tu 1 in j 2 + 207008 kalbi iey kas zodi mano uzutaykis/ne ragaus 1 By looking at the basic forms, we see all the uses of each word (Figure 8). The most convenient way to edit the basic forms and other lexical features is in this mode, because you can see all the forms of each word, the text and the sections where they are used, etc. It is therefore possible to determine the type of change of each lexicon.
