scieee AI-readable full text Open interactive document viewer

Kol nebuvo „Palemono“: kompiuteriai, lietuviški rašmenys ir šriftai

Petras Skirmantas

Full text

BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 PETRAS SKIRMANTAS Lietuvių kalbos institutas BEFORE THERE WAS PALEMONO: COMPUTERS, LITHUANIAN SCRIPTURES AND FONTS KEY WORDS: special characters, encoding, codesheet, ASCII, ANSI, unicode, text pronunciation, phonetic transcription. Talk about Lithuanian computer fonts is constantly renewed: the aim is to find out how many, what kinds and for whom they are, in which cases they are needed, how widespread they are, etc. Recently, the opportunity to talk about this occurred once again, when the State Commission for the Lithuanian Language and the Lithuanian Language Institute organized a discussion seminar Palemonas and other Lithuanian fonts: use and prospects. This paper also attempts to relate to what was said there. Practical needs, as they say – life itself, forced us to start creating computer fonts at home, or rather, to modify them… In the first half of the last decade of the last millennium, after 1990, computers of that time began to appear and become more and more abundant even in the philologists and Lithuanianists’ circles. Firstly, they became more and more affordable, secondly, many of them were brought in as charity from abroad, because foreigners bought newer, more modern ones, and gave away the older ones or sold them to us at a relatively low price. Computers then were quite different than today, much more modest capabilities. But since the computers of those days were also able to write, as we then began to say – to compile text, they very quickly pushed the typewriters out of circulation. The convenience of this was enormous and decisive: when writing or rewriting a text on a computer, the concept of draft loses its meaning, because all corrections of errors or editing changes the corrected parts of the text, and both on the screen and in print only the corrected text is visible all the time. Soon, for philologists, as well as for many people working in other fields, writing and compiling texts became the main computer-based work. Since computers came to us in various ways and often not from our first youth, there was a problem from the very beginning with the possibility to write correctly in Lithuanian on them: we wanted and needed to see all the characters used in our current spelling, first of all all the letters Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 2 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 of the current Lithuanian alphabet (ir ą, ę, į, ų, č, š, ž, ė, ū), and also other characters that occur in texts. This was needed not only by philologists or Lithuanianists, but by anyone who wanted to write or edit Lithuanian texts in their field using a computer. And the possibilities were not always... In a slight digression, it should be recalled that already at that time computers and philologists were aplinką patekdavo dažniausiai dvejopi, ir vieni nuo kitų jie gerokai skyrėsi tiek galimybėmis, as much concerned with the convenience of working (the methods of control and the nature of “communication” with human beings, the so-called interface). Some of them were designed and manufactured by Apple, others by International Business Machines (IBM). In the early 1980s, Apple began producing the Macintosh series of computers (named after an apple variety), later known as the Mac, with a mouse and a graphical user interface (GUI) that arranged almost everything needed to work on a computer into well-structured lists displayed on the screen in specially drawn graphical windows. It was enough to point the mouse to the right item in the list and the computer "understood" what to do. In the windows you can see the result of the action – the processed text or graphic document, its status after performing one or another action. The basic principles of the graphical user interface spread and remain valid to this day. At the same time, IBM began producing a similar series of PCs (abbreviation for Personal Computer) that did not have a graphical user interface for some time. All their control commands had to be entered (collected) by the keyboard, and the interface with the human was only textual, i.e. both the entered commands and the results of the computer's actions were displayed on the screen as text composed of uniform characters of standard size and shape, limited by placeholders. A checkpoint is a rectangular standard box composed of a certain number of dots divided into equal columns and rows. A keyboard and a bookmark split screen are collectively referred to as a text terminal; For some time it was the main means of exchange of information with the IBM PC-type computer. Therefore, working with such a computer is a tricky task: you have to memorize the control commands, which are quite a lot, and correctly, without mistakes, write them and the parameters necessary to pass them (data, references). A person can, so to speak, correspond with such a computer: he writes commands and data in standard text on the screen, and the computer – results obtained from executing the commands. This type of interface, based on the principle of a text terminal, is still the main way of exchanging information with a computer in some operating systems. So it is natural that it was gradually developed for IBM PC-type computers; It was given the appropriate name Windows, which profesionalų. Tačiau didžiuma žmonių vis dėlto labiau mėgsta jau minėtą grafinę sąsają, laiko refers directly to the main element of the interface – the graphical windows. By the mid-1990s, Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 3 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 Windows had grown into a standalone operating system, and since then the graphical user interface has become the main interface between humans and IBM PC-like computers. Both types of computers – both Apple Macintosh and IBM PC – the operation, in general, the work with both types of computers has become quite similar, although the fundamental "internal" differences remain. They were more accessible, cheaper, you could say, more democratic, and I didn’t have to install software Filologų aplinkoje, kaip ir daug kur Lietuvoje, labiau plito IBM PC tipo kompiuteriai. of various origins on them, first of all, the necessary applications, including even homemade ones, created by myself or ordered from known programmers. Alongside the rare original IBM PC-type computers, analogs installed in Lithuania, made in Southeast Asian countries or even from there-produced blocks, began to spread rapidly, which were called computers compatible with IBM PC… It should be noted that similar Macintosh imitations basically did not exist. Philologists, who used mainly IBM PC and compatible computers, rarely had major problems with correct, standard, Lithuanian spelling. Even the 1980s PC-like computers running on DOS (Disk Operating System) with which information was exchanged via a text terminal already used quite defined encoding schemes for the characters displayed on the screen, called code pages or code tables (CP). There were several of them, and they provided codes not only for the letters of the English alphabet, but also for the signs used in some other countries' writing systems. Along with the operating system, computer fonts were distributed – software components integrated into the operating system that allow characters to be displayed in visible, human-visible graphical images. The devices that can do this are two main ones: the display and the printer. A more problematic device that required more ingenious coordination with the computer was the printer of that time: a lot depended on its mode of operation, type, manufacturer, etc. But it is important to emphasize that both these devices do their work only under the control of the computer, and the necessary essential information is taken from computer fonts. Savaime suprantama, taisyklinga lietuviška rašyba nebuvo svarbiausias firmos Microsoft, IBM, which developed software for IBM PC compatible computers (and continues to develop it for many of today's computers). But, on the other hand, the basis of our current spelling – the Lithuanian alphabet – is neither original nor very unique, most of its letters are common, they are savo rašyboje vartoja ne tik lietuviai, bet ir kitos tautos. Tikrai lietuviškų, niekieno daugiau bent PETRAS SKIRMANTAS. Kol nebuvo Palemono: kompiuteriai, lietuviški rašmenys ir šriftai | 4 used in all languages. BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 The letters ą and ę are used in Polish, č, š and ž – in Czech, ū – in Latvian. However, the nine letters – ą, ę, į, ų, č, š, ž, ū, ė – are often distinguished from the other letters of the Lithuanian alphabet. And rightly so – all of them are double, binary, consisting of two components: a basic letter (a, e, i, u, c, s, z) and an auxiliary sign added at the bottom or top – a nose, a bird, a dash, a dot. All the other 23 of us are just one dot (dubious about such statement would be at least the small i and j, having a dot at the top, but abėcėlės raidės – a, b, c, d, e, f, g, h, i, y, j, k, l, m, n, o, p, r, s, t, u, v, z – yra viengubos, jas sudaro it is worth remembering that the large I and J do not have a dot, and both ė and Ė have it, so these letter pairs are different). It should also be noted that single letters are much more "international" than double ones, they are used in the writing systems of many European countries and all 23 are included in the English alphabet. It is very important that these two groups of letters, forming the encoding schemes of characters in computers, fell into different code groups. The single letters of our alphabet, which are also the letters of the English alphabet, were given codes first, and those that did not change later remained constant to this day. They are part of the so-called standard character set. ASCII (American Standard Code for Information Interchange) is a character set developed in the 1960s. It uses integer numbers from 0 to 127 for character encoding, and codes from 32 to 126 for script characters (letters, digits, tabs). This encoding scheme (sometimes called simply encoding) was developed even before the appearance of “small” computers, it was intended for various telecommunication devices, and computers simply took over and appropriated it when they appeared. However, with the proliferation of computers, the standard (ASCII) set of characters became insufficient, as it did not contain all the characters needed to write texts not only in English, but also in other languages. In addition to the 10-bit encoding, the 256-bit encoding is used to encode the entire number range from 0 to 255, which is used to encode the 16-bit encoding. The first 128 (0 to 127) codes regulated by ASCII remained unchanged, but additional 128 codes (numbers from 128 to 255) were added to encode other characters not included in the standard set. This extended codespace is often referred to as Extended ASCII by English speakers, but it is not an extension of ASCII (the standard or the encoding scheme regulated by ASCII) – it is an extension of the codeset to cover the unchanged ASCII encoding scheme and at the same time provide 128 additional codes for other necessary characters whose encoding is not regulated by ASCII. This resulted in a 256-code distribution model, dividing the entire code set into two equally large subsets – the fixed, regulated ASCII, covering codes from 0 to 127, and the free, covering codes from 128 to 192. Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 5 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 255. That model is exactly what the English description Extended ASCII would mean. And the free subset of codes established by it provided the prerequisites for the emergence of many different coding schemes on the basis of the model, called, as already mentioned, code sheets or code tables and markings of the most diverse markings. The constant subset of all of them, the constant part, was the same and coincided with ASCII, but the free part codes assigned to the characters varied greatly, depending on the type of character. rinkiniui kodlapis būdavo sudaromas. In the increasing variety of code sheets, there began to be code sheets in the free part of which were also encoded the composed (double) letters of the Lithuanian alphabet, more often – only those of them that are needed not only for our spelling. There are also code sheets for all letters of the Lithuanian alphabet; our computer engineers have also contributed to this. It is also possible to use a computer code to encode the text in the text editor. There was also confusion: in the first half of the last decade there were already several DOS operating system codepages with all the letters of the Lithuanian alphabet – their numbers were from 770 to 774, then added 775, 776, 777 and 778. The codes of Lithuanian letters in some of these codesheets coincided, in others they differed. With the proliferation of computer fonts representing the characters encoded in those codepages, it was necessary to be mindful and careful. It was also necessary to associate the various codes of the nine double letters of our alphabet, often called special or specific Lithuanian letters, with the keyboard, its corresponding keys (the interface of the codes of single letters covered by the standard ASCII character set with the keyboard for understandable reasons never caused any problems). But on the other hand, all this made it possible to write correctly in Lithuanian even on those computers that were controlled by the DOS operating system. The graphical user interface for IBM PC-like computers, Windows, which soon became a standalone operating system from a DOS add-on (it was the first to become Windows 95), fundamentally changed the display of characters on the screen, the position of computer fonts, and writing on computers in general. There are no signspaces left, other, new type, fonts, called vector fonts, which create the visual image of characters by means of vector computer graphics – Bezier curves marking the contours of characters. Inkjet and laser printers have become more common, which can reproduce these contours on paper very accurately, with the finest graphic details. As a result, texts and documents generated by computer and printed on such a printer have become similar to documents printed by printing houses. Many new types of fonts spread along with the Windows graphical environment, they were an integral part of it, but at the same time there was the possibility to supplement the "system" font set with other, separately purchased fonts if necessary. PETRAS SKIRMANTAS. Kol nebuvo Palemono: kompiuteriai, lietuviški rašmenys ir šriftai | 6 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 The encoding of signs has also begun to become established. Despite the diversity of local codesheets, universal international coding standards became increasingly important. Two of the most notable are ISO/IEC 8859 and ISO/IEC 10646. The first set out how to encode the characters required by the major European writing systems in code sheets with 256 codes each, and the second multiplied the number of integers that could be used as character codes by many times – it became more than two billion! Most of it is still unused, there are only about 140,000 characters encoded in this set of codes… Just next to it is another character encoding standard (not international) – Unicode. It was also designed to increase the number of character codes. The Unicode Cosorcium was established in early 1991, and the first volume of the Unicode standard was released in the autumn of 1991. According to the Unicode standard, the character codes can be 65536 integers – from 0 to 65535. Soon after, the authors of ISO/IEC 10646 and Unicode agreed to work in coordination, and the two standards were (and still are) harmonized, synchronized, and character encoding has since become completely identical, although the purposes and functions of each remain distinct. For various reasons and due to certain requirements, the total number of codes actually usable, commonly referred to as the code space, has decreased considerably and is currently slightly above 1,1 million. of which, as mentioned, about 140 thousand, or about 12.5 percent, are already assigned to signs. The combination of these two standards is commonly called unicode. The name is motivated by an effort to unify the encoding of any character used in computers, making it uniform worldwide. The above-mentioned character encoding standards are universal, oriented to the needs of information technology – they were not created only for IBM PC-type computers or their graphical interface. However, when developing computer fonts for different types and purposes of computers and for different operating systems, efforts are made to follow the character encoding prescribed by these standards as much as possible, because it is useful: unified character encoding ensures that written, character-transmitted information will be sent, received and reproduced accurately, regardless of the hardware or software used. The software manufacturer Microsoft, which created, improved and distributed the graphical user interface Windows, which evolved so rapidly that it, in a few years, literally swallowed up the DOS operating system and became itself an operating system, used the vector fonts distributed together with the Windows operating system, which differed from the DOS code sheets only in the free range of savus kodlapius, aprėpiančius po 256 ženklų kodus. Jie gavo savitus pavadinimus, bet vienas codes, i.e. the assignment of codes from 128 to 255 to characters. The constant, ASCII-regulated subset is also unchanged (by the way, it is unchanged in all parts of ISO/IEC 8859 and Unicode). In the early Windows versions, the code page 1252 was introduced. Later it was called CP-1252, Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 7 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 Windows-1252, and simply Latin 1. The number of code sheets increased, and so did the number of scripts based on both the Latin and other alphabets. In order to emphasize that they are not DOS codesheets, but Windows codesheets, they are sometimes referred to as ANSI codesheets, because the original version of the first one, CP-1252, was based on the original draft codesheet prepared by the American National Standards Institute (ANSI). This project later evolved into ISO/IEC 8859-1, but the ANSI name still became part of Microsoft's working terminology and became the collective name for Windows codepages. In addition, the code pages for DOS operating system also had a common group name – OEM. Code pages, as mentioned several times, are character encoding schemes that provide for which character to which code – an integer number, not greater than 255 – to be assigned. The code – the digital identifier of a mark in a computer system, its contractual recognition mark – is almost the same as the call sign for a radio station or the personal code for a Lithuanian citizen. In English-language computer literature, the code is now commonly referred to as a target, the successfully made word codepoint (sometimes spelled code point), thus indicating that it is a point in the set of codes, one of the numbers in that set. However, it should be emphasised that the code is not an immanent attribute of the mark, lying ‘inside’ it – it is selected and assigned to the mark freely, on the basis of a certain agreement. Consequently, the same mark, if it is included in different code sheets and/or different encoding standards, may in principle be given different codes, although this is undesirable and avoidable. Codes are not only needed to encode characters in computer fonts, they are primarily needed so that the information expressed in characters can circulate in a coordinated way within the computer system, so that the various software and hardware components of the computer that send and receive information can keep it unchanged, not distort it, not turn it into a flow of random signals, i.e. information noise, entropy. Kai kompiuteriai patobulėjo ir atsirado galimybė kodams naudoti daug kartų už 255 larger integers, the importance of the coda has somewhat faded. In Unicode, a standard that has greatly expanded the code space, the concept of code page is no longer used, different groups of characters are distinguished, related to the purpose and use of the characters, and they are called dažniau vadinamos blokais. Tačiau Microsoft nuosekliai laikosi vadinamojo suderinamumo iš scripts, and the code space is divided from the bottom up (sometimes called - historical compatibility) principle: the aim is that new, modern versions of Windows computers can also run old, used in the past programs. As a result, the current versions of Windows support dual character Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 8 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 encodings: both the modern, Unicode-compatible, and the legacy, corresponding to the previously created ANSI codesheets. The Windows API (Application Programming Interface) is a set of software functions in the system libraries, and many of its functions for working with characters and their sequences are available in two versions: A, for ANSI-coded characters, and W (wide), for Unicode-coded characters and their sequences. Since the correspondance between the ANSI codes assigned to the character of various codepages and the code assigned to Unicode is unambiguous, it is not difficult to change some codes "in both directions" to other ones - you just need to know which ANSI codepage to exchange them with. Characters are transcoded to ANSI codes when necessary due to some internal system peculiarities and because of the still existing old applications that do not recognize codes larger than 255. Modern text-based applications usually manipulate characters encoded in wide codes. Unicode character encoding is now the norm. Proper Lithuanian spelling has been ensured with the adoption of vector fonts. This is also the merit of our computer scientists. The time when computers, replacing typewriters, began to appear on the desks of our philologists, roughly coincided with the Windows 1.0 was the first version of Windows. True, normally and usefully that interface could work only in newer computers, having no less than Intel 80386 processor and VGA standard video adapter (video card) and display, but such were far from all computers purchased by philologists at that time. As Windows versions changed, the requirements for the computer hardware (so-called hardware) increased, or, more precisely, the opposite: as the hardware capabilities increased, the prerequisites for new, improved, graphical interface versions of Windows were created. Windows 3.1 already began to spread, and a little later – the font TimesLT prepared by Fotonija, which was Windows 3.11, ir jau būta vektorinių šriftų su visomis lietuviškomis raidėmis. Pirmiausia čia widely used and used by Lithuanians for a long time. His letter drawing imitates the extremely popular Times suit. Soon TimesLT counterparts appeared with letters drawings imitating different letter sets, encoding Lithuanian letters as well as Baltijos apvadas or Baltijos žiedas… It was nice, ir TimesLT. Ženklai šiuose šriftuose jau buvo koduojami pagal Lietuvos ir Latvijos kompiuterininkų tuomet kurtą bendrą kodlapį užmojo kupinu pavadinimu BalticRIM, atseit of course, that the abbreviation of the ministry was unexpectedly reflected in it. The ANSI part of tuomet veikusios ir informacijos technologijos reikalus kuravusios Ryšių ir informatikos the codesheet contained characters needed for the writing of most of the Baltic countries, but not included in the ASCII part. Of course, among them were nine distinct, double letters of ours. Starting with the Windows-95 PE version (PE – from English Pan European) this codepage was Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 9 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 actually legalized by Microsoft as the codepage Windows-1257 for the Baltic countries (called Windows Baltic, CP1257, CP1257 BalticRIM, etc.). It encodes Lithuanian letters and all Baltic letters with the same codes as were already established in Lithuania with TimesLT and other fonts corresponding to this encoding. This is how the Baltic Ring symbol set developed and became established. It has been incorporated into two more codesheets – 775 for the then still used DOS operating system and later, in 1998, the ISO/IEC 8859-13:1998 Information technology – 8-bit single-byte coded graphic character sets – Part 13: Latin alphabet No. 7, which is often referred to as BalticRIM, Baltic RIM or Baltic Rim. The character set is the same, and the encoding, except for quotation marks and apostrophes, is essentially the same as Windows-1257 and ISO/IEC 8859-13; the character encoding in OEM codesheet 775 differs significantly. Vector computer fonts have fundamentally changed the appearance of computer-generated text documents. They began to differ greatly from typewritten documents, or typefaces, and became more and more similar to prints printed by printing presses. Vector computer graphics used to create the form of writing made it possible to change the size of letters (cone) in a wide range according to the need, idea or desire without deforming their contours, it also became possible, with a suitable set of vector fonts, to change the brightness of letters of any size (the width of the lines forming them), to present the desired letters in italics, to manipulate the intervals between letters or their lines. People liked the expressive, polygraphic appearance of computer-produced documents, and philologists too, so their main tool for writing on computers almost immediately became and to this day remains one of the components of the Microsoft Office application package – the Microsoft Word text program, because it allows the mentioned vector fonts to be used relatively easily in practice. Even today there are philologists who cannot be convinced that tests can be written not only in Word, that in some cases it would be more rational to prepare them with other software tools. The iron argument – Word is the most beautiful... Lithuanian letters in computer fonts and the possibility to write Lithuanian text correctly without violating current spelling norms have satisfied and still satisfy many institutions preparing writing and many people for whom writing is a component, and often the main, for a screenwriter or a person in a similar profession, that possibility is often enough, nothing more than it is needed. However, philologists, especially linguists, use not only the usual letters of the Lithuanian alphabet in their written works, but also a lot of different signs, jų profesinės veiklos dalis, praktinius poreikius. Rašytojui, žurnalistui, literatūros kritikui, kino various scripts, sometimes Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 16 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 Fontra established itself for a long time, almost a decade – until 2004-2005, when it began to be displaced by a modern, fundamentally different Palemon. At the end of that millennium and the beginning of the new one, the number of books printed in phonetic script, using transcription marks, increased considerably, they were published one after the other, and Fontra probably had an influence on this. By the way, Professor A. Girdenis loved Fontra very much, he considered it to be a well-developed phonetic writing tool, the most suitable for printing transcribed dialect texts, and he wanted his texts to be printed in Fontra. A. Girdenis did not like Palemonas... Therefore, the second book of dialect texts prepared by the professor in 2008, Žemaičių dzūkai, was printed by Fontra, although by then it had already been replaced in many places by Palemonas. Only the third collection of Samogitian texts, Šiaurės žemaičių Skuodo zona, published after the professor’s death (†2011), was converted into Palemon... it was printed in 2013. The third set of fonts, which are used for the translation of special characters needed in linguistic work, is formed in the same way as the two previously described and is called IndoBalt. It is dedicated to peculiar characters and symbols found in works on Indo-European studies, Baltic studies and historical grammar. It greatly facilitated the preparation of the texts of Linguistics and Baltistics and was used in publishing other works of linguistics. Having three font sets – LT-Times, Fontra and IndoBalt – the Faculty of Philology of Vilnius University no longer feels the lack of special characters necessary for linguistic studies, especially since it is easy to make an auxiliary font, another Varia, if necessary... The Faculty of Philology almost did not publish writing heritage at that time, it was essentially dealt with by the Institute of Lithuanian Language, and thus the problems related to specific heritage scripts lay on its shoulders. They are solved in a similar way – by creating fonts with the necessary characters. This practice of “homemade” font making, reminiscent of the extremely popular “do-it-yourself” attitude of recent Soviet times, allowed to deal relatively successfully with the challenges of the preparation and publication of linguistic literature arising from the changes in printing technology caused by the development of computers. It became possible to prepare and publish linguistic works of almost any polygraphic complexity. However, the inconveniences also began to become apparent: the number of self-made fonts increased, the total number of special characters increased, the characters began to duplicate, repeat… Overseeing and controlling the ever-expanding farm of fonts and characters became a challenge. The recording of texts slowed down: there were no longer enough keys on the keyboard for all the characters, often I had to search through an ever-thickening folder of character tables in search of some stranger character… It was particularly problematic to send texts in which self-made “special purpose” fonts were used, transferred from computer to computer. After all, it was necessary that the computer was installed all inherited PETRAS SKIRMANTAS. Kol nebuvo Palemono: kompiuteriai, lietuviški rašmenys ir šriftai | 17 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 fonts used in the text. In fact, this meant that the text being prepared or prepared, the file of a computer document, had to be sent along with a set of fonts used in it, and that was wrong: who would want to install on his computer several or sometimes even a dozen uncertified fonts of unclear origin and quality just to be able to read another document correctly? And computer connections expanded, the Internet became more and more functional and smooth, electronic dictionaries began to appear both on media and for the Internet, and later – electronic books. Alongside traditional publications published as printed versions, electronic versions have begun to appear. Language publications also began to turn in this direction. And home-made specialized free “no taboo” encoding fonts, completely faithful to traditional publishing, began to become a great and extremely difficult to overcome obstacle. The software for working with texts has also undergone changes. The most popular and most common writing tool for philologists, Microsoft Word, since the eighth version, better known as Word 97, allowed the use of computer fonts with much more than 255 characters: they could contain hundreds and thousands of them, as foreseen by the then established, authoritative Unicode. The fonts of the self-described sets were not such, but the prospect of large fonts seemed attractive. After all, one such font could fit all the special characters needed for philologists' work! There would still be room for simple, traditional letters of various alphabets and their variants, because in one font of this format, although it is also not limitless, it is possible to put more than 65,5 thousand different characters. This perspective was made even more realistic at the turn of the millennium by Adobe's new print layout program InDesign, which also supported Unicode character encoding and Unicode-based fonts (all versions of the previously used layout program Adobe PageMaker, except the last, seventh, released after InDesign and soon disappeared, only recognized ANSI type codes). The idea of a single font containing all the signs necessary for linguistics was also pushed forward by the fact that computer-transcribed dialect texts were not only published in books, but also stored in computer repositories and databases: the abundance and variety of fonts with phonetic transcription signs began to turn into a mess, and the freedom of encoding signs led, frankly speaking, to chaos. This, of course, restricted the possibilities of developing sign sets, raised doubts about their functionality and usefulness. In short, the need for a new, multifunctional font, i.e. suitable for print, internet and databases buvo jaučiamas visur. The Lithuanian State Commission for the Language took the initiative to create such a font. In the autumn of 2002, the working group she set up to create a new font met for its first working meeting. And in autumn 2006, during the exhibition Infobalt 2006, the new font Palemonas, which has Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 18 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 been completed, tested and used for publishing books since 2004, was presented to the public. The group that created Palemonas also prepared and published in 2005 the book Lituanistinis šriftas Palemonas, the majority of which consists of specific technical information – lists of characters grouped by base letters with the unique name of each character and its coding, as well as schemes that help to quickly find the required character. Palemonas is perhaps the only Lithuanian computer font that has been documented in such detail. While working on Palemon, around 2003, a precursor appeared – a commercial font of the same format created by Fotonia, Aistika, it was publicly presented as the first Lithuanian computer font. Although it is doubtful (were the TimesLT and TimesLTAccented distributed by the same Fotonija not Lithuanian?), Aistika was still the first really successful work in this field, incorporating Lithuanian characters, not only ordinary, but also a considerable part of special ones, in a modern unicode encoding font. Aistika was intended for other products of Fotonija, its computer dictionaries, but at the same time allowed to write in both standard and pronounced Lithuanian or Russian letters, included Greek characters, a number of transcription marks and other symbols. However, all the philological publications did not have the necessary signs provided for in Aesthetics. Therefore, Palemon is still necessary even after the appearance of Aistic, especially since it was designed as a free, freely distributed font and remains such to this day. Both of these Unicode fonts are compatible with each other, and the encoding of all pronounced characters (even those that do not have Unicode codes) is also harmonized. Therefore, both plain and pronounced Lithuanian text compiled by one of them can be reproduced by another without distortion. After the creation of Palemon, the problem of special scripts needed for linguistic and Lithuanian studies was fundamentally solved. Received 15 December 2019 Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 19 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 BEFORE THERE WAS PALEMONAS: COMPUTERS, LITHUANIAN GRAPHEMES, AND FONTS Summary The article is an account of the difficulties that had to be overcome in preparing philological publications on Lithuanian linguistics, containing special graphemes (beyond the standard Lithuanian alphabet) in the last decade of the 20th century, when the typesetting technologies underwent radical changes with the advent of computer typesetting. An emphasis is placed on the importance and possibilities opened by the graphical interface in the IBM PCtype computers and the vector fonts that came with it. The topics discussed include character main focus of the article encoding schemes current at the time and the situation of the Lithuanian alphabet therein. The however is on the special linguistic characters used in Lithuanian philology (beyond the standard Lithuanian alphabet letters) and the technical challenges associated with them. The topics discussed include in-house-made special computer fonts and collections of fonts, used at the University of Vilnius Faculty of Philology at the time, for the publication of the periodicals Baltistica and Kalbotyra as well as linguistic books containing phonetically transcribed texts in Lithuanian dialects, with Lithuanian accentuation (syllable tone) marks, the so-called ‘Gerulian’ phonetic alphabet (i. e. phonetic transcription system based on the so-called Copenhagen Scheme of Transliteration and Phonetic Transcription and introduced into the study of Lithuanian dialects by a prominent scholar Jurgis Gerulis) as well as other characters necessary for this type of publications that were not available in the standard computer fonts of the time. It must be admitted that the sphere of usage of these in-housemade fonts was rather narrow and confined to publishing of linguistic texts on paper; thus, these sets of fonts were not a universal instrument of information exchange and this circumstance rendered them inferior to other, standard fonts. Difficulties arising from this non-universality and the proliferation of varieties of these privately made fonts are also discussed; the issue especially came into prominence when the time had come to publish linguistic texts on the internet and to assemble them into sortable and searchable databases. These difficulties and the new demands arising thereof are the main reason for creating a modern, universal computer font that should encompass all graphemes necessary for linguistics texts, especially those of Lithuanian linguistics. Peter Skirman. Until there was Palemon: computers, Lithuanian scripts and fonts | 20 doi.org/10.35321/bkalba.2019.92.09 BENDRINĖ KALBA 92 (2019) www.bendrinekalba.lt ISSN 2351-7204 KEYWORDS: font, special characters for linguistics, encoding, character set, code page, encoding standards, ASCII, ANSI, Unicode, accent notation in writing, phonetic transcription PETRAS SKIRMANTAS Lietuvių kalbos institutas Petro Vileišio g. 5, LT-10308 Vilnius, Lithuania [email protected]