scieee AI-readable full text Open interactive document viewer

HTR Input & Correction Manual

Meelen, Marieke; Griffiths, Rachael M.

Abstract

This is the Handwritten Text Recognition (HTR) input & correction manual developed by Marieke Meelen and Rachael Griffiths for the PaganTibet project. It provides instructions for annotators to transcribe and correct automatically transcribed Tibetan ume manuscripts. The front cover was designed by Sarah Teetor.PaganTibet is an ERC-funded Advanced Grant project (ERC, PaganTibet, 101097364) led by Professor Charles Ramble (PI) in the Horizon Europe framework, hosted by the École Pratique des Hautes Études (EPHE), PSL in Paris. Ⓒ Marieke Meelen & Rachael Griffiths (PaganTibet members), September 2025. This document is licensed under CC BY NC-SA 4.0.

Full text

HTR Input & Correction Manual pagantibet.com [email protected] Version 2.1 (public): 25 September 2025 Marieke Meelen Rachael Griffiths Funded by the European Union (ERC, Pagan Tibet, 101097364). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.” HTR input & correction manual Version 2.1 (public)1 - Marieke Meelen & Rachael Griffiths This manual is intended for those involved in the Handwritten Text Recognition (HTR) workflow of the ERC project "PaganTibet", in particular those focusing on the manual input and correction of transcriptions, i.e. the "Transcription Team" (TNI monks). Details on the project-internal side and backend of the system are not discussed here, but available elsewhere to the relevant project members and associates, i.e. the internal (ERC project) and external (BDRC & Esukhia) "HTR Teams". 1. Project goals & timeline 2. Who's who? 3. Getting started with HTR input & correction 3.1 Installing Firefox 3.2 Installing Chrome 3.3 Installing the Monlam keyboard 3.4 Getting a gmail account 3.5 Logging in to OpenPecha 3.6 Sign out and Log in again 4. HTR input and correction 4.1 Basic functionality 4.1.1 Manual input task - Team A 4.1.2 Correction task - Team B 4.1.3 Tasks for Team Leaders 4.2 Accuracy 4.3 Diplomatic transcriptions 4.3.1 Special characters? 4.3.2 Spelling mistakes? 4.3.3 Contractions? 4.3.4 Non-existing Tibetan Unicode stacks? 4.4 Different scripts 4.5 Gaps & stains 4.6 Unknown characters 5. Workflow and payment 1 PaganTibet is an ERC-funded Advanced Grant project (ERC, PaganTibet, 101097364) led by Professor Charles Ramble (PI) in the Horizon Europe framework, hosted by the École Pratique des Hautes Études (EPHE), PSL in Paris. Ⓒ Marieke Meelen & Rachael Griffiths (PaganTibet members), September 2025. This document is licensed under CC BYNC-SA 4.0. 2 Version control for shared PDFs Version 1.0 September 2023 (20230921) only sections 1-3 (MM) Version 1.1 October 2023 (20231002) first full draft (MM) Version 2.0 April 2024 (20240402) revised screenshots with new tool (RG+MM) Version 2.1 September 2025 (20250929) minor updates to timeline (RG) → first public version 3 1. Project goals & timeline The HTR tasks are part of the larger ERC project "PaganTibet" whose overall goals are to: 1. Document Leyu and Baima (and related) traditional rituals 2. Digitise ALL recordings and manuscripts → HTR part! 3. Create diplomatic & critical searchable versions of ALL texts 4. Add philological, linguistic, and other information to ALL texts 5. Create detailed, critical editions of selected texts 6. Identify features that characterise pre-Buddhist, Pagan tradition(s) 7. Use these features to reconstruct networks of Pagan tradition(s) The Handwritten Text Recognition (HTR) part highlighted in green is essential for the rest of the project, as the manuscripts all need to be digitised in a searchable way, i.e. e-texts need to be created of all images we have of the manuscripts before we can proceed with any other work steps. We originally had approximately 70,000 images to transcribe in ~9 months, but acquired more materials in the course of the year and now have access to over 100,000 images. To speed up the transcription task, we use Handwritten Text Recognition tools like Transkribus, which uses neural networks to create transcription models that can automatically transcribe our manuscripts. In order to create these models, we need to input transcriptions by hand. This is the "Manual Input" task, which Team A will focus on. The more manual input, the better the results, but the results are never 100% correct. Therefore, we also need to check the results and correct any mistakes. This is the "Correction" task, which Team B will focus on. Both tasks and the general workflow will be described in detail in section 4 below. Because the models won't have much input data in the beginning, the results will probably not be very good at first, so it will take more time to do corrections. We will therefore split our images up into smaller batches and focus on one batch per input & correction round. After every round, we will use the new input and the corrected transcriptions to improve the model to make it work better for the next round. The correction task should then become easier over time, which is good, because there is much more correction to do, so in the end both teams will spend a similar amount of time. Both tasks were completed in 12 months, starting 1 April 2024 until the final deadline of 30 April 2025. 4 2. Who's who? There are various people involved in HTR input & correction for the ERC project: ERC Project PI - Charles Ramble (CR) Internal HTR team - Marieke Meelen (MM) & Rachael Griffiths (RG) External HTR team - Élie Roux (BDRC), NT, Tashi Tsering & Tamdin (Esukhia) Transcription Team - Kemi Tsewang & TNI monks divided into two teams: - Team A 'Manual input': Sherab Chokgyal (Team Leader), Tsugphud Woeser, Tsultrim Palsang, Palgyi Wangchuk, Tsewang Drukgyal - Team B 'Correction': Tsultrim Gyaltsen (Team Leader), Tritsuk Lhundup, Tsukphud Rabsal, Tsognyi Gyatso, Sherap Woser If you have any questions, ask your Team Leader first. They can answer or pass the question on to our local contact in Kathmandu, Kemi Tsewang. Kemi will be working closely together with the internal HTR team (Marieke and Rachael) who will be on standby to help with any possible issues or they can pass on technical bugs to the external HTR team if they arise. 3. Getting started with HTR input & correction You only need the following to get started: 1. Any computer/laptop with Firefox OR Chrome with the Monlam keyboard and wifi access; 2. A Gmail account to log in to OpenPecha Tools, developed by Esukhia. 3.1 Installing Firefox Go to https://www.mozilla.org/en-GB/firefox/new/ to download Firefox and follow the instructions: 5 3.2 Installing Chrome Go to https://www.google.com/intl/en_uk/chrome/dr/download/ to download Chrome and follow the instructions: 3.3 Installing the Monlam keyboard For inputting new manual transcriptions and for correcting automatically generated transcriptions, you'll need to use the Monlam input keyboard, which works on both Windows and Unix operating systems. This keyboard is necessary as it is the only option for certain special symbols and abbreviations available in Tibetan Unicode. A full guide on the Monlam keyboard and how to install it can be found here: https://drive.google.com/file/d/1Mzzq2_l8Spi0Ntxs9eWCJl3u3Lie9CUS/view. A list of examples and how to use it, especially for contractions and other non-regular cases can be found in section 4.3 below. 3.4 Getting a Gmail account In order to use OpenPecha, where the HTR input & correction tool is hosted, you need a gmail account. Go to www.gmail.com and follow the instructions to create a free account: 6 Please let Kemi Tsewang know what your (new) Gmail address is and which team (AInput or B-Correction) you belong to so that they can grant you access to the OpenPecha Tools. If you don't know your team and/or have no preference for a team, we'll assign you to one. 3.5 Logging in to OpenPecha Once your account is activated, go to https://pecha.tools/ and click on "LOG IN": 7 Then click on the blue Google button and sign in with your Gmail account: Check if your Google image/icon/initials appear on the top right corner and then click on "HTR": 8 Check if your “user” and “group” names are displayed correctly in the top left panel after "PROJECT INFO": If all this is correct, you can now start your HTR task. If you don't see your user name and/or group name, please go back to the log in stage and make sure you are logged in with the correct email. 3.6 Sign out and log in again If you have finished all your tasks and/or the system says 'No tasks assigned', please sign out by clicking on the circle on the top right corner and then 'Sign out'. You will get back to the first page and if you log in again you will get new tasks: 15 member of the wider HTR team) to miss any mistakes, which is why we have everything checked more than once. The Team Leader working on Task A will see an image of an entire page with a transcription of the page in the text box underneath. The Team Leader needs to check this transcription line by line to make sure it is 100% accurate. If there are many mistakes, the Team Leader should make a note of the type of mistakes and click on the red cross button to reject the transcription. If there are only 1 or 2 small mistakes, they should correct these and then when the transcription is accurate, click on the green button to submit. The Team Leader working on Task B will see an image of one line with a transcription in the text box underneath. The Team Leader needs to check this transcription to make sure it is 100% accurate. If there are many mistakes, the Team Leader should make a note of the type of mistakes and click on the red cross button to reject the transcription. If there are only 1 or 2 small mistakes, they should correct these and then when the transcription is accurate, click on the green button to submit. 16 For both Team Leaders, new images and transcriptions will automatically appear after clicking the green or red button. Note that some may appear twice with or without a changed transcription. Once you've submitted, you will also see the numbers in the target process on the left panel change, keeping track of your work since you logged in today, your total amount of work and how many transcriptions you accepted or rejected: 4.2 Accuracy For both tasks it is extremely important to be precise and accurate: the transcribed text should reflect exactly what it says in the manuscript and nothing more/less. Even if you think there is a mistake in the manuscript, you should NOT fix the mistake, but instead, transcribe the syllable exactly as it is written in the manuscript WITH the mistake. For the project it is very important that we create eTexts that are exactly like the manuscripts. The sections below show how to deal with spelling mistakes, contractions, and non- 17 legitimate stacks of Tibetan characters that cannot easily be represented with the Tibetan Unicode. Note that if there are too many mistakes in your transcription, your Team Leader cannot approve it. If this is the case, your Team Leader can give you feedback on what you did wrong so you can do it correctly next time. If there are only some minor mistakes in your transcription, your Team Leader can decide to fix these and approve it. It is important for Team Leaders to make a note of any issues or regular mistakes to help their team members make their work as accurate as possible. This will mean team members can learn more quickly, which ultimately means they and the Team Leader will be able to have more approved transcriptions. 4.3 Diplomatic transcriptions For the project it is of crucial importance to get transcriptions that represent exactly what we see in the manuscript. This type of transcription is called a "diplomatic transcription". In diplomatic transcriptions, we do NOT try to fix or improve the text in the manuscript in any way. For example, if we find a common word like !ལ་པོ་ 'king' with a missing naro like !ལ་པ་, then we should transcribe this exactly like the manuscript, so without the naro: !ལ་པ་. The same goes for other spellings that may look strange or contracted forms that cannot be represented well in Tibetan Unicode: all of these need to look exactly like (or as close as possible to) the original in the manuscript. The sections below will give more examples of Special characters, Spelling mistakes, Contractions and Non-existing Tibetan Unicode stacks. Section 4.3.1. explains how to type these special characters using the Monlam Keyboard. Here's another example: The accurate way to transcribe this would be དབ(ི་ (Wylie dbalyi). We know that this is an unusual spelling or mistake for དབལ་!ི་ (Wylie dbal gyi), but we still need to transcribe it as དབ(ི་ to be 100% accurate in rendering exactly what the manuscript contains. At a later stage of the project, we will create a "normalised transcription" in addition to this diplomatic transcription. In the later "normalised transcription", we will expand contractions and regularise spelling etc. to create e-texts that are easy to search and use for other purposes. 18 4.3.1 Special characters? Most standard Tibetan Unicode characters are found on the default Monlam keyboard: Sometimes, less frequent or less standard characters are needed, however. These can be found on the Monlam keyboard by pressing either SHIFT or CAPSLOCK. The combinations with SHIFT keys look like the ones next to the capital letters, e.g. k = ཀ but K = ཁ : 19 A further set of special characters can be found pressing CAPSLOCK: Note that the above images are not exhaustive and can sometimes differ slightly depending on your operating system. You can always find the exact keyboard layout by letting your computer show the keyboard viewer: Clicking on 'Show Keyboard Viewer' will give you a default keyboard overview. If you press SHIFT, you will see the default keyboard change to the SHIFT keyboard options. Similarly, if you press CAPSLOCK, you will see the options for the CAPSLOCK key combinations. 20 On a Macbook, for example, SHIFT + 4 will yield the non-standard tsha-rtags as exemplified in ཞོ◌༹འ་ in section 4.3.4 below. A full list of examples and use cases of the Monlam keyboard with detailed explanations can be found in the Monlam keyboard manual. 4.3.2 Spelling mistakes? Since we are aiming for a diplomatic transcription, it is important to keep all spelling mistakes or unfamiliar spellings exactly as they are in the manuscript. Here are some sample cases of odd spellings and unusual symbols with their transcriptions: Note that in some cases like དཾ་, we need additional characters that may not be in the standard version of the Monlam keyboard. Section 4.3.1 above explains how to deal with these and other types of special characters; Tibetan Unicode stacks that are nonconventional are discussed in section 4.3.4. 21 4.3.3 Contractions? Similar to the odd spellings, we need to keep contracted syllables and forms in exactly the same way they appear in the manuscript: If you are unsure of how to type a contraction, there is a built-in Google sheet to look up how to type exactly what you see in the manuscript. You can access this sheet in either Task A or Task B, by clicking on the three lines (the 'hamburger icon') on the right side: When you click on this, a Google sheet will pop up in the bottom right corner with a list of images of contractions in the first column, a list of conventions for how to transcribe these in a 100% accurate way in the green middle column and, finally, a list of expanded versions of these contractions. You can scroll down this alphabetical list to find the contraction you need and then you can look at the green conventions column to learn how to type it: 22 Both annotators and Team Leaders can use this google sheet to look up conventions for contracted and non-standard forms. You can hide this Google sheet again if you like, by clicking on the red exit cross on the right: If the line or page in the manuscript you're transcribing contains the following sequence of characters, for example, then you know this is not a standard Tibetan syllable, because there are multiple vowels, not enough tshegs and there are stacks of characters that do not exist in normal Tibetan script: 23 If you don't know how to type this using the Monlam keyboard (see section 4.3.1 above) and/or if you want to make sure that you are using the correct convention for transcribing this, you look it up in the Google sheet and then you type what you find in the green 'Convention' column: If the contraction you're seeing in the image is not in this list, you should create a miniature screen shot of the syllable and add it in a new row in the shared TNI Google sheet alongside a potential convention for transcribing it and an expanded version of the contraction (if you know how to - otherwise leave these green and red columns blank, as well as the name of the batch where you found this and your own name, as shown in these two sample cases: 24 4.3.4 Non-legitimate Tibetan Unicode stacks? Non-existing stacks of consonants and/or vowels work in the same way as the conventions, spelling mistakes, and special characters above. You may need to use the SHIFT and/or CAPSLOCK keys on the Monlam keyboard to find certain elements like the tsha-rtags: 4.4 Different scripts In some images different types of scripts are used. This is often the case in ritual texts, for example, where instructions etc. are different from content. These smaller scripts or scripts in different colours should be transcribed in exactly the same way as other parts of the text. In the line above we see three different scripts, for example: 1. the regular script 2. a smaller version of the regular script 3. the regular script in red ink For the purposes of our project, these different types don’t need to be preserved, so all three types of script in the above line can be transcribed in the same way. Some texts contain words and/or lines in Zhangzhung scripts. These can be marked with an ྾. On the Monlam keyboard, you find the ྾ by typing SHIFT + 6.