scieee AI-readable full text Open interactive document viewer

Automatic Evaluation of Open-Ended Questions in MOOCs using a Named-Entity Recognition (NER) System based on the Hidden Markov Model

De Almeida, G. M.; Naukkarinen, J.; Li, C.; Matthews, S.; Jantunen, T.; Datta, S.; Kuparinen, K.; Alobaid, F.; Vakkilainen, E.

Abstract

This work investigates the automatic evaluation of open-ended questions in Massive Open Online Courses (MOOCs) for higher education. The challenge of dealing with natural language makes closed-ended questions commonly used. On the other hand, spontaneous human expression can generate additional useful feedback for teachers and students. We use a Named Entity Recognition (NER) system based on the Hidden Markov Model (HMM) statistical technique for this purpose. The data comes from a MOOC pilot test conducted with 179 engineering bachelor's students. The performance of the HMM-NER system was above 95% for both Precision and Recall metrics. In practice, all student responses in the test data were assessed correctly and confidently. Human-based steps in this process have proven to be crucial for a successful application. This highlights the role of the teacher, as technology alone cannot provide satisfactory solutions.

Full text

Practice Paper Recommended citation: De Almeida, G. M., Naukkarinen, J., Li, C., Matthews, S., Jantunen, T., Datta, S., Kuparinen, K., Alobaid, F., & Vakkilainen, E. (2025). Automatic Evaluation of Open-Ended Questions in MOOCs using a Named-Entity Recognition (NER) System based on the Hidden Markov Model. In Kangaslampi, R., Langie, G., JΓ€rvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631544. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License. AUTOMATIC EVALUATION OF OPEN-ENDED QUESTIONS IN MOOCs USING A NAMED-ENTITY RECOGNITION (NER) SYSTEM BASED ON THE HIDDEN MARKOV MODEL GM de Almeida a,1, J Naukkarinen b, C Li c, S Matthews d, T Jantunen e, S Datta f, K Kuparinen g, F Alobaid h, E Vakkilainen i a LUT University, Lappeenranta, Finland, 0000-0002-2898-5177 b LUT University, Lappeenranta, Finland, 0000-0001-6029-5515 c LUT University, Lappeenranta, Finland, 0000-0001-9400-3998 d LUT University, Lappeenranta, Finland e Lappeenranta City Hall, Lappeenranta, Finland f Digiotouch, Tallinn, Estonia, 0000-0002-2239-2194 g LUT University, Lappeenranta, Finland, 0000-0002-3373-8234 h LUT University, Lappeenranta, Finland, 0000-0003-1221-3567 i LUT University, Lappeenranta, Finland, 0000-0002-7472-3522 Conference Key Areas: Digital tools and AI in engineering education, Open and online education for engineers Keywords: MOOC, Open-ended question, Automatic evaluation, Named-Entity Recognition (NER), Hidden Markov Model (HMM) ABSTRACT This work investigates the automatic evaluation of open-ended questions in Massive Open Online Courses (MOOCs) for higher education. The challenge of dealing with natural language makes closed-ended questions commonly used. On the other hand, spontaneous human expression can generate additional useful feedback for teachers and students. We use a Named Entity Recognition (NER) system based on the Hidden Markov Model (HMM) statistical technique for this purpose. The data comes from a MOOC pilot test conducted with 179 engineering bachelor’s students. The performance of the HMM-NER system was above 95% for both Precision and Recall metrics. In practice, all student responses in the test data were assessed correctly and confidently. Human-based steps in this process have proven to be crucial for a successful application. This highlights the role of the teacher, as technology alone cannot provide satisfactory solutions. 1Corresponding Author GM de Almeida [email protected] 1 INTRODUCTION Mass Open Online Courses (MOOCs) have seen an exponential increase in higher education in recent years (Moore & Blackmon, 2022). One characteristic of this digital learning modality is the usual need for automatic assessment of student tasks. This issue often leads to the use of closed-ended questions, e.g. multiple-choice. The opportunity to use open-ended questions, such as essays, can allow for the exploration of other aspects of the teaching-learning environment (Gobbo et al., 2023). This can provide additional insights and feedback on actual learning for both teachers and students. The question that arises is how to perform the necessary automatic assessment in MOOCs, considering natural language-based responses. The answer lies in the computational area of Natural Language Processing (NLP) (Jurafsky & Martin, 2008), the branch of AI that deals with natural language data. This work explores Named Entity Recognition (NER) (Chinchor & Robinson, 1997), a type of NLP approach, for the automatic assessment of open-ended questions in the context of MOOCs. The NER system is based on the Hidden Markov Model (HMM) statistical technique, commonly used for e.g. speech recognition (Rabiner, 1990). To our knowledge, HMM-NER systems have not been used for this purpose in MOOCs. 2 CONTEXT AND PRACTICAL WORK In a recent work (Almeida et al., 2025), we proposed a framework for evaluating open-ended questions. In short, a pilot test was conducted using a developed MOOC with 179 engineering bachelor’s students. The framework was demonstrated using one of its open questions (Figure 1). First, the question is divided into parts. In this case, Part 1 ("πΆπ‘ƒπ‘†π‘†π‘‘π‘Ÿπ‘’π‘π‘‘π‘’π‘Ÿπ‘’") and Part 2 ("πΆπ‘ƒπ‘†π΄π‘π‘π‘™π‘–π‘π‘Žπ‘‘π‘–π‘œπ‘›"). Second, each part is subdivided into Sub-Questions (SQ), related to basic concepts. For Part 1: SQ1 ("π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘"), SQ2 ("𝐼𝐢𝑇"; Information and Communication Technology) and SQ3 ("π·π‘–π‘”π‘–π‘‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘"), and for Part 2: SQ1 ("πΊπ‘œπ‘Žπ‘™") and SQ2 ("πΊπ‘Žπ‘–π‘›"). For example, SQ1 refers to whether there is any mention of the "π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘" in the student’s response. The expected answer considered Bloom's Taxonomy with a matrix of rubrics (Anderson & Krathwohl, 2000), aligned with the Intended Learning Outcomes (ILOs). A positive answer scores SQ1 as β€œCorrect”; otherwise, β€œIncorrect”. This is applied to all sub-questions. Next, Part 1 is considered β€œComplete” if all of its sub-questions (SQ1-SQ3) were considered β€œCorrect”. Otherwise, β€œPartially Correct” if only some of the sub-questions, or β€œInsufficient” if none of them. This is applied to all parts. In the end, if all parts were considered β€œComplete”, the question is considered β€œCorrect”. Otherwise, β€œPartially Correct” if only some of the parts, or β€œIncorrect” if none of them. This work investigates the automatic evaluation of open questions, based on this previous framework, using the same question as a case study. Example: Sub-Question (SQ): β€œCorrect” β€œIncorrect” β€œCorrect” β€œCorrect” β€œCorrect” Part: β€œIncomplete” β€œComplete” Final assessment: β€œPartially Correct” Fig. 1. Framework for automatic evaluation of open-ended questions (previous work) 3 CONCEPTS 3.1 Named Entity Recognition (NER) Named Entity Recognition (NER) is used to locate and classify named entities mentioned, for example, in textual data (Chinchor & Robinson, 1997). For example, in the sentence: "π‘‡β„Žπ‘’π‘’π‘›π‘–π‘£π‘’π‘Ÿπ‘ π‘–π‘‘π‘¦ β„Žπ‘Žπ‘ π‘π‘’π‘’π‘›π‘€π‘œπ‘Ÿπ‘˜π‘–π‘›π‘”π‘œπ‘›πΏπ‘’π‘Žπ‘Ÿπ‘›π‘–π‘›π‘”π΄π‘›π‘Žπ‘™π‘¦π‘–π‘π‘ .", "π‘’π‘›π‘–π‘£π‘’π‘Ÿπ‘ π‘–π‘‘π‘¦" is an example of a named (mentioned) entity belonging to the entity tag (category): "πΏπ‘œπ‘π‘Žπ‘‘π‘–π‘œπ‘›". NER has been widely used in many Natural Language Processing (NLP) applications, for example, information extraction, information retrieval, machine translation, and sentiment analysis (Goyal et al, 2018). In the present work, the set of entity tags is given by {"π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘","𝐼𝐢𝑇", "π·π‘–π‘”π‘–π‘‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘","πΊπ‘œπ‘Žπ‘™","πΊπ‘Žπ‘–π‘›"}, corresponding to the five sub-questions (SQ1 to SQ5) of the case study question (section 2). If a named entity is identified for each sub-question, the student’s answer is considered β€œCorrect” (Figure 1). 3.2 Hidden Markov Model (HMM) based NER System Hidden Markov Model (HMM) is a statistical model for sequential pattern recognition (Rabiner, 1990). It can be seen as an extension of Markov chains, since in addition to the probabilistic transition between the states (of the Markov chain), the relationship between states and emissions (system outputs) is also probabilistic (that is, hidden). The choice of HMM in this work was precisely due to its explicit probabilistic modeling that favors interpretability, essential for feedback purposes to students and teachers. The initial applications were in speech recognition around the 1970s. Since then, it has been applied in many other areas such as human activity recognition, bioinformatics, and network analysis (Mor et al., 2021). More information about the theory of HMMs can be found, for example, in Rabiner (1990). A discrete HMM is defined by five elements (𝑆,𝑂,πœ‹,𝐴,𝐡), as shown in Table 1 (1st and 2nd columns). The 3rd column contextualizes them in NER applications using the case study of this work. Since the system outputs in this case are given by a sequence of words (student response), which are discrete symbols belonging to a finite set (vocabulary), the matrix 𝐡 refers to discrete probability distributions, which characterizes a discrete HMM. Table 1. Elements of a discrete HMM in the context of a NER application Element Description An e xample in this work States ( 𝑆 ) Each state of the Markov chain represents an entity tag. 𝑆 =  " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " , " 𝐼𝐢𝑇 " , … " π·π‘–π‘”π‘–π‘‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " , " πΊπ‘œπ‘Žπ‘™ " , " πΊπ‘Žπ‘–π‘› "  Observations ( 𝑂 )Sequence of words (student response). 𝑂 =  " 𝐴𝑛 " , " 𝑒π‘₯π‘Žπ‘šπ‘π‘™π‘’ " , " π‘œπ‘“ " , " π‘Ž " , " π‘π‘Ÿπ‘œπ‘π‘’π‘ π‘  " , … " 𝑖𝑠 " , " 𝑑 β„Ž 𝑒 " , " 𝑝𝑒𝑙𝑝 " , " π‘π‘Ÿπ‘œπ‘π‘’π‘ π‘  "  Transition probability matrix ( 𝐴 ) Probability of moving from one state to another. 𝐴 ( 𝑖 , 𝑗 ) = 𝑃 ( 𝑆 𝑑 = 𝑗 | 𝑆 𝑑 βˆ’ 1 = 𝑖 ) 𝐴 ( " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " , " 𝐼𝐢𝑇 " ) = 𝑃 ( " 𝐼𝐢𝑇 " | " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " ) Emission probability matrix ( 𝐡 ) Probability of a state generating (that is, emitting) a specific observation (word). 𝐡 ( 𝑖 , 𝑀 ) = 𝑃 ( π‘Š 𝑑 = 𝑀 | 𝑆 𝑑 = 𝑖 ) 𝐡 ( " 𝐼𝐢𝑇 " , " π‘‘π‘–π‘”π‘–π‘‘π‘Žπ‘™ " ) = 𝑃 ( " π‘‘π‘–π‘”π‘–π‘‘π‘Žπ‘™ " | " 𝐼𝐢𝑇 " ) Initial probability vector (  ) Probability of a state occurring at the beginning of a sentence. πœ‹ ( 𝑖 ) = 𝑃 ( 𝑆 1 = 𝑖 ) πœ‹ ( " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " ) = 𝑃 ( " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " ) One use of HMM as a generative model is to seek to reveal the underlying process (sequence of states) responsible for generating the system outputs (sequence of observations). In this work, the observations are given by a sequence of words (𝑂= {π‘ π‘‘π‘’π‘‘π‘’π‘›π‘‘π‘Ÿπ‘’π‘ π‘π‘œπ‘›π‘ π‘’}) and each state is associated with an entity tag in particular (𝑆= {"π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘","𝐼𝐢𝑇","π·π‘–π‘”π‘–π‘‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘","πΊπ‘œπ‘Žπ‘™","πΊπ‘Žπ‘–π‘›"}). Considering the objective of this work to automatically evaluate open-ended questions in MOOCs, an HMM (  ,𝐴,𝐡) will compute a sequence of entity tags (model output) for a sequence of words (model input). In other words, the idea with an HMM-based NER system is to find the most likely sequence of entity tags given a sequence of words (student response), with one entity tag for each word in the sequence. That is, in the end, each word in the student's response is classified into an entity tag. Table 2 (1st and 2nd columns) shows an example of the input-output association of an HMM in the context of an HMM-based NER system. The resulting NER tag sequence can be directly used for automatic evaluation of student responses using the proposed framework for automatic evaluation of open-ended questions (section 2). The term "𝑂𝑒𝑑𝑠𝑖𝑑𝑒" is used when a word does not match any previously defined entity tag. Table 2. Input-output association in the context of an HMM-based NER system HMM input (in original values, that is, words) HMM output HMM input (in index values, one for each word in the 1 st column of this table) Sequence of words Sequence of entity tags S equence of indices " 𝐴𝑛 " " 𝑒π‘₯π‘Žπ‘šπ‘π‘™π‘’ " "π‘œπ‘“" "π‘Ž" "π‘π‘Ÿπ‘œπ‘π‘’π‘ π‘ " "𝑖𝑠" "π‘‘β„Žπ‘’" " 𝑝𝑒𝑙𝑝 " " π‘π‘Ÿπ‘œπ‘π‘’π‘ π‘  " " 𝑂𝑒𝑑𝑠𝑖𝑑𝑒 " " 𝑂𝑒𝑑𝑠𝑖𝑑𝑒 " "𝑂𝑒𝑑𝑠𝑖𝑑𝑒" "𝑂𝑒𝑑𝑠𝑖𝑑𝑒" "π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘" "𝑂𝑒𝑑𝑠𝑖𝑑𝑒" "𝑂𝑒𝑑𝑠𝑖𝑑𝑒" " 𝑂𝑒𝑑𝑠𝑖𝑑𝑒 " " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " 1 2 3 4 5 6 7 8 5 4 METHODOLOGY Figure 2 shows the methodology adopted in this work, where the starting point is the original set of student responses. There are four main tasks. (1) Text preparation cleans the original textual data using word normalization, punctuation removal, and word lowercasing. (2) Reference definition manually evaluates all student responses for use as a reference and then establishes sets of named entities (that is, vocabularies, one for each sub-question; section 2) that are used to generate the sequences of entity tags for the sequences of words (cleaned student responses). In this sense, an additional state for words not associated with any entity tag is categorized as "𝑂𝑒𝑑𝑠𝑖𝑑𝑒" (that is, outside of any vocabulary). (3) Model training obtains the discrete HMM model, that is, the set of parameters ((  ,𝐴,𝐡); Table 1) through an iterative estimation process using the Baum-Welch algorithm (Rabiner, 1990). Using the hold-out approach, 85% of the student responses are used for this purpose, after a word-to-index conversion. (4) Model testing uses the remaining 15% of the student responses. Words not present in the training data and in turn in the emission probability matrix (𝐡) need to be handled in some way (shown in Table 3). The resulting trained HMM-based NER system is then used to generate the most likely sequences of entity tags for the input sequences of words (student responses in the test data not used during model training). This computation usually employs the Viterbi algorithm (Rabiner, 1990). The obtained result is compared to the target sequences of entity tags to evaluate the quality of the system using the confusion matrix and derived metrics. Once validated, the HMM-based NER system can be used for automatic evaluation of open-ended questions, as proposed in this work. Fig. 2. Methodology steps. Table 3 lists the set of hyperparameters that are combined to obtain candidates for the HMM-based NER system. Eight candidates are generated, of which one is selected by performance comparison using the test data. The first hyperparameter refers to using representatives of all words (provided by the text corpus, that is, all student responses) in the training data. This is similar to the handling of minimum and maximum values in prediction problems in machine learning. Second, two approaches were used to deal with unseen words, that is, words that appear in the test data but not in the training data. Laplace Smoothing addresses the counting problem, more specifically, the zero frequency (probability) issue. In the context of NER, a general observation called "π‘’π‘›π‘˜π‘›π‘œπ‘€π‘›" is created and given a small probability value, being included in the emission probability matrix (𝐡). When an unseen word appears, it is then assigned as "π‘’π‘›π‘˜π‘›π‘œπ‘€π‘›".Word Embedding maps words to vectors. A pre-trained model provided by the fastText library (Mikolov et al., 2018) was adopted, which returns a 300-dimension real vector. The "π‘’π‘›π‘˜π‘›π‘œπ‘€π‘›" word vector is then compared to all word vectors in the vocabularies and replaced by the closest word, considering semantic relations. Regarding the third hyperparameter, two distance measures were used for such comparison, Euclidean distance and cosine similarity. The Matlab computing environment was used in this work. Reference definition Table 3. Hyperparameter set explored during HMM training Task Hyperparameter Value [1] Model training Presence of representatives of all words in the training data (1) N o verification (2) At least one representative of each vocabulary word (named entity) is present in the training data [2] Model training (in case of (1)) and Model testing How to deal with unseen words, that is, words that are not present in the predefined vocabularies (1) Laplace smoothing (2) Word embeddings [3] Model testing (given the use of word embedding s in [2] ) Distance metric between unseen words (test data) and vocabulary words (training data) (1) Euclidean distance (2) Cosine similarity 5 RESULTSAND INSIGHTS 5.1 Text preparation The pilot test (section 2) allowed students to choose between different learning paths. Of the total of 179 students who carried out the activity, 97 answered the question used as a case study in this work (Figure 1). This step had the main objective of reducing the words to a root form, so that the inflected words could be analysed as a single term. We adopted the process of lemmatization, which reduces the words to their dictionary forms e.g. "𝑠𝑒𝑛𝑑" to "𝑠𝑒𝑛𝑑" and "π‘Žπ‘π‘π‘œπ‘Ÿπ‘‘π‘–π‘›π‘”" to "π‘Žπ‘π‘π‘œπ‘Ÿπ‘‘". 5.2 Reference definition First, all 97 students’ responses were manually scored according to the framework shown in Figure 1. This human-based evaluation was used as a reference labelling to assess the quality of the HMM-NER system. This initial analysis was then used to construct five vocabularies, one for each sub-question (SQ1-SQ5, which refers to {"π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘","𝐼𝐢𝑇","π·π‘–π‘”π‘–π‘‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘","πΊπ‘œπ‘Žπ‘™","πΊπ‘Žπ‘–π‘›"}, respectively) (section 2 and Figure 1). For example, the vocabulary for the entity tag: "π‘ƒβ„Žπ‘¦π‘ π‘–π‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘", contains twelve words: {"π‘Žπ‘ π‘ π‘’π‘šπ‘π‘™π‘¦","π‘’π‘žπ‘’π‘–π‘π‘šπ‘’π‘›π‘‘","π‘“π‘Žπ‘π‘–π‘™π‘–π‘‘π‘¦",…,"π‘‘π‘œπ‘œπ‘™"}. Both tasks are crucial and completely teacher-dependent and their settings should be aligned with the Intended Learning Outcomes (ILO). In this work, we consider the first level of Bloom’s taxonomy: β€œRemember”, combined with a set of rubrics: {β€œCorrect”, β€œPartially Correct”, β€œIncorrect”}. The set of sub-questions defines the set of entity tags. In this way, we construct the sequences of entity tags, which will be used as input information to train the HMM models along with the sequences of words (student responses). Table 2 shows an example of a sequence of entity tags (2nd column) derived from a sequence of words (1st column). 5.3 Model training The training dataset consists of 82 (85% of 97) student responses of varying lengths. Given the discrete nature of the HMM model, the words were converted to indices (3rd column of Table 2). As an illustration, the word "π‘π‘Ÿπ‘œπ‘π‘’π‘ π‘ " that appears twice receives the same index (β€œ5”). Then, both the integer-valued word sequences and the corresponding entity tag sequences are ready for the model training phase. Eight candidate HMM models are obtained by varying the hyperparameter set (Table 3). 5.4 Model testing This section presents the results for the 15 student responses of varying lengths in the test dataset. It is crucial to check the behaviour of a model with data not used during its training. Table 4 summarizes the results for all eight candidate HMM-based NER systems (hyperparameter values specified by the gray cells). Values refer to averages of the ten runs for each candidate model. Two usual metrics are shown, Precision and Recall, which are calculated based on the confusion matrix. Given the class imbalance issue, where classes are given by entity tags with different sizes, the overall accuracy metric is not appropriate. Precision [=𝑇𝑃 (𝑇𝑃+𝐹𝑃) ⁄] refers to the rate of true positives out of all positive predictions and Recall [=𝑇𝑃 (𝑇𝑃+𝐹𝑁) ⁄] refers to the rate of positive predictions out of all true positives, where 𝑇𝑃,𝐹𝑃, and 𝐹𝑁 are the amounts of true positives, false positives, and false negatives, respectively. The range of both metrics is [0,1] and the closer to 1, the better. Values below 0.95 are shown in red. Recall values are generally close to 1, meaning that true entity tags are usually classified correctly, as desired. However, this is not the case for Precision, whose lowest value is 0.32 (model 5 and entity tag "𝐼𝐢𝑇") meaning a relatively higher rate of false positives. That is, a portion of the predictions in one entity tag actually belong to another entity tag. For example, considering model 5, of all the "𝐼𝐢𝑇" predictions, only 32% are actually true. This also occurs for "π·π‘–π‘”π‘–π‘‘π‘Žπ‘™π‘Šπ‘œπ‘Ÿπ‘™π‘‘". Unlike the Recall metric, for Precision, there are improvements in the model (i.e., reduction in the false positive rate) for some hyperparameter values. The best result was given by model 7 (bold values in Table 4), which uses [1] representatives of all words in the training data, [2] word embeddings to deal with unseen words (the most significant factor among all three), and [3] the Euclidean distance metric to select the representative word for unseen words. All values for the Precision and Recall metrics in this case are considerably high, above 0.95, as desired. An example of using the word embedding approach is given by replacing the unseen word "𝑏𝑒𝑖𝑙𝑑" (which appeared in the test data but not in the training data) with the vocabulary word "π‘‘π‘’π‘£π‘’π‘™π‘œπ‘" (present in the training data) in one of the model runs. Other examples are "π‘ π‘’π‘π‘’π‘Ÿπ‘–π‘‘π‘¦" with "π‘ π‘Žπ‘“π‘’π‘‘π‘¦" and "π‘π‘Žπ‘‘β„Ž" with "π‘Ÿπ‘œπ‘’π‘‘π‘’". The Macro-Precision and Macro-Recall metrics are also shown, which summarize the corresponding individual values into a single one. For example, Macro-Precision is given by βˆ‘π‘ƒπ‘Ÿπ‘’π‘π‘–π‘ π‘–π‘œπ‘›π‘–π‘– π‘›π‘’π‘šπ‘π‘’π‘Ÿπ‘œπ‘“π‘’π‘›π‘‘π‘–π‘‘π‘¦π‘‘π‘Žπ‘”π‘  ⁄, with 𝑖=1,2,…, π‘›π‘’π‘šπ‘π‘’π‘Ÿπ‘œπ‘“π‘’π‘›π‘‘π‘–π‘‘π‘¦π‘‘π‘Žπ‘”π‘ . Both metrics are close to 1 for model 7, as desired. Therefore, the final HMM-NER system (model 7) was able to address both issues, namely, correctly predicting entity tags (high Recall values) and at the same time, given an entity tag prediction, ensure high confidence that it is correct (high Precision values). Both metrics are crucial to achieve a reliable model for automating open question evaluation. In practice, this means that all student responses in the test data were assessed correctly and confidently, compared to the reference (teacherbased) labelling (Figure 2). This also shows the benefit of a hyperparameter grid search. Similar results were obtained for other open questions (not shown) in the same pilot test (section 2), contributing to the validation of the proposal of this work. One current issue regards to cheating with the use of AI writers. Together with pedagogical strategies, automatic detection tools can also be used. As an example in this work, entities in students’ answers can be compared with those in the vocabularies (section 4). This can be measured by the probability of generating the sequence of entities (𝑃(𝑂|π»π‘€π‘€βˆ’π‘πΈπ‘…π‘ π‘¦π‘ π‘‘π‘’π‘š)) for the student’s answer (𝑂). Relatively low values may suggest the use of AI writers. Table 4. Final results in average values (as a result of ten runs): Comparison between all eighth candidate HMM-based NER systems (best model highlighted in bold; values below 0.95 in red; standard deviation shown only for the best model due to lack of space) HMM-based NER system Hyperparameter set (Table 3) Precision (by entity tag) Recall (by entity tag) Macro-Precision (by model) Macro-Recall (by model) Representative words in the training data Handling of unseen words Distance measure for word embeddings No Yes Laplace smoothing Word embedding Euclidean distance Cosine similarity " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " " 𝐼𝐢𝑇 " " π·π‘–π‘”π‘–π‘‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " " πΊπ‘œπ‘Žπ‘™ " " πΊπ‘Žπ‘–π‘› " " 𝑂𝑒𝑑𝑠𝑖𝑑𝑒 " " 𝑃 β„Ž π‘¦π‘ π‘–π‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " " 𝐼𝐢𝑇 " " π·π‘–π‘”π‘–π‘‘π‘Žπ‘™ π‘Šπ‘œπ‘Ÿπ‘™π‘‘ " " πΊπ‘œπ‘Žπ‘™ " " πΊπ‘Žπ‘–π‘› " " 𝑂𝑒𝑑𝑠𝑖𝑑𝑒 " 1 0.90 0.40 0.72 0.94 0.96 1.00 0.99 0.98 0.99 0.97 0.95 0.93 0.82 0.97 2 0.90 0.37 0.64 0.94 0.94 1.00 1.00 0.98 0.98 0.97 0.96 0.91 0.80 0.97 3 0.99 0.95 0.97 0.96 0.96 0.99 0.99 0.98 0.98 0.98 0.97 0.99 0.97 0.98 4 0.97 0.96 0.95 0.97 0.97 0.99 1.00 0.99 0.99 0.99 0.97 0.99 0.97 0.99 5 0.88 0.32 0.70 0.96 0.94 1.00 1.00 1.00 1.00 1.00 0.99 0.92 0.80 0.98 6 0.89 0.34 0.74 0.96 0.94 1.00 1.00 1.00 1.00 1.00 0.99 0.92 0.81 0.99 7 0.99 Β±0.01 1.00 Β±0.00 1.00 Β±0.00 0.96 Β±0.03 0.96 Β±0.02 1.00 Β±0.00 1.00 Β±0.00 1.00 Β±0.00 1.00 Β±0.00 1.00 Β±0.00 0.99 Β±0.01 0.99 Β±0.00 0.98 Β±0.01 1.00 Β±0.00 8 0.97 0.98 1.00 0.95 0.98 1.00 1.00 1.00 1.00 1.00 0.99 0.99 0.98 1.00 6 CONCLUSIONS AND IMPLICATIONS This work investigates the automatic evaluation of open-ended questions using a Named Entity Recognition (NER) system based on the Hidden Markov Model (HMM) statistical technique. The performance of the final HMM-NER system was above 95% for both Precision and Recall metrics, showing great potential for further educational use. An advantage of HMM is its explicit probabilistic nature, which favors interpretability and therefore feedback to teachers and students. For a broader use of this work, a Learning Management System (LMS) could be used or a web platform could be developed as a front-end for the teacher. One issue with data-driven applications is maintaining their accuracy over time. In this work, this means constant review of the vocabularies (section 4) by the teacher, as well as retraining the HMM model with new student responses periodically, for example. This approach makes it possible to use open questions in assessment even with large courses with limited teacher resources. As part of formative assessment, the feedback enabled by the HMM-NER system, can support and direct student learning, for example by clarifying misunderstandings and knowledge gaps or adapting learning paths. The use of open questions in summative assessment focuses students’ attention on deeper learning rather than rote learning or developing guessing techniques, which are usually encouraged by multiple-choice questions. In conclusion, it is important to highlight the pedagogical role of the teacher throughout this process, as technology alone cannot provide satisfactory solutions. 7 ACKNOWLEDGEMENTS This work was developed within the scope of the THREADING-CO2 project, funded by the EU Horizon Europe research and innovation programme (grant 101092257).