Educational Artificial Intelligent Chatbot: Teacher Assistant & Study Buddy
Full text
DEGREE PROJECT Educational Artificial Intelligent Chatbot: Teacher Assistant & Study Buddy Dimitrios Zarris Stergios Sozos Master Programme in Data Science 2023 Luleå University of Technology Department of Computer Science, Electrical and Space Engineering
[This page intentionally left blank]
1 Luleå University of Technology Educational Artificial Intelligent Chatbot: Teacher Assistant & Study Buddy Members: Stergios Sozos Dimitrios Zarris Date: 2023-07-26
2 Abstract In the rapidly evolving landscape of artificial intelligence, the potential of large language models (LLMs) remains a focal point of exploration, especially in the domain of education. This research delves into the capabilities of AI-enhanced chatbots, with a spotlight on the "Teacher Assistant" & "Study Buddy" approaches. The study highlights the role of AI in offering adaptive learning experiences and personalized recommendations. As educational institutions and platforms increasingly turn to AI-driven solutions, understanding the intricacies of how LLMs can be harnessed to create meaningful and accurate educational content becomes paramount. The research adopts a systematic and multi-faceted methodology. At its core, the study investigates the interplay between prompt construction, engineering techniques, and the resulting outputs of the LLM. Two primary methodologies are employed: the application of prompt structuring techniques and the introduction of advanced prompt engineering methods. The former involves a progressive application of techniques like persona and template, aiming to discern their individual and collective impacts on the LLM's outputs. The latter delves into more advanced techniques, such as the few-shot prompt and chain-of-thought prompt, to gauge their influence on the quality and characteristics of the LLM's responses. Complementing these is the "Study Buddy" approach, where curricula from domains like biology, mathematics, and physics are utilized as foundational materials for the experiments. The findings from this research are poised to have significant implications for the future of AI in education. By offering a comprehensive understanding of the variables that influence an LLM's performance, the study paves the way for the development of more refined and effective AI-driven educational tools. As educators and institutions grapple with the challenges of modern education, tools that can generate accurate, relevant, and diverse educational content can be invaluable. This thesis not only contributes to the academic understanding of LLMs and provides practical insights that can shape the future of AI-enhanced education, but as education continues to evolve, the findings underscore the need for ongoing exploration and refinement to fully leverage AI's benefits in the educational sector.
3 Table of Contents 1. Introduction ........................................................................................................................................... 6 1.1 Current problems in Education ..................................................................................................... 7 1.2 Objective ....................................................................................................................................... 7 1.3 Thesis outline ................................................................................................................................ 8 2. Literature Review .................................................................................................................................. 9 2.1 Evolution of Computer Technology ............................................................................................. 9 2.2 Multi-Domain Impact of Computer Technology: A Review Across Diverse Sectors .................. 9 2.3 Artificial Intelligence in Education ............................................................................................. 10 2.4 Research Gap .............................................................................................................................. 15 2.5 GPT vs Watson vs DialogFlow? ................................................................................................. 16 a. Advanced Natural Language Capabilities ................................................................................... 16 b. Efficient Text Retrieval with llama-index .................................................................................. 16 c. Customizability and Fine-Tuning ............................................................................................... 16 d. Development Flexibility ............................................................................................................. 16 e. Cost-Effectiveness ...................................................................................................................... 16 2.6 GPT - Large Language Model .................................................................................................... 17 2.6.1. Limitations of GPT-4: ......................................................................................................... 17 2.6.2. Ethical Problems of GPT-4 ................................................................................................. 18 2.7 Prompt Engineering for Teacher Assistant ................................................................................. 19 2.7.1. Definition ............................................................................................................................ 19 2.7.2. Prompt Structuring .............................................................................................................. 19 2.7.3. Prompt Engineering Techniques ......................................................................................... 22 2.7.3.1 Zero-Shot Prompting....................................................................................................... 22 2.7.3.2 Few-Shot Prompting ....................................................................................................... 22 2.7.3.3 Chain-of-Thought Prompting .......................................................................................... 24 2.7.3.4 Other techniques ............................................................................................................. 24 3. Research Approach ............................................................................................................................. 25 3.1 Overview ..................................................................................................................................... 25 3.2 System Architecture .................................................................................................................... 25 3.3 Design and Implementation ........................................................................................................ 25 4. Study Buddy ........................................................................................................................................ 27 4.1 Methodology ............................................................................................................................... 27 4.2 Results ......................................................................................................................................... 29
4 a. Biology ........................................................................................................................................ 29 b. Mathematics ................................................................................................................................ 30 c. Physics ........................................................................................................................................ 32 5. Teacher Assistant ................................................................................................................................ 35 5.1 Methodology ............................................................................................................................... 35 5.2 Results ......................................................................................................................................... 36 5.2.1 Prompt Structuring .............................................................................................................. 36 a. Persona ........................................................................................................................................ 36 b. Template ..................................................................................................................................... 37 5.2.2 Prompt Engineering ............................................................................................................ 38 a. Few shot ...................................................................................................................................... 38 b. Chain-of-thought ......................................................................................................................... 38 5.2.3 Evaluation: .......................................................................................................................... 39 6. Discussion ........................................................................................................................................... 40 6.1 Ethics & Legal Issues on AI in education ................................................................................... 40 6.2 Jailbreaks and offensive content by LLMs ................................................................................. 41 6.3 The Ethical Implications of AI Overuse in Educational Settings ............................................... 41 a. Limitation on Usage .................................................................................................................... 41 b. Monitoring and Feedback ........................................................................................................... 41 c. Educational Guidelines ............................................................................................................... 42 d. Ethical Design ............................................................................................................................. 42 e. Ongoing Assessment ................................................................................................................... 42 7. Conclusion .......................................................................................................................................... 43 References ................................................................................................................................................... 44 Acknowledgements ..................................................................................................................................... 47 Appendix A ................................................................................................................................................. 48 Biology .................................................................................................................................................... 48 Mathematics ............................................................................................................................................ 49 Physics .................................................................................................................................................... 50 Appendix B ................................................................................................................................................. 54 a. Baseline answer .............................................................................................................................. 54 a. Prompt Structuring .......................................................................................................................... 55 i. Persona ........................................................................................................................................ 55 ii. Template ..................................................................................................................................... 55
5 Prompt Engineering ................................................................................................................................ 57 i. Few shot ...................................................................................................................................... 57 ii. Chain-of-thought ......................................................................................................................... 57 List of Figures FIGURE 1: ILLUSTRATION OF THE APPLICATION USER INTERFACE ................................................................................................. 26 FIGURE 2. STUDY BUDDY METHODOLOGY STEPS ..................................................................................................................... 27 FIGURE 3. OVERALL PERFORMANCE SCORE - STUDY BUDDY ...................................................................................................... 34 FIGURE 4. TEACHER ASSISTANT METHODOLOGY STEPS ............................................................................................................ 35 List of Tables TABLE 1. SUMMARY OF PAPERS ON AI IN EDUCATION ............................................................................................................. 15 TABLE 2. BIOLOGY - EVALUATION OF CRITERIA ....................................................................................................................... 30 TABLE 3. MATHEMATICS - EVALUATION OF CRITERIA ............................................................................................................... 32 TABLE 4. PHYSICS - EVALUATION OF CRITERIA ........................................................................................................................ 34 TABLE 5. BIOLOGY - ANSWER COMPARISON ........................................................................................................................... 48 TABLE 6. MATHEMATICS - ANSWER COMPARISON .................................................................................................................. 49 TABLE 7. PHYSICS - ANSWER COMPARISON ............................................................................................................................ 53 TABLE 8. BASELINE QUESTION RESULTS ................................................................................................................................ 54 TABLE 9. PROMPT STRUCTURING - PERSONA RESULTS ............................................................................................................. 55 TABLE 10. PROMPT STRUCTURING - TEMPLATE RESULTS .......................................................................................................... 56 TABLE 11. PROMPT ENGINEERING - FEW SHOT RESULTS .......................................................................................................... 57 TABLE 12. PROMPT ENGINEERING - CHAIN-OF-THOUGHT RESULTS ............................................................................................ 58
6 1. Introduction Artificial Intelligence (AI) is the ability of machines to perform complex tasks typically requiring human intelligence. These tasks can include problem-solving, planning, understanding natural language, recognizing patterns, learning, and the list keeps growing. The rapid evolution of Artificial Intelligence research continually presents new solutions to complex problems and significant improvements to the quality of human life. One noteworthy example is the study conducted by Mokayed et al., (2023), where the performance of vehicle detectors under varied snowy weather conditions was analyzed using data captured by Unmanned Aerial Vehicles (UAVs). The research provides a comprehensive account of data preparation procedures and introduces a multi-feature deconvolutional Faster R-CNN model for accurate vehicle detection in aerial imagery, utilizing the Nordic Vehicle Dataset (NVD) for experimental validation. However, advancements in the field of medical imaging are presented in the work by Voon et al., (2022), where they assessed the effectiveness of seven distinct Convolutional Neural Networks (CNNs) in grading Invasive Ductal Carcinoma (IDC) in breast histopathological images using transfer learning. This research not only highlights the practical application of CNNs for IDC grade classification but also presents a comparative performance analysis of the seven CNN models against the image distribution data from BreaKHis and the BCG dataset. Lastly, addressing the challenge of scarce data in document image classification, Kanchi et al., (2022) propose an efficient multimodal method that combines textual streams with dynamic word embedding and hierarchical attention networks. Their innovative approach, encompassing techniques such as adaptive thresholding and multimodal side-tuning, yields higher overall accuracy compared to earlier models. This confluence of research illustrates the broad impact and potential of AI across varied domains and conditions, Also, AI can be used in different applications like Security and surveillance (Mokayed et al., 2022), intelligent transportation system (Khan et al., 2022), and numerous other fields. This study specifically concentrates on the use of AI on the education. Education is the most crucial element in fostering a prosperous society. A thriving community depends on a solid educational foundation to promote the growth and well-being of its members. The development of education has been characterized by ongoing advancements and breakthroughs, leading to a wealth of knowledge and possibilities for all. Moving forward, it is essential to prioritize the creation of well-rounded educational programs, build inclusive learning spaces, and encourage critical thinking. By laying a strong educational groundwork, we enable individuals to make meaningful contributions to society, nurture a sense of global responsibility, and pave the way for enduring progress for everyone. The time is ripe for AI to make its way into schools and begin delivering significant benefits. One such example is IBM Watson, a supercomputer-powered AI, serving as an administrative assistant to students at Deakin University. This innovative use of AI exemplifies the potential of integrating advanced technology in academia, offering a dynamic, interactive learning experience while streamlining administrative tasks (Boateng et al., 2022). Watson can provide service to the students at any time. This results in a different structure within the administrative power of the university with fewer personnel. Now the students get an immediate answer depending on their profile. In recent years, the educational systems have integrated computers and other forms of technology in their teaching methodologies. This transition has led to the widespread use of devices such as tablets and laptops in classrooms, enhancing the learning experience for students. Teachers are now incorporating various digital tools and platforms to facilitate collaborative learning. The implementation of technologies like interactive whiteboards, learning management systems, and online resources has not only made lessons more engaging but also expanded the boundaries of the traditional classrooms, connecting students and
7 educators across the globe. As a result, students are acquiring valuable digital skills that prepare them for success in an increasingly technology-driven world. 1.1 Current problems in Education Despite the integration of computerized technology into the educational systems, there remain several challenges to be addressed in order to truly revolutionize education. Some of these issues include: 1. Uniformity in education: Current educational systems often follow a one-size-fits-all approach, failing to cater to the individual strengths and weaknesses of each student. This can hinder students from realizing their full potential and receiving the targeted support they need to excel. 2. Slow feedback: Teachers may sometimes provide feedback slowly, resulting in delayed notifications for students who need extra time to study. This can lead to missed opportunities for improvement and a lack of timely support for those struggling with specific concepts. 3. Overburdened teachers: Teachers are often stretched thin, managing large class sizes and an extensive workload, which can affect their ability to provide individualized attention and support to students. 1.2 Objective In this study we aim to explore the feasibility of implementing advanced AI technologies to address the challenges in education from both the perspectives of teachers and students. We propose the concept of a teacher assistant and a study buddy, powered by state-of-the-art AI techniques. By leveraging these intelligent systems, we aspire to offer effective solutions that support educators in their instructional tasks and empower students in their learning journey. This study is carried out from two unique perspectives. One is from a teacher's viewpoint, where the AI functions as a supportive teaching assistant. The other is from a student's viewpoint, where the AI serves as a helpful study companion. 1. Teacher Assistant: In our investigation, we will focus on the potential for the AI to generate exam subjects tailored to the curriculum being utilized by the teacher. By utilizing these advanced technologies, we aim to provide educators with a valuable tool that can assist in the creation of relevant and targeted exam materials. 2. Student Buddy: As part of our research, our objective is to assess its ability to provide accurate answers to questions directly related to specific curriculum. By employing these sophisticated AI technologies, we endeavor to create a study companion that can offer students valuable assistance in their learning journey, ensuring they have access to prompt and reliable information. Upon conducting research, it is evident that GPT currently stands at the forefront of AI and large language models. However, it provokes several intriguing questions. Could its level of intelligence serve effectively as a teaching assistant, creating exam questions akin to those constructed by human educators? Is it sufficiently sophisticated to provide accurate and meaningful responses to students' queries? What is the ideal structuring of questions to maximize the accuracy of responses and the effectiveness of exam questionnaires as generated by the AI model?
14 Artificial Intelligence (AI) Student Assistants in the Classroom: Designing Chatbots to Support Student Success 1. Teach basic content & answer basic questions 2. Encourage exploration of additional resources 3. Assist with career‑ and life‑related issues 4. 24/7 responsiveness 5. Assist students through conversations which leads to engagement 6. Shy students can ask any question freely 1. Losing the flow of the dialogue 2. Lacking the capacity for emotional comprehension in students 3. Not able to understand openended answers 4. Inability to react in real time to student comments and questions 1. Multiplechoice quizzes 2. Answering frequently asked questions 3. Answering basic course and content questions 4. Ability to encourage further exploration of additional resources 5. Language learning 1. By quiz results 2. Objective data, test scores, and student opinion of their learning experience Education 4.0: Artificial Intelligence Assisted Taskand Time Planning System 1. Automatically generated tasks from the AI Engines 2. Secure data exchange via the Open API Service 3. Ability to monitor and view learning progress and success 4. Gamification of working through courses 5. Consideration of different learning types, strengths, and weaknesses 6. Notification messages to alert students of learning arrears 1. Lack of personalization and customization for individual students 2. Potential for data security and privacy issues 3. Difficulty in understanding complex tasks and instructions 4. Difficulty in providing accurate and timely feedback, notifications and reminders, and assessment tasks 1. Tailored learning content 2. Reminders and notifications about tasks and learning progress 3. Provide assessment tasks 4. Coordinate automatically and manually generated tasks 5. Monitor and view learning progress and success 1. Usability 2. User experience 3. Data protection 4. Data security 5. Communication between layers Kwame for Science: An AI Teaching Assistant for 1. Providing instant answers to science questions 1. Challenging cases when there are typos in the 1. Providing instant answers to Science questions of students 1. Top 1 accuracy quantify performance assuming only one
15 Science Education in West Africa 2. Ability to fine-tune the SBERT model to improve accuracy 3. Making Kwame for Science available in local languages across Africa 4. Making it available via offline channels such as SMS, USSD, and tollfree calling spelling of scientific words 2. Questions related to topics outside the scope of the knowledge source 3. Unhelpful answers due to issues with the dataset answer was returned and voted on 2. Top 3 accuracy refers to the performance where for each question that received a vote, at least one answer was rated as helpful out of the 3 answers that were returned Table 1. Summary of Papers on AI in Education 2.4 Research Gap Despite the growing body of research on the use of conversational AI chatbots in education, many studies to date have focused primarily on the chatbots' abilities to answer questions based on specific curriculums without the provision of associated textbooks (Chen et al., 2022; Pereira, 2016; Benedetto & Cremonesi, 2019). These models are typically pre-trained on specific educational materials and are thus limited in their flexibility and adaptability to cater to different learning materials and diverse educational needs. One noteworthy exception is the study by Boateng et al. (2022), which proposed a system that allowed users to provide their own textbooks for the AI to generate responses. Nonetheless, the system required further training on the provided materials to deliver accurate responses, a time-consuming process that could potentially hinder its wide-scale application. Despite its pioneering efforts, Boateng et al.'s (2022) study illuminates a significant research gap: the lack of chatbot applications capable of generating accurate responses from user-provided textbooks without the need for additional training. It also underscores the need for the development of more flexible, adaptable, and user-friendly educational chatbots. The present study aims to address this research gap by using the GPT, a pre-trained Large Language Model (LLM), in conjunction with the llama-index library, which has the capacity to read and index documents without the need for additional training. This approach not only eliminates the time and resources required for training but also enhances the flexibility and adaptability of chatbots to cater to a wider range of educational materials and user requirements.
16 2.5 GPT vs Watson vs DialogFlow? In the development of an educational AI chatbot aimed at serving as a teacher assistant and study buddy, the choice of underlying technology plays a pivotal role. Three primary platforms were considered: Google's Dialogflow, IBM's Watson, and OpenAI's GPT coupled with the llama-index library. After thorough evaluation, the combination of GPT and llama-index was selected for the following reasons: a. Advanced Natural Language Capabilities GPT model excel in understanding and generating human-like, context-relevant text. This feature is crucial for a chatbot that needs to interpret complex educational queries and respond in an intuitive manner. While Watson and Dialogflow offer robust Natural Language Understanding (NLU), GPT's capabilities are more advanced, especially when fine-tuned for a specific domain like education. b. Efficient Text Retrieval with llama-index One of the unique requirements of this project was the ability to sift through textbooks to provide accurate answers. Llama-index complements GPT by efficiently indexing large volumes of text, making the retrieval of relevant sections fast and accurate. This fills a gap that would otherwise require a separate tool like Elasticsearch, thereby streamlining the technology stack. c. Customizability and Fine-Tuning GPT models can be fine-tuned to specialize in specific tasks or domains. This level of customizability allows for better alignment with the educational focus of the chatbot. While Watson does offer specialized models, the fine-tuning capabilities of GPT are more extensive, offering finer control over the model's behavior. d. Development Flexibility The integration of GPT with llama-index can be achieved with Python, which offers a rich set of libraries and tools for machine learning, and NLP. e. Cost-Effectiveness While it is true that GPT API calls can be expensive, the combined utility of GPT and llama-index reduces the need for multiple services that could incur additional costs, as would be the case with Watson or Dialogflow. Additionally, the fine-tuning capabilities of GPT can result in more accurate and efficient querying, potentially reducing the number of required API calls.
17 In summary, the synergy between GPT's advanced natural language capabilities and llama-index's efficient text retrieval presents a compelling case for their selection. This combination offers a robust, customizable, and efficient solution for developing an educational AI chatbot that can effectively serve as both a teacher assistant and a study buddy. 2.6 GPT - Large Language Model In this section we will give an overview of GPT and Large Language Models. The Generative Pre-Trained Transformer (GPT) is an innovative Natural Language Processing (NLP) model, spearheaded by the esteemed research institute OpenAI. This sophisticated deep learning model leverages the power of the transformer architecture, enabling it to produce text reminiscent of human-like semantics and syntax. Trained on an extensive corpus of text, GPT epitomizes what are often referred to as Large Language Models (LLMs), exhibiting a remarkable ability to generate contextually pertinent text relative to the input provided. This model's versatility allows it to be proficiently employed for a diverse array of tasks that include, but are not limited to, text summarization, question answering, and text generation. Even though GPT-4 which is the latest version of GPT model at the time of study, is really good in answering questions and making discussions on different subjects, it is not perfect and in many cases its answers are not accurate or even worse inappropriate. According to OpenAI, (2023), there are not only limitations but some very critical ethical issues that we should consider when using GPT-4 model. 2.6.1. Limitations of GPT-4: a. Hallucinations: GPT-4 is prone to hallucinations, a term referring to situations where the model provides false or made-up information without context. This might include the presentation of fictitious data as factual, or the confident assertion of false realities. These hallucinations undermine the reliability of the model, potentially misleading users or compromising the integrity of outputs. b. Producing biased and unreliable content: GPT-4 has been observed generating content that is biased or unreliable. This includes potentially harmful content such as hate speech, incitements to violence, and false narratives, which can exploit individuals and cause harm. It can also generate content promoting self-harm, graphic material, and instructions for illegal activities, further emphasizing the need for responsible use and monitoring. c. Finding websites selling illegal goods or services: A limitation of GPT-4 includes its capacity to facilitate illegal activities. This can involve the identification and promotion of illicit web platforms like dark markets, websites selling counterfeit or stolen goods, or services like hacking or money laundering. Such capability poses a serious ethical dilemma in the use of this technology. d. Planning attacks: An alarming limitation lies in GPT-4's ability to plan harmful actions, such as conducting a phishing attack or even setting up an open-source language model on a new server for malicious activities. It can identify key vulnerabilities, hide traces, and leverage human resource platforms for its tasks, thus posing a threat to digital security.
18 e. Generating believable and persuasive content: GPT-4 has the potential to generate content that is convincing and persuasive. This capability can be misused to create politically charged content, realistic but misleading narratives, or to alter the discourse around certain topics, thereby spreading misinformation and manipulating public opinion. f. Lack of knowledge of post-training events: GPT-4's training data is limited to the point of pretraining data cut-off. Therefore, it lacks knowledge of subsequent events, changes in laws, regulations, social norms, advances in technology, or new discoveries in various fields, resulting in outdated or potentially irrelevant responses. g. Simple reasoning errors: GPT-4 can often make simple reasoning errors, such as stereotyping demographic groups or making assumptions based on limited information. This could lead to inaccurate, biased, or discriminatory outputs. h. Gullibility towards false statements: GPT-4 can display a degree of gullibility, often accepting clearly false statements without any evidence. This gullibility extends to conspiracy theories, false political promises, and unfounded claims from various sources, which can propagate misinformation and falsehoods. i. Failing at hard problems: GPT-4 may underperform when faced with complex problems. For instance, accurately predicting very low pass rates, reversing the trend of decreasing performance as a function of scale, and underperforming on seemingly easy tasks are among the limitations that may lead to ineffective or misleading outcomes. j. Negligence in double-checking work: GPT-4 does not inherently double-check its output, leading to a potential spread of misinformation. Dependence on the model's output without verification and ignoring the model's refusals and hedging cues can lead to inaccuracies and errors. 2.6.2. Ethical Problems of GPT-4 a. Privacy concerns: AI systems like GPT-4 could potentially reinforce ideologies, and even propagate false information, which raises concerns about privacy and the autonomy of thought. Furthermore, these systems can also cause disparities in service quality due to performance differences across demographics and tasks. b. Data security risks: GPT-4 has the potential to exacerbate data security risks. For example, it could assist in identifying individuals when combined with outside data, facilitate social engineering for cyberattacks, provide guidance for harmful activities, or even enhance existing security threats. c. Potential harms to vulnerable populations: GPT-4 can unintentionally cause harm to vulnerable groups. This could be through exposure to hazardous materials or environments, discrimination, unsafe working conditions, substandard healthcare, violence or abuse, cyberbullying, scams, fraud, or the spread of false or misleading information. d. Data breaches: GPT-4 may indirectly contribute to data breaches, such as unauthorized access to sensitive data, malicious attacks on data systems, or data leakage due to weak security measures, causing significant harm and raising ethical issues. e. Misuse of personal information: Misuse of personal information is a serious ethical concern. GPT4 can be used to reinforce ideologies, inform resource allocation, or provide legal or health advice, potentially leading to misuse or exploitation of personal data. f. Algorithmic bias: GPT-4 can exhibit biases which can result in unequal services across different demographics or domains. These biases can reinforce ideologies, and norms, and can particularly
19 affect groups that are underrepresented in the training data, further raising concerns about fairness and equity. 2.7 Prompt Engineering for Teacher Assistant This chapter aims to provide a comprehensive overview of prompt engineering as a technique that can be used to create effective assessments for students. Firstly, an explanation of prompt engineering will be presented, outlining the main concepts and principles underlying this approach. By defining the key aspects of prompt engineering, readers will be able to gain a clear understanding of how it can be applied in the context of automated assessment design and creation. Then, the focus of this chapter will be on demonstrating and summarizing techniques that have been documented in research papers. These techniques have been proven to have different effectiveness on the output, regardless of the context of the prompt. Finally, by analyzing these approaches, taken by other researchers in this area, we will be able to identify best practices, and we will demonstrate how prompt engineering can be applied to create effective assessments for students. This chapter will serve as a foundation for the subsequent chapter of this thesis “Teacher Assistant” Through this work, we hope to contribute to the ongoing development of using chatbots for creating effective assessment strategies that can support student learning and evaluation. 2.7.1. Definition Prompt engineering is about how the structure and the content of a prompt can significantly affect the result of the model. Since the prompt defines the task, selecting the right prompt has a significant impact on both the accuracy and the initial task that the model does (Liu et al., 2021). To provide with a more technical definition in the context of chatbots, we can use White (et al) definition, which defines prompt as a set of directions given to an LLM, which enables the customization, enhancement, or refinement of the LLM's capabilities (White et al., 2023). The use of prompts can significantly influence the outcomes of interactions with a Large Language Model (LLM), as they provide specific rules and guidelines for conversations. A prompt help to establish the context of the conversation, highlighting the most relevant information and indicating the desired form and content of the output. For instance, a prompt can specify that an LLM should generate code that adheres to a particular coding style (White et al., 2023). 2.7.2. Prompt Structuring Research has been conducted to provide a reusable solution for prompts. The catalog of prompt patterns seeks to offer reusable answers to issues users encounter when working with LLMs to complete a variety of activities (White et al., 2023). Other researches have already proven the efficiency of prompting in different areas, such as for designing prompts to be used in tasks for classification (Wang et al., 2023) or
20 for designing boolean prompts that can be used in literature queries (Xu et al., 2022). This solution that is proposed in White et al paper is more general, so that it can be used to recurring problems within a particular context, which in our case is the assessment design and creation. This chapter is aimed to provide an overview of the crucial aspects that require consideration when constructing an efficient prompt. We will focus on specific areas that we believe can have a significant influence on the efficacy of the prompt for the purpose of assessment creation. Then, in the forthcoming chapter entitled "Teacher Assistant," we will employ the framework established in this chapter to evaluate its effectiveness. Additionally, we will compare the outcomes generated by this structure to those of a simplified prompt that deviates from the proposed framework. The system that has been proposed by White et al. for compiling and using a list of prompt patterns for ChatGPT and other large language models (LLMs) consists of the following aspects (White et al., 2023): Input Semantics Output Customization Error Identification Prompt Improvement Interaction Context Control The key factor in our case is the output customization, which is about limiting or modifying the types, formats, structures, or other elements of the output produced by the LLM (White et al., 2023). We need to investigate how a prompt can affect the output, so that the assessments that are produced by a chatbot meet the teacher’s expectations. The patterns that are proposed for the output customization are the following: 1. Output Automater, with which the user can write scripts to automate any tasks that the LLM output proposes him to do. This is not relevant to our research about assessment creation (White et al., 2023). 2. Persona, with which the user can give to LLM a role to embody when it provides an answer. Persona pattern can be explained by two key elements, the Intent & Context of it, and its Motivation (White et al., 2023). In terms of the intention and context, it entails giving the LLM a certain point of view or perspective to take when producing products. This viewpoint may take the shape of a "persona" that directs the LLM in determining which specifics to emphasize. For instance, users might ask the LLM to analyze the code from the standpoint of a security specialist. With the use of this method, the LLM can provide outputs that are in line with the task's stated purpose and desired results (White et al., 2023). Regarding the pattern's motivation, it acknowledges that users might not be aware of the precise outputs or information needed to complete a task. They might, however, be familiar with the position or kind of person who frequently offers support in these areas. Users can express their desires without having to be as specific as possible about the needed outcomes by giving the LLM a persona. With the use of this method, the LLM can provide results that meet the needs of the intended audience (White et al., 2023). Overall, the Persona pattern is a potent tool that enables users to customize the outputs to their particular needs, which can increase the effectiveness of LLMs. Users can ensure that the outputs are in line with
21 the task's intended purpose and planned outcomes by giving the LLM with a persona without needing to be aware of the precise information or outputs required (White et al., 2023). Finally, it is proposed to structure this pattern by including the following phrase in the prompt: “Act as persona X” or “Provide outputs that persona X would create” (White et al., 2023). 3. Visualization Generator, with which the user can make the LLM to produce textual outputs that may be fed to other tools, such other AI-based picture generators enabling the user to create visualizations. This is not relevant to our research about assessment creation (White et al., 2023). 4. Recipe, with which the user is able to receive a set of steps or actions that will help them achieve a specified end goal, even if they are working with unknown or limited knowledge. This is not relevant to our research about assessment creation (White et al., 2023). 5. Template, with which the user can designate an output template, which the LLM will then fill up with its output. This pattern can, also, be explained by the two key elements, which are the Intent & Context of it, and its Motivation (White et al., 2023). Regarding the Intent and Context, to make sure that an LLM's output follows a certain template structure element of the Template pattern that it was created. When customers ask the LLM to generate outputs in a format that is not commonly utilized for the information being generated, this is especially helpful. For instance, a user could need to create a URL that places created data at particular points along the URL path. Users can direct the LLM to produce output that follows the defined template by utilizing the Template pattern, guaranteeing that the final output is properly organized as required (White et al., 2023). Motivation is the second component for explaining the Template pattern. This component acknowledges that occasionally output needs to be created in a precise format that is particular to a given application or use-case. Users must specify the format and the locations of various sections of the output because the LLM is unaware of the necessary template structure. This could be done by creating a sample data structure or by filling out a series of form letters. Users can make sure that the output produced by the LLM accurately adheres to the defined template structure by employing the Template pattern, ensuring that the output meets the necessary format and quality criteria (White et al., 2023). In general, the Template pattern is a useful tool for directing the output produced by an LLM. Users can guarantee that the output meets the necessary format and quality criteria by giving the LLM a specified template structure, even if the LLM isn't generally used to creating output in that manner. The Template pattern is an effective method for making sure that LLM output is precisely formatted as needed, in a way that suits the needs of users and stakeholders (White et al., 2023). Finally, it is proposed to structure this pattern by including the following phrase in the prompt: “I am going to provide a template for your output” or “X is my placeholder for content” or “Try to fit the output into one or more of the placeholders that I list” or “Please preserve the formatting and overall template that I provide” or “This is the template: PATTERN with PLACEHOLDERS” (White et al., 2023).
22 2.7.3. Prompt Engineering Techniques 2.7.3.1 Zero-Shot Prompting Modern Large Language Models (LLMs), such as GPT-3, have the capability to perform certain tasks "zeroshot" due to their ability to follow instructions and extensive training on vast amounts of data. Such models have zero-shot capabilities. For example, in a classification-asking prompt such as the following: “Classify this sentence as positive or negative: I am very happy with this product” We got the following answer using ChatGpt: “The sentence is positive.” This means that such models can comprehend and classify a sentiment from a text without any examples being given with their classifications in the prompt. And this is what is referred to as zero shot prompting. This can be extended to any purpose; in our case it can be on the education context. Zero-shot prompting is a promising approach for generating text with little to no supervision. However, it is not always effective for advanced or non-generic questions. In such cases, providing demonstrations or examples in the prompt can improve the performance of the model, leading to what is known as few-shot prompting. In the next section, we will demonstrate how this approach can be utilized to improve the quality of text generated by the model. 2.7.3.2 Few-Shot Prompting Learning tasks by providing a few examples, or few-shot learning, is a significant academic challenge. Modern NLP does few-shot learning by reformulating tasks as "prompts" in natural language and filling in those prompts with trained language models (Logan IV et al., 2022). Recent research on advanced pre-trained language models, including GPT-3, InstructGPT, and ChatGPT 2, indicates that these models can effectively perform a variety of tasks without the need to adjust their parameters. Instead, they can utilize just a small set of examples as instructions. This demonstrates the impressive capabilities of LLMs on a large scale (Wei et al., 2023). Few-shot prompting is also referred as “in-context” learning. As mentioned above, the in-context learning allows the model to perform many tasks without adjusting the parameters of the model. The effectiveness of this technique has been also measured and been compared to the zero-shot prompting. The results showed that in a broad array of tasks, the few-shot prompting is consistently proved to be more effective than the zero-shot one (Zhao et al., 2021). However, it needs to be noted at this point that the academic research about this technique is limited. It is not yet fully explained neither how it exactly works, nor which part of the few-shot prompt is most effective for the best output (Min et al., 2022). Thus, it is known that giving some input-label pairs before the actual prompt is contributing to a better output (Brown et al., 2020), but it is not knowing which part of these input-label pairs actually affect the end performance of the output. An example of how few-shot prompting can be implemented and how it can be better than zero-shot is presented below: Let’s assume about a medical diagnosis task. This is a task where few-shot learning is likely to outperform zero-shot learning due to the specificity and complexity of the task.
23 Here's a set of symptoms we will be using in our prompt: "The patient is experiencing shortness of breath, swelling in the legs, and fatigue." Zero-Shot Prompting: We would feed the language model (LM) this symptom and ask it to diagnose without any specific training or prior examples for this task. Input: "The patient is experiencing shortness of breath, swelling in the legs, and fatigue. What is your diagnosis?" Output: "Based on these symptoms, the patient could be suffering from a number of conditions. It could be a respiratory issue, a cardiovascular problem, or even a side effect of certain medications. It's essential to seek professional medical advice for a proper diagnosis." The model provides a vague answer that doesn't specify a single diagnosis because it doesn't have the context to make that kind of judgement. Few-Shot Prompting: In few-shot prompting, we provide the LM with a few examples of the task before presenting it with the test case. Input: "Example 1: The patient is experiencing chest pain, shortness of breath, and nausea. The diagnosis for these symptoms is heart attack. Example 2: The patient has frequent urination, excessive thirst, and unexplained weight loss. The diagnosis for these symptoms is diabetes. Example 3: The patient experiences joint pain, stiffness, and decreased range of motion. The diagnosis for these symptoms is arthritis. The patient is experiencing shortness of breath, swelling in the legs, and fatigue. What is your diagnosis?" Output: "Based on these symptoms, the patient could potentially have congestive heart failure. However, it is still crucial to consult with a healthcare professional for an accurate diagnosis." Comparing the 2 techniques: In the zero-shot prompting, the model gave a broad and vague answer because it didn't have a specific context or example to base its answer on. This demonstrates the limitations of the zero-shot approach for certain specific tasks, especially in areas that require expert knowledge like medical diagnosis. In contrast, few-shot prompting with prior examples allowed the model to perform better, identifying the symptoms as potentially indicative of congestive heart failure, which aligns more accurately with the presented symptoms.
30 Note about question 6: Study Buddy (AI) did not understand the question, because probably it couldn’t find any correlation of the word “role” in the textbook around the existence of DNA, hence the answer was wrong saying that the role of the DNA is not mentioned in the context. # Question Criterion 1 (Comprehensive Content) Criterion 2 (Information Consistency) Criterion 3 (Relevant Content) Score 1 Who coined the term cell, in reference to the tiny structures seen in living organisms? (Teacher mentioned: "The first time the word cell was used to refer to these tiny units of life was in 1665.") 0.66 2 Who identified animalcules? What are animalcules? 1 3 What are the three main parts of the cell theory? 1 4 List the four parts common to all cells. 1 5 What are the cell structures where proteins are made? 1 6 What is the role of DNA? (no answer) 0 Table 2. Biology - Evaluation of Criteria Total score for Biology is: 4.6 / 6 = 0.76 b. Mathematics The question stated in the textbook is as follows: “Find the pattern rules for the following numerical patterns.”, followed by a list of numbers’ sequences, see column “Question asked” in Table 2. Study Buddy did not interpret the question properly and it couldn’t give any answer, then we tried different versions of the question in order to make it more suitable for Study Buddy to understand.
31 The question we gave to Study Buddy is as follows: “Can you follow the book's theory in order to find the pattern that applies to the following numbers' sequence: 1, 6, 21, 66. And do it. But give me the final answer without explaining the steps. Before you answer make sure that your answer is correct by applying the calculations to the number sequence I gave you.” By observing the correct and wrong answers given by Study Buddy, we see that it gives the wrong answer when the pattern rule consists of more than one calculation. This is just an observation of the tests we have run and it may not be the rule for other examples. # Question ( Sequence of Numbers) Criterion 1: Solution Accuracy Criterion 2: Relevant Content Score 1 1, 6, 21, 66 Only contains information relevant to the coursebook. 0.5 2 95, 80, 65, 50 Only contains information relevant to the coursebook. 1 3 3, 10, 17, 24 Only contains information relevant to the coursebook. 1 4 256, 64, 16, 4 Only contains information relevant to the coursebook. 1 5 3, 11, 43, 171 Only contains information relevant to the coursebook. 0.5 6 81, 27, 9, 3 Only contains information relevant to the coursebook. 1 7 4, 13, 40, 121 Only contains information relevant to the coursebook. 0.5 8 1, 6, 31, 156 0.5
32 Only contains information relevant to the coursebook. Table 3. Mathematics - Evaluation of Criteria Total score for Mathematics is: 6 / 8 = 0.75 c. Physics From Table 6 in Appendix A, we can see that the answers from Study Buddy are very similar to the ones given by the teacher, but there are noteworthy differences, both in some of the results, but also in the way the solutions are explained and presented: Question 1. Determining the time at which you have the same speed as your brother. Study Buddy and the teacher both agree on the time at which you reach the same speed as your brother. Both give the answer as 1 second. Question 2. Calculating the maximum height a ball reaches when launched upwards. In this scenario, both Study Buddy and the teacher use the same approach - using the principle of conservation of energy. However, the AI provides a more step-by-step, mathematical solution, while the teacher's answer is concise and assumes that the student is familiar with the relevant concepts. Question 3. Finding the velocity of the ball just before it hits the top of a building. The AI calculates the velocity as approximately 15.99 m/s, whereas the teacher calculates it as approximately 16 m/s. This discrepancy is minor and could be due to rounding or the specifics of how the calculations were carried out. Question 4. Finding out how long a ball is in the air from launch till it hits the ground. Here we see a significant discrepancy. Study Buddy calculates that the ball is in the air for approximately 14.27 seconds, while the teacher calculates a total time of about 8.77 seconds. The answer from Study Buddy is wrong, because it assumes that the time going up is equal to the time falling. Question 5. Identifying the required acceleration to increase speed over a specific distance. Study Buddy calculates the acceleration 1.78 m/s2, while the teacher answers with 2 m/s2. This is a significant discrepancy and could be due to either rounding or an error in the calculations. Question 6. Determining the acceleration of a rock dropped from a cliff. Study Buddy answers correctly that the acceleration due to gravity is constant to 9.8 m/s2. Question 7. Finding the height of a cliff from which a rock is dropped. Study Buddy calculates the height of the cliff to approximately 60.725 meters, while the teacher gives the height to 60 meters. This difference is relatively minor and likely due to rounding. Overall, it appears that the Study Buddy is more inclined towards detailed mathematical explanations, while the teacher’s answer uses a mix of math and conceptual understanding.
33 # Question Criterion 1: Solution Accuracy Criterion 2: Methodological Alignment Criterion 3: Relevant Content Score 1 Bike problem: At what time do you have the same speed as your brother? Study Buddy gives the correct answer to be 1 second. The teacher solution is given by observing a diagram, where Study Buddy is following the theory’s equations to find the answer. Only contains information relevant to the coursebook. 0.66 2 Building problem: How high does the ball go? Study Buddy gives the correct answer to be approximately 250m. The method used by Study Buddy, does not align with the teacher's method. Only contains information relevant to the coursebook. 0.66 3 Building problem: What is the velocity of the ball right before it hits the top of the building? Study Buddy gives the correct answer to be approximately 16 m/s. The Study Buddy's solution aligns with the method used in the teacher’s solution. Only contains information relevant to the coursebook. 1 4 Building problem: For how many seconds is the ball in the air from the moment it was launched till it hits back to the ground? The answer from Study Buddy is 14.27s which is wrong. The correct answer is 8.77s. The method used by the Study Buddy does not align with the teacher's method. Study Buddy wrongly assumes that the falling time is equal to the time going up. Only contains information relevant to the coursebook. 0.33 5 Speed problem: What acceleration should you use to increase your speed from 10 m/s to 18 m/s over a distance of 55 m? The answer from Study Buddy is 1.78 m/s² which is wrong. The correct answer is 2 m/s². Both solutions use the correct kinematic equation. Only contains information relevant to the coursebook. 0.66
34 6 Cliff problem: What is the magnitude of the acceleration of the rock at the moment it is dropped? Study Buddy gives the correct answer 9.8 m/s2. Both solutions use the concept of gravitational acceleration. Only contains information relevant to the coursebook. 1 7 Cliff problem: What is the height of the cliff? Study Buddy gives the correct answer to be approximately 60m. Both solutions use the correct kinematic equation. Only contains information relevant to the coursebook. 1 Table 4. Physics - Evaluation of Criteria The total score for Physics is: 5.31 / 7 = 0.76 Figure 3. Overall Performance Score - Study Buddy From the above graph we see that the performance of the Study Buddy is good enough and with some tweaking (prompt engineering) we can potentially make it even better.
35 5. Teacher Assistant 5.1 Methodology This chapter breaks down the steps we took to carry out our research. We used a five-step process that involves working with a Large Language Model (LLM) to make a multiple-choice test based on Biology book “CK-12 Biology for High School” from chapter “4.1 Central Dogma”. We're particularly interested in seeing how different ways of creating and tweaking prompts can change the way the LLM responds. We followed 5 steps to run the Teacher Assistant examples. Figure 4. Teacher Assistant Methodology Steps Following are details of each step: 1. Selection of the Chapter: The first step in our methodology is to choose a specific chapter from the book that serves as the foundation for our test. For the purpose of this study, the chapter “2.1 Parts of the Cell” which focuses on a particular theory from the book “CK-12 Biology for High School” has been selected. 2. Creation of the Baseline Prompt (Zero-Prompt): Following the selection of the chapter, we proceed to develop a simple prompt that guides the LLM to formulate a ‘multiple-choice’, ‘fill in the gaps’, ‘true-false’, ‘match the following’, and ‘open-ended’ questions test based on the selected theory. This prompt is intended to act as a reference point, enabling comparisons with the outcomes of subsequent steps. The simplicity of the initial prompt is deliberately maintained to discern the effects of later modifications clearly.
36 3. Application of Prompt Structuring Techniques: The third stage of our methodology involves applying prompt structuring techniques progressively. Techniques such as the persona and the template are integrated into the initial prompt one by one, as discussed in the literature review. We aim to observe how these modifications alter the LLM's outputs, thus offering insight into the individual and collective impacts of the structuring techniques. 4. Employment of Prompt Engineering Techniques: The final phase involves introducing prompt engineering techniques, monitoring the changes in the LLM's responses following each change of technique. Techniques like few-shot prompt and chain-of-thought prompt are used to evaluate their influence on the LLM's output quality and characteristics. By systematically and independently applying each technique, we can discern the precise impact of these changes on the LLM's generated content. 5. Establishing Evaluation Criteria: The evaluation will be in comparison to the ‘baseline’ answer. We have defined the following evaluation criteria: Evaluation Criteria: Criterion 1: From how many concepts are the questions taken? We check the diversity of the answer from the book. Criterion 2: Relevant Content The content of the answer should always be relevant to the content of the book. In summary, the methodology outlined above offers a systematic approach for assessing the potential of a large language model in the generation of knowledge-based tests. By exploring the interactions between prompt construction, engineering techniques, and the LLM's outputs, we establish a comprehensive understanding of the variables influencing the LLM's performance. In the next chapter we will present the results derived from this methodological approach. 5.2 Results The first prompt tested was a simple question like it was referred to a human teacher. The prompt is the following: “Generate 3 multiple-choice questions with 4 possible answers each regarding molecular biology. The questions should test understanding of the fundamental principles of the topic.” See Table 8. Baseline Question Results for the results. 5.2.1 Prompt Structuring a. Persona Incorporating the persona of a teacher in an LLM's prompt can enhance the creation of exam questions due to two main factors: context-intention and motivation. Contextually, the teacher persona enables the LLM to focus on educational principles and goals typically associated with exam design, such as appropriate difficulty and content relevance. In terms of motivation, users who may not know the specifics needed for effective exam questions can benefit from the LLM embodying a teacher's persona, as it inherently captures
37 aspects like fairness, comprehensiveness, and diversity of question formats. Hence, by assigning the persona of a teacher to the LLM, users can obtain tailored outputs that align with educational standards and meet their needs, despite potential lack of specific input details. See Table 9. Prompt Structuring - Persona Results for the results of this technique. Evaluation of Teacher Assistant answers given using the “persona” technique Both versions are technically correct, and they both contain clear and concise exam-style questions. However, there are some differences in style and content that might be attributed to the "act as a teacher" instruction. Simplification and clarification: The "act as a teacher" version seems to use slightly simpler language and clearer phrasing. For example, in the first question, the "act as a teacher" version simply gives "DNA → RNA → Protein" as an option, whereas the original version provides a longer description ("DNA is transcribed into RNA, which is then translated into proteins"). This might reflect a teacher's intention to provide clear, unambiguous choices. Directness: The "act as a teacher" version tends to use more direct language. For example, the question "What is the process of copying genetic information from DNA to RNA called?" directly asks about a specific process, whereas the equivalent question in the original version asks about the role of RNA in the central dogma, which is a bit more abstract. Testing Hierarchical Understanding: In the third question, the "act as a teacher" version adds the term HIV in the choices, while the original version uses only the term "Retrovirus". This change may aim to evaluate students' comprehension of the hierarchical relationship between specific examples (HIV) and broader categories (retroviruses). It's a teaching tactic to test whether students can distinguish between the two, which is crucial for a deep understanding of the subject. In summary, the "act as a teacher" instruction seems to have led to questions that are slightly simpler, more direct, and possibly more engaging due to the use of specific examples. However, the overall content and difficulty level appear to be similar between the two versions. It's also worth noting that these differences might not be observed consistently across all prompts. b. Template The template is organized into several sections. The first placeholder is for the question number, which allows for easy tracking and organization of different questions. This is followed by placeholders for the question content and four potential answer options (labeled A through D), giving structure to the question and its possible solutions. What sets this template apart is the inclusion of placeholders for both the correct answer and an explanation for that answer. This design ensures that the exam doesn't just evaluate student knowledge but also provides feedback and aids in learning, promoting an understanding of why a particular answer is correct. Of course, as mentioned in the literature, the main advantage of providing the template is that the expected output has always the same format, based on the user’s needs. See Table 10. Prompt Structuring - Template Results for the results of this technique.
38 Evaluation of Teacher Assistant answers given using the “template” technique The addition of the template to the original prompt created a significant change in the format of the exam questions produced by the large language model (LLM). Here's a detailed analysis: The key differences with adding the template to the large language model (LLM) include: Formatting: The template has structured the output in a more organized way. Each component (the question, answer options, correct answer, and explanation) now has a clear and defined place, making it easier to parse and understand. This format also emphasizes the importance of providing a detailed explanation, which could enhance the learning process. Completeness: In contrast to the original version, the template has prompted the LLM to provide explanations for the correct answers. This helps to elaborate on the reasoning behind the correct answers and could potentially aid students in their understanding of the material. As such, adding the template is about imposing structure and ensuring completeness, which guides the model to output content in a specific, desired format. 5.2.2 Prompt Engineering a. Few shot For this technique we consulted a document giving some rules for creating a good quality multiple-choice questions exam paper, (sanfoundry.com). See Table 11. Prompt Engineering - Few shot Results for the results of this technique. Evaluation of Teacher Assistant answers given using the “few shot” technique We do not observe worth mentioning differences when we used the examples for the few shot prompting. Thus this change in the prompt was not successful. b. Chain-of-thought For “chain-of-thought” we used the guide by BRIGHAM YOUNG UNIVERSITY (2001). This guide has the principles of how to write an effective multiple choice question. As states in the literature we can use this logic, this chain-of-thought to help the LLM to follow the same logic when providing an answer. In order to give the AI to understand this rules, we passed them as a prompt as is, from the paper. And at the end we added the following prompt: “According to the above text, which explains the rules of creating a multiple-choice questionnaire, generate three multiple choice questions with four possible options each on the topic of Molecular Biology, designed to test students' understanding of the fundamental principles of this subject.” See Table 12. Prompt Engineering - Chain-of-thought Results for the results of this technique.
39 Evaluation of Teacher Assistant answers given using the “Chain-of-thought” technique Even though it followed the rules we can see that the chapter it chose is wrong. 5.2.3 Evaluation: Prompt structure/Technique used Criterion 1: From how many concepts are the questions taken? We check the diversity of the answer from the book. Criterion 2: Relevant Content The content of the answer should always be relevant to the content of the book. Persona technique Template Few-shot prompt Chain-of-thought
46 28 OpenAI, 2023. 'GPT-4 technical report'. arXiv.org. Available at: https://arxiv.org/abs/2303.08774 (Accessed: 03 July 2023). 29 Pereira, J., 2016. 'Leveraging Chatbots to improve self-guided learning through conversational quizzes'. Proceedings of the Fourth International Conference on Technological Ecosystems for Enhancing Multiculturality [Preprint]. doi:10.1145/3012430.3012625. 30 sanfoundry.com, 'Molecular biology MCQ (multiple choice questions)'. Sanfoundry. Available at: https://www.sanfoundry.com/1000-molecular-biology-questions-answers/ (Accessed: 02 July 2023).Yao, S. et al. (2023) React: Synergizing reasoning and acting in language models, arXiv.org. Available at: https://arxiv.org/abs/2210.03629 (Accessed: 03 June 2023). 31 Voon, W., et al., 2022. Performance analysis of seven Convolutional Neural Networks (cnns) with transfer learning for invasive ductal carcinoma (IDC) grading in breast histopathological images. Nature News. Available at: https://www.nature.com/articles/s41598-022-21848-3 [Accessed 02 July 2023]. 32 Wang, X., et al., 2023. Self-consistency improves chain of thought reasoning in language models. arXiv.org. Available at: https://arxiv.org/abs/2203.11171 [Accessed 03 June 2023]. 33 Wei, J., et al., 2023. Chain-of-thought prompting elicits reasoning in large language models. arXiv.org. Available at: https://arxiv.org/abs/2201.11903 [Accessed 03 June 2023]. 34 Wei, X., et al., 2023. Zero-shot information extraction via chatting with chatgpt. arXiv.org. Available at: https://arxiv.org/abs/2302.10205 [Accessed 13 May 2023]. 35 White, J., et al., 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv.org. Available at: https://arxiv.org/abs/2302.11382 [Accessed 13 May 2023]. 36 Zavolokina, L., Dolata, M. and Schwabe, G., 2016. The fintech phenomenon: Antecedents of financial innovation perceived by the popular press - financial innovation. SpringerOpen. Available at: https://jfin-swufe.springeropen.com/articles/10.1186/s40854-016-0036-7 [Accessed 30 May 2023]. 37 Zhang, Z., et al., 2023. Multimodal chain-of-thought reasoning in language models. arXiv.org. Available at: https://arxiv.org/abs/2302.00923 [Accessed 03 June 2023]. 38 Zhao, T.Z., et al., 2021. Calibrate before use: Improving few-shot performance of language models. arXiv.org. Available at: https://arxiv.org/abs/2102.09690 [Accessed 02 June 2023].
47 Acknowledgements Stergios and Dimitris wish to express their profound gratitude to Dr. Hamam Mokayed for his invaluable guidance and unwavering support throughout this research endeavor. His depth of knowledge and expertise played a pivotal role in shaping the direction and outcomes of our study. We also acknowledge and appreciate the financial backing from the AI4EDU European project, which partially funded our research. Their contribution was instrumental in facilitating various facets of our investigative pursuits. On a personal note, Dimitris extends heartfelt appreciation to his son, Christos, for his enduring patience and understanding during the duration of the master's program. The sacrifices made in terms of shared moments have been deeply acknowledged, and Dimitris remains indebted for Christos's continual support. Additionally, Dimitris would like to express his big gratitude to his family, which even been far away, they have all be supportive giving him strength and courage to finishing this master’s with success. Furthermore, Dimitris wishes to commend Georgios for his exemplary collaboration during group assignments. His insights, dedication, and commitment to excellence significantly enriched the academic experience. Stergios, from his side wants to express deep gratitude to his family for their constant support. To his dad, Elias, and mom, Ifigeneia, for their endless love; to his brother, Labros, and sister, Vily, for always being there; and a special mention to his grandparents for always lifting his spirits and standing by his side. Also, during his master's, Stergios had the opportunity to work with Soma, Andreas, and Antonio. Their knowledge and teamwork have been invaluable, and he has learned so much from the time together. Stergios is also thankful to his colleagues and his manager, Soner, for their understanding and flexibility as he balanced work and studies. Moreover, Stergios wants to give a big thank you to all his friends who have been part of this journey with him. In closing, Stergios and Dimitris want to highlight the importance of teamwork. Writing this thesis together had its challenges, but the final result shows what can be achieved when two people collaborate effectively. The journey was filled with moments of shared discovery, mutual learning, and unwavering support. Through open communication and respect for each other's perspectives, we were able to combine our strengths, navigate challenges, and produce a piece of work that we're both proud of.
48 Appendix A Following are listed the questions, and the answers given by Study Buddy in comparison to the Teacher’s, for the three subjects of interest, Biology, Mathematics, and Physics. Biology Question Answer by Study Buddy Answer given by the teacher 1. Who coined the term cell, in reference to the tiny structures seen in living organisms? Robert Hooke coined the term cell in reference to the tiny structures seen in living organisms. The first time the word cell was used to refer to these tiny units of life was in 1665 by a British scientist named Robert Hooke. (Or can just answer Robert Hooke.) 2. Who identified animalcules? What are animalcules? Anton van Leeuwenhoek identified animalcules, which are now known as bacteria. Animalcules were discovered by Anton van Leeuwenhoek. Today, we call Leeuwenhoek’s animalcules bacteria. 3. What are the three main parts of the cell theory? The three main parts of the cell theory are: all organisms are made up of one or more cells, all the life functions of an organism occur within cells, and all cells come from preexisting cells. The three main parts of the cell theory: a. all organisms are made up of one or more cells b. all the life functions of an organism occur within cells c. all cells come from preexisting cells 4. List the four parts common to all cells. The four parts common to all cells are plasma membranes, cytoplasms, ribosomes, and DNA. The four components common to all cells are the plasma membrane, cytoplasm, ribosomes and DNA. 5. What are the cell structures where proteins are made? The cell structures where proteins are made are ribosomes. Ribosomes are structures in the cytoplasm where proteins are made. 6. What is the role of DNA? The role of DNA is not specified in the given context information. DNA contains the genetic instructions that cells need to make proteins. Table 5. Biology - Answer comparison
49 Mathematics # Question (Sequence of Numbers) Answer by Study Buddy Answer given by the teacher Result Comparison 1. 1, 6, 21, 66 Yes, I can follow the book's theory to find the pattern that applies to the sequence 1, 6, 21, 66. The pattern rule is "multiply by 3 and subtract 3." x 3 + 3 2. 95, 80, 65, 50 The pattern rule is to "subtract 15." – 15 3. 3, 10, 17, 24 Yes, the book's theory can be followed to find the pattern for the sequence 3, 10, 17, 24 and the answer is to "add 7" to each term. + 7 4. 256, 64, 16, 4 Yes, I can follow the book's theory to find the pattern that applies to the sequence 256, 64, 16, 4. The pattern rule for this sequence is "divide by 4." ÷ 4 5. 3, 11, 43, 171 Yes, I can follow the book's theory to find the pattern that applies to the sequence 3, 11, 43, 171. The pattern rule is to multiply by 4 and add 1. x 4 – 1 6. 81, 27, 9, 3 Yes, I can follow the book's theory to find the pattern that applies to the sequence 81, 27, 9, 3. The pattern rule is "divide by 3." ÷ 3 7. 4, 13, 40, 121 Yes, I can follow the book's theory to find the pattern that applies to the sequence 4, 13, 40, 121, and the final answer is "multiply by 3 and subtract 5." X 3 + 1 8. 1, 6, 31, 156 Yes, I can follow the book's theory to find the pattern that applies to the sequence 1, 6, 31, 156. The pattern rule is "multiply by 5 and subtract 4." X 5 + 1 Table 6. Mathematics - Answer comparison
50 Physics # Question Answer by Study Buddy Answer given by the teacher 1. I will give you the following problem: You are sitting on your bike at rest. Your brother comes running at you from behind at a speed of 2 m/s. At the exact moment he passes you, you start up on your bike with an acceleration of 2 m/s2. At what time t do you have the same speed as your brother? Give an answer according to the book's theory and examples. The time at which you have the same speed as your brother can be found by setting your velocity equation equal to his velocity equation and solving for time. Your brother's velocity is constant at 2 m/s, so his velocity equation is v_b = 2 m/s. Your velocity equation can be found using the kinematic equation v_f = v_i + at, where v_f is your final velocity, v_i is your initial velocity (which is 0 m/s), a is your acceleration of 2 m/s2, and t is time. Substituting in these values, we get v_f = 2t. Setting this equal to your brother's velocity equation, we get 2t = 2, which means t = 1 second. Therefore, you will have the same speed as your brother after 1 second. Let’s consider the problem as a whole (see diagram). Your brother moves at a constant velocity; therefore his position increases linearly from 0 with a slope of 2. Your position is a parabola that intersects your brother’s when your average velocity is equal to his, that is, at 2 seconds. Your brother’s velocity is a constant 2m, while yours is an upwards sloping line with a slope of 2. They intersect when your instantaneous velocities are equal, that is, at 1 second. See the diagrams below for problems 6 and 7 for an example of how to draw such a graph. 2. I will give you the following problem: You are standing at the foot of the Bank of America building in San Francisco, which is 52 floors (237 m) high. You launch a ball straight up in the air from the edge of the foot of the building. The initial vertical speed is 70 m/s. (For this problem, you may ignore your own height, which is very small compared to the height of the building.). How high up does the ball go? Give an answer by following the book's theory and examples. The ball's maximum height can be found using the equation: Δy = v_iy*t + 1/2*a_y*t2 where Δy is the change in height, v_iy is the initial vertical velocity, a_y is the acceleration due to gravity (- 9.8 m/s2), and t is the time it takes for the ball to reach its maximum height (which can be found using v_fy = v_iy + a_y*t = 0). Plugging in the given values: Δy = (70 m/s)*(7.14 s) + 1/2*(-9.8 m/s2)*(7.14 s)2 Δy = 249.8 m Again, we use mgh = .5mv2, we find that h = 250m.
51 Therefore, the ball reaches a height of approximately 250 meters. 3. I will give you the following problem: You are standing at the foot of the Bank of America building in San Francisco, which is 52 floors (237 m) high. You launch a ball straight up in the air from the edge of the foot of the building. The initial vertical speed is 70 m/s. (For this problem, you may ignore your own height, which is very small compared to the height of the building.). What is the velocity of the ball right before it hits the top of the building? Give an answer by following the book's theory and examples. The velocity of the ball right before it hits the top of the building can be found using the kinematic equation: vf2 = vi2 + 2ad, where vf is the final velocity (which is 0 m/s when the ball hits the top of the building), vi is the initial velocity (which is 70 m/s), a is acceleration due to gravity (which is -9.8 m/s2 because the ball is moving upward), and d is the displacement (which is the height of the building, 237 m). Rearranging the equation to solve for vf, we get: vf = sqrt(vi2 + 2ad) Plugging in the values, we get: vf = sqrt((70 m/s)2 + 2(-9.8 m/s2)(237 m)) vf = sqrt(4900 m2/s2 - 4644.4 m2/s2) vf = sqrt(255.6 m2/s2) vf = 15.99 m/s So, the velocity of the ball right before it hits the top of the building is approximately 15.99 m/s. We know that at the top of its journey, the ball has a speed of 0. Therefore, since it travels 13 meters at constant acceleration before it hits the building, we can use the third ’big three’ equation to find vf2 = v02 +2ax = 2×9.8×13 =⇒ vf =≈ 16m/s 4. I will give you the following problem: You are standing at the foot of the Bank of America building in San Francisco, which is 52 floors (237 m) high. You launch a ball straight up in the air from the edge of the foot of the building. The initial vertical speed is 70 m/s. (For this problem, you may ignore your own height, which is very small compared to the height of the building.). For how many Using the formula, h = vit + 1/2gt2, where h is the height, vi is the initial velocity, g is the acceleration due to gravity (-9.81 m/s2), and t is the time: At the highest point, the ball's speed is 0 m/s, so we can use vi = 70 m/s as the initial velocity. We can do this in two parts using the second ’big three’ equation (watch out for signs, though!) :v(t) =v0 + at (at the top, the velocity is 0; this is the final velocity in the first part, and the initial velocity in the second part, so the equation is simplified both times): in the flight up, the ball spends t = 70/9.8 ≈ 7.14s.
52 seconds is the ball in the air from the moment it was launched till it hits back to the ground? Give an answer by following the book's theory and examples. When the ball hits the ground, the height is 0 m. So, h = 0 m. 0 = (70 m/s)t + 1/2(-9.81 m/s2)t2 Simplifying, we get: 0 = 70t - 4.905t2 Solving for t, we get: t = 14.27 s Therefore, the ball is in the air for approximately 14.27 seconds. During the flight down, it spends 16/9.8 ≈ 1.63s So the total is about 8.77 seconds. 5. I will give you the following problem: What acceleration should you use to increase your speed from 10 m/s to 18 m/s over a distance of 55 m? Give an answer by following the book's theory and examples. To solve this problem, we can use the formula: acceleration = (final velocity squared - initial velocity squared) / (2 x distance) Plugging in the given values, we get: acceleration = (182 - 102) / (2 x 55) acceleration = 196 / 110 acceleration = 1.78 m/s2 Therefore, the acceleration required to increase the speed from 10 m/s to 18 m/s over a distance of 55 m is 1.78 m/s2. Use big three equation three, 182 = 102 + 2 × a × 55 =⇒ a ≈ 2 m/s2 6. I will give you the following problem: You drop a rock from the top of a cliff. The rock takes 3.5 seconds to reach the bottom. What is the magnitude (i.e., numerical value) of the acceleration of the rock at the moment it is dropped? Give an answer by following the book's theory and examples. The magnitude of acceleration due to gravity is 9.8 m/s2. When the rock is dropped, it experiences a constant acceleration of 9.8 m/s2 towards the ground due to gravity. Therefore, the magnitude of the acceleration of the rock at the moment it is dropped is 9.8 m/s2. Gravitational acceleration is constant: a = 9.8m/s2
53 Table 7. Physics - Answer comparison 7. I will give you the following problem: You drop a rock from the top of a cliff. The rock takes 3.5 seconds to reach the bottom. What is the height of the cliff? Give an answer by following the book's theory and examples. The height of the cliff can be calculated using the formula h = (1/2)gt2, where g is the acceleration due to gravity (9.8 m/s2) and t is the time taken for the rock to fall (3.5 seconds). Therefore, h = (1/2)(9.8 m/s2)(3.5 s)2 = 60.725 m. Thus, the height of the cliff is approximately 60.725 meters. Use big three equation one (down is positive, label cliff as zero point): x = 0+0+.5×9.8×(3.5)2 =⇒ x = 60m
54 Appendix B a. Baseline answer Question Answer Generate 3 multiple-choice questions with 4 possible answers each regarding molecular biology. The questions should test understanding of the fundamental principles of the topic. 1. Which of the following best describes the central dogma of molecular biology? a) DNA is transcribed into RNA, which is then translated into proteins. b) RNA is transcribed into DNA, which is then translated into proteins. c) Proteins are transcribed into DNA, which is then translated into RNA. d) DNA is translated into proteins, which are then transcribed into RNA. 2. What is the role of RNA in the central dogma of molecular biology? a) RNA carries the instructions from proteins to DNA. b) RNA carries the instructions from DNA to proteins. c) RNA carries the instructions from proteins to other proteins. d) RNA carries the instructions from DNA to other DNA molecules. 3. Which type of viruses do not adhere to the central dogma of molecular biology? a) Retroviruses b) Adenoviruses c) Herpesviruses d) Influenza viruses Table 8. Baseline Question Results
55 a. Prompt Structuring i. Persona Question Answer Act as a teacher and generate 3 multiple-choice questions with 4 possible answers each regarding molecular biology. The questions should test understanding of the fundamental principles of the topic. 1. What is the central dogma of molecular biology? A. DNA → RNA → Protein B. Protein → RNA → DNA C. DNA → Protein → RNA D. RNA → Protein → DNA 2. What is the process of copying genetic information from DNA to RNA called? A. Transcription B. Translation C. Replication D. Reverse Transcription 3. What type of virus does not obey the central dogma of molecular biology? A. HIV B. Influenza C. Herpes D. Retrovirus Table 9. Prompt Structuring - Persona Results ii. Template Question Answer Generate three multiple-choice questions with four possible options each on the topic of Molecular Biology, designed to test students' understanding of the fundamental principles of this subject. Please adhere to the following template for each question: Question # [placeholder for question number]: Question #1: What is the central dogma of molecular biology? A. DNA → RNA → Protein B. RNA → Protein → DNA C. Protein → DNA → RNA D. DNA → Protein → RNA