Full text
DEGREE PROJECT Chatting Over Course Material. The Role of Retrieval Augmented Generation Systems in Enhancing Academic Chatbots Hélder Monteiro Master Programme in Applied Artificial Intelligence 2024 Luleå University of Technology Department of Computer Science, Electrical and Space Engineering
[This page intentionally left blank]
Abstract Large Language Models (LLMs) have the potential to enhance learning among students. These tools can be used in chatbot systems allowing students to ask questions about course material, in particular when plugged with the so-called Retrieval Augmented Systems (RAGs). RAGs allow LLMs to access external knowledge, which improves tailored responses when used in a chatbot system. This thesis studies different RAGs through an experimentation approach where each RAG is constructed using different sets of parameters and tools, including small and large language models. We conclude by suggesting which of the RAGs best adapts to high school courses in Physics and undergraduate courses in Mathematics, such that the retrieval systems together with the LLMs are able to return the most relevant answers from provided course material. We conclude with two RAG-powered LLM with different configurations performing over 64% accuracy in physics and 66% in mathematics.
Preface In this thesis, I explore retrieval-augmented generation systems (RAGs), which is an exciting technique for those working with or interested in large language models (LLMs) and are keen on augmenting their chatbots with external knowledge. Throughout the document, I walk you through the rationale for the experiments that I conducted and what they entail, and I conclude with some remarks on the results. In the project, local LLMs were used: i) to generate synthetic question answer pairs on publicly available educational material from MIT’s OpenCourseWare in order to experiment with different RAGs; ii) to run different RAGs, and iii) to evaluate the results. The hope is that the results presented here are meaningful and can be used to further provide an understanding of RAGs, tailored for education material, such as class notes, videos, and audio.
Contents 1 Introduction 1 1.1 Goals ...................................... 3 1.2 Outline ..................................... 3 2 Background and related work 4 2.1 FromNLPtoLLMs .............................. 4 2.2 Open-sourceLLMs............................... 5 2.3 ChatbotsinEducation............................. 6 2.4 RAGTechniques ................................ 8 2.5 Synthetic Data Generation . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3 Materials and Methods 10 3.1 Synthetic Data Generation . . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.2 ToolsandLanguages.............................. 11 3.3 ExperimentDesign............................... 12 4 Results 14 4.1 SyntheticQAdata............................... 14 4.2 Retrievalcapability............................... 15 4.3 Q&AEvaluation ................................ 16 5 Discussion and Conclusion 19 5.1 Discussion.................................... 19 5.2 Conclusion ................................... 20 5.3 Futurework................................... 20 5.4 Ethical considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 Bibliography 21 A Appendices 31 A.1 RAGScenarios ................................. 31 A.2 Short ground-truth answers in Maths . . . . . . . . . . . . . . . . . . . . . 32
B Extra figures 35 B.1 Maths RAG on high performing physics RAG . . . . . . . . . . . . . . . . 35 B.2 Physics RAG on high performing maths RAG . . . . . . . . . . . . . . . . 36
List of Figures 3.1 Schematic of the pipeline to generate synthetic data. . . . . . . . . . . . . 10 4.1 Count of QA pairs generated. . . . . . . . . . . . . . . . . . . . . . . . . . 14 4.2 Accuracy per each Physics RAG (See table A.1 for detailed configuration of the RAGs shown the figure). . . . . . . . . . . . . . . . . . . . . . . . . 15 4.3 Accuracy per each Maths RAG (See table A.1 for detailed configuration of the RAGs shown the figure). . . . . . . . . . . . . . . . . . . . . . . . . 16 4.4 Count of characters (log-scale) of question, answer (RAG) and groundtruth answer for the high performing Physics RAG (#2). . . . . . . . . . 17 4.5 Count of characters (log-scale) of question, answer (RAG) and groundtruth answer for the high performing Mathematics RAG (#57). . . . . . . 18 B.1 Count of characters (log-scale) of question, answer (RAG) and groundtruth answer for Mathematics RAG (#2). . . . . . . . . . . . . . . . . . . 35 B.2 Count of characters (log-scale) of question, answer (RAG) and groundtruth answer for Physics RAG (#57). . . . . . . . . . . . . . . . . . . . . 36
List of Tables 3.1 Experiment parameters used in the study . . . . . . . . . . . . . . . . . . 12 A.1 Retrieval-augmented generation (RAG) scenarios used in the experimentation....................................... 32 A.2 Questions, Answers, and Ground Truths . . . . . . . . . . . . . . . . . . . 34
1 Introduction Since the rise of ChatGPT, there has been a lot of hype around Large Language Models (LLMs) and chatbot systems. These technologies have enabled us to improve our workflow (e.g. GitHub Copilot as code completion tool), and even being used as a study companion. LLMs are often seen as giant statistical models (Rosenfeld, 2000) trained on billions of texts from the internet made available through projects like Common Crawl1. They are capable of not only generating text in multiple natural languages but also code, do machine translation and even summarize texts. There has been different research conducted within the use of LLMs in education settings (Vacalopoulou et al., 2024; Alexandra Farazouli and McGrath, 2024; Yu, 2023; Latif et al., 2024; Xiao et al., 2023; Nechakhin, D’Souza, and Eger, 2024; Yen and Hsu, 2023) with different focuses including mathematical learning (Yen and Hsu, 2023) and impact in teachers’ assessments (Alexandra Farazouli and McGrath, 2024). In some cases, the use of LLMs is encouraged, such as at Stanford University’s “Creativity and Design Thinking Program”2course, where students submit the prompt they used that gave rise to the solution, thereby evaluating their creativity in writing prompts (Klebahn and Krakowski, 2023; Leung1and Lo, 2024). Different universities and teachers see the tool differently, either as an enabler for better education (Gaˇsevi´c, Siemens, and Sadiq, 2023) or as a detractor for effective learning (Srishti, 2024), whereby students become dependent on the tools. The research in this field has progressed at a steady pace, and has seen the rise of open-source LLMs and techniques to augment their knowledge using domain specific data. The availability of such models made its adoption on consumer hardware much easier, thanks to affordable Graphics Processing Unit (GPU) cards and techniques aimed at compressing them for inference on Central Processing Units (CPUs) only. This means that anyone with a decent computer can have locally run LLMs and chatbots that are powerful enough to generate text and allow customization for different tasks. When it comes to augmenting the knowledge of an LLM, one idea is to use what is called retrieval augmented generation (RAG) system, which simply put is a way to integrate external knowledge into the LLM using different tools. This knowledge can come from different sources: documents, databases, media files, the internet, etc., such that the model can access them beforehand, in a preprocessed form, and create the 1https://commoncrawl.org/ 2https://online.stanford.edu/how-you-can-use-chatgpt-increase-your-creative-output 1
2.4 RAG Techniques The Retrieval-Augmented Generation (RAG) system is an essential component to enhance knowledge to a large language model (LLM). Having a RAG-powered LLM addresses issues such as outdated information and hallucinations (Ding et al., 2024) which would limit the LLM in providing relevant answers to the user, particularly in the context of a chatbot. The RAG system works through a combination of data indexing, retrieval and generation (Gao et al., 2024) in an end-to-end fashion. In their survey paper, Gao et al., 2024 categorize RAGs into three types: Naive, Advanced and Modular. The naive consists solely in common pipelines for indexing, retrieving and generation. The advanced, optimizes the user query by rewriting it (Peng et al., 2024), so that relevant information is retrieved. This is useful when the query from the user is too wordy or unclear that may affect the system retrieval capability when searching for similar texts. The modular RAG is more customizable to integrate with different components within the RAG pipeline. Es et al., 2023 proposes Ragas, a framework to evaluate RAG pipelines using metrics such as faithfulness, answer and context relevance. The framework can be used to generate synthetic data that can then be used to test RAG systems. Through their documentation3, Ragas seem to require OpenAI’s REST API to generate the data, there is no mention of usage with local LLMs but the framework gives an idea of what is possible, when it comes to evaluating RAG systems. Salemi and Zamani, 2024 introduces eRAG, another framework to evaluate RAG pipelines. This tool evaluates returning answers from the RAG system against their ground-truth, meaning that if the retrieval system returns k responses, each of them is evaluated against the ground-truth and assigned a label, before the final answer is returned to the user. This differs from a more direct evaluation approach for instance, the one from Roucher, n.d. which evaluates solely on the final response from the RAG and not on each response from the retrieval system. Nonetheless, the authors argue that their framework is more efficient that existing approaches. 2.5 Synthetic Data Generation Synthetic data generation is a common practice within the AI practice (Puri et al., 2020; Shakeri et al., 2020; Riabi et al., 2020; Alberti et al., 2019; Wang et al., 2022; Roucher, n.d.) to reduce reliance on human annotations which can be expensive. Riabi et al., 2020 introduces an approach to cross-lingual synthetic data generation, making use of English question and answer (QA) model and translate the generated pairs into multiple languages. The data is then used to train better multilingual QA models. The authors say that their approach outperforms English-only baseline models (Riabi et al., 2020). Shakeri et al., 2020 build on top of the SQuAD (Rajpurkar et al., 2016) dataset and generate additional synthetic QA pairs using a transformer-based model. The model 3https://docs.ragas.io/en/stable/concepts/testset_generation.html 8
not only generates the data but also filters the best candidates using likelihood score (Shakeri et al., 2020). Puri et al., 2020 uses GPT-2 model to generate synthetic QA pairs. They breakdown text into paragraphs and pass those to an LLM to generate QA pairs. Quality checks are done using BERT (Devlin et al., 2019; Javed et al., 2022), which filters irrelevant couples (Puri et al., 2020). Similar approach is done by Roucher, n.d. which uses Mixtral8x7B-Instruct-v0.1 (Jiang et al., 2024) model having 56 billion parameters. The work of Roucher, n.d. slightly differs on that of Puri et al., 2020, where Mixtral model is used for everything: question and answer generation and evaluation, both of which done through prompts. The evaluation consists of three metrics: groundedness, relevance and standalone scores. The groundedness evaluates the truthfulness of the generated pair given the retrieved context. The relevance relates to the domain of the data, for instance if we want to evaluate the relevance for physics and mathematics, the model will be asked to assign a score based on the relevance for these fields. The standalone metric scores the QA pair on whether there is implicit mention of context in the question, which might indicate low quality of question generated. All these metrics by Roucher, n.d. take on the values between 1 to 5, which are then filtered out for a minimum of 4 across the three metrics. These studies are important for our project, specially the work of Roucher, n.d. as it can be adapted to run with slightly smaller language models within a consumer hardware, for instance Llama 3 8B4(AI@Meta, 2024) to generate and evaluate synthetic data, as well as test various RAG systems. 4https://en.wikipedia.org/wiki/Llama_(language_model) 9
3 Materials and Methods 3.1 Synthetic Data Generation Since our goal is to experiment with different RAGs, we should have question and answer pairs to evaluate each RAG that we construct. In figure 3.1 we show the schematic of the pipeline to generate synthetic data. DirectoryLoader (PyPDFLoader) Recursive Character Text Splitter Chunk size: 2000 Chunk overlap: 200 Separators: ["\n\n", "\n", ".", " ", ""] "./courses/RES.8-009/" "./courses/18.01/" Prompt Generator LLM (Llama3-8B) Unfiltered question and answer pairs Evaluator LLM (Llama3-8B) > Groundedness (1-5) > Relevance (1-5) > Standalone (1-5) filtered question and answer pairs >=4 Figure 3.1: Schematic of the pipeline to generate synthetic data. We follow the work of Roucher, n.d. that uses an open-source LLM to generate synthetic data. The author uses Mixtral-8x7B-Instruct-v0.1 (Jiang et al., 2024) model whereas in our project we use Llama3-8B model as it has shown better performance compared to earlier variants of the series with comparable size (AI@Meta, 2024) and it can be run on the available computational resources that we have. We modify the code to allow loading PDF documents from a directory using DirectoryLoader module from 10
Langchain1where we pass PyPDFLoader module as class to guide DirectoryLoader that the expected files are of PDF type and that PyPDFLoader should be used as a parser. As seen from figure 3.1, after each course folder is loaded separately, the next step is to split the data. We use Langchain’s Recursive Character splitter with the same parameters as Roucher, n.d., that is, chunk size of 2000 and chunk overlap of 200. The list of separators are default. These parameters were kept to ensure enough text is retrieved (e.g. 2000 characters per chunk, with overlap between chunks of 200 characters). The separators are the means to split the text, first starting with double newlines, followed by newline, full-stop, space and character level (no space). Each of the separators are such that the splits have the most text within the maximum allowed chunk size, and the splitter iterates over the separators to find the best that keeps relevant chunks together. For the generator LLM which generates the QA pairs, and the evaluator LLM which provide scores to the QA pairs across three metrics (groundedness, relevance and standalone) were used with Llama3-8B model using the same prompts of Roucher, n.d. However, the prompt for relevance, we modified to include that the scoring should ensure that the synthetic questions are relevant for physics and mathematics. •Relevance score (Physics) The relevance score is given depending on how useful this question can be to high school seniors taking the course Introduction To Oscillations And Waves. •Relevance score (Mathematics) The relevance score is given depending on how useful this question can be to undergraduate students taking the course Single Variable Calculus. When the question pairs are scored, we filter them to only retain those with score greatet than or equal to 4. 3.2 Tools and Languages Throughout the project we use Python2programming language to generate synthetic data and test multiple RAG pipelines. Orchestration tools are required to allow us to make use of the LLMs for building applications. Most components needed for building RAG-powered LLMs are provided by the orchestration tool. These include components for text splitting, retrieval and storage in semantic databases. Usually these components are third-party libraries that are integrated into the orchestration tool. For our project, we use LangChain as orchestration tool as it is one of the easiest and comprehensive tools that current exists besides LlamaIndex, for programmatically interact with LLMs. In order to run local LLMs we use Ollama as it best optimizes running local LLMs for different hardwares (Zimmermann and Rohrer, 2024) with or without GPU, and for 1https://en.wikipedia.org/wiki/LangChain 2https://www.python.org/ 11
its simplicity and ease of integration with different orchestration tools like LangChain. Ollama works in a similar manner as Docker, allowing to “pull” models from a repository and running them as local LLM REST APIs. 3.3 Experiment Design Major part in the project consist of experimentation. First by generation of synthetic QA data and then evaluation of different RAG systems that we carefully designed considering time and resource constraints. To generate the synthetic data and test our RAG, we focus on subjects related to undergraduate Mathematics course on Single Variable Calculus (Jerison, 2006) and high school Physics course on Introduction to Oscillation and Waves (Williams, 2017) both from MIT’s OpenCourseWare. We picked these subjects as we find that designing RAG-powered LLMs for mathematics and natural science subjects like Physics are more interesting from application point-of-view as these subjects are hard to master and thus it would be useful to students in general to enhance their learning with a RAG-powered LLM tailored for this type of educational material. In table 3.1 we have the parameters used in our experiments. Considering Cartesian product of the count of each parameter, we have a total of 64 scenarios for RAGs that we test each per course subject. For detailed combinations, see table A.1 in the appendix. Parameter Values Chunk Sizes 500, 1000 Overlaps 50, 100 Vector stores Chroma, FAISS Models Phi3, Llama3 Embedding Models mxbai-embed-large, llama3 Text Splitters CharacterTextSplitter, RecursiveCharacterTextSplitter Table 3.1: Experiment parameters used in the study The choice of chunk sizes, the idea was to have a balance between small (500) and large (1000). The overlaps were chosen in similar fashion, small (50) and large (100). The vector stores is where we store the external knowledge to do semantic search. We used two that are popular, Chroma3and Facebook AI Similarity Search (FAISS) (Douze et al., 2024). These are used without any customization, meaning that we use default parameters when using them in our experimentation pipeline, as we are focusing on finding the best RAG solely using default parameters as they are, as these can be finetuned later on once the RAG is in use. The models we use are Phi-3 (Abdin et al., 2024) from Microsoft and Llama-3 (AI@Meta, 2024) from Meta AI. These two models provide a good balance between small (3.8B parameters in Phi-3) and large (8B parameters in Llama-3) so we can study if model size affects the generation quality within the RAG-powered LLM. The same can 3https://docs.trychroma.com/ 12
be studied with the embedding models which are responsable for creating a vector space for which semantic search can be carried out. We test two models mxbai-embed-large (Sean Lee, 2024; Li and Li, 2023) from MixedBread AI that has only 335M parameters and Llama-3. The LLMs should be passed with prompts to guide them through the task. The following prompt was used in the experiments: Answer the question using only on the provided context . Only respond to what was asked without repeating the question . The response should be concise and ’ straight to the point ’. If you are unable to answer the question , say "I don ’t know ". Context: {context} Question: { question } For text splitting, we are using character-level and recursive character-based text splitters. The difference between them is that the character-level splitter splits chunks of text on each character, whereas recursive splitter allow us to decide on a list of text separators to consider, where each of one is tried until chunk size usage is maximized. If a recursive character splitter has empty string separator, it becomes a character splitter. Evaluation of the RAGs are done using the prompt with the five scores proposed by Roucher, n.d., ranging from 1 for completely incorrect/inaccurate to 5 for completely correct/accurate. The prompt are passed to the local LLM, in our case LLama3-8B that acts like a judge on the generation quality of the RAGs compared against the groundtruth answers. In order to calculate accuracy, we choose score of 3 as the cut-off for accurate results as the the score implies somewhat correct/accurate response from the RAG. 13
4 Results In this section, we present the results of the experiments, where we ran a pipeline to test different RAG combinations. 4.1 Synthetic QA data We generated a total of 183 QA pairs for physics and 200 for mathematics. After filtering for relevance, groundedness and standalone scores greater than equal to 4, we obtained 119 QA pairs for physics and 137 pairs for mathematics, see figure 4.1 for details. unfiltered filtered unfiltered filtered Subjects 0 25 50 75 100 125 150 175 200 Count (QA pairs) 183 119 200 137 Physics Mathematics Figure 4.1: Count of QA pairs generated. This generated data was then used to test the 128 RAGs that we created with different parameters. 14
4.2 Retrieval capability After running the generated data for each of the 128 RAGs, we obtained interesting results. For Physics RAGs (see figure 4.2), the maximum accuracy was 64% and that was achieved with RAG #2 having chunk size of 500, overlap 50, chroma as vector store, CharacterTextSplitter as text splitter, and embedding model was mxbai-embed-large and main LLM was Llama-3. 123456789 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 Scenario / RAG (Physics) 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Accuracy 0.64 Figure 4.2: Accuracy per each Physics RAG (See table A.1 for detailed configuration of the RAGs shown the figure). For mathematics RAGs (see figure 4.3), 66% maximum accuracy was achieved for RAG #57 having chunk size of 1000, overlap 100, Chroma as vector store, RecursiveCharacter as text splitter, and embedding model was mxbai-embed-large and main LLM was Phi-3. This was completely opposed to what we obtained for the Physics RAGs. 15
123456789 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 Scenario / RAG (Mathematics) 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Accuracy 0.66 Figure 4.3: Accuracy per each Maths RAG (See table A.1 for detailed configuration of the RAGs shown the figure). 4.3 Q&A Evaluation For the highly performing RAGs for physics and mathematics we plotted the character count for the question created during synthetic data generation process, and compare the count of the ground-truth answer and the answer returned by the individual RAGs. From figure 4.4, we see that the RAG answers are of comparable size to the ground truth, even though there are some picks in either of them. On the other hand, the character count for mathematics RAG (see figure 4.5) that performed best, the number of characters count is highest in most cases for the generated answer. The ground-truth answers were relatively short. 16
0 20 40 60 80 100 120 Question/Answer index 100 101 102 Character count (log scale) Count of characters // RAG 2 // Physics Question Answer Ground Truth Figure 4.4: Count of characters (log-scale) of question, answer (RAG) and ground-truth answer for the high performing Physics RAG (#2). 17
Gunes, Yasin Celal and Turay Cesur (2024). “A Comparative Study: Diagnostic Performance of ChatGPT 3.5, Google Bard, Microsoft Bing, and Radiologists in Thoracic Radiology Cases”. In: medRxiv, pp. 2024–01. Gwon, Yong Nam, Jae Heon Kim, Hyun Soo Chung, Eun Jee Jung, Joey Chun, Serin Lee, and Sung Ryul Shim (2024). “The Use of Generative AI for Scientific Literature Searches for Systematic Reviews: ChatGPT and Microsoft Bing AI Performance Evaluation”. In: JMIR Medical Informatics 12, e51187. Hochreiter, Sepp and J¨urgen Schmidhuber (1997). “Long short-term memory”. In: Neural computation 9.8, pp. 1735–1780. Hum, Yan Chai, Yee Kai Tee, Wun-She Yap, Hamam Mokayed, Tian Swee Tan, Maheza Irna Mohamad Salim, and Khin Wee Lai (2022). “A contrast enhancement framework under uncontrolled environments based on just noticeable difference”. In: Signal Processing: Image Communication 103, p. 116657. Ide, Nancy and Jean V´eronis (1998). “Introduction to the special issue on word sense disambiguation: the state of the art”. In: Computational linguistics 24.1, pp. 1–40. Javed, Saleha, Fredrik Sandin, Hamam Mokayed, Jerker Delsing, and Marcus Liwicki (2022). “Deep Ontology Alignment with BERT INT: Improvements and Industrial Internet of Things (IIoT) Case Study”. In. Javed, Saleha, Muhammad Usman, Fredrik Sandin, Marcus Liwicki, and Hamam Mokayed (2023a). “Deep Ontology Alignment Using a Natural Language Processing Approach for Automatic M2M Translation in IIoT”. In: Sensors 23.20, p. 8427. Javed, Salman, Aparajita Tripathy, Jan van Deventer, Hamam Mokayed, Cristina Paniagua, and Jerker Delsing (2023b). “An approach towards demand response optimization at the edge in smart energy systems using local clouds”. In: Smart Energy 12, p. 100123. Jerison, David (2006). Single Variable Calculus. MIT OpenCourseWare: Massachusetts Institute of Technology. 18.01 (Fall 2006). Licensed under CC BY-NC-SA 4.0. Available at https://ocw.mit.edu/courses/18-01single-variablecalculusfall-2006/. Jiang, Albert Q., Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L´elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timoth´ee Lacroix, and William El Sayed (2023). Mistral 7B. arXiv: 2310.06825 [cs.CL]. Jiang, Albert Q, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. (2024). “Mixtral of experts”. In: arXiv preprint arXiv:2401.04088. 24
Jiao, Xiaoqi, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu (2019). “Tinybert: Distilling bert for natural language understanding”. In: arXiv preprint arXiv:1909.10351. Jurafsky, Daniel and James H Martin (n.d.). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Khalid, Marzuki, Rubiyah Yusof, and Hamam Mokayed (2011). “Fusion of multi-classifiers for online signature verification using fuzzy logic inference”. In: International Journal of Innovative Computing 7.5, pp. 2709–2726. Klebahn, Perry and Sebastian Krakowski (2023). How You Can Use ChatGPT to Increase Your Creative Output.https://online.stanford.edu/how-you-can-usechatgpt-increase-your-creative-output. Accessed: 2024-04-10. Lan, Zhenzhong, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut (2020). ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. arXiv: 1909.11942 [cs.CL]. Latif, Ehsan, Luyang Fang, Ping Ma, and Xiaoming Zhai (2024). Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments. arXiv: 2312.15842 [cs.CL]. Leung1, Rosanna and Iris Sheungting Lo (2024). “Check for Can ChatGPT Inspire Me? Evaluate Students’ Questioning Techniques on AI Tool for Overcoming Fixation Rosanna Leung1and Iris Sheungting Lo2”. In: Information and Communication Technologies in Tourism 2024: ENTER 2024 International eTourism Conference, Izmir, T¨urkiye, January 17–19. Springer Nature, p. 75. Li, Xianming and Jing Li (2023). “AnglE-optimized Text Embeddings”. In: arXiv preprint arXiv:2309.12871. Li, Yuhua, David McLean, Zuhair A Bandar, James D O’shea, and Keeley Crockett (2006). “Sentence similarity based on semantic nets and corpus statistics”. In: IEEE transactions on knowledge and data engineering 18.8, pp. 1138–1150. Lieb, Anna and Toshali Goel (2024). “Student Interaction with NewtBot: An LLM-astutor Chatbot for Secondary Physics Education”. In: Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems. CHI EA ’24. ¡conf-loc¿ ¡city¿Honolulu¡/city¿ ¡state¿HI¡/state¿ ¡country¿USA¡/country¿ ¡/conf-loc¿: Association for Computing Machinery. isbn: 9798400703317. doi:10 . 1145 / 3613905 . 3647957.url:https://doi.org/10.1145/3613905.3647957. Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv: 1907.11692 [cs.CL]. 25
Luo, Ziyang, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang (2023). “Wizardcoder: Empowering code large language models with evol-instruct”. In: arXiv preprint arXiv:2306.08568. Maryamah, Maryamah, Muhammad Maula Irfani, Edric Boby Tri Raharjo, Netri Alia Rahmi, Mohammad Ghani, and Indra Kharisma Raharjana (2024). “Chatbots in Academia: A Retrieval-Augmented Generation Approach for Improved Efficient Information Access”. In: 2024 16th International Conference on Knowledge and Smart Technology (KST), pp. 259–264. doi:10.1109/KST61284.2024.10499652. Meta Platforms, Inc. (2023). Llama 2 License Agreement.https://github.com/metallama/llama/blob/main/LICENSE. Version Release Date: July 18, 2023. Ireland and USA: Meta Platforms. Minaee, Shervin, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao (2024). “Large language models: A survey”. In: arXiv preprint arXiv:2402.06196. Mokayed, Hamam, Liang Kim Meng, Hon Hock Woon, and Ng Hooi Sin (2014). “Car plate detection engine based on conventional edge detection technique”. In: The International Conference on Computer Graphics, Multimedia and Image Processing (CGMIP2014). The Society of Digital Information and Wireless Communication. Mokayed, Hamam, Amirhossein Nayebiastaneh, Lama Alkhaled, Stergios Sozos, Olle Hagner, and Bj¨orn Backe (2024). “Challenging YOLO and Faster RCNN in Snowy Conditions: UAV Nordic Vehicle Dataset (NVD) as an Example”. In: 2024 2nd International Conference on Unmanned Vehicle Systems-Oman (UVS). IEEE, pp. 1– 6. Mokayed, Hamam, Amirhossein Nayebiastaneh, Kanjar De, Stergios Sozos, Olle Hagner, and Bj¨orn Backe (2023). “Nordic Vehicle Dataset (NVD): Performance of vehicle detectors using newly captured NVD from UAV in different snowy weather conditions.” In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5313–5321. Mokayed, Hamam, Shivakumara Palaiahnakote, Lama Alkhaled, and Ahmed N ALMasri (2022). “License plate number detection in drone images”. In: Artificial Intelligence and Applications. Mokayed, Hamam, Palaiahnakote Shivakumara, Marcus Liwicki, and Umapada Pal (2020). “A new defect detection method for improving text detection and Recognition performances in natural scene images”. In: 2020 Swedish Workshop on Data Science (SweDS). IEEE, pp. 1–7. Mokayed, Hamam, Palaiahnakote Shivakumara, Rajkumar Saini, Marcus Liwicki, Loo Chee Hin, and Umapada Pal (2021). “Anomaly detection in natural scene images based on enhanced fine-grained saliency and fuzzy logic”. In: IEEE Access 9, pp. 129102– 129109. 26
Nasukawa, Tetsuya and Jeonghee Yi (2003). “Sentiment analysis: Capturing favorability using natural language processing”. In: Proceedings of the 2nd international conference on Knowledge capture, pp. 70–77. Navigli, Roberto (2009). “Word sense disambiguation: A survey”. In: ACM computing surveys (CSUR) 41.2, pp. 1–69. Nechakhin, Vladyslav, Jennifer D’Souza, and Steffen Eger (2024). Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph. arXiv: 2405.02105 [cs.AI]. Neupane, Subash, Elias Hossain, Jason Keith, Himanshu Tripathi, Farbod Ghiasi, Noorbakhsh Amiri Golilarz, Amin Amirlatifi, Sudip Mittal, and Shahram Rahimi (2024). From Questions to Insightful Answers: Building an Informed Chatbot for University Resources. arXiv: 2405.08120 [cs.ET]. Nikolaidou, Konstantina, George Retsinas, Vincent Christlein, Mathias Seuret, Giorgos Sfikas, Elisa Barney Smith, Hamam Mokayed, and Marcus Liwicki (2023). “Wordstylist: styled verbatim handwritten text generation with latent diffusion models”. In: International Conference on Document Analysis and Recognition. Springer Nature Switzerland Cham, pp. 384–401. OpenAI et al. (2024). GPT-4 Technical Report. arXiv: 2303.08774 [cs.CL]. Peng, Wenjun, Guiyang Li, Yue Jiang, Zilong Wang, Dan Ou, Xiaoyi Zeng, Derong Xu, Tong Xu, and Enhong Chen (2024). “Large language model based long-tail query rewriting in taobao search”. In: Companion Proceedings of the ACM on Web Conference 2024, pp. 20–28. Pickering, Martin J and Roger PG Van Gompel (2006). “Syntactic parsing”. In: Handbook of psycholinguistics. Elsevier, pp. 455–503. Puri, Raul, Ryan Spring, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro (2020). “Training question answering models from synthetic data”. In: arXiv preprint arXiv:2002.09599. Radev, Dragomir, Eduard Hovy, and Kathleen McKeown (2002). “Introduction to the special issue on summarization”. In: Computational linguistics 28.4, pp. 399–408. Radford, Alec, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. (2018). “Improving language understanding by generative pre-training”. In. Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang (2016). “Squad: 100,000+ questions for machine comprehension of text”. In: arXiv preprint arXiv:1606.05250. Ram, Bal and Pratima Verma (2023). “Artificial intelligence AI-based Chatbot study of ChatGPT, Google AI Bard and Baidu AI”. In: World Journal of Advanced Engineering Technology and Sciences 8.01, pp. 258–261. 27
Resnik, Philip (1999). “Semantic similarity in a taxonomy: An information-based measure and its application to problems of ambiguity in natural language”. In: Journal of artificial intelligence research 11, pp. 95–130. Riabi, Arij, Thomas Scialom, Rachel Keraron, Benoˆıt Sagot, Djam´e Seddah, and Jacopo Staiano (2020). “Synthetic data augmentation for zero-shot cross-lingual question answering”. In: arXiv preprint arXiv:2010.12643. Rosenfeld, Ronald (2000). “Incorporating linguistic structure into statistical language models”. In: Philosophical Transactions of the Royal Society of London. Series A: Mathematical, Physical and Engineering Sciences 358.1769, pp. 1311–1324. Roucher, Aymeric (n.d.). RAG Evaluation.https://huggingface.co/learn/cookbook/ en/rag_evaluation. Accessed: 2024-04-04. Roumeliotis, Konstantinos I, Nikolaos D Tselikas, and Dimitrios K Nasiopoulos (2023). “Llama 2: Early Adopters’ Utilization of Meta’s New Open-Source Pretrained Model”. In. Roy, Ayush, Palaiahnakote Shivakumara, Umapada Pal, Hamam Mokayed, and Marcus Liwicki (2023). “Fourier feature-based CBAM and vision transformer for text detection in drone images”. In: International Conference on Document Analysis and Recognition. Springer Nature Switzerland Cham, pp. 257–271. Salemi, Alireza and Hamed Zamani (2024). “Evaluating Retrieval Quality in RetrievalAugmented Generation”. In: arXiv preprint arXiv:2404.13781. Sanh, Victor, Lysandre Debut, Julien Chaumond, and Thomas Wolf (2019). “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”. In: arXiv preprint arXiv:1910.01108. Sean Lee Aamir Shakir, Darius Koenig-Julius Lipp (2024). Open Source Strikes Bread - New Fluffy Embeddings Model.url:https://www.mixedbread.ai/blog/mxbaiembed-large-v1. Shakeri, Siamak, Cicero dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang (2020). “End-to-end synthetic data generation for domain adaptation of question answering systems”. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5445– 5460. Shoufan, Abdulhadi (2023). “Exploring Students’ Perceptions of ChatGPT: Thematic Analysis and Follow-Up Survey”. In: IEEE Access 11, pp. 38805–38818. doi:10. 1109/ACCESS.2023.3268224. Srishti, Richa (2024). “ChatGPT in Education: Augmenting Learning Experience or Dehumanizing Education?” In: Educational Perspectives on Digital Technologies in Modeling and Management. IGI Global, pp. 114–128. 28
Sutskever, Ilya (2013). Training recurrent neural networks. University of Toronto Toronto, ON, Canada. Sutskever, Ilya, James Martens, and Geoffrey E Hinton (2011). “Generating text with recurrent neural networks”. In: Proceedings of the 28th international conference on machine learning (ICML-11), pp. 1017–1024. Taori, Rohan, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto (2023). “Alpaca: A strong, replicable instruction-following model”. In: Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html 3.6, p. 7. Team, Gemma, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi`ere, Mihir Sanjay Kale, Juliette Love, et al. (2024). “Gemma: Open models based on gemini research and technology”. In: arXiv preprint arXiv:2403.08295. Touvron, Hugo, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. (2023). “Llama: Open and efficient foundation language models”. In: arXiv preprint arXiv:2302.13971. Vacalopoulou, Anna, Viktor Gardelli, Theodoris Karafyllidis, Foteini Liwicki, Hamam Mokayed, Marios Papaevripidou, George Paraskevopoulos, Spyridoula Stamouli, Athanasios Katsamanis, and Vassilis Katsouros (2024). “AI4EDU: An Innovative Conversational Ai Assistant For Teaching And Learning”. In: INTED2024 Proceedings. IATED, pp. 7119–7127. Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin (2017). “Attention is all you need”. In: Advances in neural information processing systems 30. Wang, Jiayi, David Ifeoluwa Adelani, Sweta Agrawal, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Marek Masiak, Xuanli He, Sofia Bourhim, Andiswa Bukula, et al. (2023). “AfriMTE and AfriCOMET: Empowering COMET to Embrace Underresourced African Languages”. In: arXiv preprint arXiv:2311.09828. Wang, Yizhong, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi (2022). “Self-instruct: Aligning language models with self-generated instructions”. In: arXiv preprint arXiv:2212.10560. Williams, Mobolaji (2017). Introduction to Oscillations and Waves. MIT OpenCourseWare: Massachusetts Institute of Technology. RES.8-009 (Summer 2017). Licensed under CC BY-NC-SA 4.0. Available at https://ocw.mit.edu/courses/res-8009-introduction-to-oscillations-and-waves-summer-2017/. Xiao, Changrong, Sean Xin Xu, Kunpeng Zhang, Yufang Wang, and Lei Xia (July 2023). “Evaluating Reading Comprehension Exercises Generated by LLMs: A Showcase 29
of ChatGPT in Education Applications”. In: Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023). Ed. by Ekaterina Kochmar, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, Nitin Madnani, Ana¨ıs Tack, Victoria Yaneva, Zheng Yuan, and Torsten Zesch. Toronto, Canada: Association for Computational Linguistics, pp. 610–625. doi:10.18653/ v1/2023.bea-1.52.url:https://aclanthology.org/2023.bea-1.52. Yen, An-Zi and Wei-Ling Hsu (Dec. 2023). “Three Questions Concerning the Use of Large Language Models to Facilitate Mathematics Learning”. In: Findings of the Association for Computational Linguistics: EMNLP 2023. Ed. by Houda Bouamor, Juan Pino, and Kalika Bali. Singapore: Association for Computational Linguistics, pp. 3055–3069. doi:10.18653/v1/2023.findingsemnlp.201.url:https:// aclanthology.org/2023.findings-emnlp.201. Yu, Hao (2023). “Reflection on whether Chat GPT should be banned by academia from the perspective of education and teaching”. In: Frontiers in Psychology 14, p. 1181712. Zhang, Peiyuan, Guangtao Zeng, Tianduo Wang, and Wei Lu (2024). “Tinyllama: An open-source small language model”. In: arXiv preprint arXiv:2401.02385. Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. (2024). “Judging llmas-a-judge with mt-bench and chatbot arena”. In: Advances in Neural Information Processing Systems 36. Zimmermann, Lucien and Florian Rohrer (2024). Study Buddy. OST-Ostschweizer Fachhochschule. url:https://eprints.ost.ch/id/eprint/1176/. 30
Appendices A Appendices A.1 RAG Scenarios The table A.1 shows the full list of RA scenarios used in the experimentation. There are a total of 64 RAG combinations. Scenario Chunk Size Overlap Text Splitter Vector Store Embedding Model Model 1 500 50 Character Chroma mxbai-embed-large phi3 2 500 50 Character Chroma mxbai-embed-large llama3 3 500 50 Character Chroma llama3 phi3 4 500 50 Character Chroma llama3 llama3 5 500 50 Character FAISS mxbai-embed-large phi3 6 500 50 Character FAISS mxbai-embed-large llama3 7 500 50 Character FAISS llama3 phi3 8 500 50 Character FAISS llama3 llama3 9 500 50 RecursiveCharacter Chroma mxbai-embed-large phi3 10 500 50 RecursiveCharacter Chroma mxbai-embed-large llama3 11 500 50 RecursiveCharacter Chroma llama3 phi3 12 500 50 RecursiveCharacter Chroma llama3 llama3 13 500 50 RecursiveCharacter FAISS mxbai-embed-large phi3 14 500 50 RecursiveCharacter FAISS mxbai-embed-large llama3 15 500 50 RecursiveCharacter FAISS llama3 phi3 16 500 50 RecursiveCharacter FAISS llama3 llama3 17 500 100 Character Chroma mxbai-embed-large phi3 18 500 100 Character Chroma mxbai-embed-large llama3 19 500 100 Character Chroma llama3 phi3 20 500 100 Character Chroma llama3 llama3 21 500 100 Character FAISS mxbai-embed-large phi3 22 500 100 Character FAISS mxbai-embed-large llama3 23 500 100 Character FAISS llama3 phi3 24 500 100 Character FAISS llama3 llama3 31
25 500 100 RecursiveCharacter Chroma mxbai-embed-large phi3 26 500 100 RecursiveCharacter Chroma mxbai-embed-large llama3 27 500 100 RecursiveCharacter Chroma llama3 phi3 28 500 100 RecursiveCharacter Chroma llama3 llama3 29 500 100 RecursiveCharacter FAISS mxbai-embed-large phi3 30 500 100 RecursiveCharacter FAISS mxbai-embed-large llama3 31 500 100 RecursiveCharacter FAISS llama3 phi3 32 500 100 RecursiveCharacter FAISS llama3 llama3 33 1000 50 Character Chroma mxbai-embed-large phi3 34 1000 50 Character Chroma mxbai-embed-large llama3 35 1000 50 Character Chroma llama3 phi3 36 1000 50 Character Chroma llama3 llama3 37 1000 50 Character FAISS mxbai-embed-large phi3 38 1000 50 Character FAISS mxbai-embed-large llama3 39 1000 50 Character FAISS llama3 phi3 40 1000 50 Character FAISS llama3 llama3 41 1000 50 RecursiveCharacter Chroma mxbai-embed-large phi3 42 1000 50 RecursiveCharacter Chroma mxbai-embed-large llama3 43 1000 50 RecursiveCharacter Chroma llama3 phi3 44 1000 50 RecursiveCharacter Chroma llama3 llama3 45 1000 50 RecursiveCharacter FAISS mxbai-embed-large phi3 46 1000 50 RecursiveCharacter FAISS mxbai-embed-large llama3 47 1000 50 RecursiveCharacter FAISS llama3 phi3 48 1000 50 RecursiveCharacter FAISS llama3 llama3 49 1000 100 Character Chroma mxbai-embed-large phi3 50 1000 100 Character Chroma mxbai-embed-large llama3 51 1000 100 Character Chroma llama3 phi3 52 1000 100 Character Chroma llama3 llama3 53 1000 100 Character FAISS mxbai-embed-large phi3 54 1000 100 Character FAISS mxbai-embed-large llama3 55 1000 100 Character FAISS llama3 phi3 56 1000 100 Character FAISS llama3 llama3 57 1000 100 RecursiveCharacter Chroma mxbai-embed-large phi3 58 1000 100 RecursiveCharacter Chroma mxbai-embed-large llama3 59 1000 100 RecursiveCharacter Chroma llama3 phi3 60 1000 100 RecursiveCharacter Chroma llama3 llama3 61 1000 100 RecursiveCharacter FAISS mxbai-embed-large phi3 62 1000 100 RecursiveCharacter FAISS mxbai-embed-large llama3 63 1000 100 RecursiveCharacter FAISS llama3 phi3 64 1000 100 RecursiveCharacter FAISS llama3 llama3 Table A.1: Retrieval-augmented generation (RAG) scenarios used in the experimentation. A.2 Short ground-truth answers in Maths The table A.2 shows the top 20 synthetic question and ground-truth pairs and the answer generated by maths RAG having configuration 2 (the configuration that Physics 32
had highest accuracy). 33