scieee AI-readable full text Open interactive document viewer

A CMS AI to Improve Paper Publishing

Álvarez-Cascos, Annunziata; Rehm, Florian

Abstract

This project explores the potential of AI-powered automation in the peer review process for CMS experiment publications at CERN, using the LLaMA 3.1 language model. The research initially employed a fine-tuning approach through Parameter-Efficient Fine-Tuning (PEFT), which, while effective in automating certain aspects such as ensuring consistency with formatting and style guidelines, required significant resources and time. To address these limitations, a prompt engineering approach was later adopted, significantly reducing the time required while improving the accuracy and consistency of reviews. By embedding a summary of CMS internal publication guidelines directly within the prompts, the model was able to perform without additional task-specific training. Both approaches demonstrated the potential for AI to streamline the peer review process, with the prompt engineering method showing particular promise in efficiency. The findings lay a strong foundation for future improvements, including further customization of training processes, dataset expansion, and the development of a graphical user interface (GUI), while also contributing valuable insights into the current capabilities and limitations of AI models in specialized domains.

Full text

A CMS AI to Improve Paper Publishing AUGUST 2024 AUTHOR(S): Annunziata Álvarez-Cascos Hervías BE-CSS DSB SUPERVISOR(S): Florian Rehm CERN OpenLab Report // 2024 3 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías ABSTRACT ç This project explores the potential of AI-powered automation in the peer review process for CMS experiment publications at CERN, using the LLaMA 3.1 language model. The research initially employed a fine-tuning approach through Parameter-Efficient Fine-Tuning (PEFT), which, while effective in automating certain aspects such as ensuring consistency with formatting and style guidelines, required significant resources and time. To address these limitations, a prompt engineering approach was later adopted, significantly reducing the time required while improving the accuracy and consistency of reviews. By embedding a summary of CMS internal publication guidelines directly within the prompts, the model was able to perform without additional task-specific training. Both approaches demonstrated the potential for AI to streamline the peer review process, with the prompt engineering method showing particular promise in efficiency. The findings lay a strong foundation for future improvements, including further customization of training processes, dataset expansion, and the development of a graphical user interface (GUI), while also contributing valuable insights into the current capabilities and limitations of AI models in specialized domains. CERN OpenLab Report // 2024 4 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías TABLE OF CONTENTS INTRODUCTION OBJECTIVES BACKGROUND AI-ASSISTED PEER-REVIEW LARGE LANGUAGE MODELS PART 1: FINE-TUNING APPROACH METHODOLOGY DATA COLLECTION INITIAL SET UP OF LLAMA3.1 MODEL FINE-TUNING PEFTs HYPERPARAMETERS PROMPT ENGINEERING EXPERIMENTAL DESIGN RESULTS AND ANALYSIS CONCLUSION AND FUTURE STEPS PART 2: PROMPT ENGINEERING APPROACH METHODOLOGY RESULTS FUTURE WORK APPENDICES 05 05 06 1 10 10 1 12 1 13 1 14 15 16 CERN OpenLab Report // 2024 5 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías 1. INTRODUCTION Scientific paper publishing, especially at CERN and its experiments, follows a rigorous process where each paper must undergo a peer review before it can be made public. At CERN, scientists from over 100 nationalities work together, bringing a variety of writing styles to their papers. This diversity in writing styles makes the peer review process not only time-consuming but also challenging, as reviewers must carefully evaluate each article's style. To alleviate the burden and reduce the time required for peer review, the need for AI assistance in this process has become increasingly important. Automated peer reviewing tasks would not only involve correcting the entire text but also providing detailed feedback on the changes made. Recent advancements in Natural Language Processing (NLP) and Large Language Models (LLMs) have proven to be more efficient and accurate, making them suitable for handling these time-intensive tasks effectively. In this study, we present an AI-powered peer reviewer framework that utilizes LLaMA3.1 (LLM from Meta AI) to automate the peer review process for CMS experiment paper publishing. To tackle the computational challenges associated with fine-tuning a LLM, we will also integrate Parameter-Efficient Fine-Tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA). a. OBJECTIVES The main objectives of this project are: • Streamline the Peer Review Process: Use LLMs, specifically a fine-tuned Llama 3.1 model, to automate and enhance the peer review process, making it more efficient and consistent. • Reduce Time and Effort: Decrease the time and manual effort required for peer reviewing by automating parts of the process, allowing for quicker turnaround times without sacrificing quality. • Optimize Model Training: Use techniques like PEFT and LoRA to fine-tune the model efficiently, reducing computational resources while maintaining high performance. 2. BACKGROUND a. AI-ASSITED PEER REVIEW Peer reviewing is the process of assessing the quality, validity and originality of a paper before is published. It is a crucial but time-consuming process. The diversity in writing styles and the inconsistencies in following guidelines often slow down the review process, leading to misunderstandings and delays. This is particularly challenging in the context of CMS experiment papers, where precision and clarity are paramount. Therefore, automatizing this process in the most accurate way would be crucial for ensuring that reviews are consistent, efficient, and maintain the high standards required for CMS experiment papers. b. LARGE LANGUAGE MODELS Large Language Models (LLMs) are large neural networks trained on enormous datasets, enabling them to perform a wide range of natural language generation tasks with high quality. The main purpose of LLMs is to understand, generate, and interact with human language in a meaningful and contextually relevant way. CERN OpenLab Report // 2024 6 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías LLMs are trained using self-supervised learning, where models learn from unlabeled data by predicting the next word in a sentence. This allows them to leverage vast amounts of raw text. The Transformer architecture, with its self-attention mechanism, enables models to focus on the context of all words in a sentence, making word associations efficiently. This parallel processing capability makes Transformers highly effective for training large-scale language models. Meta's LLaMA, for instance, focuses on efficient language model training with fewer parameters but strong performance, making it suitable for research and applications where resources are limited. It’s used for tasks like text generation, translation, and summarization. Additionally, this models are open-sources, which facilitates broad accessibility and communitydriven improvements. 3. PART 1: FINE-TUNING APPROACH 3.1 METHODOLOGY The process of automating peer review begins by selecting a pre-trained model, which in this case is Llama3.1-8B-Instruct. Since pre-training large language models (LLMs) for peer review can be expensive, we opted to use fine-tuning and prompt engineering as practical and costeffective approaches to adapt the model for this specific task. The next step involves defining the objective of the automation. I then prepared a dataset consisting of over 1,300 peer-reviewed LaTeX documents and proceeded with fine-tuning and prompt engineering approaches. Following this, the model's performance was evaluated, and based on the results, adjustments and iterations were made as necessary before deploying the model for actual use. In the following sections, I will provide a more detailed explanation of the most important steps. A. DATA COLLECTION The data collection process involved web scraping over 1,300 peer-reviewed LaTeX documents from reputable scientific sources. The documents were divided into paragraphs, with each paragraph forming a row in the dataset. During preprocessing, empty rows and incomplete documents were removed to ensure data quality. We ensured compliance with copyright laws by only including open-access documents or those with proper permissions. This clean and ethically sourced dataset was then used to train the model in understanding the structure and content of scientific papers. B. INITIAL SET UP OF LLAMA3.1 The initial setup of Llama3.1-8B-Instruct involved loading the pre-trained model and tokenizer. The model was sourced from the specified directory, and the tokenizer was configured to handle padding appropriately by assigning the end-of-sequence (EOS) token as the padding token. The dataset was then pre-processed to tokenize the text data, ensuring that it was truncated and padded to a maximum length of 128 tokens. The tokenized input was prepared with corresponding labels to enable effective training for the causal language modelling task. All experiments were conducted on NVIDIA A100-PCIE-40GB GPU platforms, which provided the necessary computational power to handle the large model and extensive fine-tuning process. This setup ensured that the model was properly initialized and ready for fine-tuning with the given datasets. CERN OpenLab Report // 2024 7 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías C. MODEL FINE-TUNING One of the approaches I followed to try to achieve an accurate model for peer-review automation was fine-tuning. Fine-tuning refers to the practice of adapting a pre-trained model to a related, but not identical, task or dataset from the one it was originally trained on. i. PEFTs To enhance the model's efficiency and accuracy, we employed a PEFT approach, specifically using LoRA. PEFT is important because it allows us to fine-tune large models like Llama3.1-8BInstruct with fewer trainable parameters. This is crucial due to the model's size and the need to optimize computational resources. LoRA achieves this by introducing low-rank matrices to specific model components, rather than modifying the entire model. By focusing on key modules such as q_proj and v_proj, which are integral to the model's attention mechanisms, LoRA allows for targeted adjustments. This focused adaptation helps improve the model’s performance on specialized tasks while minimizing computational demands. As a result, we can efficiently enhance the model’s capabilities without the need for extensive resources or compromising its core pre-trained knowledge. Key LoRA parameters included: • lora_alpha: Scales the impact of LoRA layers. We set it to 64 to ensure significant adaptation. • lora_dropout: Applies dropout to LoRA layers to prevent overfitting. We used 0.05 for balanced regularization. • lora_r: Defines the rank of low-rank matrices, set to 64 to capture sufficient complexity without excessive resource use. • target_modules: Focused on q_proj and v_proj modules, crucial for model attention mechanisms. ii. HYPERPARAMETERS During the fine-tuning of the Llama3.1-8B-Instruct model, we carefully selected and adjusted several training hyperparameters to optimize model performance. These training arguments significantly impact how effectively the model learns and generalizes from the data. We experimented with various configurations to identify the most effective settings for our specific task. Key adjustments included setting the token length limit to 2048 tokens and using a batch size of 2. We employed the adamw_hf optimizer, which is well-suited for large models and provides effective weight decay regularization. Additionally, we trained the model for various numbers of epochs—5, 10, 15, and 20—to observe how the training and validation loss converged. Through these experiments, we found that training for 20 epochs provided the best results, with the most effective convergence of both training and validation loss. Each parameter was adjusted iteratively, and different values were tested to find the optimal settings that balanced performance, efficiency, and resource usage. This thorough tuning process ensured that the fine-tuning was both effective and computationally efficient, leading to improved model performance on the peer-review task. CERN OpenLab Report // 2024 8 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías D. PROMPT ENGINEERING The second approach to enhancing the model for peer-review automation involved prompt engineering, which provides explicit instructions to guide the model’s generation process. Unlike fine-tuning, which adjusts the model's weights, prompt engineering leverages the pre-trained model's capabilities by crafting input prompts that effectively direct its behavior. We employed two primary techniques for prompt engineering: • Zero-Shot Learning: This technique involves designing prompts that enable the model to perform the peer review task without any additional task-specific training. This approach was used to assess the model's ability to generalize its pre-existing knowledge to the new task with no additional examples. An example of a zero-shot prompt template is presented in figure 1. • Few-Shot Learning with Input–Output Examples: In this approach, we provided the model with 4 examples within the prompt to guide its responses more effectively. These examples illustrated the type of feedback or review comments expected, helping the model generate more relevant and accurate peer review outputs. For clarity, the template used for this approach is illustrated in Figure 2. • Few-Shot Learning with CMS Writing Guidelines: In this approach, we incorporated a summary of the CMS writing guidelines into the prompt. These guidelines outline the standards for writing paper publications related to CMS experiments. By providing the model with a condensed version of these guidelines, we aimed to help the model generate peer review comments that align with the expected format and content. We experimented with various numbers of examples to identify the optimal amount that enhanced performance while ensuring the model was not overwhelmed with excessive information. The template used for integrating these guidelines is shown in Figure 3. Through these experiments, we aimed to determine how different prompting strategies impacted the model's ability to perform peer review tasks. The effectiveness of each approach was evaluated based on the relevance and quality of the generated reviews, as well as the model's ability to generalize from the prompts provided. This exploration of prompt engineering techniques allowed us to leverage the model’s capabilities effectively while minimizing the need for extensive task-specific training. CERN OpenLab Report // 2024 9 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías Figure 1. Example of Zero-Shot Learning Figure 2. Example of Few-Shot Learning with Input–Output Examples. CERN OpenLab Report // 2024 10 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías Figure 3. Example of Few-Shot Learning with CMS Writing Guidelines. 3.2 EXPERIMENTAL DESIGN To evaluate our model’s performance, we designed a comprehensive experimental framework, which involved creating an inference script to process and assess LaTeX documents. The inference process begins with initializing the pre-trained model and tokenizer, which are loaded from a specified directory. For correcting paragraphs, the script formats the text using a chat template that includes detailed instructions for the model. This formatted input is then passed through the model, which generates corrections based on the provided parameters, such as token limits and sampling methods. The output from the model is decoded to retrieve the corrected text. Additionally, the script processes LaTeX documents by reading and correcting each section, while carefully preserving the document’s formatting and handling specific elements like figures and tables. This approach ensures that the LaTeX documents are accurately reviewed and corrected using the model. For a detailed view of the code used for model inference, please refer to the repository https://gitlab.cern.ch/dsb/proofreading-llm.git 3.3 RESULTS AND ANALYSIS Our primary quantitative metrics for evaluating the model's performance were the training loss and validation loss. Although these metrics provide an overview of the model’s convergence, they have limitations in reflecting the model's effectiveness in real-world peer review tasks. • Training Loss: This metric decreased steadily throughout the training process, indicating that the model was learning and fitting the training data well. • Validation Loss: We observed a similar downward trend in the validation loss, but with some fluctuations. Lowering the learning rate could potentially smooth these fluctuations and improve validation loss further. CERN OpenLab Report // 2024 17 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías APPENDIX B: IMPLEMENTING A GRAPHICAL USER INTERFACE This interface enhances the usability of the model by allowing peer reviewers to evaluate and decide on individual AI-recommended changes as they read through a paragraph. By presenting suggested modifications clearly, the tool empowers users to make informed decisions on whether to accept or reject each recommendation. This approach ensures that the AI is used as a supportive tool rather than as the sole authority in the editing process. Below, you will find a draft of the GUI design that illustrates this concept. The interface highlights suggested changes by underlining the relevant words or sentences. The peer reviewer can then choose to accept or reject each change. Depending on the reviewer’s decisions, the final output will be a LaTeX document with the approved corrections. This draft provides a visual representation of how the solution is structured and how it facilitates a collaborative and controlled editing process. CERN OpenLab Report // 2024 18 A CMS AI to Improve Paper Publishing – Annunziata Álvarez-Cascos Hervías Figure 8. Draft Design of the Graphical User Interface (GUI) for Peer Review Automation