scieee AI-readable full text Open interactive document viewer

AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL

Caysar, Zeynep; Hollis, Dominic; Roun, Tomáš; Mönnich, Adrian

Abstract

This project presents a plugin developed for Indico, CERN’s open-source event management platform, designed to automatically summarize meeting minutes using large language models (LLMs). The goal is to reduce the manual effort of post-meeting documentation by leveraging open-source, efficient, and accurate AI models. The plugin integrates directly into the Indico interface, enabling users to generate structured summaries by selecting meetings, editing prompts, and receiving results. A range of open-source LLMs were evaluated, from lightweight models to large quantized ones, based on output quality, accuracy, inference time, and consistency. While smaller models proved insufficient, more powerful quantized models deployed via llama.cpp and eventually a GPU-backed model on CERN’s Kubeflow infrastructure achieved significantly better results. Human evaluations were used to assess summary quality and inform model selection. The outcome is a working prototype capable of producing clear and concise summaries, laying the foundation for future improvements such as automatic metric-based evaluation, support for more complex prompts, and context-aware summarization.

Full text

AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL August 2025 AUTHOR: Zeynep Çaysar University of Zurich SUPERVISORS: Dominic Hollis Tomáš Roun Adrian Mönnich CERN openlab Report 2025 PROJECT SPECIFICATION The objective of this project was to implement a plugin for Indico, CERN’s open-source event management platform, that automatically generates summaries of meeting minutes using open-source large language models (LLMs). The plugin was intended to reduce the manual workload of post-meeting documentation by integrating AI-driven summarization directly into the Indico workflow. •Objectives –Evaluate different open-source LLMs (small, large, and quantized) with respect to output quality, inference time, and usability. –Develop a proof-of-concept Indico plugin prototype capable of requesting and displaying automatically generated summaries. •Scope Included: Plugin development, backend–frontend integration, testing of multiple LLMs, human-based evaluation of summaries. Excluded: Automatic metric evaluation (e.g., ROUGE/BLEU), and full-scale production deployment of large models beyond proof of concept. •Requirements Functional Requirements: –Allow users to select meetings and trigger summarization directly from Indico. –Support prompt customization for flexible summarization outputs. –Display results as structured summaries in the Indico interface. Non-Functional Requirements: –Ensure reasonable inference times (under 1 minute for typical meeting notes) –Provide readable and consistent summaries across different meeting types. –Maintain modular design to allow independent updates to summarization models. •Constraints – Time: 9-week internship duration. – Resources: Limited access to GPU resources; reliance on CPU and quantized models during early stages. AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 1 CERN openlab Report 2025 ABSTRACT This project presents a plugin developed for Indico, CERN’s open-source event management platform, designed to automatically summarize meeting minutes using large language models (LLMs). The goal is to reduce the manual effort of post-meeting documentation by leveraging open-source, efficient, and accurate AI models. The plugin integrates directly into the Indico interface, enabling users to generate structured summaries by selecting meetings, editing prompts, and receiving results. A range of open-source LLMs were evaluated, from lightweight models to large quantized ones, based on output quality, accuracy, inference time, and consistency. While smaller models proved insufficient, more powerful quantized models deployed via llama.cpp and eventually a GPU-backed model on CERN’s Kubeflow infrastructure achieved significantly better results. Human evaluations were used to assess summary quality and inform model selection. The outcome is a working prototype capable of producing clear and concise summaries, laying the foundation for future improvements such as automatic metric-based evaluation, support for more complex prompts, and context-aware summarization. AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 2 CERN openlab Report 2025 TABLE OF CONTENTS 1 Introduction 4 1.1 Objectives....................................... 4 1.2 LLMKeyRequirements ............................... 4 1.3 Plugin Workflow (UI in Indico) . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2 Methodology and implementation 7 2.1 ToolsandTechnologies................................ 7 2.2 Model selection and evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 2.3 Human-reviewevaluation............................... 8 3 Prototype in action 10 4 Future work 16 5 Acknowledgements 17 6 References 17 AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 3 CERN openlab Report 2025 1 Introduction This project focuses on developing a plugin for Indico[10], CERN’s[2] open-source event and meeting management platform, to automatically generate summaries of meeting minutes using an open-source Large Language Model (LLM). Manual summarization of meetings is time-consuming, error-prone, and inconsistent. The goal is to streamline this process by integrating LLMs directly into Indico’s user interface, enabling users to generate accurate, readable summaries. LLMs are transformer-based neural networks trained on massive datasets to comprehend and generate human-like text. They enable a wide range of natural language processing tasks, including translation, summarization, question answering, and content generation[15]. While closed-source models such as GPT-4[7] demonstrate strong performance, they also present limitations, including restricted transparency, limited reproducibility, and dependency on proprietary third-party services. In contrast, open-source LLMs, such as LLaMA[13] or Gemma[5], promote transparency, facilitate community-driven development, and offer greater adaptability to diverse deployment scenarios[11]. Platforms like Hugging Face play a central role in this ecosystem by providing a collaborative hub for accessing, evaluating, and deploying open-source models. The Hugging Face Model Hub hosts thousands of pre-trained LLMs and provides tools for inference, fine-tuning, and integration, making it an essential resource for building scalable and customizable summarization solutions[9]. This project leverages open-source models from Hugging Face[9] and also quantized models deployed locally (via ‘llama.cpp‘[14] and Hugging Face[20]) or on high-performance inference platforms (CERN’s Kubeflow[12] and Groq[8]) and implements a minimal prototype interface designed to align with Indico’s native design. 1.1 Objectives •Develop an intuitive plugin to summarize meeting minutes within Indico. •Use fully open-source, CPU-compatible models. •Test more powerful cloud-deployed models for scalable inference. •Minimize hallucinations and ensure deterministic, reproducible summaries. 1.2 LLM Key Requirements •Fast inference •Open-source and license-compatible models •High-quality outputs •Support for prompt customization AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 4 CERN openlab Report 2025 User on Indico event page Click Summarize Modal: preview area + prompt selection/editor Click Generate summary GET /ai-summary/summarize-event/$eventId {event_id, prompt, token} Preview summary (read-only) Save to event? POST /api/note {event_id, source, render_mode, revision_id} Write to Event notes Edit/Change prompt or Clear summary Errors: Timeout, empty minutes, permission denied Modal shows error + keeps prompt Yes No Figure 1: Plugin/UI flowchart 1.3 Plugin Workflow (UI in Indico) Goal: let a user generate, review, and optionally save a summary of meeting minutes. 1. Entry point. On the event page, the user opens the “Event manage button” dropdown menu and clicks Summarize. 2. Prompt setup. A modal opens with: •a dropdown of predefined prompts (incl. “Dev review”, “Grumpy”, “LOL”, “TL;DR”, “Default”, “Old and angry engineer”) and a “Custom” option, •aprompt editor to tweak or write a custom prompt (with Save/Delete for reuse). AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 5 CERN openlab Report 2025 3. Preview request. The user hits Generate summary. The plugin: (a) disables the form and shows a spinner, (b) sends a GET to the backend /ai-summary/summarize-event with event_id,prompt, and token. 4. Preview rendering. On success, the modal displays a read-only preview plus: •Save to event (writes to Event/Notes via backend), •Clear summary Clears the summary from the preview area. 5. Save flow. If the user clicks Save to event, the plugin calls /event/$eventId/api/note with the source(summary text), render mode(html) and revision id(current revision id). When user goes back to the event page and refreshes the page, the summary is added to the event note. 6. Failure handling. For timeouts, auth errors, or empty minutes: •the modal shows a concise error banner or error message in the summary preview area, •preserves the user’s prompt text AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 6 CERN openlab Report 2025 2 Methodology and implementation 2.1 Tools and Technologies •Frontend: HTML, JavaScript, React(Semantic UI)[19] integrated into the Indico meeting interface. •Backend APIs: Python-based, interacts with model server. •LLMs tested: –‘llama.cpp‘ for quantized local models –Hugging Face Transformers pipeline for experimentation –CERN Kubeflow for Qwen-32B inference –Groq API for ultra-fast testing of LLaMA-3 70B, Gemma 9B, etc. 2.2 Model selection and evaluation Initially, smaller CPU-compatible models such as google/flan-t5-large,flan-t5-small, bart-large-cnn, and meeting-focused models like knkarthick/meeting_summary were tested. While these models were efficient in terms of computational resource usage, the generated summaries were often of low quality. In particular, the outputs contained duplicated sentences, omitted relevant information, and occasionally introduced hallucinated content. To improve consistency across runs, several inference parameters (e.g., temperature,top_k,top_p, and do_sample) were tuned, but the overall performance remained inadequate due to model size limitations. To improve the quality of summaries while maintaining reasonable inference times, quantized models such as qwen2.5-7b,llama-3-8b, and mistral-7b were evaluated using the llama.cpp framework. These models provided a better balance between summary quality and computational efficiency; however, their performance was still limited compared to larger-scale models. Subsequently, several high-performance cloud-based models were evaluated on the Groq platform, including llama-3-70b,qwen-3-32b,llama-3.1-8b-instant,gemma-2-9b-it, and moonshotai-kimi-k2-instruct. These models achieved high-quality summaries with very low latency (approximately 2 seconds), establishing a performance benchmark for subsequent testing. Finally, the Qwen2.5-32B model, hosted on CERN’s Kubeflow ML platform, was evaluated. It demonstrated strong performance, producing coherent and contextually accurate summaries while maintaining competitive inference times. Given its balance of quality, speed, and integration feasibility, the Qwen2.5-32B model was selected for deployment in the current prototype. AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 7 CERN openlab Report 2025 Table 1: Comparison of Tested Models for Meeting Summarization Model Type / Size Avg. Inference Time Evaluation Summary flan-t5-small / large Encoderdecoder (small) ∼5–7s Fast but produced low-quality summaries, missing context and introducing duplicates. bart-large-cnn Encoderdecoder (medium) ∼6–8s Better than Flan but still struggled with coherence and hallucinations. knkarthick/meeting_summaryFine-tuned for meetings ∼5s Optimized for meeting data but underperformed in accuracy and completeness. qwen2.5-7b /llama-3-8b /mistral-7b Quantized models ∼3–5s Balanced efficiency and quality but lacked depth compared to larger models. llama-3-70b /qwen-3-32b /llama-3.1-8b-instant /gemma-2-9b-it / moonshotai-kimi-k2-instruct Large cloudhosted models (Groq) ∼2s Delivered high-quality outputs and very low latency; established performance benchmark. Qwen2.5-32B Large model (Kubeflow) ∼2–3s Produced coherent, accurate, and contextaware summaries; currently used in the prototype. 2.3 Human-review evaluation An internal form was created for team-based evaluation of summaries. Criteria included: •Accuracy •Summary and format quality •Avoidance of redundancy The form remains open; however, models Qwen3-32B, Qwen2.5-32B and moonshotai-kimi-k2instruct have received the most positive feedback so far. AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 8 CERN openlab Report 2025 Figure 12: Hallucination example Figure 13: Another hallucination example: sudden switch to Mandarin although the context is correct (Used LOL prompt - maybe it was a joke) AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 15 CERN openlab Report 2025 4 Future work Although the current implementation successfully delivers an end-to-end summarization workflow within Indico, several areas remain open for future improvements to enhance performance, scalability, and overall impact. The following directions outline possible next steps: •Evaluation with larger models While smaller Hugging Face models were prioritized during the internship due to limited compute resources and the need for a fast proof of concept, larger models should be reevaluated. These models may deliver significantly higher-quality summaries, particularly for long and complex meeting minutes. •Systematic evaluation metrics At present, evaluation was largely manual and qualitative. Future work should adopt standardized, automatic evaluation metrics such as ROUGE[18], BLEU[1], and F1[3] to provide objective and reproducible measures of summary quality. This would allow more rigorous model comparisons and benchmarking. •Advanced prompting techniques The current summarization relies on relatively simple prompts. An important future step is to introduce few-shot examples[4] within prompts to guide the models towards more structured, domain-relevant summaries. This could reduce hallucinations and improve coverage of critical meeting content. •Multi-turn context summarization For long meetings spanning multiple sections or topics, implementing multi-turn summarization (e.g., retrieval-augmented generation(RAG)[17]) would preserve context across segments. Instead of truncating or splitting blindly, this approach would allow summaries to remain coherent and informed by earlier discussion. •Deployability and scalability Future iterations should focus on building GPU-enabled deployment pipelines, enabling faster inference for larger models. Incorporating these pipelines into platforms such as Kubeflow[12] would ensure scalability, load balancing, and smoother integration with production environments. •Microservice architecture The initial design envisioned the summarization logic as a separate microservice decoupled from Indico. Moving further in this direction would keep Indico responsible only for user interaction and request handling, while the summarization backend could evolve independently. This modularity would make the system easier to maintain, upgrade, or replace in the future. •User Experience Enhancements In addition to backend improvements, several user-facing features could be introduced to enhance usability: –Provide options for users to configure the summarization style (e.g., concise, detailed, or bullet-point format). AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 16 CERN openlab Report 2025 –Support section-specific summarization, allowing users to focus on relevant sections, such as Dev review, while excluding less important content. –Integrate feedback mechanisms, enabling users to rate summaries and facilitate iterative refinement of the system. 5 Acknowledgements My nine weeks at CERN have been an incredible journey. I learned so much, from working with open-source LLMs and the Hugging Face platform to exploring plugin architecture and Semantic UI React. CERN is a truly unique place to do an internship, and being part of Indico was very special to me. Indico plays such an important role in communication at CERN, and it was very exciting to contribute to a platform that is not only used here but also shared with the wider community as open source. I want to thank my supervisors, Tomáš, Dom, and Adrian, for all their guidance and support. A big thank you as well to my office mates-Ajob, Marina, and Michel-and the entire Indico team for welcoming me, teaching me, and making me feel part of the team from day one. I honestly couldn’t have asked for a better group of people to spend these weeks with. I truly enjoyed every day of this internship, and I’m very grateful for the experience. Thank you! 6 References [1] Bleu. Bleu.url:https://en.wikipedia.org/wiki/BLEU (visited on 08/27/2025). [2] CERN. CERN: the European Organization for Nuclear Research.url:https://home. cern/fr (visited on 08/21/2025). [3] F1. F1.url:https://datasciencedojo.com/blog/understanding-f1-score/ (visited on 08/27/2025). [4] Fewshotprompting. Fewshotprompting.url:https://www.promptingguide.ai/techniques/ fewshot (visited on 08/27/2025). [5] Gemma. Gemma.url:https://deepmind.google/models/gemma/ (visited on 08/27/2025). [6] Github. Github.url:https://github.com/zeynepcaysar/indico-plugins/tree/ summary-rebase/ai_summary (visited on 08/27/2025). [7] GPT-4. GPT-4.url:https : / / openai . com / fr - FR / index / gpt - 4/ (visited on 08/27/2025). [8] Groq. GROQ: company that provides fast AI inference in the cloud and in on-prem AI compute centers.url:https : / / console . groq . com / docs / models (visited on 08/21/2025). [9] Hugging Face. Hugging Face: open-source model hub.url:https://huggingface.co/ (visited on 08/21/2025). [10] Indico. Indico: open-source event management platform.url:https://indico.cern.ch (visited on 08/21/2025). AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 17 CERN openlab Report 2025 [11] Jiya Manchanda et al. opensourceLLM.url:https://www.researchgate.net/publication/ 387140760_The_Open_Source_Advantage_in_Large_Language_Models_LLMs (visited on 08/27/2025). [12] Kubeflow. CERN Kubeflow ML platform.url:https : / / ml . cern . ch/ (visited on 08/21/2025). [13] Llama. Llama.url:https://www.llama.com/ (visited on 08/27/2025). [14] Llamacpp. Llamacpp: the engine that runs AI models locally.url:https://en.wikipedia. org/wiki/Llama.cpp (visited on 08/21/2025). [15] LLMwiki. LLM.url:https:/ /en . wikipedia. org/ wiki / Large _language _ model (visited on 08/27/2025). [16] Qwen-Hugging Face. Qwen-HF.url:https : / / huggingface . co / Qwen (visited on 08/27/2025). [17] RAG. RAG.url:https://en.wikipedia.org/wiki/Retrieval-augmented_generation (visited on 08/27/2025). [18] Rouge. Rouge.url:https://en.wikipedia.org/wiki/ROUGE_(metric) (visited on 08/27/2025). [19] Semantic UI. Semantic UI React is the official React integration for Semantic UI.url: https://react.semantic-ui.com/ (visited on 08/21/2025). [20] TheBlokeHF. TheBloke: quantized model using GGUF format for open-source models provided in Hugging Face platform.url:https://huggingface.co/TheBloke/Mistral7B-Instruct-v0.2-GGUF (visited on 08/21/2025). AUTOMATIC MINUTES SUMMARIZATION IN INDICO USING AN OPEN-SOURCE LARGE LANGUAGE MODEL 18