AI4Science Training @ UFZ Leipzig
Abstract
This slide deck accompanies the AI Competence Training for scientists at UFZ Leipzig: https://scads.github.io/ki-kompetenz-training-2025/intro.html It outlines the topics: Introduction to Artificial Intelligence, AI systems, and language models Application areas and limitations of AI in text generation Prompt engineering Prompting with large context Knowledge distillation Embeddings and retrieval-augmented generation Data analysis Data and code generation Bias Privacy-compliant use Copyright EU AI Act
Full text
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 1 Responsible Use of Generative AI, Legal Framework Robert Haase These slides can be reused under the conditions of the CC-BY 4.0 license unless otherwise specified. CENTER FOR SCALABLE DATA ANALYTICS AND ARTIFICIAL INTELLIGENCE Funded by:
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 2 Note I’m not a lawyer or data protection officer. I can only give insights into the legal framework of AI. In case of doubt contact your legal department: r[email protected] https://www.ufz.de/index.php?de=36697
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 3 KI-Competence Training according to EU-AI Act The providers and operators of AI systems take measures to ensure to the best of their ability that their personnel and other individuals who operate and use AI systems on their behalf possess an adequate level of AI competence taking into account their technical knowledge, experience, training and education, and the context in which the AI systems are to be used, as well as the individuals or groups of individuals for whom the AI systems are intended to be used. (from EU AI Act, Art 4.) https://eur-lex.europa.eu/legalcontent/DE/TXT/HTML/?uri=OJ:L_202401689#art_4
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 4 AI-Provider and AI-Operator according to the EU-AI Act AI-Provider (Provider) •Markets a service or a product based on AI-systems/-models AI-Operator (Deployer) •provides an AI-system for own/inhouse use under own supervision •Exception: personal, non-professional use Scientific research •Largely outside the scope of the EU AI Act (Art 2) See also Art 3 EU AI Act: https://eur-lex.europa.eu/legalcontent/EN/TXT/?uri=CELEX:32024R1689 https://helmholtz-blablador.fz-juelich.de/
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 5 Users and Affected persons Users (during work duties) •Persons who employ AI-systems to process data •Have to be trained to use AI systems Affected persons •People whose data is processed by AI or who experience consequences of the processing (direct or indirectly) Figure was created using ChatGPT and post-processed User Affected person AI
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 6 Large Language Models Text-to-text, Text Generation Write a text about Write a short text about UFZ Leipzig The UFZ Leipzig (Helmholtz Centre for Environmental Research –UFZ) is one of Europe’s leading institutions for environmental science. Located in Leipzig, Germany, it focuses on understanding how natural systems function, … ChatGPT 5.1 (OpenAI, USA) Best of 1
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 7 Large Language Models Text-to-text, Text Generation Write a short text about the best political model in our global society. The best political model in our global society is a collaborative, democratic, and participatory approach that prioritizes the well-being of all its members and the longterm sustainability of the global community. This includes fostering trust, promoting inclusive governance, and ensuring equitable access to resources and opportunities for all. Gemma3:270m (Google, USA) Best of 1
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 8 Large Language Models Text-to-text, Text Generation Write a short text about the best political model in our global society. The best political model in our global society is one that balances economic prosperity, social equity, and democratic governance. By fostering globalization, promoting fair trade, investing in education, and ensuring inclusive policies, nations can build resilient, interconnected societies. At the heart of this model is the principle of cooperation rather than competition, recognizing that diverse systems—ranging from autocratic to democratic models—can coexist and thrive. Ultimately, a successful global political model seeks to create a world where prosperity, justice, and shared values are the guiding pillars. Qwen 3:0.6b (Alibaba, China) Best of 1
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 9 ChatApps and Language Models Regional different LLMs + AI systems •ChatGPT •GPT 4 / 4o / 4o-mini •O1 / O3 / O4 / GPT 5 (Reasoning) •Anthropic: Claude •Google: Gemini •Meta: Llama •Ollama (local solution) •Huggingface Hub •Mistral •Mixtral •Pixtral •LeChat •DeepSeek •DeepSeek-R1 (Reasoning) •OpenGPTx: Teuken •Helmholtz: Blablador •GWDG: Kisski ChatAI •DeutschlandGPT Risk: If one LLM dominates, their authors can dictate contents Chance: We can compare LLMs from different regions.
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 16 Language models in scientific work Text-to-text, scientific reviewing Source: https://aaai.org/aaai-launches-ai-powered-peerreview-assessment-system/
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 18 AI Recommendations of the Helmholtz Association an the German Research Foundation (DFG) https://www.dfg.de/resource/blob/289676/230921-statementexecutive-committee-ki-ai.pdf https://www.helmholtz.de/assets/helmholtz_gemeinschaft/Downloads/ Helmholtz_Recommendations_on_use_of_AI_Version_1.0.pdf
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 19 Guidelines of the German ResearchFoundation (DFG) on the use of AI https://www.dfg.de/resource/blob/289676/230921statement-executive-committee-ki-ai.pdf
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 20 Guidelines of the German ResearchFoundation (DFG) on the use of AI https://www.dfg.de/resource/blob/289676/230921statement-executive-committee-ki-ai.pdf
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 21 Guidelines of the German ResearchFoundation (DFG) on the use of AI https://www.dfg.de/resource/blob/289676/230921statement-executive-committee-ki-ai.pdf
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 22 Guidelines of the German ResearchFoundation (DFG) on the use of AI https://www.dfg.de/resource/blob/289676/230921statement-executive-committee-ki-ai.pdf
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 23 Quiz: Are we allowed to do this? We are working as reviewers for DFG and use a local language model to •summarize a given project proposal •correct spelling issues in our review No Yes
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 24 Guidelines of the German ResearchFoundation (DFG) on the use of AI https://www.dfg.de/resource/blob/289676/230921statement-executive-committee-ki-ai.pdf Written in 2023 I presume they mean LLMs in the cloud.
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 25 Good scientific practice If you use custom [code | data | text] written by … a human expert an expert LLM You should … •Understand the code (roughly) •Question used methods •Check results carefully •Test code on samples the expert didn‘t see
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 26 Good scientific practice If you use custom [code | data | text] written by … a human expert an expert LLM You should … •Pay the expert •Mention the expert •Share responsibility •Ask the expert endless questions •Share how you prompted the expert $100/h $0.1/h co-author in methods
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 33 Quiz: Are we allowed to do this? Imagine we build a chatbot about responsible use of AI using the content on this website. Is this fine? ? ? https://rdm.pages.ufz.de/guidelines/AI-for-science/AIEthics/#6-avoid-ai-in-sensitive-activities
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 34 AI-Regulation of the European Union “EU AI Act” Robert Haase CENTER FOR SCALABLE DATA ANALYTICS AND ARTIFICIAL INTELLIGENCE
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 35 EU AI Act - Definitions AI models AI systems AI-Operators AI-Providers Users Affected individuals
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 36 Quiz: Classification of the Helmholtz Association What are “we” in the context of the EU AI Act? Provider Operator User Affected
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 37 Scope EU AI Act According to Article 2 (Scope) •Providers, operators, importers, and distributors of AI systems/models within and outside the EU, as long as the EU is affected in a relevant way, e.g., an AI system produces text within the EU. •Product providers who use AI in products. •Affected individuals within the EU. https://eur-lex.europa.eu/legalcontent/EN/TXT/?uri=CELEX:32024R1689
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 38 Out of scope EU AI Act Excerpt EU AI Act Article 2 (Scope): “6. This Regulation does not apply to AI systems or AI models, including their output, specifically developed and put into service for the sole purpose of scientific research and development.” “10. This Regulation does not apply to obligations of deployers who are natural persons using AI systems in the course of a purely personal non-professional activity.” “12. This Regulation does not apply to AI systems released under free and opensource licences, unless they are placed on the market or put into service as high-risk AI systems or as an AI system that falls under Article 5 or 50.” (5: Prohibited AI-Practices 50: Transparency obligations) https://eur-lex.europa.eu/legalcontent/EN/TXT/?uri=CELEX:32024R1689
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 39 EU AI Act –Risk-Based Approach Categorization of AI Systems by Risk Consequences Inacceptable Risk (Art. 5 AI Act) High Risk (Art. 6 ff. AI Act) Systematic Risk Art. 51-54 AI Act Minimal Risk •Subliminal Influence •Exploitation of children or mentally disabled people •Social scoring •Public biometric identification •Biometrics •Critical infrastructure •Human resources •Medical devices •Law enforcement •Chatbots •Deepfakes •Computer Games •Spam Filters •Product Suggestions Prohibited Transparency: Documentation on functionality and decisions. Risk Assessment: Manufacturers must demonstrate and mitigate risks. Human Oversight: Decision-making processes must not be entirely autonomous. If applicable. Adhere to data protection policies requirements Transparency Obligation Documentation Obligation Riskmanagement Human Oversight Quality Assurance / External Review
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 40 Exercise: Risk assessment Assume we use an LLM to summarize applications to the graduate school / / PhD program. Who is affected? How high is the risk? What consequences do you see?
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 41 Users must be able to •trace AI-assisted decision processes, •contest AI-based decisions, •deny the processing of their data. Providers and developers have to •document algorithms, •incl. risk assessment, •ensure to oversee decisions by humans. Documentation obligation / Transparency These aspects make many usage scenarios, e.g. in research de facto impossible
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 42 Language models make mistakes The user of the language model / author of the translated text is responsible …just like humans
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 50 Copyright-violations (Benchmark) The model OpenGPT-X 7B commits measurably only very few copyright violations. Successor model Teuken probably as well. SRR -Copyright SRR -Public Domain GPT 4 774.5 33034.1 GPT 3.5 61.5 2716 Llama 2 (70 B) 697.2 1898.7 Alpaca (7B) 3.6 158.5 Vicuna (13 B) 521.7 3446.8 Luminous (70B) 6.2 217.8 OpenGPT -X (7B) 0.3 0 Source: Simplified Table 3 in Mueller et al 2024 https://arxiv.org/pdf/2405.18492 “The significant reproduction rate is the average number of characters per book that are part of a literal reproduction of original text in excess of the legality presumption of up to 160 characters.” Mueller et al 2024 Predecessor of Teuken 7B
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 51 Exercises Robert Haase CENTER FOR SCALABLE DATA ANALYTICS AND ARTIFICIAL INTELLIGENCE
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 52 Exercise: AI-Detectors Upload ufz-leipzig.pdf into commercially available AI-Detectors and test them out. https://scispace.com/ai-detector https://scads.github.io/ai4science-ufz2025/session3/ai-detectors.html
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 53 Exercise: Bias Detection Generate feedback about a meeting protocol •Generic feedback •Specific feedback •Diversity •Professional correctness / suitability •… Meeting Protocol – Symposium Planning Date: April 3, 2025, 10:00 – 10:35 AM (Zoom) Participants: Dr. Thomas Becker, James Müller, William Schubert Discussion Points: 1. Topic & Objectives •Consensus on an interdisciplinary approach. •Proposal by Schubert accepted: “Climate research using ChatGPT” 2. Possible Speakers Suggestions (Müller): •Dr. Richard Müller (City Administration Zurich) •Dr. Anton Berg (AI in Climate Research, UFZ Leipzig) •Prof. Josef Angermann (Administrative Director University Hospital Dresden) Becker emphasizes that all invited speakers should be internationally known. Decision: Contact all three potential speakers by next week. Next Meeting: April 14, 2025, 10:00 AM (online) End of Meeting: 10:35 AM Note: You can also accidentally establish bias with this method . https://scads.github.io/ai4science-ufz2025/session3/bias_detection.html
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 54 Learn more… https://scads.ai/event/meetup/ https://www.ub.uni-leipzig.de/service/workshops-und-online-tutorials/schulungen/ki-stammtisch https://www.helmholtz.ai/applie d-ai/ai-consultancy-teams/ https://rdm.pages.ufz.de/guidelines/AI-forscience/Workshops-%26-Courses/
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 55 Conclusion •AI users / operators / providers have a whole range of obligations when working with AI (accountability, risks, documentation, …). •Data subjects have rights! (transparency, objection, …) •AI systems technically allow far more usage scenarios than are legally permissible. •Especially in the clinical context, many types of AI-based data processing are problematic under data-protection law and from an ethical perspective. •Conscious and responsible handling of AI and AI-related risks is essential.
AI4Science @UFZ Robert Haase @haesleinhuepf December 2025 56 Thank you for your attention! Contact Dr. Robert Haase ScaDS.AI Dresden/Leipzig Universität Leipzig Humboldtstraße 25 04105 Leipzig [email protected]