Toward LLM-enabled business process coherence checking based on multi-level process documentation
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Schulte, Marek; Franzoi, Sandro; Köhne, Frank; vom Brocke, Jan Article — Published Version Toward LLM-enabled business process coherence checking based on multi-level process documentation Process Science Suggested Citation: Schulte, Marek; Franzoi, Sandro; Köhne, Frank; vom Brocke, Jan (2025) : Toward LLM-enabled business process coherence checking based on multi-level process documentation, Process Science, ISSN 2948-2178, Springer International Publishing, Cham, Vol. 2, Iss. 1, https://doi.org/10.1007/s44311-025-00024-6 This Version is available at: https://hdl.handle.net/10419/333235 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
RESEARCH Open Access © The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. Schulte et al. Process Science (2025) 2:22 https://doi.org/10.1007/s44311-025-00024-6 *Correspondence: Sandro Franzoi [email protected] 1University of Münster, Münster, Germany 2European Research Center for Information Systems, Münster, Germany 3viadee Unternehmensberatung AG, Münster, Germany Toward LLM-enabled business process coherence checking based on multi-level process documentation MarekSchulte1, SandroFranzoi1,2*, FrankKöhne3 and Janvom Brocke1,2 Introduction Deviations from expected process behavior persist in virtually every organization. As such, understanding and addressing process deviance is a critical concern for both business process management (BPM) research and practice (König et al. 2019). In the context of today’s dynamic business environments, deviations from established processes are not merely challenges; they represent potential opportunities for significant process improvements and innovations (Bartelheimer et al. 2023). When effectively identified and managed, these deviations can enhance operational efficiency and adaptability, providing substantial value to businesses (Setiawan and Sadiq 2013; Delias 2017). Abstract In this paper, aProCheCk, an Autonomous Process Coherence Checking method, is developed. aProCheCk leverages large language models (LLMs) to enhance the coherence checking of multi-level process documentation within business process management (BPM). This research addresses the need for automated ways of managing incoherencies in process documentation. The development of the artifact was guided by a design science research approach, which involved iterative development and refinement. This was achieved through expert interviews with researchers and practitioners, iterative experimental benchmarking, and focus group validation based on demonstrations of a prototypical implementation with naturalistic data from diverse industries. aProCheCk can dynamically analyze and assess changes in BPM documentation, detect incoherencies, and provide actionable insights for maintaining process coherence. The findings reveal significant potential for improving operational efficiency, reducing manual effort, and detecting negative and positive process variation early to support continuous process innovation. This research contributes to the field of BPM by integrating LLMs into the BPM lifecycle, enhancing generative AI-based applications within BPM practices, introducing the Business Process Change Classification Framework, and providing an open-source dataset that can serve as a foundation for future research and development. Keywords Large language model, Process coherence, Artificial intelligence, Process deviance, Business process management Process Science
Page 2 of 33Schulte et al. Process Science (2025) 2:22 In the past, inconsistencies in process documentation were typically solely regarded as errors that needed to be rectified (van der Aa et al. 2017). However, there is a growing recognition that these deviations can be viewed as opportunities for improvement (Galperin 2012; Mertens and Recker 2017). This is illustrated in the shift toward an understanding that BPM needs to balance the enforcement of process compliance with the identification of positive deviance (Setiawan and Sadiq 2013; Mendling et al. 2020). Identifying and analyzing these deviations can provide insights that lead to more effective business outcomes, innovative process improvements, or necessary corrective actions. This shift in perspective underscores the need for sophisticated tools that can efficiently detect, analyze, and leverage both positive and negative deviations to drive business growth and innovation (König et al. 2019). The emergence of generative artificial intelligence (AI), particularly large language models (LLMs), signifies a transformative development in this domain. LLMs and their capability to generate and understand human-like text render them particularly suitable for interpreting complex business documentation associated with BPM (Vidgof et al. 2023; Fahland et al. 2024). The potential role of AI in detecting and managing process deviations becomes crucial when considering the automation and enhancement of these tasks (Weinzierl et al. 2024). In particular, LLMs are capable of analyzing extensive sets of process documentation to identify inconsistencies and changes that may indicate deviations (Vidgof et al. 2023). Recent research shows that LLMs are also increasingly capable of making complex subjective decisions based on the given context (Binz and Schulz 2023; Mcintosh et al. 2024). This characteristic enables the expansion of the scope from detecting quantitative deviations, for instance, based on event logs and process diagrams, to encompass the more complex and subjective problem of checking for incoherencies in text-based business process documentation. By detecting deviations through the analysis of changes in the process documentation, LLMs enable organizations to not only maintain the integrity of their processes but also to take advantage of these deviations to drive innovation and improve efficiency (Vidgof et al. 2023). This approach addresses a critical research gap identified by Feuerriegel et al. (2024), which emphasizes the potential of generative AI methods to reveal opportunities for process innovation and support process redesign initiatives. This is especially important because organizations frequently maintain multiple, interrelated documents, often resulting in incoherencies between them – a problem magnified by changes made to individual documents that are not reflected consistently across all related documents (Martin-Toral et al. 2010). This issue is particularly evident in the context of business process documentation, as expressed, for example, in the presence of content inconsistencies between process descriptions and the corresponding process models (van der Aa et al. 2017). This suggests that analyzing changes in process documentation and comparing them with other related process documents may reveal incoherencies, which could potentially be leveraged to detect deviance and improve processes faster and more efficiently in organizations. To address this problem, we formulate the following two research objectives, which aim to explore the integration of LLMs within process coherence checking based on process documentation:
Page 3 of 33Schulte et al. Process Science (2025) 2:22 RO1: Define the essential design objectives and specifications for an LLM-enabled process coherence checking method based on changes to process-related documentation. RO2: Develop and evaluate a solution for employing LLMs to automaticallyevaluate the consistency and coherence across different versions of multipledocumentation types for business processes. Addressing these research objectives, we followed a design science research (DSR) approach and developed an artifact called aProCheCk, an acronym for Autonomous Process Coherence Checking. This artifact represents a method that is designed to leverage LLMs to enhance the enactment and evaluation phases of the BPM lifecycle by assessing the consistency and coherence of multi-level process documentation. The core functionality of aProCheCk is to automate process coherence checking, thereby enabling organizations to identify and manage both negative and positive deviance within their processes. Specifically, aProCheCk compares two entities of process documentation (e.g., process model and process description or process model and regulatory document, among others) and their change over time to identify incoherencies and subsequently create notifications (see notification creation in Fig.2), presented in the form of management summaries, to inform process owners. We instantiate a software prototype to demonstrate and evaluate the utility of our artifact in practice. As such, our research marks a fundamental step toward coherence checking based on multi-level process documentation. The identification and assessment of incoherencies allows organizations to not only rectify errors but also to leverage positive deviations as opportunities for strategic advancement and innovation. Following established guidelines on structuring DSR (Gregor and Hevner 2013), we present the remainder of the paper as follows. First, we introduce the research background on process coherence checking as well as generative AI in BPM. Next, we outline our methodological approach and its respective phases. Then we describe our developed artifact and provides information on its design and development. Afterwards, we present our experimental and focus group evaluations. Finally, we discuss theoretical and practical implications as well as the limitations of our work before we conclude the paper. Research background To provide a comprehensive foundation for our work, the research background section is structured into two subsections. First, we introduce the concept of process coherence checking, highlighting the challenges of maintaining semantic consistency across multilevel process documentation and the limitations of traditional approaches. Second, we shift the focus to generative AI in business process management, outlining how recent advances, particularly in LLMs, are opening up new opportunities for addressing challenges in BPM. Process coherence checking Process coherence represents a vital concept for ensuring logical consistency across business processes and their related documentation. In the context of information systems, the antonym of coherence, incoherence, provides a useful point of reference. Saint-Dizier (2018), for instance, defines incoherence as linguistic discrepancies between documents. Martin-Toral et al. (2008) define content incoherence as “[…] the weakness
Page 4 of 33Schulte et al. Process Science (2025) 2:22 of consistency amongst related documents, or amongst different pieces of the same document, or the lack or excess of information in a document” (p. 283). In a broader sense, coherence is essential for preserving logical consistency across systems, processes, and documentation (Martin-Toral et al. 2010). It ensures that various components interrelate coherently, enabling clear communication and operational efficiency. This principle of coherence is particularly important given the extensive and complex documentation typical of business processes. The concept of process deviance (Mertens and Recker 2017) is closely related. It refers to user actions that deviate from the intended process flow and can be studied through various methods, including process mining (Di Francescomarino et al. 2025). These deviations illustrate how users interact with systems in unpredictable or unintended ways, leading to positive or negative effects (Delias 2017). These user-driven deviations may result in updates to operational process documents, such as training materials, to reflect actual practices. However, other documents that are less user-centric, such as process models, may not be updated synchronously, resulting in potential inconsistencies. In contrast to process deviance, incoherence is more closely associated with general inconsistencies within the documentation of processes, which may occur independently of any user-driven actions (Martin-Toral et al. 2010). Examples of such incoherencies that do not originate from user-driven actions include changes to overarching documents such as regulations or guidelines, organizational updates such as role changes, and software or hardware upgrades (Rosemann et al. 2008). The introduction of multi-level process documentation adds further complexity. The various types of documentation, including process models, textual descriptions, guidelines, and policies, not only differ in form but also in perspective, level of abstraction, and state of knowledge (Polyvyanyy et al. 2015; Rosemann and vom Brocke 2015; Nwankpa et al. 2022). Some documents may address only specific aspects of a process, whereas others may deal with particular abstraction layers, such as the database or strategic abstraction layers. Similarly, internal or external regulations, as well as modeling conventions, often span multiple processes and documentation types, placing them at a higher abstraction level. Ensuring coherence across these diverse documents is essential for effective BPM. Based on these insights, process coherence in BPM can be defined as semantic consistency across multi-level process documentation. Maintaining this coherence ensures that documents related to a process remain logically aligned, despite differences in perspective and granularity. This alignment is critical to avoid miscommunication and inefficiencies and to support a cohesive working environment. To address these challenges, the deployment of advanced technologies, particularly LLMs, enables organizations to automate the coherence checking of their process documentation, thereby capitalizing on the technological advancements over existing methods. For example, Martin-Toral et al. (2008) investigated the detection of incoherencies in a corpus of technical and regulatory documents using traditional methods such as extraction techniques involving text and document mining. Recent advances, as exemplified by the work of Sai et al. (2023), highlight the importance of bridging the gaps between organizational documentation and regulatory texts through the application of advanced NLP techniques. Their study focused on the comparison of general business documents with regulatory frameworks without addressing the specifics of BPM-related documents or the recent developments in LLM technology. The integration of LLMs in
Page 5 of 33Schulte et al. Process Science (2025) 2:22 this context not only supports regulatory compliance but also improves overall process coherence and reliability. Generative AI in business process management The progression of generative AI has redefined the potential applications of AI in a multitude of domains, including BPM (Kampik et al. 2024). By enabling more intuitive, adaptive, and creative forms of process interaction and design, generative AI extends the traditional boundaries of BPM beyond automation toward conversational, autonomous, and sophisticated process capabilities (Rosemann et al. 2024). For example, van Dun et al. (2023) leveraged generative adversarial networks to help process designers in the creation of business process improvement ideas. In another paper, Harl et al. (2024) showcase how generative machine learning can be used for automated business process redesign at runtime. Research has also investigated the opportunities and challenges of natural language processing for BPM (van der Aa et al. 2018) and explored its application, for instance, in predictive business process monitoring (Teinemaa et al. 2016). More recently, the evolution of transformer architectures has marked another key shift in the field of generative AI in BPM. Here, a key development was the emergence of tools such as ChatGPT, which employ LLMs trained on extensive datasets to generate contextually relevant responses to user queries (Vidgof et al. 2023). In the context of future developments, the potential applications of generative AI in BPM are numerous and diverse (Kampik et al. 2024). One area of significant growth is the development of new generations of process guidance systems (Feuerriegel et al. 2024). In contrast to traditional systems that rely on static, manually crafted knowledge bases, new systems powered by generative AI can dynamically retrieve information from a wide array of structured and unstructured data sources, including emails, manuals, and corporate documents (Morana et al. 2019; Feuerriegel et al. 2024). Such systems can guide a wide range of business process management tasks, such as modeling (Kourani et al. 2024), knowledge management (Franzoi et al. 2025b), or analysis support in process mining (Brützke et al. 2025). This enables the provision of real-time, context-sensitive guidance, making these systems more adaptive and intelligent. This shift from static to dynamic guidance systems indicates a broader move toward more responsive and intelligent BPM tools that greatly enhance decision-making and optimization across organizational levels. As highlighted by Feuerriegel et al. (2024), there is a continuing need to explore how generative AI can unveil new opportunities for process innovation. Their work suggests that further investigation into the capabilities of generative AI could reveal transformative approaches to BPM, redefining operational strategies and competitive dynamics in numerous industries. As such, there is still ample potential to employ LLMs to assess process deviance or process coherence across documents. Research methodology To address our research objectives, we employ a multifaceted methodology, centering on the applicability of an emerging technology to address a practical issue. The foundational methodological approach adopted is DSR, which aims at identifying a problem and developing an artifact designed to mitigate the issue within predefined parameters (Hevner et al. 2004).
Page 6 of 33Schulte et al. Process Science (2025) 2:22 Specifically, we follow the guidelines proposed by Peffers et al. (2007) and Tuunanen et al. (2024), which provide a comprehensive framework for carrying out research that is centered around the creation and evaluation of IT artifacts intended to solve identified problems. The process proposed by Peffers et al. (2007) is divided into six key phases: problem identification, definition of objectives, and design and development as Build phases, followed by demonstration, evaluation, and communication of research findings as Evaluate phases. The approach is illustrated in Section A in the online appendix1. In adapting the DSR framework for this research, specific modifications were made to tailor the approach to the unique requirements and constraints of this study. The aim was to enhance the focus on the practical application of emerging technology in solving real-world problems. This customized approach allowed for a more targeted development and evaluation of the artifact, ensuring that the research outcomes are both practical and theoretically sound. Additionally, we also incorporated principles of an iterative approach, including frequent evaluation and adaptation (Tuunanen et al. 2024). Our instantiation of the DSR approach is depicted in Fig.1. The evaluation phase in the DSR approach is critical for validating that the developed artifact meets both theoretical expectations and practical utility in real-world applications. This duality ensures that artifacts are robust in their theoretical underpinnings and highly functional in practical scenarios, indicating a successful bridging of theory and practice (Hevner et al. 2004). Our evaluation strategy employs a structured approach, utilizing various methodologies such as reviewing literature, expert interviews, focus groups, or experimental benchmarking (Sonnenberg and vom Brocke 2012). This strategy is designed to align with various evaluation types as identified in the DSR framework. Importantly, we emphasize the iterative nature of our DSR approach by highlighting the recurring cycles between design and development, as well as demonstration and evaluation. Here, phase 3.a, 3.b, and 3.c represent design and development activities, and 4.a, 4.b, and 4.c represent demonstration and evaluation activities (see Fig.1). The following paragraphs outline the specific steps of our approach. 1 The full online appendix with all relevant documents is available here: h t t p s : / / g i t h u b . c o m / v i a d e e / p r o c e s s - d o c u m e n t - c o h e r e n c e - c h e c k e r. Fig. 1 Applied Design Science Research Approach
Page 7 of 33Schulte et al. Process Science (2025) 2:22 Phase 1. Identify problem and motivate (Eval 1) In the initial phase, the underlying problem is carved out and discussed. For this purpose, existing literature is examined to identify relevant research, establish the theoretical foundation for the work, and justify the problem statement. Phase 2. Define objectives of a solution Based on the identification of the problem and the relevant research identified in the previous phase, the scope of the targeted artifact is defined, and provisional design objectives are derived from the literature. Phase 3. Design and Development Phase 3.a design specifications and initial proof of concept (PoC) Guided by the established provisional design objectives, an exemplary application is composed utilizing existing technologies. Furthermore, provisional design specifications are derived from the design objectives. Phase 3.b design and development of prototype In accordance with the refined design specifications, an initial working prototype is created. Phase 3.c adaptation of the prototype The artifact is refined through an iterative process based on the results of the experimental iterations. Phase 4. Demonstration and Evaluation Phase 4.a demonstration and evaluation in Semi-structured expert interviews (Eval 2) The PoC is demonstrated, and the design objectives and specifications are discussed in six semi-structured interviews with three practitioners in BPM-related roles and three researchers in the BPM field. This phase validates the design objectives and refines the specifications, focusing on understandability, feasibility, applicability, and operationality. Phase 4.b experimental evaluation of prototype (Eval 3) The applicability of the artifact is demonstrated through iterative experiments based on a dataset sourced from literature and enriched by a BPM expert. This phase focuses on a quantitative evaluation of the artifact’s efficiency, effectiveness, and robustness. Phase 4.c demonstration and evaluation in focus groups (Eval 4) The usefulness of the artifact is evaluated through demonstrations of an instantiation in two focus groups, each comprising four to five BPM and AI consultants. Following the demonstration, a semi-structured discussion assesses practical applicability, usability, and real-world integration. Phase 5. Communication The contribution to the knowledge base will be achieved through the publication of the research.
Page 8 of 33Schulte et al. Process Science (2025) 2:22 This transparent, structured, and iterative evaluation approach ensures that the artifact adheres to rigorous academic standards while enhancing BPM practices in diverse operational contexts (Hevner et al. 2024). Both formative and summative evaluation methods are applied in artificial as well as naturalistic settings (Venable et al. 2016), strengthening confidence in the overall evaluation methodology (vom Brocke et al. 2020). By aligning the interests and feedback of both researchers and practitioners, the artifact is refined into a robust tool that offers significant theoretical and practical contributions to the field of business process management (Sonnenberg and vom Brocke 2012). Artifact description This section outlines the developed artifact, aProCheCk, which represents a method for LLM-enabled process coherence checking. We first present an overview of the final artifact and its three main stages –preprocessing, content comparison, and coherence checking– followed by the design rationale, including objectives, proof of concept, expert evaluation, and the Business Process Change Classification Framework. We then detail the development and operating logic, highlighting prompt engineering techniques and the implementation strategy. Artifact overview: aProCheCk Drawing on the presented DSR approach, this section presents an overview of the final developed artifact, aProCheCk, which comprises three primary stages: (1) preprocessing, (2) content comparison, and (3) coherence checking. Each stage can terminate early if the documents are found to be coherent, thereby optimizing efficiency and reducing hallucination risks. Figure2 depicts an overview of aProCheCk, from identifying and inputting process documents to the optional creation of notifications. The darker grey boxes denote LLM interactions, while dotted lines indicate modifiable elements. Preprocessing The preprocessing phase focuses on noise reduction by removing non-relevant visual data from XML representations of BPMN files, which were identified as uncritical for coherence checking through expert interviews and testing. An equality check then determines if the preprocessed documents are identical; if so, the workflow terminates early, indicating coherence and conserving computational resources. Content comparison In the content comparison phase, the content of the two process document versions is compared to identify changes. Identified changes are aggregated into a JSON element and classified according to the Business Process Change Dimensions: Task, Data, Fig. 2 Overview of aProCheCk
Page 15 of 33Schulte et al. Process Science (2025) 2:22 process owner is primarily interested in all changes to the control flow, but the production planner may be more interested in the resource dimension”. Section C of the online appendix contains a more detailed explanation of the Business Process Change Dimensions and Change Relevance Categories, presenting a framework that enhances the efficiency and accuracy of the coherence checking mechanism. An overview of this framework is depicted in Table1. By addressing the subjective nature of change classification, leveraging the decision-making capabilities of LLMs, and maintaining human-centric oversight, the artifact aims to more effectively manage process coherence and provide accurate and actionable notifications to users. Artifact development and operating logic This section describes the transition from theoretical design to practical implementation of aProCheCk for LLM-enabled business process coherence checking. Importantly, while aProCheCk constitutes a general method for LLM-enabled process coherence checking based process documentation, we also instantiate a software artifact to demonstrate and evaluate aProCheCk in a real-world scenario. The code for the instantiation is available in Section G of the online appendix. The development phase is based on the insights and refined design specifications gained from the expert interviews. This section focuses on prompt engineering techniques that are critical to optimizing the capabilities and performance of the artifact. Section D of the online appendix provides in-depth descriptions of the technology stack underlying the instantiated artifact, the prototype development process, and the software architecture, including its structure, components, and connections. There, we also highlight the iterative nature of the software architecture and the various aspects required to build a robust, scalable artifact based on LLMs. Prompt engineering represents a vital technique for optimizing the utility of LLMs across a range of domains, evident in its use in the context of BPM (Busch et al.2023). Despite its rising significance, prompt engineering remains a relatively new area of research. Only recently established and validated techniques have begun to emerge in the literature, significantly increasing the potential for the use of LLMs (Sahoo et al.2024). The comprehensive taxonomy proposed by Schulhoff et al. (2024) classifies 58 general text-only prompting techniques into six clusters: Zero-Shot, Few-Shot, Thought Generation, Ensembling, Self-Criticism, and Decomposition. Among these approaches, Few-Shot Prompting and Chain of Thought prompting have proven to be particularly effective (Schulhoff et al.2024). Few-Shot prompting involves providing the model with a limited number of examples from which to learn and perform specific tasks (Sahoo et al.2024). This technique is particularly valuable in contexts where extensive training data is not available, allowing the model to generalize effectively from minimal Table 1 Business process change classification framework overview
Page 16 of 33Schulte et al. Process Science (2025) 2:22 examples. Chain of Thought prompting, the only technique introduced in the Thought Generation category, encourages the LLM to articulate its reasoning process step by step before delivering a final answer (Sahoo et al.2024). This encourages responses that are more accurate and logically structured. Both techniques are integrated into the iterative experimental optimization process described in the section Experiment Composition. The operating logic of the artifact involves a two-stage process of logic-based coherence checking that serves as the foundation for the method, as illustrated in Fig.4. The first stage of the process is to compare the content of the modified and original documents. During this stage, individual semantic changes are identified and categorized according to the Business Process Change Dimensions. The second step, the coherence checking step, consists of comparing the previously identified changes with the content of the related document. This comparison includes the description of any required changes to the related document that are necessary to restore process coherence. Here, the identified changes are classified into change categories as described in the Business Process Change Classification Framework. The process can be terminated at any stage if coherence is confirmed. The final step, which is part of our instantiation, involves the generation of the notification, which is constructed in the form of a textual notification based on the structured output. In addition, a notification title and severity indicator are added based on the results of the coherence check. Figure4 shows a conceptual breakdown of the three stages of the operating logic rather than isolated substeps. Since each stage is executed within a single LLM prompt, measuring the performance of each step individually is not feasible. Therefore, we evaluate the accuracy of the coherence-checking logic holistically, as detailed in Appendix B as well as Section E of the online appendix. Adherence to general prompt engineering best practices is crucial in the development of an LLM-based artifact (Lo 2023; Marvin et al. 2024). Hence, throughout the development of aProCheCk, we incorporated established prompting techniques to optimize performance. The prompts were structured to place key information at the beginning, ensuring clarity from the outset and facilitating processing by the model. Clear structural divisions within prompts allowed for logical flow and coherence, while the strategic use of keywords highlighted critical points and guided the model’s focus. Background information was provided after the task description to provide the necessary context without overwhelming the initial instructions. Step-by-step instructions were used to maintain a logical sequence, with constraints listed at the end to avoid distraction from Fig. 4 Operating Logic of the Artifact
Page 17 of 33Schulte et al. Process Science (2025) 2:22 task performance. Importantly, as part of the experimental iterations, a system message was set to assign the role of an experienced BPM expert to the LLM, helping to improve its contextual understanding and decision accuracy. The final prompts as well as the changes made in each experimental iteration are available in the code base in Section G of the online appendix. Evaluation The evaluation of aProCheCk was conducted in two complementary stages to ensure both technical rigor and practical applicability. First, an experimental evaluation systematically assessed the artifact’s effectiveness, robustness, and efficiency through an iterative refinement process. Second, a focus group evaluation examined an instantiation of the artifact in naturalistic settings with real tasks, systems, and users, providing confirmatory insights into its utility, integration potential, and areas for further enhancement. Together, these evaluations establish the artifact’s readiness for real-world application and identify directions for future improvements. Experimental evaluation Experiment composition Experimental evaluation is a crucial step in the validation and benchmarking of aProCheCk, in line with the Eval 3 phase of Sonnenberg and vom Brocke (2012). To accomplish this, we designed the evaluation as a reverse ablation study, that is, we began with a baseline prompt and incrementally added techniques in each iteration. This approach allowed us to isolate and measure the individual impact of each technique. The goal of the evaluation is to iteratively refine the artifact and rigorously assess its performance in a controlled environment, thereby ensuring its readiness for practical application. Conducting such experimental evaluations iteratively is important for the systematic optimization of the artifact, closely following the principles of DSR, which emphasizes iterative development and continuous improvement. Each iteration refines a specific aspect of the artifact, ensuring both practical utility and theoretical soundness. This dynamic approach is consistent with the problem-solving cycle inherent in DSR, which involves problem identification, solution design, implementation, evaluation, and ongoing refinement (Peffers et al. 2007; Tuunanen et al. 2024). This framework ensures that the artifact evolves progressively, incorporating feedback and benchmarking results to improve its performance. The procedure of our experiment iterations is visualized in Fig.5. The dotted arrows visualize the prospect of rolling back the changes made in the most recent iteration if the benchmarking results of the iterations are not satisfactory. Fig. 5 Experiment Iterations
Page 18 of 33Schulte et al. Process Science (2025) 2:22 For the empirical validation, specifically during the experiment phase as defined in the DSR evaluation process, a comprehensive dataset from multiple process repositories was used. This novel Process Coherence Checking Dataset, which is published under the GNU General Public License, is central to assessing the effectiveness and robustness of LLMs in identifying and verifying the coherence of multi-level process documentation. The initial dataset, derived from the research of Sànchez-Ferreres et al. (2018), includes a wide range of process models originally compiled by the BPM Academic Initiative and further detailed by Eid-Sabbagh et al. (2012). These models, enriched with textual descriptions compiled by expert process modelers, form the basis of a dataset for checking the coherence of documented business processes, as described in more detail in Section E of the online appendix. The criteria used for the evaluation of the experiment are also introduced in detail in Section E of the online appendix. The first iteration focuses on establishing a baseline using basic prompting guidelines. This baseline uses structured prompting without any advanced prompt engineering or extensive data preparation. This simple approach forms the baseline against which subsequent iterations are compared. The second iteration aims to improve performance by reducing noise in the dataset. Importantly, this iteration focuses solely on data preprocessing. The prompt itself remains unchanged from the structured baseline prompt. Specifically, irrelevant visual data is removed from BPMN files, significantly reducing the file size and, therefore, the number of input tokens required for processing. This step addresses the needle-inthe-haystack challenge by reducing extraneous data that may divert attention from the incoherence of the content (Nelson et al. 2024). By preserving all the important connections and process flows while removing the noise, the artifact is expected to deliver more accurate results. In addition, an extra check ensures that documents that are identical after the noise reduction step are identified beforehand, saving resources and minimizing the risk of hallucinations in the output. In the third iteration, few-shot prompting is introduced. This refinement provides the LLM with more comprehensive examples of input and output scenarios, both specific to BPMN models and general process documentation, appended to the basic structured prompting template. The inclusion of more examples is intended to better guide the LLM’s decision-making process and ensure consistent and accurate categorization of changes. The enriched context is expected to improve the artifact’s ability to reliably detect and categorize inconsistencies. The fourth iteration incorporates advanced prompt engineering techniques. A system message is introduced to frame the LLM as an experienced BPM expert, which is expected to improve its contextual understanding and decision-making. In addition, the chain of thought method is implemented, which prompts the LLM to generate a detailed reasoning process before making categorizations and decisions. This approach aims to improve the reasoning capabilities of the artifact, leading to more accurate and consistent categorization of changes. The iterative refinement process reinforces the principles of DSR by continuously enhancing the artifact through empirical evaluation and feedback integration. Each iteration builds on the previous one, incrementally improving the functionality and reliability of aProCheCk. To ensure full transparency, the final prompt developed in this last iteration, including comments on where content was
Page 19 of 33Schulte et al. Process Science (2025) 2:22 added in previous iterations, is available in the supplementary repository in Section G of the online appendix. A higher level anatomy of the final prompts used can be found in Appendix C. Experiment findings The experimental results provide comprehensive insights into the artifact’s performance across multiple evaluation criteria, confirming its robustness and readiness for practical application. Through a series of iterations, the accuracy, consistency, and cost-effectiveness of aProCheCk were systematically evaluated and refined. The results are summarized in Table2, where accuracy and consistency are expressed as a value from 0 to 1, with 1 representing the optimal result. Appendix B provides a detailed explanation of how accuracy and consistency are calculated (this is further extended in Section E of the online appendix). The average API cost is expressed on a monetary scale, with the lowest values indicating the best results. The data demonstrates that the mean accuracy and consistency have increased with each iterative refinement, indicating highly effective performance. Moreover, the costs per process run were reduced by over 50% from the first to the second iteration, indicating a significant optimization. Although there was a slight increase in cost in the third and fourth iterations due to the incorporation of additional sophisticated prompting techniques, the overall cost-effectiveness remained significantly better than the baseline established in the first iteration. A more in-depth analysis of the results of the experiment iterations, the statistical significance, and the respective scores for accuracy, robustness, and efficiency can be found in Section E of the online appendix. Overall, aProCheCk demonstrated exceptional performance, given the complexity and subjectivity of the problem domain. While perfect accuracy is unattainable due to the inherent subjectivity of the problem at hand, the structured and iterative optimization process proved effective in significantly improving robustness, efficiency, and effectiveness, validating the artifact’s readiness for real-world BPM applications. Focus group evaluation The focus group evaluation is anchored in the three realities proposed by Sun and Kantor (2006): real tasks, real systems, and real users. These dimensions are central to assessing the practical applicability of the artifact and ensuring its alignment with the requirements of real business processes. Real tasks assess the performance of the artifact in the context of real business processes. Real systems evaluate the artifact’s integration into established BPM ecosystems. Real users highlight the practical utility and user experience of the artifact, ensuring alignment with business needs. The use of naturalistic data from different industries underlines aProCheCk’s ability to effectively address real-world scenarios. This data not only provides a robust dataset for validation but also highlights the versatility and practicality of the artifact. By illustrating the coherence Table 2 Experiment result summary Experiment Iteration Average Accuracy (Effectiveness) Average Consistency (Robustness) Average API Costs per Process Run (€) (Efficiency) 1 0,751 0,891 0,125 2 0,815 0,924 0,054 3 0,823 0,932 0,069 4 0,835 0,936 0,075
Page 20 of 33Schulte et al. Process Science (2025) 2:22 checking capabilities of the artifact’s instantiation in a variety of settings, the readiness for use in real-world BPM environments is demonstrated. Demonstrating the artifact with naturalistic data An instantiation of aProCheCk is demonstrated and evaluated in a naturalistic setting, following the Eval 4 phase introduced by Sonnenberg and vom Brocke (2012). This naturalistic demonstration showcases the functionality of the artifact in a wider range of real-world scenarios with a prototypical implementation. The naturalistic demonstration employed for the focus groups includes checking the coherence of changes to a process model with a related BPMN diagram and vice versa. In addition, this demonstration introduces an overarching document coherence check, where process models are compared against updated process modeling guidelines, and a coherence check on a changing process diagram and a related process diagram of a variant of the process. It further generalizes the applicability of aProCheCk by incorporating different file formats, such as SVG for process models. The data sources for the naturalistic data demonstration consist of real process documents from three German companies of diverse sizes and industries. These companies range from 200 to 300 employees to 10,000–15,000 employees, and the sectors represented include consulting, industrial manufacturing, and the energy and fuel sectors. Table3 summarizes the demonstration data for the four use cases. To ensure confidentiality, parts of the documents have been anonymized, and only the anonymized results generated by the aProCheCk tool are provided. The notification generated by the initial naturalistic example process is illustrated in Fig.6. The remaining three examples can be found in Section F of the online appendix, along with detailed descriptions of the exemplary use cases. These examples demonstrate aProCheCk’s ability to identify and manage changes across multiple levels of process documentation. The artifact effectively fulfills its design specifications by providing actionable insights, ensuring process coherence, and operating efficiently in diverse real-world contexts. These demonstrations underline the robustness and applicability of the artifact, highlighting its potential to significantly improve BPM practices in naturalistic scenarios. Confirmatory Artifact Evaluation The confirmatory evaluation of the developed artifact was conducted through two focus groups consisting of 4 and 5 BPM IT consultants, respectively. Each session lasted approximately one hour and aimed to validate aProCheCk in a naturalistic setting, Table 3 Summary of naturalistic artifact demonstration data Industry of Company Process Name Changed Original Documentation Related Process Document Consulting Conducting Job Interviews BPMN Models (XML Depiction) Process Description (Text) Energy/Fuel Sector Register Incoming Orders (Variants: Manual and Automated) BPMN Model (SVG Depiction) BPMN Model (SVG Depiction) of Process Variant Energy/Fuel Sector Register Incoming Orders (Variants: Manual and Automated) Training Document (Presentation Text) BPMN Models (SVG Depiction) of two Process Variants Insurance Vehicle Insurance Process Overarching Process Modelling Convention (Text) BPMN Model (XML Depiction)
Page 21 of 33Schulte et al. Process Science (2025) 2:22 employing the use cases introduced in the previous section. Both focus groups were conducted for the same purpose and using the same methodology. Conducting two separate sessions enabled us to collect more extensive feedback and validate our artifact more robustly. This approach follows the guidelines established by Tremblay et al. (2010) and is aligned with the Eval 4 phase of the evaluation process proposed by Sonnenberg and vom Brocke (2012). The focus groups were designed to evaluate the artifact in the context of the three realities proposed by Sun and Kantor (2006): real tasks, real systems, and real users, ensuring a comprehensive assessment of the artifact’s performance in realistic conditions. Participants were asked to evaluate the artifact based on three specific criteria derived from the Eval 4 phase: fidelity with real-world phenomenon, impact Fig. 6 aProCheCk Demonstration with Naturalistic Data
Page 22 of 33Schulte et al. Process Science (2025) 2:22 on artifact environment and user, and applicability (Sonnenberg and vom Brocke 2012). Similar to the conducted expert interviews, each criterion was accompanied by a Guiding question to anchor the discussions and allow for a comprehensive evaluation, and participants were instructed to rate each criterion on a Likert scale from 1 to 7. Fidelity with real-world phenomenon was explored with the guiding question, “Could the developed artifact be used in a realistic working environment?” to assess its applicability and practicality in real-world settings. This criterion assesses the artifact’s capacity to manage the complexity of authentic BPM operations and to integrate seamlessly into existing workflows. Impact on artifact environment and user was evaluated through the question “How do you assess the potential influence of the artifact on the working environment and the user?” to understand the impact of the artifact on the existing working environment and user interaction. Participants rated this criterion on a scale from 1, indicating “no positive influence at all”, to 7, indicating “completely positive influence”. Applicability was assessed using the guiding question, “Would the system’s notifications be more of a burden, or would you find them helpful?” to determine the functional usefulness and practicality of the artifact. This criterion determines whether the artifact’s notifications are actionable and beneficial in improving process coherence while reducing manual effort. The boxplot diagram in Fig.7 depicts the distribution of scores across the specified criteria, derived from the nine responses provided by the participants. The results of the evaluation showed a consistently high rating for the criterion of impact on artifact environment and user, reflecting a strong positive perception of the artifact’s potential influence. Moreover, the fidelity with real-world phenomenon was rated favorably, although some reservations were expressed regarding the quality of the data present in practice. Such factors have the potential to impact the artifact’s performance in real-world scenarios. The applicability of the artifact was also met with favorable responses, although some raised concerns regarding the verbosity of the email notifications. Fig. 7 Focus Group Evaluation Results
Page 23 of 33Schulte et al. Process Science (2025) 2:22 Key takeaways from the focus groups highlight various strengths and suggested improvements to aProCheCk. The experts consistently found the artifact impressive and highly relevant, identifying numerous use cases in their respective client organizations, with I 9 (FG1) stating that it is “Highly relevant, because […] in so many customer settings this issue somehow so quickly and easily leads to uncontrolled documentation”. Focus group participants emphasized the practical applicability of the artifact and provided insights into how its functionality could be enhanced. They suggested incorporating deeper process knowledge by embedding more of the organization’s process documentation in the LLM to identify interrelated changes across different processes, which could significantly enhance the usefulness of aProCheCk. The experts emphasized the considerable potential of aProCheCk in managing overarching process documents, particularly in reducing the necessity for manual work, exemplified by I 11 (FG2) saying, “something like this would extremely reduce the manual effort.“ They highlighted that when internal guidelines are updated, the artifact could automatically confirm coherence across all related documents. This capability eliminates the time-consuming task of manually checking each document individually, thereby ensuring organizational consistency and compliance with new policies or regulations. By automating these checks, the artifact significantly enhances operational efficiency and utility within the organization. While acknowledging the current limitations, the experts were optimistic that future iterations of LLMs would further enhance the capabilities of the artifact. They suggested that future versions could enable automated process model generation and other advanced features. Data quality issues, prevalent in many organizations, were identified as a challenge, but the subjective decision-making capabilities of LLMs based on given contexts were seen as a promising solution for maintaining process coherence. In summary, the focus groups provided valuable feedback highlighting the significant potential of aProCheCk, its practical utility, and areas for further improvement. By incorporating these insights into future iterations, the artifact can be refined and optimized to better meet the needs of diverse BPM environments. This confirmatory evaluation highlights both the current strengths of aProCheCk and the promising directions for its continuous improvement. Discussion In order to address our research objectives, we employed a comprehensive and multifaceted DSR approach involving multiple iterations and experts from both research and practice. First, design objectives were derived from the existing literature. Subsequently, provisional design specifications were developed based on the design objectives and evaluated and refined through expert interviews with researchers and practitioners. The evaluation was informed by a preliminary PoC demonstration. aProCheCk was developed iteratively, refined through experimental benchmarking, and then validated through focus groups using naturalistic data and an instantiation of the artifact. In addition, a Business Process Change Classification Framework and an open-source business process coherence checking dataset were developed. As such, our research has important implications for research and practice.
Page 24 of 33Schulte et al. Process Science (2025) 2:22 Theoretical implications Our work represents a significant advancement in the field of BPM by applying generative AI, specifically LLMs, to improve process management practices. In particular, we address the research gap identified by Feuerriegel et al. (2024) concerning the detection of positive process deviance through the use of generative AI. Our research makes several important contributions: First, we contribute to the field of business process coherence checking by introducing a novel approach that utilizes LLMs for the dynamic analysis of diverse process documentation, with the objective of identifying incoherencies. The field is currently dominated by static approaches to inconsistency detection utilizing structured process documentation, such as event logs (Ko and Comuzzi 2023). Existing research on unstructured text-based business process documents has largely focused on static pattern-matching approaches (Martin-Toral et al. 2010; van der Aa et al. 2017). Our LLMbased approach advances this line of work by enabling the autonomous identification of incoherencies in multi-level process documentation, thus moving beyond traditional static analyses. Second, integrating LLMs into the BPM lifecycle, particularly in the process implementation and the monitoring phase, transforms traditional methods that rely on static data and manual reviews (Vidgof et al. 2023). While first studies have started to investigate the potential of LLMs in BPM (Franzoi et al. 2025b), specific applications, such as process coherence checking, remain scarce. Here, our work provides an important starting point by rigorously developing and evaluating an LLM-based artifact to continuously assess the coherence of multi-level process documentation. By enabling dynamic, AIdriven evaluations, the artifact facilitates the detection of negative deviations indicating inefficiencies and positive deviations suggesting innovation opportunities, thereby supporting more context-sensitive decision-making (Franzoi et al. 2025a) and demonstrating the importance of leveraging generative AI to move from static BPM methods to adaptive, proactive management systems (Feuerriegel et al. 2024). Third, we contribute to the field by establishing detailed design specifications for coherence checking based on multi-level process documentation. These specifications balance functional requirements with the complexities of maintaining BPM documentation coherence, providing a robust foundation for future research. Fourth, we introduce the Business Process Change Classification Framework, which comprises Business Process Change Dimensions and Change Relevance Categories. Developed through engagement with established BPM literature and extensive expert interviews, the framework deals with changes in text-based, multi-level business process documentation. By systematically categorizing and quantifying changes, the framework provides a structured mechanism for managing and interpreting the nuanced nature of BPM documentation. This advance improves the theoretical understanding of change management within BPM. Traditional methods often prove inadequate in addressing these complexities and subjectivities, whereas the developed artifact employs LLMs to effectively overcome these challenges. Fifth, we contribute a robust, open-source dataset based on the established work of Sànchez-Ferreres et al. (2018), which provides further support for empirical BPM research. Enriched with expert insights, this dataset is structured around the introduced Business Process Change Classification Framework, providing a valuable resource for
Page 31 of 33Schulte et al. Process Science (2025) 2:22 Table 6 Structure of content comparison and coherence check prompts Prompt Element Short Description Example from Software Instantiation 4. JSON key-value pairs. - changed in Reasoning Structuring iteration to include chain of thought Lists required JSON keys for the output and details what each should contain. “‘[…] ‘technical_comparison’: A technical comparison of the two BPMN diagrams of the same process, including change IDs and technical detail. The Change ID should start at c01 and then count up. […]” 5. Business Process Change Clarification Framework Details Defines relevant elements of the BP Change Clarification Framework with BPMN-specific examples. “BPM Change Dimensions Clarifications: Task: Changes to the fundamental elements of the workflows, including the introduction of new tasks, the modification of existing tasks or the deletion of obsolete tasks. […]” 6. General Clarifications Lists what changes to disregard (e.g. visual changes, syntax changes, unrelated pools, irrelevant IDs). “Changes that are only visual can be neglected, like the size of pools/swim lanes or the coordinates of elements do not matter […]” 7. JSON Format Details - changed in Reasoning Structuring iteration to include chain of thought Shows a complete example JSON output in the correct structure. “[…] \“content_comparison\“: [ {{\“id\“: \“c01\“, \“detail\“: \“content comparison detail 1\“}}, {{\“id\“: \“c02\“, […] }” 8. Few Shot prompting examples - introduced in Data Enrichment iteration Provides concrete change examples with correct classifications to guide the model. “[…] Example 2: Two tasks swap places and therefore all incoming and outcoming control flows need to be adopted as well. Correct Classification: Only one relevant change of type ‘control flow’ […]” Table 7 Structure of notification creation prompt Prompt Element Short Description Example from Software Instantiation 1. Task description Details the main purpose of the task of creating the management summary “[…] Please first describe the changes made to the original document and then describe, how those changes are inconsistent with the related document. […]” 2. Task Data Includes structured comparison, related process document, and most recent BPMN version. “[…] Corresponding related process document only for reference: {txt_filename}: \n{txt_content} […] ” 3. Respond in the following JSON format Return JSON with management summary and email title including urgency indicator. “[…]’, ‘Email title’: { ‘Urgency and Importance Rating’: ‘ 🟢 / 🟡 / 🟠 / 🔴 ’, ‘Email Title’: ‘Short and comprehensive title […]’ } }” 4. Task clarification Outlines what to include/omit, writing style, structure, and formatting rules. “[…] c. Write the text in the same language as the process documents. […] ” Authors’ contributions We describe the authors contributions by describing the respective contribution roles according to the CRediT Taxonomy: Marek Schulte: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Validation, Visualization, Writing – original draft, Writing – review & editing; Sandro Franzoi: Conceptualization, Investigation, Methodology, Project administration, Validation, Writing – original draft, Writing – review & editing; Frank Kühne: Conceptualization, Funding acquisition, Resources, Supervision, Validation, Writing – review & editing; Jan vom Brocke: Conceptualization, Resources, Supervision, Validation, Writing – review & editing. Funding Open Access funding enabled and organized by Projekt DEAL. As part of the Change.WorkAROUND project (promotion sign 02J21C166), this research was funded by the German Federal Ministry of Education and Research. Data availability Supplementary materials for the paper titled ‘Toward LLM-Enabled Business Process Coherence Checking Based on Multi-Level Process Documentation’ by Schulte, M.; Franzoi*, S.; Köhne, F.; vom Brocke, J. submitted for publication to the journal Process Science can be accessed here: https://git hub.com/via dee/process -documen t-coherence-checker.
Page 32 of 33Schulte et al. Process Science (2025) 2:22 Declarations Competing interests The authors declare no competing interests. Received: 30 April 2025 / Accepted: 7 September 2025 References Bartelheimer C, Wolf V, Beverungen D (2023) Workarounds as generative mechanisms for bottom-up process innovation— insights from a multiple case study. Inform Syst J 33:1085–1150 Becker J, Bergener P, Delfmann P, Eggert M, Weiß B (2011) Supporting Business Process Compliance in Financial Institutions - A Model-Driven Approach. In: Bernstein A (ed) Proceedings of the 10th International Conference on Wirtschaftsinformatik: 16–18 February 2011 Zurich, Switzerland, vol 10, Zürich, pp 355–364 Binz M, Schulz E (2023) Using cognitive psychology to understand GPT-3. Proc Natl Acad Sci U S A 120:1–10 Bose RPJC, van der Aalst WMP, Žliobaitė I, Pechenizkiy M (2011) Handling concept drift in process mining. In: Mouratidis H, Rolland C (eds) Advanced information systems engineering, vol 141, 23rd edn. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 391–405 Brützke P, Killewald R, Franzoi S, vom Brocke J (2025) AI-assisted Process Mining for Context-sensitive Analysis Support. Proceedings of the European Conference on Information Systems (ECIS) Busch K, Rochlitzer A, Sola D, Leopold H (2023) Just tell me: prompt engineering in business process management. In: van der Aa H, Bork D, Proper HA, SchmidtR (eds) Lecture notes in business information processing. Enterprise, business-process and information systems modeling, vol. 479. Springer Nature, Switzerland,pp 3–11. h t t p s : / / d o i . o r g / 1 0 . 1 0 0 7 / 9 7 8 - 3 - 0 3 1 - 3 4 2 4 1 - 7 _ 1 Delias P (2017) A positive deviance approach to eliminate wastes in business processes. Ind Manag Data Syst 117:1323–1339 Di Francescomarino C, Donadello I, Ghidini C, Maggi FM, Puura J (2025) Business process deviance mining with sequential and declarative patterns. Bus Inf Syst Eng Eid-Sabbagh R-H, Kunze M, Meyer A, Weske M (2012) A platform for research on process model collections. In: van der Aalst W, Mylopoulos J, Rosemann M, Shaw MJ, Szyperski C, Mendling J, Weidlich M (eds) Business process model and notation, vol 125. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 8–22 Fahland D, Fournier F, Limonad L, Skarbovsky I, Swevels AJE (2024) How well can large language models explain business processes? Feuerriegel S, Hartmann J, Janiesch C, Zschech P (2024) Generative AI. Bus Inf Syst Eng 66:111–126 Franzoi S, Hartl S, Grisold T, van der Aa H, Mendling J, vom Brocke J (2025a) Explaining process dynamics: a process mining context taxonomy for sense-making. Process Sci. https://doi.org/10.1007/s44311-025-00008-6 Franzoi S, Delwaulle M, Dyong J, Schaffner J, Burger M, vom Brocke J (2025b) Using large Language models to generate process knowledge from enterprise content. In: Gdowska K, Gómez-López MT, Rehse J-R (eds) Business process management workshops, vol 534. Springer Nature Switzerland, Cham, pp 247–258 Friedrich F, Mendling J, Puhlmann F (2011) Process model generation from natural Language text. In: Mouratidis H, Rolland C (eds) Advanced information systems engineering, 23rd edn. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 482–496 Galperin BL (2012) Exploring the nomological network of workplace deviance: developing and validating a measure of constructive deviance. J Appl Soc Psychol 42:2988–3025 Gregor S, Hevner AR (2013) Positioning and presenting design science research for maximum impact. MIS Q 37:337–355 Marvin G, Hellen N, Jjingo D, Nakatumba-Nabende J (2024) Prompt engineering in large Language models. In: Jacob Ij, Piramuthu S, Falkowski-Gilski P (eds) Data intelligence and cognitive informatics. Springer Nature Singapore, Singapore, pp 387–402 Grisold T, van der Aa H, Franzoi S, Hartl S, Mendling J, vom Brocke J (2024) A Context Framework for Sense-making of Process Mining Results. In: 2024 6th International Conference on Process Mining (ICPM). IEEE, pp 57–64 Harl M, Zilker S, Weinzierl S (2024) Towards automated business process redesign in runtime using generative machine learning. Proceedings of the European Conference on Information Systems (ECIS) Hevner AR, March ST, Park J, Ram S (2004) Design science in information systems research. MIS Q 28:75–105 Hevner AR, Parsons J, Brendel AB, Lukyanenko R, Tiefenbeck V, Tremblay MC, vom Brocke J (2024) Transparency in design science research. Decis Support Syst 182:1–11 Kampik T, Warmuth C, Rebmann A, Agam R, Egger LNP, Gerber A, Hoffart J, Kolk J, Herzig P, Decker G, van der Aa H, Polyvyanyy A, Rinderle-Ma S, Weber I, Weidlich M (2024) Large Process Models: A Vision for Business Process Management in the Age of Generative AI. KI - Künstliche Intelligenz:1–15 Ko J, Comuzzi M (2023) A systematic review of anomaly detection for business process event logs. Bus Inf Syst Eng 65:441–462 König UM, Linhart A, Röglinger M (2019) Why do business processes deviate? Results from a Delphi study. Bus Res 12:425–453 Kourani H, Berti A, Schuster D, van der Aalst WMP, van der Aa H, Bork D, Schmidt R, Sturm A (2024) Process modeling with large Language models. Enterprise, Business-Process and information systems modeling, vol 511. Springer Nature Switzerland, Cham, pp 229–244 Leopold H, Eid-Sabbagh R-H, Mendling J, Azevedo LG, Baião FA (2013) Detection of naming convention violations in process models for different languages. Decis Support Syst 56:310–325 Lo LS (2023) The art and science of prompt engineering: a new literacy in the information age. Internet Ref Serv Q 27:203–210 Martin-Toral S, Sainz-Palmero G, Dimitriadis Y (2008) Detection Of Incoherences In A Technical And Normative Document Corpus. In: Cordeiro J, Filipe J (eds) Proceedings of the Tenth International Conference on Enterprise Information Systems. SciTePress - Science and and Technology Publications, pp 282–287 Martin-Toral S, Sainz-Palmero G, Dimitriadis Y (2010) Hybrid Approach for Incoherence Detection Based on Neuro-fuzzy Systems and Expert Knowledge. In: Cordeiro J, Filipe J (eds) Proceedings of the 12th International Conference on Enterprise Information Systems. SciTePress - Science and and Technology Publications, pp 408–413
Page 33 of 33Schulte et al. Process Science (2025) 2:22 Mcintosh TR, Liu T, Susnjak T, Watters P, Halgamuge MN (2024) A reasoning and value alignment test to assess advanced GPT reasoning. ACM Trans Interact Intell Syst 14:1–37 Mendling J, Pentland BT, Recker J (2020) Building a complementary agenda for business process management and digital innovation. Eur J Inform Syst 29:208–219 Mertens W, Recker J (2017) Positive Deviance and Leadership: An Exploratory Field Study. In: Sprague R, Bui TX (eds) Proceedings of the 50th Hawaii International Conference on System Sciences (2017). Hawaii International Conference on System Sciences Morana S, Kroenung J, Maedche A, Schacht S (2019) Designing process guidance systems. JAIS 20:499–535 Nelson E, Kollias G, Das P, Chaudhury S, Dan S (2024) Needle in the haystack for memory based large language models. h t t p s : / / d o i . o r g / 1 0 . 4 8 5 5 0 / a r X i v . 2 4 0 7 . 0 1 4 3 7 Nwankpa JK, Roumani Y, Datta P (2022) Process innovation in the digital age of business: the role of digital business intensity and knowledge management. JKM 26:1319–1341 Peffers K, Tuunanen T, Rothenberger MA, Chatterjee S (2007) A design science research methodology for information systems research. J Manage Inf Syst 24:45–77 Polyvyanyy A, Smirnov S, Weske M (2015) Business process model abstraction. In: vom Brocke J, Rosemann M (eds) Handbook on business process management 1, vol 1. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 147–165 Rosemann M, vom Brocke J (2015) The six core elements of business process management. In: vom Brocke J, Rosemann M (eds) Handbook on business process management 1, vol 1. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 105–122 Rosemann M, Recker J, Flender C (2008) Contextualisation of business processes. IJBPIM 3:47 Rosemann M, vom Brocke J, van Looy A, Santoro F (2024) Business process management in the age of AI – three essential drifts. Inf Syst E-Bus Manage. https://doi.org/10.1007/s10257-024-00689-9 Sai C, Winter K, Fernanda E, Rinderle-Ma S (2023) Detecting deviations between external and internal regulatory requirements for improved process compliance assessment. In: Indulska M, Reinhartz-Berger I, Cetina C, Pastor O (eds) Advanced information systems engineering, vol 13901. Springer Nature Switzerland, Cham, pp 401–416 Sahoo PK, Datta R, Rahman MM, Sarkar D. Sustainable environmental technologies: recent development, opportunities, and key challenges. Applied Sciences. 2024;14(23):10956. Saint-Dizier P (2018) Mining incoherent requirements in technical specifications: analysis and implementation. Data Knowl Eng 117:290–306 Sànchez-Ferreres J, van der Aa H, Carmona J, Padró L (2018) Aligning textual and model-based process descriptions. Data Knowl Eng 118:25–40 Schulhoff S, Ilie M, Balepur N, Kahadze K, Liu A, Si C, Li Y, Gupta A, Han H, Schulhoff S [Sevien], Dulepet PS, Vidyadhara S, Ki D, Agrawal S, Pham C, Kroiz G, Li F, Tao H, Srivastava A, . . . Resnik P (2024) The prompt report: a systematic survey of prompting techniques. https://doi.org/10.48550/arXiv.2406.06608 Schulhoff S, Ilie M, Balepur N, Kahadze K, Liu A, Si C, Li Y, Gupta A, Han H, Schulhoff S [Sevien], Dulepet PS, Vidyadhara S, Ki D, Agrawal S, Pham C, Kroiz G, Li F, Tao H, Srivastava A, . . . Resnik P (2024) The prompt report: a systematic survey of prompting techniques. https://doi.org/10.48550/arXiv.2406.06608 Setiawan MA, Sadiq S (2013) A methodology for improving business process performance through positive deviance. Int J Inf Syst Model Des 4:1–22 Sonnenberg C, vom Brocke J (2012) Evaluations in the science of the Artificial – Reconsidering the Build-Evaluate pattern in design science research. In: Hutchison D, Kanade T, Kittler J, Kleinberg JM, Mattern F, Mitchell JC, Naor M, Nierstrasz O, Pandu Rangan C, Steffen B, Sudan M, Terzopoulos D, Tygar D, Vardi MY, Weikum G, Peffers K, Rothenberger M, Kuechler B (eds) Design science research in information systems. Advances in theory and practice, vol 7286. Springer Berlin Heidelberg, Berlin, Heidelberg, pp 381–397 Sun Y, Kantor PB (2006) Cross-evaluation: a new model for information system evaluation. J Am Soc Inf Sci 57:614–628 Teinemaa I, Dumas M, Maggi FM, Di Francescomarino C (2016) Predictive business process monitoring with structured and unstructured data. In: La Rosa M, Loos P, Pastor O (eds) Business process management, vol 9850. Springer International Publishing, Cham, pp 401–417 Tuunanen T, Winter R, vom Brocke J (2024) Dealing with complexity in design science research: a methodology using design echelons. MIS Q 48:427–458 van der Aa H, Leopold H, Reijers HA (2017) Comparing textual descriptions to process models – the automatic detection of inconsistencies. Inf Syst 64:447–460 van der Aa H, Carmona J, Leopold H, Mendling J, Padró L (2018) Challenges and opportunities of applying natural language processing in business process management. In: Bender EM (ed) The 27th International Conference on Computational Linguistics - proceedings of the conference: August 20–26, 2018, Santa Fe, New Mexico, USA: COLING 2018. Association for Computational Linguistics, Stroudsburg, PA, pp 2791–2801 van Dun C, Moder L, Kratsch W, Röglinger M (2023) ProcessGAN: supporting the creation of business process improvement ideas through generative machine learning. Decis Support Syst 165:113880 Venable J, Pries-Heje J, Baskerville R (2016) FEDS: a framework for evaluation in design science research. Eur J Inf Syst 25:77–89 Vidgof M, Bachhofner S, Mendling J (2023) Large Language models for business process management. Opportunities and Challenges vom Brocke J, Winter R, Hevner A, Maedche A (2020) Special issue editorial –accumulation and evolution of design knowledge in design science research: a journey through time and space. JAIS 21:520–544 Weinzierl S, Zilker S, Dunzer S, Matzner M (2024) Machine learning in business process management: A systematic literature review. Expert Syst Appl :1–43 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
