scieee AI-readable full text Open interactive document viewer

Robust Lazy Conflict Detection via Multi-Conflict Extraction and Genetic Diversity Control: Technical Appendix

Le, Viet-Man; Feldgrill, Lukas André; Felfernig, Alexander

Abstract

This document serves as the Technical Appendix to the paper "Robust Lazy Conflict Detection via Multi-Conflict Extraction and Genetic Diversity Control", accepted for publication at the 40th Annual AAAI Conference on Artificial Intelligence (AAAI-26). It provides the full experimental results, extended figures, and detailed analyses that underpin the empirical findings reported in the main paper. All tables, plots, and metrics presented here are consistent with the methodology and evaluation setup described in the main manuscript.

Full text

Robust Lazy Conflict Detection via Multi-Conflict Extraction and Genetic Diversity Control: Technical Appendix Viet-Man Le, Lukas Andr´ e Feldgrill, Alexander Felfernig Graz University of Technology, Graz, Austria v[email protected], [email protected], alexander[email protected] This document serves as the Technical Appendix to the paper “Robust Lazy Conflict Detection via Multi-Conflict Extraction and Genetic Diversity Control”, accepted for publication at the 40th Annual AAAI Conference on Artificial Intelligence (AAAI-26). It provides the full experimental results, extended figures, and detailed analyses that underpin the empirical findings reported in the main paper. All tables, plots, and metrics presented here are consistent with the methodology and evaluation setup described in the main manuscript. This appendix is published as a standalone artifact to ensure transparency, reproducibility, and easy access to all supplementary data referenced in the main paper. The following sections present the complete results across all five configuration knowledge bases (CKBs), with all metrics computed over 20 independent runs per configuration to ensure statistical robustness. Each quantitative claim in the analysis is directly supported by the corresponding tables (Tables 1–5) and visualizations (Figures 1–2), ensuring internal consistency. Coverage and efficiency results are aligned with trade-off and convergence trends, while conclusions about robustness are backed by variance statistics and normalized scores. Furthermore, all overall configuration rankings are computed using a transparent and reproducible scoring formula: Overall = 0.4×NormCoverage + 0.4×NormEfficiency + 0.2× NormStability. This weighted aggregation avoids overemphasizing unstable but high-performing configurations and reflects a realistic assessment of solution quality. Taken together, the presented data, metrics, and derived insights form a coherent and verifiable basis for all claims made in the main paper. The following sections report detailed evaluation results per CKB. Full Evaluation Results For each configuration and CKB, we report the following metrics: •Generations: The average number of generations completed per run, indicating where configurations with extinction mechanisms may terminate earlier. Copyright © 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. •Found CS: The average number of conflict sets discovered. •Std Dev: The standard deviation of the number of conflict sets found, reflecting result stability. •CS Improv.: The improvement in conflict sets found relative to the Baseline. •Runtime: The average wall-clock time per run (formatted as minutes:seconds:milliseconds). •CS/min: The average number of conflict sets identified per minute, indicating discovery throughput. •Checks/CS: The average number of consistency checks required to discover a unique conflict set. •CC Improv.: The improvement in solver effort (consistency checks per conflict set) relative to the Baseline. For the key metrics such as Found CS,CS Improv., Checks/CS, and CC Improv., we highlight the top two configurations per column (with ties equally highlighted). Note that higher values are preferred for Found CS,CS Improv., and CC Improv., while lower values are better for Checks/CS. In the Arcade model (Table 1), All Features delivers the best overall performance, combining the highest coverage (123.25) with a substantial 58.91% reduction in consistency checks. Ext+Multi-CS achieves the highest efficiency (Checks/CS = 39.15), while Full Pop offers the most stable results (Std Dev = 1.16). Multi-CS and Full Pop also improve coverage individually (both over 16%), and Extinction significantly reduces solver effort (–49.15%). In contrast, Baseline remains the weakest across all metrics, and Weighted configurations show only moderate improvements. For the larger FQA model (Table 2), Ext+Multi-CS achieves the highest coverage (288.65) but incurs a higher solver cost. In contrast, All Features provides a more balanced trade-off, reaching nearly the same coverage (286.40) while reducing consistency checks by 38.28%, highlighting its scalability under solver-intensive conditions. MultiCS and Extinction deliver moderate performance across all metrics, without leading in any dimension. Interestingly, the Baseline configuration shows the lowest variance (Std Dev = 4.96), indicating stability despite overall low effectiveness. In the medium-sized B2C model (Table 3), Ext+MultiCS achieves the highest coverage (197.05) and best effi- ciency (Checks/CS = 37.81), confirming the benefit of combining early pruning with multi-conflict extraction. Extinction alone also improves throughput (+15.05% in CC Improv.), while All Features yields the most balanced performance, increasing coverage by 19.87% and reducing solver calls by 35.26%. Full Pop shows the lowest variance (Std Dev = 2.93), suggesting reliable convergence. In contrast, Weighted Basic and Baseline remain less effective across all dimensions. For the highly constrained BusyBox model (Table 4), Ext+Multi-CS achieves the best trade-off between coverage (78.50) and efficiency (Checks/CS = 62.29), boosting conflict discovery by 85.58% while reducing solver calls by over 73%. Extinction alone also improves coverage significantly (+59.10%). In contrast, most standalone enhancements degrade performance. All Features delivers near-peak results with improved robustness, while Weighted Basic offers the highest stability (Std Dev = 3.55). Full Pop shows moderate variance but incurs slightly higher solver costs. The Baseline configuration ranks lowest in both coverage and efficiency. In the largest and sparsest model, EA (Table 5), MultiCS stands out with the highest coverage (2096.50) and best efficiency (Checks/CS = 4.77), nearly doubling the baseline coverage (+85.63%) while cutting solver calls nearly in half (–46.10%). All Features performs comparably well (+85.32% coverage, –45.99% solver calls), confirming its robustness across model scales. Extinction achieves the best stability (Std Dev = 29.74) but contributes little to coverage or efficiency here, suggesting early convergence is not a bottleneck in sparse settings. Full Pop also yields competitive coverage, while Weighted configurations lag behind across all metrics. Across the five evaluated models, clear patterns emerge regarding the strengths and trade-offs of different enhancement strategies: 1. Multi-CS consistently improves conflict coverage and efficiency, particularly when combined with Extinction. This combination (Ext+Multi-CS) yields the best or nearbest results in B2C,BusyBox,FQA, and EA models of varying sizes and constraint densities. This demonstrates its adaptability and effectiveness across diverse settings. 2. All Features consistently delivers strong and balanced performance. It dominates in Arcade and performs nearoptimally in FQA,B2C, and EA, showing that holistic configurations scale well across problem domains. In several cases, it achieves comparable coverage to Ext+Multi-CS but with significantly reduced solver effort. 3. Extinction alone offers significant gains in models prone to early convergence (e.g., BusyBox and B2C), substantially reducing solver calls (up to –73%) and increasing conflict throughput, though its effect is more limited in sparse models like EA. 4. Full Pop stands out for its stability. It achieves the lowest variance in B2C and Arcade, and contributes to high coverage in EA, confirming its role in preserving genetic diversity and ensuring convergence reliability. 5. Weighted configurations show inconsistent or moderate improvements, offering little benefit over more advanced strategies. They often trail behind in both efficiency and coverage, especially in dense or highly constrained models. The combination of multi-conflict extraction (via MultiCS) and diversity preservation (via Extinction and Full Pop) leads to the most effective and scalable improvements across all metrics. While individual techniques provide targeted benefits, the All Features configuration consistently delivers the best balance between coverage, efficiency, and stability, making it the most robust choice across varied model complexities. Trade-off Analysis (Expanded Figure 1) Figure 1 presents the complete version of the trade-off visualization, showing the relationship between Found CS and Checks/CS across all configurations and CKBs. Each subfigure corresponds to a CKB and includes all configurations retained after filtering. Configurations closer to the top-right are preferred, indicating higher effectiveness and efficiency. Across all CKBs, Ext+Multi-CS and All Features consistently demonstrate strong performance, frequently appearing near the Pareto frontier by combining high coverage with relatively low solver effort. Multi-CS excels in the EA model, achieving the highest coverage (2096.50) and efficiency (4.77), clearly dominating that setting. The Extinction configuration effectively reduces solver effort in dense models like BusyBox and B2C, but yields limited benefits on sparser models such as EA. As expected, Baseline performs the weakest across most dimensions. These findings highlight the robustness and scalability of combining enhancements in Lazy conflict detection. In isolation, Multi-CS improves coverage but delivers only modest solver-effort reductions unless integrated with other mechanisms. However, Figure 1 does not capture stability, which is critical for assessing robustness. We refer to Table 4 in the main paper for complete rankings. Convergence and Stability (Expanded Figure 2) Figure 2 provides convergence curves over 100 generations, reporting the average Found CS per generation for selected configurations (Baseline,Multi-CS,Extinction,All Features). These plots illustrate both the convergence rate and the variance across generations. In the Arcade and B2C models (Figures 2a and 2c), convergence is achieved within 40 generations, with both All Features and Extinction rapidly identifying almost all possible conflict sets. In contrast, Baseline stagnates early, while Multi-CS offers moderate improvement but with slower growth. For the larger FQA and EA models (Figures 2b and 2e), All Features maintains a steady discovery rate, ultimately doubling the number of conflict sets compared to Baseline. In EA, the most complex model, All Features approaches near-optimal coverage (2093.05 CS) through gradual improvement, whereas Extinction stabilizes early with Configuration Generations Found CS Std Dev CS Improv. Runtime CS/min Checks/CS CC Improv. [%] [m:s:ms] [%] Baseline 100.00 104.75 2.17 0.00 2:25:979 43.05 95.47 0.00 Multi-CS 100.00 122.80 1.64 17.23 5:32:102 22.19 81.43 14.71 Full Pop 100.00 122.25 1.16 16.71 5:46:18 21.20 81.80 14.32 Weighted Basic 100.00 119.35 1.42 13.94 5:3:948 23.56 83.79 12.23 Weighted Advanced 100.00 121.25 1.37 15.75 5:5:645 23.80 82.47 13.62 Weighted Full 100.00 120.60 2.44 15.13 5:7:109 23.56 82.92 13.15 Extinction 46.15 95.05 4.38 -9.26 1:10:44 81.42 48.55 49.15 Extinction & Multi-CS 48.15 123.00 1.65 17.42 4:16:293 28.80 39.15 58.99 All Features 48.35 123.25 1.58 17.66 4:5:629 30.11 39.23 58.91 Table 1: Performance metrics evaluated with the Arcade knowledge base Configuration Generations Found CS Std Dev CS Improv. Runtime CS/min Checks/CS CC Improv. [%] [m:s:ms] [%] Baseline 100.00 219.35 4.96 0.00 5:53:23 37.28 45.59 0.00 Multi-CS 100.00 241.05 7.29 9.89 9:19:863 25.83 41.49 8.99 Full Pop 100.00 243.75 7.76 11.12 8:56:998 27.23 41.03 10.00 Weighted Basic 100.00 243.85 6.49 11.17 9:51:37 24.75 41.01 10.05 Weighted Advanced 100.00 242.60 7.07 10.60 11:48:793 20.54 41.22 9.59 Weighted Full 100.00 241.15 8.25 9.94 9:36:160 25.11 41.47 9.04 Extinction 94.80 237.55 6.71 8.30 5:39:524 41.98 39.91 12.46 Extinction & Multi-CS 82.45 288.65 6.10 31.59 17:48:387 16.21 28.56 37.35 All Features 80.60 286.40 5.37 30.57 17:16:628 16.58 28.14 38.28 Table 2: Performance metrics evaluated with the FQA knowledge base limited gains, confirming its diminished impact in nonstagnating search spaces. In the highly constrained BusyBox model (Figure 2d), diversity control via Extinction and All Features leads to substantial improvements in both convergence speed and coverage. Overall, the convergence plots reinforce the importance of diversity-preserving mechanisms, especially in dense or stagnation-prone knowledge bases. The combined strategy in All Features enables consistent progress and late-stage discovery across all settings. These results are reflected in the stability dimension of the Overall Score. Rankings by Metric The following steps were used to determine the Best Overall configuration per CKB in Table 4 of the main paper: •Normalized Scores: Each metric, including coverage, efficiency, and stability, is scaled to the [0, 1] range prior to aggregation. •Overall Score: Computed as a weighted sum: Overall = 0.4×NormCoverage + 0.4×NormEfficiency + 0.2× NormStability •Overall Rank: Derived by ranking configurations based on their overall scores. This weighted aggregation scheme ensures that configurations with high variance do not dominate the ranking solely due to strong coverage. By assigning greater importance to coverage and efficiency (0.4 each) while still accounting for stability (0.2), the ranking provides a more balanced and realistic assessment of overall performance. For example, while Ext+Multi-CS often excels in both coverage and efficiency, its higher variance can reduce its overall rank. In contrast, configurations like Full Pop (in EA) and Extinction (in BusyBox) may be preferred due to their superior stability in more sensitive or variance-prone scenarios. This justifies the selection of configurations in Table 4 of the main paper. Configuration Generations Found CS Std Dev CS Improv. Runtime CS/min Checks/CS CC Improv. [%] [m:s:ms] [%] Baseline 100.00 164.30 4.88 0.00 5:36:701 29.28 60.86 0.00 Multi-CS 100.00 174.50 3.43 6.21 6:29:979 26.85 57.31 5.83 Full Pop 100.00 168.60 2.93 2.62 6:9:84 27.41 59.31 2.55 Weighted Basic 100.00 170.35 4.69 3.68 5:49:553 29.24 58.70 3.55 Weighted Advanced 100.00 170.75 3.77 3.93 5:58:723 28.56 58.57 3.76 Weighted Full 100.00 169.40 4.15 3.10 6:21:528 26.64 59.03 3.01 Extinction 95.15 184.05 4.24 12.02 5:9:895 35.63 51.70 15.05 Extinction & Multi-CS 74.50 197.05 4.61 19.93 7:37:162 25.86 37.81 37.87 All Features 77.60 196.95 3.95 19.87 7:37:51 25.85 39.40 35.26 Table 3: Performance metrics evaluated with the B2C knowledge base Configuration Generations Found CS Std Dev CS Improv. Runtime CS/min Checks/CS CC Improv. [%] [m:s:ms] [%] Baseline 100.00 42.30 6.88 0.00 5:26:222 7.78 236.41 0.00 Multi-CS 100.00 44.35 6.52 4.85 5:56:248 7.47 225.48 4.62 Full Pop 100.00 39.50 3.90 -6.62 6:10:4 6.41 253.16 -7.09 Weighted Basic 100.00 37.75 3.55 -10.76 5:58:395 6.32 264.90 -12.05 Weighted Advanced 100.00 39.10 3.67 -7.57 5:55:162 6.61 255.75 -8.18 Weighted Full 100.00 39.65 3.57 -6.26 5:54:328 6.71 252.21 -6.68 Extinction 50.10 67.30 3.69 59.10 3:19:99 20.28 74.44 68.51 Extinction & Multi-CS 48.90 78.50 6.49 85.58 4:10:255 18.82 62.29 73.65 All Features 48.85 77.45 6.96 83.10 4:16:36 18.15 63.07 73.32 Table 4: Performance metrics evaluated with the Busybox knowledge base Configuration Generations Found CS Std Dev CS Improv. Runtime CS/min Checks/CS CC Improv. [%] [m:s:ms] [%] Baseline 100.00 1129.40 29.88 0.00 74:32:479 15.15 8.85 0.00 Multi-CS 100.00 2096.50 55.32 85.63 120:5:385 17.46 4.77 46.10 Full Pop 100.00 2062.55 42.17 82.62 116:46:875 17.66 4.85 45.20 Weighted Basic 100.00 2063.80 71.99 82.73 118:42:896 17.38 4.85 45.20 Weighted Advanced 100.00 2058.65 64.62 82.28 117:49:81 17.47 4.86 45.08 Weighted Full 100.00 2084.15 55.56 84.54 116:36:535 17.87 4.80 45.76 Extinction 100.00 1130.25 29.74 0.08 74:44:590 15.12 8.85 0.00 Extinction & Multi-CS 100.00 2078.30 55.42 84.02 115:58:297 17.92 4.81 45.65 All Features 100.00 2093.05 58.82 85.32 114:31:573 18.28 4.78 45.99 Table 5: Performance metrics evaluated with the EA knowledge base 405060708090 95 100 105 110 115 120 Found CS ( ↑ better) CKB = Arcade 30354045 220 230 240 250 260 270 280 290 CKB = FQA 4045505560 Checks per CS ( ↓ better) 165 170 175 180 185 190 195 CKB = B2C 100150200 Checks per CS ( ↓ better) 45 50 55 60 65 70 75 80 Found CS ( ↑ better) CKB = BusyBox 56789 Checks per CS ( ↓ better) 1200 1400 1600 1800 2000 CKB = EA Baseline Multi-CS Extinction Ext+Multi-CS All Features Figure 1: Trade-off between conflict set coverage (Found CS) and solver effort (Checks/CS) across five configuration knowledge bases. Configurations closer to the top-right are preferred. 40 60 80 100 120 Found CS CKB = Arcade 50 100 150 200 250 300 CKB = FQA 50 100 150 200 CKB = B2C 20 30 40 50 60 70 80 CKB = BusyBox 0 20 40 60 80 100 Generation 0 500 1000 1500 2000 CKB = EA Baseline Baseline Std Multi-CS Multi-CS Std Extinction Extinction Std All Features All Features Std Figure 2: Convergence of conflict set discovery over generations for five CKBs. Lines denote means; shaded areas show standard deviations.