scieee AI-readable full text Open interactive document viewer

The AutoCorrelation Integral Drill (ACID) Test Set

Toraman, Gözdenur; Fauconnier, Dieter; Verstraelen, Toon

Abstract

This repository contains the scripts and StepUp workflows to regenerate the "AutoCorrelation Integral Drill" (ACID) test set. The ACID test set comprises a diverse collection of algorithmically generated time series designed to evaluate the performance of algorithms that compute the autocorrelation integral. The set contains in total 15360 test cases, and each case consists of one or more time series. The cases differ in the kernel characterizing the time correlations, the number of time series, and the length of the time series. For each combination of kernel, number of sequences and sequence length, 64 test cases are generated with different random seeds to allow for a systematic validation of uncertainty estimates. The total dataset, once generated, is about 80 GB in size. In addition to the ACID test set, this repository also contains scripts and workflows to validate STACIE, a software package for the computation of the autocorrelation integral. The results of this analysis are discussed in the following paper: Gözdenur Toraman, Dieter Fauconnier, and Toon Verstraelen "STable AutoCorrelation Integral Estimator (STACIE): Robust and accurate transport properties from molecular dynamics simulations" Journal of Chemical Information and Modeling 2025, 65 (19), 10445–10464, doi:10.1021/acs.jcim.5c01475, arXiv:2506.20438. This dataset is distributed under a choice of license: either the Creative Commons Attribution-ShareAlike 4.0 International license (CC BY-SA 4.0) or the GNU Lesser General Public License, version 3 or later (LGPL-v3+). The SPDX License Expression for the documentation is CC-BY-SA-4.0 OR LGPL-3.0-or-later. You should have received a copy of the CC BY-SA 4.0 and LGPL-v3+ licenses along with the data set. If not, see: https://creativecommons.org/licenses/by-sa/4.0/ https://www.gnu.org/licenses/

Full text

Validation of the STable AutoCorrelationIntegral Estimator (STACIE) Using TheAutoCorrelation Integral Drill (ACID) TestSetmodel: exppoly(0, 2)Gözdenur Toraman,† Dieter Fauconnier,†‡ and Toon Verstraelen✶¶† Soete Laboratory, Ghent University, Technologiepark-Zwijnaarde 46, 9052 Ghent, Belgium‡ FlandersMake@UGent, Core Lab EEDT-MP, 3001 Leuven, Belgium¶ Center for Molecular Modeling (CMM), Ghent University, Technologiepark-Zwijnaarde 46, B-9052,Ghent, Belgium✶E-mail: [email protected]Version 2025-06-25 (bf3ff8a)Contents1.Description of Figures and tables ........................................................................... 22.Kernel exp1p ................................................................................................. 33.Kernel exp1w ................................................................................................ 54.Kernel exp2 .................................................................................................. 75.Kernel sho1pcrit ............................................................................................. 96.Kernel sho1pover ........................................................................................... 117.Kernel sho1punder ......................................................................................... 138.Kernel sho1wcrit ........................................................................................... 159.Kernel sho1wover .......................................................................................... 1710.Kernel sho1wunder ........................................................................................ 1911.Kernel sho2crit ............................................................................................. 2112.Kernel sho2over ............................................................................................ 2313.Kernel sho2under .......................................................................................... 251 1.Description of Figures and tablesThe following sections contain figures and tables with the same type of results in each section, butcomputed for different kernels. All figures and tables are labeled with a letter and are explained below. Fora full discussion of the results, we refer to the STACIE paper: TODO ADD CITATION.(a)Illustration of input data.•Left: an example input sequences (first 100 steps).•Center: the sampling autocorrelation function (ACF) of the input data (𝑁=1024, 𝑀=256, purpleline) and the analytical ACF (dashed line).•Right: the sampling power spectral density (PSD) of the input data (𝑁=1024, 𝑀=256, turquoiseline) and the analytical PSD (dashed line).(b)Scaling of uncertainty of the autocorrelation integral with input data.•The slope of the slanted gray lines indicates the ideal scaling of the uncertainty (proportional to1√𝑁𝑀). The spacing between the lines corresponds to a factor of 2 in the uncertainty, the ideal casewhen changing 𝑁 by a factor of 4.•A square represents the standard deviations over 64 repetitions of STACIE’s estimate of theautocorrelation integral for a specific combination of 𝑁 and 𝑀.•The dotted lines represent the corresponding predicted uncertainties.(c)Assessment of the error estimate of the autocorrelation integral.•The square blocks show the ratio of the standard deviation of the STACIE estimate and the RMSvalue of the predicted uncertainty, over 64 repetitions. This value is ideally 100%.•The dots show the ratio of the mean error and the RMS value of the predicted uncertainty, over 64repetitions. This value is ideally 0%.(d)Validation of the Maximum A Posteriori (MAP) estimate.•The MAP estimate for the autocorrelation integral (blue, cross is the maximizer, ellipse is the 2-sigma volume).•The Monte Carlo samples of the posterior distribution of the model parameters (black points).•The mean and covariance of the Monte Carlo samples (red, cross is the mean, ellipse is the 2-sigmavolume).(e)Sensitivity of the autocorrelation integral to the cutoff frequency.•This plot shows how the autocorrelation integral correlates with the effective number of points usedin the fit (top) and the cutoff frequency (bottom).•Results are shown only for the 𝑀=64.•The color code for different 𝑁 corresponds to the legends shown in figures (b), (c), (d) and (e).(f)Number of successful test cases (Failures are typically due to not finding any cutoff frequency withacceptable results.)(g)Sanity check counts for the effective number of points•Number of test cases for each combination of 𝑁 and 𝑀 where the effective number of points used inthe fit is below 20𝑃=40.(h)Sanity check counts for the regression cost z-score•Number of test cases for each combination of 𝑁 and 𝑀 where the z-score of the regression costexceeds 2.(i)Sanity check counts for the cutoff criterion z-score•Number of test cases for each combination of 𝑁 and 𝑀 where the z-score of the cutoff criterionexceeds 2.2 2.Kernel exp1p(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency3 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10245564646464𝑁=4096002519𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102400000𝑁=409601001𝑁=1638400002𝑁=6553600200(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102411325𝑁=409600111𝑁=1638400002𝑁=65536001004 3.Kernel exp1w(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency5 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10245163646464𝑁=4096011420𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102402000𝑁=409600201𝑁=1638401112𝑁=6553600112(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102412544𝑁=409601001𝑁=1638401011𝑁=65536002106 4.Kernel exp2(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency7 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10243148646464𝑁=4096003313𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102400032𝑁=409600111𝑁=1638400010𝑁=6553600112(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102432001𝑁=409601012𝑁=1638400011𝑁=65536000108 5.Kernel sho1pcrit(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency9 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246264646464𝑁=4096103724𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102400200𝑁=409600003𝑁=1638401111𝑁=6553600401(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102442134𝑁=409610000𝑁=1638400000𝑁=655360011116 9.Kernel sho1wover(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency17 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10245064646464𝑁=4096002924𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102401211𝑁=409600020𝑁=1638401103𝑁=6553601111(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102441425𝑁=409601101𝑁=1638400002𝑁=655361010018 10.Kernel sho1wunder(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency19 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40961072345𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102400120𝑁=409600111𝑁=1638400041𝑁=6553600134(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102422337𝑁=409620001𝑁=1638400000𝑁=655360001020 11.Kernel sho2crit(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency21 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246264646464𝑁=4096011618𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102402111𝑁=409601220𝑁=1638400231𝑁=6553600001(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102412323𝑁=409610111𝑁=1638400120𝑁=655360010022 12.Kernel sho2over(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency23 (f) Number of successful test cases𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10246464646464𝑁=40966464646464𝑁=163846464646464𝑁=655366464646464(g) Sanity check counts for the effective number of points𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=10242955646464𝑁=409600007𝑁=1638400000𝑁=6553600000(h) Sanity check counts for the regression cost z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102402203𝑁=409600001𝑁=1638401120𝑁=6553601010(i) Sanity check counts for the cutoff criterion z-score𝑀=1𝑀=4𝑀=16𝑀=64𝑀=256𝑁=102401332𝑁=409613000𝑁=1638400100𝑁=655360100024 13.Kernel sho2under(a) Illustration of input data(b) Scaling of uncertainty of the autocorrelationintegral with input data(c) Assessment of the error estimate of theautocorrelation integral(d) Validation of the Maximum A Posteriori(MAP) estimate(e) Sensitivity of the autocorrelation integral tothe cutoff frequency25