scieee AI-readable full text Open interactive document viewer

Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation

Ahmad, Hadi; Liem, Radita; Lofstead, Jay

Full text

Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation Hadi Ahmad*, Radita Liem*, Jay Lofstead+ *Chair for High Performance Computing, IT Center, RWTH Aachen University +Sandia National Laboratory Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 2 Background •Storage and I/O are critical components of HPC clusters. •We are looking into the IO500 benchmark. The current de facto standard for I/O system benchmarking, combining IOR, mdtest, and find to evaluate bandwidth, metadata, and search performance. •In the IO500 list, the benchmark results are published. It is modeled after TOP500 and Green500 rankings, enabling comparisons across data centers. •However, despite its widespread adoption, IO500 tuning remains complex, often undocumented, and difficult to reproduce. •There is a clear research gap: published studies are limited while there’s a need for more transparent and systematic investigation that can guide data centers to tune the benchmark. •This work tries to address the gap through a literature-guided tuning study on CLAIX23, combining benchmark parameter adjustments with filesystem configuration experiments. Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 3 IO500 Benchmark Overview [1] •IO500 is a benchmark suite designed to evaluate the system’s I/O performance of a cluster •Developed in 2017 and provides a standardized evaluation of both metadata and bandwidth performance •IO500 runs a series of tests that measure I/O bandwidth, metadata performance, and small file handling. •It consists of three main benchmarks: −IOR : Measures read and write bandwidth for parallel I/O. −mdtest : Evaluates metadata operations like file creation, deletion, and stat calls. −Find : Finding relevant objects based on patterns. •The IOR and mdtest benchmarks are configured to represent the ‘worst-case’ and ‘best-case’ I/O operation scenarios in the storage system. Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 4 •Scenarios of the IO500 benchmarks Component Tests Metric Explanation IOR 'easy' ior_easy_write , ior_easy_read GiB /s Free to tune IOR parameters. Typically file -per-process, large, aligned chunks to get the best possible bandwidth performance IOR 'hard' ior_hard_write , ior_hard_read GiB/s Limited options to tune. Forced to use small unaligned I/O to a single shared file for the worst possible bandwidth performance mdtest 'easy' mdtest_easy_delete, mdtest_easy_stat , mdtest_easy_write KIOPS Free to tune mdtest parameters with zero size files in separate directory per process to represent best case scenario for metadata rate mdtest 'hard' mdtest_hard_delete, mdtest_hard_stat, mdtest_hard_write, mdtest_hard_read, KIOPS Limited options to tune. Forced all processes to write on a single shared directory. Representing worst case scenario for metadata rate Find find KIOPS Finding specific subset of files from those created by four scenarios. IO500 Benchmark Overview [2] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 5 IOR •The benchmark used in IO500 to evaluate read and write bandwidth of parallel filesystem. •User can specifies parameters as input to the program to adjust the configuration for their runs •Sample run: ./ ior -t 2m –b 9920000m -a POSIX –s 100000 -F •IOR uses MPI to coordinate the processes to read and write concurrently to ensure they start and running in a coordinated manner. •To measure bandwidth, IOR synchronizes all processes ensuring they start simultaneously. It records start and end times, then divides it by the aggregate amount of data written. Flag Parameter -a API(POSIX,HDF5,MPIIO,…) -b Block Size/Chunk Size -F File per process -s Segment Count -t Transfer Size Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 6 IOR in IO500 •IO500 benchmark’s IOR scenarios are divided into ‘easy’ and ‘hard’ ▪Worst case scenario comes from IOR ‘hard’ scenario ▪Best case scenario comes from IOR ‘easy’ scenario •Sample configuration of IOR easy in IO500: ./ ior -- dataPacketType = timestamp -C -Q 1 -g -G -309386941 -k -e -t 2 m –b 9920000m -F -r -R -a POSIX •Differences between ‘hard’ and ‘easy’ scenarios: Feature Easy Hard File Access Independent Shared Data Pattern/Segment(s) Contiguous Non -contiguous Block/Chunk Size Large (customizable) Small (47008 Bytes) Transfer Size Customizable Small (47008 Bytes) Expected Throughput Higher Lower Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 7 mdtest •It is used to evaluate metadata performance of parallel file systems by measuring how effectively a system handles create, status, and delete operations. •User specifies parameters as input to the program to adjust the configuration for their runs ▪Input example: ./ mdtest -n 1000000 -t -w 1k -e 1K -N 1 -a POSIX •According to the input configuration, mdtest generates a directory tree, then fills the tree with the desired number of files and directories. Different configurations result in different tree depths, and file locations (at leaf only, distributed throughout, etc.). •After creation, directories and files have their metadata retrieved and once the stat operations are complete, the contents of the tree are removed. Users are also given the option to read from these files. •Similar to IOR, mdtest uses MPI to coordinate the processes to create, stat and remove concurrently to ensure they start and proceed in a coordinated manner. •To measure IOPS, it records start and end times, then divides by the number of actions performed. Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 8 mdtest in IO500 •IO500 benchmark’s Mdtest scenarios used to evaluate metadata operation performance. ▪Worst case scenario is coming from Mdtest ‘hard’ scenario ▪Best case scenario is coming from Mdtest ‘easy’ scenario •Sample configuration of mdtest hard in IO500: ./ mdtest -- dataPacketType = timestamp -n 1000000 -t -w 3901 -e 3901 -P -G =577035642 -N 1 -F -C -Y -W 300 -a POSIX •Differences between ‘easy’ and ‘hard’ scenarios: Feature Easy Hard Directory Depth Flat Hierarchical/Deep Operations Create/Stat/Remove Create(Write)/Stat/Read/Remove Write Size 0 3901 Bytes Read Size 0 3901 Bytes Expected Throughput Higher Lower Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 9 Literature Study Publication Task count Chunk Size Transfer Size Stripe Count Filesystem Benchmark Boito et al. BeeGFS IOR Borkar et al. BeeGFS IOR Brzenski et al. BeeGFS IOR Carns et al. PVFS Mdtest Chowdhury et al. BeeGFS IOR Hennecke DAOS IOR, mdtest (IO500) Reed et al. Lustre IOR Saini et al. Lustre IOR Shan et al. Lustre , GPFS IOR Sung et al. Lustre IOR Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 16 •mdtest easy: ~1200% peak performance of mdtest hard. •mdtest easy ~40% of manufacturer listed metadata performance. •mdtest easy performance stable from 240 to 720 tasks. •Drop in performance at 960 tasks. •mdtest hard stable in range of 20-720 tasks. Task Counts : mdtest –Experiment Results [1] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 17 •mdtest ‘hard’ performance declines at 960 tasks. •Large increase in variability of performance at 960 tasks. •Repeated testing for 960 tasks shows that result is not an outlier. •Possible resource contention or context switching affecting performance as IOPS rise and are more stable with decreased core count: Tasks Mean Write (KIOPS) Std. dev Mean stat (KIOPS) Std. dev 920 9.68 0.22 69.01 2.97 940 3.6 3 0. 21 42. 27 3. 15 960 1.59 0.36 27.46 22.11 Task Counts : mdtest –Experiment Results [2] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 18 Storage Targets/Stripe Count •Storage targets refers to the number of stripes or chunks of data that can be simultaneously written or read from a parallel filesystem. •Larger number of storage targets increases performance due to parallelization and load balancing (Boito et al. Reed et al.) In the work of Boito et al. (left image), increases in stripe counts will increases BW from ~1750MiB/s to ~8000MiB/s The work of Reed et al. (right image) shows that increase in stripe count and varying stripe count for different files affects performance. IOR1-3 have static stripe count, IOR 4-6 are altered for different file sizes in experiment. Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 19 Storage Targets/Stripe Count: IOR –Experiments Results •IOR easy performs best at lower stripe counts. •IOR easy has more variability at lower stripe counts. •IOR hard has best performance at 4 targets 8 Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 20 •mdtest easy shows little to no improvement across all metrics with increased stripe count. •Best mdtest easy performance at 8 stripe count albeit with highest variation in stat. •mdtest hard shows lower performance than mdtest easy, as well as greater normalized variation. •Drop in performance in mdtest hard at stripe counts becomes greater than 4. Storage Targets/Stripe Count: mdtest –Experiment Results Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 21 Chunk Size •Chunk size refers to the unit of data stored per object or target in parallel filesystems. •Larger chunks can result in increased performance for larger file sizes as data is fragmented less frequently according to the work of Saini et al. •Smaller chunks can sometimes benefit when frequent access to smaller files is required. IOR benchmark with variable block size. (left) Different file sizes. (right) Different number of OSTs. It shows larger stripe counts benefit from increased stripe size as well. Images from Saini et al.’s paper Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 22 •IOR easy best performance at lower chunk sizes. •IOR hard performs best at 512KiB. •Best performance for IOR easy at 128KiB. •Lower chunk size configuration than 128KiB resulted in errors in IOR easy and a worse performance than 128KiB in IOR hard (R mean: 8277, W mean: 2294) Chunk Size: IOR –Experiment Results [1] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 23 •Peak at 512KiB for read operations. •Similar performance at 256 and 512 KiB for write operations. •Initial increase till 512KiB, drop at larger chunk size for both read and write. Chunk Size: IOR –Experiment Results [2] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 24 •mdtest easy ~900% peak performance of mdtest hard. •mdtest easy performance best at 512KiB. •mdtest easy shows similar performance for write and delete. Chunk Size: mdtest –Experiment Results [1] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 25 •mdtest ‘hard’ performance consistent throughout. •Variation in performance from 118K (best) to 110K (worst) in stat operation. •Read, write and delete perform consistently across range with no distinguishable shift in variation . •No improvement detected that cannot be attributed to variation between the runs. Chunk Size: mdtest –Experiment Results [2] Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 32 Chunk/Block/Stripe Size - mdtest - Setup •BeeOND setup: −Storage targets: 4 −Chunk size: Variable −Core Count: 240 •For mdtest-easy: ./mdtest '-n' '1000000' '-u' '-L' '-F' '-P’ '-G' '1583163012' '-N' '1' '-C' '-Y' '-W' ‘300’ '-a' 'POSIX' •Variable: −Chunk size •For mdtest-hard: ./mdtest '-n' '1000000' '-t' '-w' '3901' '-e' '3901' '-P' '-G=1583177082' '-N' '1' '-F' '-C' '-Y' '-W' ‘300' '-a' 'POSIX' •Differences: Easy scenario has no bytes written to/read from the files. Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 33 Storage Targets/Stripe Count - IOR - Setup •BeeOND setup: −Storage targets: Variable −Chunk size: 512 −Core Count: 240 •For IOR-easy: •Variable: −Storage Targets •For IOR-hard: Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation REX-IO - Cluster 2025 34 Storage Targets/Stripe Count - mdtest - Setup •BeeOND setup: −Storage targets: Variable −Chunk size: 512KiB −Core Count: 240 •For mdtest-easy: ./mdtest '-n' '1000000' '-u' '-L' '-F' '-P’ '-G' '1583163012' '-N' '1' '-C' '-Y' '-W' ‘300’ '-a' 'POSIX' •Variable: −Storage Targets •For mdtest-hard: ./mdtest '-n' '1000000' '-t' '-w' '3901' '-e' '3901' '-P' '-G=1583177082' '-N' '1' '-F' '-C' '-Y' '-W' ‘300' '-a' 'POSIX' •Differences: Easy scenario has no bytes written to/read from the files