scieee AI-readable full text Open interactive document viewer

File access patterns of distributed Deep Learning applications

Parraga, Edixon,Leon, Betzabeth,Mendez, Sandra,Rexachs, Dolores,Luque, Emilio

Abstract

Nowadays, Deep Learning (DL) applications have become a necessary solution for analyzing and making predictions with big data in several areas. However, DL applications introduce heavy input/output (I/O) loads on computer systems. These types of applications, when running on distributed systems or distributed memory parallel systems, handle a large amount of information that must be read in the training stage. Inherently parallel and distributed systems and persistent file accesses can easily overwhelm traditional shared file systems and negatively impact application performance. In this way, the management of these applications constitutes a constant challenge due to their popularity in HPC systems. Scientific applications or simulators have traditionally been executed and are optimized for this type systems. Therefore, it is essential to identify the key factors involved in the I/O of a DL application to find the most appropriate form of configuration to minimize the impact of I/O on the performance of this type of application. In the present work, we present an analysis of the behavior of the patterns generated by I/O operations in the training stage of distributed deep learning applications. We selected two well-known datasets such as CIFAR and MNIST to describe file access patterns

Full text

File Access Patterns of Distributed Deep Learning Applications Edixon Parraga1(B), Betzabeth Leon1, Sandra Mendez1,2 , Dolores Rexachs1, and Emilio Luque1 1Computer Architecture and Operating Systems Department, Universitat Aut´onoma de Barcelona, 08193 Bellaterra, Barcelona, Spain {edixon.parraga,betzabeth.leon,sandra.mendez, dolores.rexachs,emilio.luque}@uab.es 2Computer Sciences Department, Barcelona Supercomputing Center (BSC), 08034 Barcelona, Spain [email protected] Abstract. Nowadays, Deep Learning (DL) applications have become a necessary solution for analyzing and making predictions with big data in several areas. However, DL applications introduce heavy input/output (I/O) loads on computer systems. These types of applications, when running on distributed systems or distributed memory parallel systems, handle a large amount of information that must be read in the training stage. Inherently parallel and distributed systems and persistent file accesses can easily overwhelm traditional shared file systems and negatively impact application performance. In this way, the management of these applications constitutes a constant challenge due to their popularity in HPC systems. Scientific applications or simulators have traditionally been executed and are optimized for this type systems. Therefore, it is essential to identify the key factors involved in the I/O of a DL application to find the most appropriate form of configuration to minimize the impact of I/O on the performance of this type of application. In the present work, we present an analysis of the behavior of the patterns generated by I/O operations in the training stage of distributed deep learning applications. We selected two well-known datasets such as CIFAR and MNIST to describe file access patterns. Keywords: Distributed deep learning ·High performance computing · File input/Output ·Parallel I/O 1 Introduction In recent years, Deep Learning (DL) has drawn much attention due to its potential usefulness in different types of applications in the real world. Basic paramThis publication is supported under contract PID2020-112496GB-I00, funded by the Agencia Estatal de Investigaci´on (AEI), Spain and the Fondo Europeo de Desarrollo Regional (FEDER) UE and partially funded by a research collaboration agreement with the Fundaci´on Escuelas Universitarias Gimbernat (EUG). c The Author(s), under exclusive license to Springer Nature Switzerland AG 2022 E. Rucci et al. (Eds.): JCC-BD&ET 2022, CCIS 1634, pp. 3–19, 2022. https://doi.org/10.1007/978-3-031-14599-5_1 4 E. Parraga et al. eters about the data are configured, and the computer is trained to learn on its own. It recognizes patterns by using many layers of processing, managing large volumes of data for the training of a model, through which predictions can be made. The ingestion of this massive amount of information needed to train networks means that high-performance computing (HPC) requires handling large amounts of data and it is not limited to traditional workloads such as simulation [14]. DL applications introduce heavy I/O loads with long, highly simultaneous, and persistent file access, which can overwhelm the conventional file system. Thus, for data-intensive DL applications, I/O performance can become a bottleneck that generates high training latency and introduces CPU and memory overhead, especially when datasets are too large to fit in memory. Some of the most popular deep-learning uses include speech recognition, natural language processing, image recognition, recommendation systems, amongst other things. This vast difference in the information that feeds these systems results in a diverse set of I/O patterns that differ from traditional HPC I/O behavior [14]. Therefore, the way to treat this information must also be different, adapted to the needs of variability, randomness, frequency, and repetitive use that these deep learning applications and training from datasets with these characteristics naturally have. Since the training stage requires a large number of computing resources, there are several proposals in the literature to improve the performance of DL applications [13]. However, there are still many areas to be developed, especially regarding I/O, due to its complexity. In this way, in the present work, we present a methodology for analyzing the behavior of input/output file access patterns of DL applications in distributed systems. We will focus on two well-known datasets such as CIFAR and MNIST, which can be handled with few resources, in order to carry out a detailed analysis at small scale of the behavior of the patterns generated in the I/O operations in the training stage. This paper is structured as follows: Sect. 2presents related work. Section 3 characterizes the I/O Patterns Models of DL applications. Section 4shows experimental data-extraction for file I/O pattern modelling characterization. Finally, in Sect. 5, we explain our conclusions and future work. 2 Related Work In the literature, several studies have been carried out around the characterization of the I/O from different perspectives. For example, in [14], the authors provided a systematic I/O description of ML I/O jobs to understand both how I/O behavior differs across several scientific domains and the scale of workloads. They studied the use of the parallel and burst file system per ML I/O jobs. Furthermore, in [8,20], the authors carried out an analysis of the bottleneck generated by the I/O in the training phase of machine learning. This analysis includes access patterns and a performance model that gives us an overview of storage strategies and their influence on I/O. In the first one, the authors introduced NoPFS, a machine learning I/O middleware and in the second one the File Access Patterns of Distributed Deep Learning Applications 5 authors presented a design DeepIO, an I/O framework for training deep neural networks. Likewise, in [18], the authors designed and built prediction models to estimate performance degradation due to workload placement. On the other hand, in [17], the authors formulated the prediction of I/O performance in production HPC systems as a classification problem and used machine learning techniques to address it. This paper states that storage subsystems are complex, and the I/O operations of running applications can have irregular patterns. In [21] and [9], the authors also dealt with the analysis of access patterns. In the first paper, the I/O patterns of deep neural networks were examined, and in the second, a deep recurrent neural network was proposed that learns the patterns of I/O requests and predicts the next ones. Regarding the behavior of HPC applications using different file systems, formats, and libraries, [5,22] compared the I/O load performance of an HPC application on several file systems. Parallel libraries like HDF5, PnetCDF, and MPIIO were used to attain their I/O patterns. Likewise, in [19], the authors showed the essential characteristics that file systems must have so that they can meet the needs of DL applications. Moreover, in [3], the authors studied some of the most used frameworks, because they indicate that deep learning systems suffer from scalability limitations, particularly for data I/O. These studies characterized the I/O performance and scaling of different frameworks. Likewise, the authors carried out an analysis of the I/O subsystem to understand the cause of its inefficiency. As can be seen from the literature cited in this section, the authors proposed solutions to introduce new middle-ware, I/O frameworks, file system, file format analysis, burst buffer usage, and various prediction models. With our work in this first phase of the research, we intend to analyze the file access patterns of distributed deep learning applications. Compared to the research cited in this section, our contribution is that we intend to determine the influence of the variability of various parameters and techniques that could impact the I/O of DL applications. This is an analysis of the temporal and spatial behavior, taking into account the phases of I/O as groupings of the behavior. With this information, a user could recognize the problem and decide to use configuration strategies that allow I/O to have less impact on DL applications. 3 Characterizing the I/O Patterns Models of DDL Applications 3.1 Software Stack DL A software stack consists of all the software components necessary to support the execution of the application. The components of a software stack work together to deliver application services to the end-user efficiently. The DL application I/O software stack is different from the traditional I/O software stack used in HPC. As shown in Fig. 1, the applications are located at the top, which will go 6 E. Parraga et al. through other layers necessary for its operation, in which the frameworks and libraries are found. The most used frameworks are Tensorflow [2], Pytorch [1], Caffe [10] and, for distributed DL, Horovod [16]. Then there is the file system with all the elements associated with it, which can directly impact the pattern of access to files concerning the behavior of I/O operations. This lower layers are related to the logical relationship between hardware and software, such as I/O buffers, controllers, and the disks on which the information is stored. It is important to know the I/O software stack on which we are working to know which monitoring tools should be used. In addition, it is essential to know at which level of the stack to observe the specific behavior of each layer that can affect the access patterns to files of the DL applications. Fig. 1. Deep learning - IO software stack 3.2 File Access Pattern In the present work, we will focus on the training phase of DL applications. Thus, this section briefly describes the training stage, specifying its close relationship with the I/O when reading the datasets. The purpose is to improve the understanding of the moment in which the I/O operations occur to a greater extent and, therefore, where the I/O pattern of file access can be analyzed. File Access Patterns of Distributed Deep Learning Applications 7 DL applications work through deep neural networks, grouped into input layers that receive the data and hidden layers that perform the mathematical calculations. The model is fitted to the training dataset; this function trains the model for a fixed number of epochs. In this phase, the dataset information is read, and checkpoints [15] are generated according to the configuration used. After this training, an evaluation is performed to confirm that the model is working as desired. Once the model has been successfully evaluated, predictions can be made. There are two essential elements to consider: a large amount of computing power and a large amount of data to train the network correctly. Deep learning applications introduce heavy I/O loads on computing systems regarding a large amount of training data. Distributed DL file data access is highly concurrent and persists throughout the training process. The features of files vary from many small files or one very large shared file [19]. This situation can easily overwhelm traditional shared file systems and negatively impact application performance. When considering the types of parallel I/O, there are two types, as shown in Fig. 2. One file per process, which is more straightforward, and where each process generates its own file. All access to the files is done independently, so no synchronization is needed between I/O operations. In the case of the shared file, all the processes access a single file, which has advantages in disk storage space. Fig. 2. Training stage - Dataset access In addition to the types of file I/O access, it should be noted that if the datasets are small, the system cache or local node storage, such as an SSD, might hold the entire dataset and not affect performance. However, if they are substantial datasets that do not fit in the file system cache or local node storage, they could affect the performance of the learning workload due to the resulting 8 E. Parraga et al. bottleneck [4]. In this way, new technologies and methodologies have had to be incorporated to manage all the information of these applications. The file access pattern is helpful to being able to make predictions and thus being able to optimize the execution of the application by configuring several parameters. The access pattern can be spatial and temporal. – Spatial pattern: this indicates how the file is being accessed taking into account the file offset and the request size. – Temporal pattern: this shows how processes are accessing the file during the application’s execution. Figure 3depicts the key I/O characteristics that we consider to model the file access patterns of DL applications running on HPC. We divide this characteristics in two main groups as follows: – Data Access: it is mainly related with the information needed to represent the spatial and temporal I/O patterns for each file open by a DL application. – Meta-information: it related with the information need to understand the usefulness of the file for the whole application and to determine the I/O resources required by the files of a DL application. The information related with the data access is described as follows: – Type of Operation: this mainly related with the I/O operations done after opening the file, such as read and write operations. – Number of Operations: count of I/O operations done to/from a file. – Size of Operations: bytes wrote to or read from a file by an I/O operation. – Offset: it refers to the displacement within the file, it indicates a reading or writing process where the operations on the file begin. The open subroutine resets it to 0. We should keep in mind that when we monitor a system, depending on where we observe the behavior, the number of operations and the size of the operations also depend on the file system. To understand the I/O context of an application we also consider the metainformation (See Fig. 3) that is described as follows: – A number of files and file size: Depending on the dataset’s characteristics, the data format, and the application used, there may be one or many files, and each of these files has a size that can vary between them. This is useful to know the I/O resources required by an application. – Type of Access: The type of access can be read-only, write-only, or open for read/write. – File Type: When the file type is shared, all processes access a single file. A different kind of file is when each process accesses it independently, so there is no need for synchronization between I/O operations; one file is accessed per process. – Access mode: it can be random, sequential, or strided. File Access Patterns of Distributed Deep Learning Applications 9 Fig. 3. File access pattern and meta information The behavior does not necessarily have to be unique throughout the entire application’s execution, but taking into account the temporal pattern, we can distinguish phases with different spatial behavior. A phase is temporarily identified during a time interval and it is described by its spatial pattern. This behavior can be repeated throughout the application. In the next phase, it can have a different spatial pattern. 4 Experimental Data-extraction for File Access Pattern Modelling Characterization In this section, we present the experimental Data-extraction for file access pattern modelling characterization. A benchmark that trains a neural network for pattern recognition was used. It was executed with 1, 4, and 8 processes in a single node. 4.1 Experimental Environment The technical description of experimental environments is as follows: – Compute nodes - Processor Haswell 2680v3, Intel(R) Xeon(R) CPU E5-2680 v3 @ 2.50 GHz, 24 CORES, 128 GB of RAM. – I/O systems: LUSTRE, NFS 10 E. Parraga et al. – Software: Tensorflow 2.0 [2], Keras [7], Horovod, Python 3.7, Dataset: CIFAR10 [11], MNIST [12]. – HPC I/O Characterization Tool: Darshan. 4.2 Mechanisms Used to Characterize File Access Patterns We instrument the I/O of the deep learning applications by using the Darshan [6] tool. It is a tool to trace and profile the I/O behavior of HPC applications. This tool provides detailed information about the I/O operations performed by serial or parallel applications. Steps to perform the I/O instrumentation: 1. Deploying Darshan: this tool has two module, one is to do the instrumentation (darshan-runtime), and the other is to analyze the trace logs (darshan-util). 2. Instrumentation: Darshan instruments applications via either compile time wrappers or dynamic library preloading. We apply the second option by using the LD PRELOAD environment variable to insert instrumentation at runtime without modifying the application binary. About the placement of darshan logs, we set up the environment variable DARSHAN LOG DIR PATH to indicate the location where Darshan will store the log files. 3. Darshan report generation and analysis: To have the whole I/O information of the application, we use the profile provided by Darshan. Part of the information obtained per job when run an application with Darshan is: total numbers of files, total read and write operations, maximum displacement, operation size grouped by ranges, time spent in operations, and so on. To have the information to represent the spatial and temporal pattern per file we enable the Darshan eXtended Tracing (DXT) module, because we need detailed information about the operation order, offset and request size. 4.3 Characterization of File Access Patterns to the CIFAR-10 Dataset CIFAR-10 (a natural image data set with ten categories) consists of 60,000 32 ×32 color images in 10 classes, with 6,000 images per class. There are 50,000 training images and 10,000 test images. The dataset is divided into five training batches and one test batch containing 10,000 images. The test batch contains exactly 1,000 randomly selected images from each class. The training batches contain the remaining images in random order; the training batches contain exactly 5000 images of each class. The classes are mutually exclusive; the same picture cannot belong to more than one class. The results of the experimentation with CIFAR-10 with four processes in a node on two file systems LUSTRE and NFS will be shown below, since the behavior varies according to the number and size of the operations. Figure 4shows the spatial pattern of access to a file (data batch 1) on LUSTRE. Thus, we observe the following: – Type of Operation: read-only. File Access Patterns of Distributed Deep Learning Applications 11 Fig. 4. Spatial pattern of access to 1 file (data batch 1). Shared file (Dataset: Cifar10, File size: 29.6 MiB) File System:LUSTRE – Number of Read Operations: 4. – Size of read operations: most of the file was read in a single read (28.50 MiB); the rest of the reads were smaller, one at the beginning and two at the end. As can be seen in Fig. 5, all the files of the dataset are accessed by one process. The small reads are shown at the beginning of the first operation of 512 kiB. Then, a large operation of 29184 kiB follows and ends with two small reads, one of 512 KiB and the last one 100.30 KiB. Let us remember that the CIFAR10 dataset is made up of five files called data batch and a test batch file, where we can see that this pattern is repeated for all the dataset files. Fig. 5. Spatial and temporal pattern for a process. Each process open files in order starting from data batch 1todatabatch 5 and finally the test batch. Each process read whole files following the same I/O pattern. Dataset: Cifar10, File size: 29.6MiB, and File System: LUSTRE. Figure 6shows the spatial pattern of access to a file (data batch 1) on NFS. Thus, we observe the following: – Type of Operation: read-only. 18 E. Parraga et al. different; in CIFAR, the same file is shared for all processes, and with MNIST, the file is replicated per process. This behavior can have advantages and disadvantages, such as the need for synchronization, performance, and storage space. We hope to analyze other types of DL applications for future work to see their impact on file access patterns. References 1. Pytorch. https://pytorch.org/docs/stable/index.html/. Accessed 24 Mar 2021 2. Abadi, M., et al.: TensorFlow: large-scale machine learning on heterogeneous systems (2015). https://www.tensorflow.org/, software available from tensorflow.org 3. Bae, M., Jeong, M., Yeo, S., Oh, S., Kwon, O.K.: I/O performance evaluation of large-scale deep learning on an HPC system. In: 2019 International Conference on High Performance Computing & Simulation (HPCS), pp. 436–439. IEEE (2019) 4. Brinkmann, A., et al.: Ad hoc file systems for high-performance computing. J. Comput. Sci. Technol. 35(1), 4–26 (2020) 5. Byna, S., et al.: Exahdf5: delivering efficient parallel I/O on exascale computing systems. J. Comput. Sci. Technol. 35(1), 145–160 (2020) 6. Carns, P., et al.: Understanding and improving computational science storage access through continuous characterization. Trans. Storage 7(3), 8:1–8:26 (2011). https://doi.org/10.1145/2027066.2027068 7. Chollet, F., et al.: Keras. https://github.com/fchollet/keras (2015) 8. Dryden, N., B¨ohringer, R., Ben-Nun, T., Hoefler, T.: Clairvoyant prefetching for distributed machine learning I/O. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC 2021, Association for Computing Machinery, New York (2021). https://doi.org/10.1145/ 3458817.3476181 9. Farhangi, A., Bian, J., Wang, J., Guo, Z.: Work-in-progress: a deep learning strategy for I/O scheduling in storage systems. In: 2019 IEEE Real-Time Systems Symposium (RTSS), pp. 568–571. IEEE (2019) 10. Jia, Y., et al.: Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093 (2014) 11. Krizhevsky, A.: Learning multiple layers of features from tiny images. Technical Report (2009) 12. LeCun, Y., Cortes, C., Burges, C.: MNIST. http://yann.lecun.com/exdb/mnist/. Accessed 24 Mar 2021 13. Mittal, S., Rajput, P., Subramoney, S.: A survey of deep learning on CPUs: opportunities and co-optimizations. IEEE Trans. Neural Netw. Learn. Syst. 1–21 (2021). https://doi.org/10.1109/TNNLS.2021.3071762 14. Paul, A.K., Karimi, A.M., Wang, F.: Characterizing machine learning I/O workloads on leadership scale HPC Systems. In: 2021 29th International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS), pp. 1–8 (2021). https://doi.org/10.1109/MASCOTS53633. 2021.9614303 15. Rojas, E., Kahira, A.N., Meneses, E., Gomez, L.B., Badia, R.M.: A study of checkpointing in large scale training of deep neural networks. arXiv preprint arXiv:2012.00825 (2020) 16. Sergeev, A., Balso, M.D.: Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 (2018) File Access Patterns of Distributed Deep Learning Applications 19 17. Wan, L., et al.: I/O performance characterization and prediction through machine learning on hpc systems. In: CUG2020 Proceedings (2020) 18. Zacarias, F.V., Petrucci, V., Nishtala, R., Carpenter, P., Moss´e, D.: Intelligent colocation of HPC workloads. J. Parall. Distrib. Comput. 151, 125–137 (2021) 19. Zhang, Z., Huang, L., Pauloski, J.G., Foster, I.: Aggregating local storage for scalable deep learning I/O. In: 2019 IEEE/ACM Third Workshop on Deep Learning on Supercomputers (DLS), pp. 69–75. IEEE (2019) 20. Zhu, Y., et al.: Entropy-aware I/o pipelining for large-scale deep learning on HPC systems. In: 2018 IEEE 26th International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS), pp. 145– 156 (2018). https://doi.org/10.1109/MASCOTS.2018.00023 21. Zhu, Y., Yu, W., Jiao, B., Mohror, K., Moody, A., Chowdhury, F.: Efficient userlevel storage disaggregation for deep learning. In: 2019 IEEE International Conference on Cluster Computing (CLUSTER), pp. 1–12 (2019). https://doi.org/10. 1109/CLUSTER.2019.8891023 22. Zhu, Z., Tan, L., Li, Y., Ji, C.: PHDFS: optimizing I/O performance of HDFS in deep learning cloud computing platform. J. Syst. Archit. 109, 101810 (2020)