Full text
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 184 نيزختلا يف ةيتوصلا تافلملل ةلاعف ةرركم تانايب ءاغلإ ةيمزراوخ ميمصت يباحسلا .أقوزقز رامع .د 1 .منسح نسح 2 :صخلملا ةحاسم كلهتست يهف ،ةمخضلا تانايبلا نيزخت ةمظنأ يف ًامهم ًايدحت ةرركملا تانايبلا لكشت ميظنت ىلع رثؤي ام ،ةينيزخت تايمزراوخ مدختست ،ةلكشملا هذه لحل .اهتجلاعمو اهترادإو تانايبلا نمز دادزي ةباحسلا يف ةنزخملا تانايبلا ةدايز عم نكل .تافلملل لازتخا حيتافم ديلوتل لازتخلاا ةقباطملاو ثحبلا ةيلمع. ىمسي ام اذهو ،ةفلتخم تافلمل لازتخلاا حاتفم قباطتي نأ نكمي امك عتي يذلاو مداصتلاب ثودح ةيلامتحا تّلق حاتفملا لوط داز املكف .لزتخملا حاتفملا لوطب قل .مداصت ةيتوصلا تانايبلا نيزخت ءاغللإ فلملا ىوتسم ىلع ةرركم تانايب ءاغلإ ةينقت مدقن ،ةقرولا هذه يف ميق نع ثحبلا نمز ليلقت ىلإ ةحرتقملا ةينقتلا فدهت .يباحسلا نيزختلا ماظن يف ةرركملا لازتخلاا ىلع دامتعلااب سراهفلا هذه ممصُت .ةددعتم سراهفب لازتخا لودج ءاشنإ قيرط نع فلم ةغيص .توصلا لجأ نم سرهف لكو ،ةددعتم ًاسراهف لازتخلاا لودج نمضتي يلاتلابو ةيمزراوخ مادختسا متي ،مداصت ثودح ةيلامتحا ليلقتل .ةددحم ةغيص MD6 ًاحاتفم جتنت يتلاو ، لوطب512 ًاتب. :ةيحاتفملا تاملكلا ةرركملا تانايبلا ءاغلإةئزجتلا لودجMD6 - ةيتوصلا تافلملا- .يباحسلا نيزختلا ________________________ 1 ،ذاتسأساحلا ةسدنه مسق و ،بةسدنهلا ةيلكامح ،ةصاخلا ةينطولا ةعماجلا ،ةايروس ،. 2 يروس ،صمح ،ةيئابرهكلاو ةيكيناكيملا ةسدنهلا ةيلك ،بيساوحلاو يلَلاا مكحتلا ةسدنه مسق ،هاروتكد بلاط.ة
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 185 Designing an Efficient Deduplication Algorithm for Audio Files in Cloud Storage 2 Hasan Hasan Eng. 1 Dr. Ammar Zakzouk ABSTRACT: Duplicate data poses a significant challenge in big data storage systems as it consumes storage space, affecting data organization, management, and processing. To solve this problem, hash algorithms are used to generate hash keys for files. However, as the amount of data stored in the cloud increases, the search and matching process takes longer. Additionally, hash keys can match different files, known as collisions, which are related to the length of the hash key. The longer the key, the less likely collisions will occur. In this paper, we present a technique for eliminating duplicate data at the file level to reduce storage of duplicate audio data in the cloud storage system. The proposed technique aims to reduce the search time for hash values by creating a reduction table with multiple indexes. These indexes are designed based on the audio file format. Therefore, the hash table includes multiple indexes, each for a specific format. To minimize the probability of collisions, MD6 algorithm is used, which produces a key with a length of 512 bits. KEYWORDS: Deduplication - Hash Table - MD6 - Audio Files - Cloud Storage. ________________________ 1 Professor, Computer Engineering Department, Engineering Faculty, Al-Wataniya Private University, Hama, Syria. 2 Ph.D. Student, Automatic Control and Computers Engineering Department, Mechanical and Electrical Engineering Faculty, Al-Baath University, Homs, Syria.
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 186 1. Introduction The cloud storage system handles a large volume of data in various formats (audio, text, video, images) from different sources. Some of data are duplicated, which is considered unnecessary [1]. To optimize the storage space and improve the performance of the cloud system in terms of storage and processing, techniques for eliminating duplicate data have been employed. There are two types of duplicate data detection methods: source-based detection (client-side) and target-based detection (serverside) [2]. In general, the deduplication process involves generating a hash key for each file and comparing it with the hash keys of other files to determine if it should be stored. A hash table is used to store all the hash keys of the stored files [3]. However, as the size of the data stored in the cloud storage system increases, the number of hash values stored in the hash table also increases. This results in longer search times for hash values. To mitigate this issue, multiple tables or indexes can be created, reducing the time required for comparing the hash key of an incoming file with the existing data in the cloud. This approach also helps address collision problems. By utilizing multiple indexes or tables, the number of comparison operations is reduced, thereby decreasing the probability of collisions [4]. The choice of hash algorithm also plays a crucial role in the deduplication process, as the length of the key directly impacts collision rates [5]. Based on these considerations, we propose an algorithm that involves creating a hash table specifically for audio files on the server. This table consists of multiple indexes, each containing hash keys generated using the MD6 algorithm. MD6 is a cryptographic hash function that employs a Merkle tree-like structure, enabling efficient parallel computation of hashes for extremely long inputs [6].
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 187 2. Audio File Formats There are three main types of audio files formats: Uncompressed Audio Formats, Audio Formats with Lossy Compression, Audio Formats with Lossless Compression [7]. 2.1 Uncompressed Audio Formats Uncompressed audio formats (UAF) consist of real sound waves captured and converted to digital format without further processing. Although, uncompressed audio files tend to be the most accurate but take up a lot of disk space. The Most Common Uncompressed Audio Formats are: 1Pulse-Code Modulation (PCM). 2Waveform Audio File Format (WAV). 3Audio Interchange File Format (AIFF). 2.2 Audio Formats with Lossy Compression Lossy compression is characterized by some data loss during the compression process. However, compression is vital because uncompressed audio takes up a considerable amount of the disk capacity. The Most Common Audio Formats with Lossy Compression are: 1MPEG-1 Audio Layer 3 (MP3). 2Advanced Audio Coding (AAC). 3Windows Media Audio (WMA). 2.3 Audio Formats with Lossless Compression Opposite lossy compression is lossless compression, a method that reduces an audio file's size without ANY data loss between the source and the compressed audio files. The Most Common Audio Formats with Lossless Compression are: 1Free Lossless Audio Codec (FLAC). 2Apple Lossless Audio Codec (ALAC).
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 188 Figure (1): the most common audio files formats 3. Related Work In the cloud, File-Level Deduplication generally depends on generating a hash value for the incoming file using a hash algorithm and comparing it with the hash values already stored in the cloud. If this value doesn’t exist in the hash table, the file is stored in the cloud and its hash value in the hash table. We'll show some of the cloud deduplication techniques: 1In 2015, Naveen A N and V Ravi, proposed a technique to detect duplicate user's files, and then the unique data is stored in the server. This technique is characterized by a low time-complexity, since the process works with small amount of data. However, this technique is considered the least effective because the final user's data may match files on the server [8.] 2In 2016, V. Radia and D. Dingh, made a study of data deduplication techniques: file level, block level, inline post process, source based and target based. The study concluded that source-based deduplication
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 189 technique is the best as it optimizes the uploading bandwidth and storage space over cloud. Distributed deduplication provides security. Both approaches together provide reliability [9.] 3In 2016, Parth Shah et al, proposed a technique to detect duplication between files of users. The technique involves detecting duplication not only within a user's files but also across files from different users. Once the unique files have been identified. This technique is considered more effective in saving storage space and has a medium time complexity since the process takes place at the level of users’ data and not at the server or client. However, user files coming into the storage system may match files that already exist [10.] 4In 2017, Ishita Vaidya and Prof. Rajender Nath, proposed a technique to generate a hash key for the file using MD5 algorithm. Then, this key is compared with the stored keys in Hadoop [11 .] 5In 2018, Manjunath R. Hudagi and Sachin A. Urabinahatti, proposed a technique to deduplicate data on file-level. This technique is based on building a hash table in the Hbase that contains the hash keys for the files stored within the system. To process each incoming file, a hash value is generated using a specific algorithm. This hash value is then compared with the values stored in the hash table. If there is no match between the generated hash and any of the stored values, the file is considered unique and subsequently stored in the system [12.] 6In 2020, Weiqi ZHANG et al. proposed a technique for deduplication in Hadoop. Hash table was designed in namenode. For every block of files, a hash key using SHA-512 was used. To determine the uniqueness of each block, the generated hash key is compared to the existing keys in the system. If there is no match, it indicates a unique block, which gets stored in the Hadoop Distributed File System (HDFS). SHA-512 algorithm
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 190 is used to generate a hash value with 512 bits, which results in a lower collision rate compared to previous algorithms. This ensures a higher level of confidence in identifying and storing unique blocks of data [13 .] 7In 2021, Niteesha Sharma and Dr. A. V Krishna Prasad, proposed a technique to solve the storage issues and deduplication in Hadoop. A table with hash keys is built in Hbase and SHA-256 is used. The process of reading from HBase is faster than Hadoop Distributed File System (HDFS) [14.] 8In 2021, G. Sujatha and Dr. Jeberson Retna Raj. Proposed approach to improve the searching time of duplicated data. Dedicated hash tables were designed, each of which is used for each digital data type. When a file is received, its hash key is compared with the hash table corresponding to its type. Thus, the time required for the matching process is reduced compared to previous techniques. However, this technique did not give importance to the type of hash algorithm [15.] Thus, the previous studies and their solutions as presented above are not sufficient because one of the parameters in the process of eliminating duplicate data is the time required to implement the technique, which is mainly related to the number of comparison operations for the hash key of the incoming file with the stored keys. The fewer the number of operations, the less the execution time. For any incoming audio file, it will be compared with all stored files (text, images, videos, audio) if a single indexing table is used. However, it will be compared with all audio files if multiple indexing tables are used. This requires a long execution time, which increases with the increase in the stored files. 4. Proposed Algorithm We developed an algorithm to deduplicate audio files. The algorithm consists mainly of two phases:
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 191 1building the indexing system. 2deduplication. 4.1 Index System Index System consists of a table with multi-indexes. Nine indexes for the most common audio files formats and one index for other formats. The first index is for Pulse-Code Modulation index. The second index is for Advanced Audio Coding. The third index is for Audio Interchange File Format. The fourth index is for MPEG-1 Audio Layer3. The fifth index is for Advanced Audio Coding. The sixth index is for Windows Media Audio and. The seventh index is for Free Lossless Audio Codec. The eighth index is for Apple Lossless Audio Codec. The last index is for other formats. Figure (2) shows index system and it contains examples for hash values: Figure (2): index system 4.2 Deduplication In this approach, when a file is uploaded to the cloud storage, the metadata of the audio file is read to determine its format. Then, a hash key is calculated using the MD6 algorithm. Based on the format of the file, the file's hash value is compared with the hash values stored in the index specific to that format. If the hash value already exists in the corresponding index, it indicates that the file is a duplicate and therefore should not be
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 192 stored again. However, if the hash value is not found in the index, it means the file is unique and can be stored in the cloud storage. Additionally, the hash value of the file is added to the corresponding index for future reference. Figure (3) shows the routing of hash values. Figure (3): routing system 4.3 Technique Scheme Based on the provided information, the stages of the technique can be arranged as follows: 1 .File Upload: The user uploads the file to the cloud storage system. 2 .Read Metadata: The metadata of the uploaded audio file is read to determine its format. 3 .Calculate Hash Key: MD6 algorithm is used to calculate the hash key of the file. 4 .Index Comparison: Based on the file format, the calculated hash key is compared with the hash values stored in the corresponding index. 5 .Determine Storage Decision: If the hash value is found in the index, it indicates that the file is a duplicate and should not be stored. If the hash value is not found, it means the file is unique and can be stored in the cloud storage.
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 199 reducing the number of comparisons needed. That means that the proposed algorithm eliminates the need for unnecessary comparisons, resulting in improved efficiency and faster processing times when handling different audio formats in the cloud. By reducing the number of comparison operations, the proposed technique not only speeds up the search process but also minimizes the probability of data collisions. Using the proposed technique based on multi-tables, all files are compared with each other, there is a higher probability of two files having similar keys, resulting in collisions. However, with the proposed technique, the chances of collisions are reduced since the comparison is limited to files within the same format index. 7. Conclusion With the increasing volume of data received by the cloud system, which includes duplicate data, it is not enough to design techniques that only focus on efficient duplicate detection. In the face of this massive amount of data, it is necessary to minimize the time required to process duplicate data as much as possible. In this research, we have built a multi-indexing system based on metadata (file format) to reduce the number of comparisons and thus accelerate the search process for the partition key. Reducing the number of comparisons also reduces the probability of collisions between partition keys. In this research, we also used the MD6 algorithm to calculate the hash value for the files, which produces a 512bit key. The longer the key length, the lower the collision rate. In the future, it is possible to build a duplicate data elimination system that relies on multiple tables and multiple indexes to minimize the time complexity as much as possible.
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 200 8. References 1. T. Hashem, I. Yaqoob, N. B. Anuar, S. Mokhtar, A. Gani, S. U. Khan, “The rise of “big data” on cloud computing: Review and open research issue”, Information Systems 47, pp. 98-115, 2015. 2. N. Sharma, A. V. Krishna Prasad, V. Kakulapati, “Data Deduplication Techniques for Big Data Storage Systems”, International Journal of Innovative Technology and Exploring Engineering (IJITEE), ISSN: 2278-3075, Volume-8 Issue-10, pp. 1145-1150, 2019. 3. M. Muniswamaiah, T. Agerwala, C. Tappert, “BIG DATA IN CLOUD COMPUTING REVIEW AND OPPORTUNITIES”, International Journal of Computer Science & Information Technology (IJCSIT), Vol. 11, No. 4, pp. 43-57, 2019. 4. V. Schmitt, J. Jordaan, “Establishing the Validity of Md5 and Sha-1 Hashing in Digital Forensic Practice in Light of Recent Research Demonstrating Cryptographic Weaknesses in these Algorithms”, International Journal of Computer Applications, Volume 68– No.23, pp. 40-43, 2013. 5. M. Eichlseder, F. Mendel, M. Schl¨affer, “Branching Heuristics in Differential Collision Search with Applications to SHA-512”, Conference: Fast Software Encryption - FSE 2014, P. 16. 6. Ronald L. Rivest, “The MD6 hash function A proposal to NIST for SHA-3”, Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA 02139, P. 237, 2008. 7. Audio file format, Audio file format - Wikipedia, 2023. 8. Naveen A N and V Ravi, “Client Side Deduplication Scheme for Secured Data Storage in Cloud Environments”, International Journal
ةصاخلا ةينطولا ةعماجلا ةلجم– دلجملا1 – ددعلا1 -2023 .أقوزقز رامع .د 1 .منسح نسح 2 201 of Engineering Research & Technology (IJERT), Vol. 4 Issue 05, pp. 1465-1467, 2015. 9. V. S. R. and D. K. Singh, “Secure Deduplication Techniques: A Study,” Int. J. Comput. Appl., vol. 137, no. 8, pp. 41–43, 2016, doi: 10.5120/ijca2016908874. 10. Parth. S, Amit. G, Sandipkumar. P, Priteshkumar. P,” Efficient Cross User Client Side Data Deduplication in Hadoop”, Journal of Computers, DOI: 10.17706/jcp.12.4, pp. 362-370, 2016. 11. I. Vaidya, Prof. R. Nath, “An Improved De-Duplication Technique for Small Files in Hadoop”, International Research Journal of Engineering and Technology (IRJET), Volume: 04 Issue: 07, pp. 2040-2045, 2017. 12. Manjunath R. Hudagi and Sachin A. Urabinahatti, “EFFICIENT DEDUPLICATION USING HADOOP”, International Journal of Latest Trends in Engineering and Technology, Vol. (10) Issue (3), pp.236238, 2018. 13. w. Zhang, B. Shao, G. Bian, Q. He, “Research on Multifeature Data Routing Strategy in Deduplication”, Scientific Programming, Volume 2020, Article ID 8869237, P. 11, 2-020. 14. N. Sharma, Dr. A. V Krishna Prasad, “File-level Deduplication by using text files – Hive integration”, International Conference on Computer Communication and Informatics (ICCCI -2021), P. 6, 2021. 15. G. Sujatha, Dr. Jeberson Retna Raj, “Improving the Efficiency of Deduplication Process by Dedicated Hash Table for each Digital Data Type in Cloud Storage System”, Webology, Volume 18, Special Issue on Artificial Intelligence in Cloud Computing, pp. 288-301, 2021.