Application of Big Data in Agriculture: A Technical Review Paper
Full text
Application of Big Data In Agriculture Saurav Upadhyaya Louisiana State University, LA, United States I. ABSTRACT In recent years, I have seen the extreme use of Big Data and Deep Learning in the Agriculture sector. However, I noticed that it is very hard to create a robust system that works in an agriculture setting. In this paper, I am analyzing and evaluating the performance of big data frameworks and deep learning algorithms on plant leaves and rice disease detection. I learned how distributed computing platforms such as Apache Spark, Hadoop, and Hive are used with the Convolutional Neural Network (CNN) for detecting diseases in plant leaves. Overall, in this report, I present how researchers are using big data with deep learning to solve the farmers problems, such as detecting tip burn diseases in strawberry leaves using Random Forest Classifier and PySpark, using modern methods for automating the detection of the fungal diseases in apple and cotton plants using CNN, PySpark, computer vision and Hadoop, identifying rice disease using Apache Spark and deep learning algorithms, using Internet of Things (IOT) and big data frameworks in agriculture for smart and sustainable farming, and detecting the quality of the seeds using probabilistic stochastic process. II. LITERATURE REVIEW In “Detecting Plant Diseases at Scale: A Distributed CNN Approach with PySpark and Hadoop” (published in 2024) paper [1], I understood the underlying concept and working of the distributed Convolution Neural Network (CNN), computer vision, PySpark, Tkinter and Hadoop in developing the robust and user-friendly plant disease detection system. I learned how vast amounts of apple and cotton plant image data are used for detecting fungal diseases such as Anthracnose and Apple Scab. I liked how different data augmentation techniques, such as resizing, rotating, and zooming, are used in 2000 images for increasing variability in the datasets. I learned that they are resizing their image to 150 pixels and using those images in a batch using PySpark for distributed processing. Furthermore, I appreciated how they are using machine learning and a big data framework to automate the disease detection process and help farmers detect the disease at an early stage. I liked the detailed explanation on the configuration of Hadoop, Hive, CNN, and Python. Also, I understood how the system detects healthy and unhealthy plants using CNN, PySpark, computer vision, and Hadoop. Likewise, I appreciated how their developed system automated the damage detection early and improved the yields and productivity. In the “Assessing the Feasibility and Scalability of Using Spark for Identifying Tip Burn Diseases in Strawberry Leaves” (2023) paper [2], I understood the importance of machine learning for identifying and detecting tip burn diseases in strawberry leaves with automation. In this paper, I learned how they utilized 1431 diseased and unhealthy leaves of strawberries. I appreciated how they are using Random Forest Classifier and PySpark for distributed computing, improving accuracy and execution speed. I liked the detailed explanation of how they are creating a reliable system using big data and deep learning to detect tip burn diseases in strawberry leaves. In this paper, I understood how large-scale strawberries is possible to process using Apache Spark, which processes the image data in multiple nodes. This makes the model more efficient and scalable. In the “RDI-SD: An Efficient Rice Disease Identification based on Apache Spark and Deep Learning Technique” paper [3], I understood the underlying concept of using Apache Spark and deep learning techniques in developing a model for detecting rice diseases with automation. I noticed that they are using 3,472 images that include 4 different types of rice diseases. I learned that the DenseNet-121 CNN model is integrated with Apache Spark for distributed computing to handle massive dataset collected from Kaggle and Dehradun. Furthermore, I appreciated the detailed explanation of using a big data framework and deep learning for identifying and detecting rich diseases, such as brown spots, tungro, bacterial blight, and leaf smuts. I noticed how the model uses images captured by using phones, drones, and CCTV. I learned that this model could detect four different types of rice diseases with an F1-score of 99.56%. In this paper, I understood the importance of using deep learning models and big data in agriculture settings. Likewise, I also learned about the modular architecture used for identifying and detecting four categories of rice diseases, making the model more robust, efficient, scalable, and adaptable. In “An Overview of Data Science and Big Data Applications in Agriculture for Smart and Sustainable Farming” paper [4], I understood the importance of Deep Learning and Big Data in smart and sustainable farming. I learned how IOT, machine learning, deep learning, and Big Data frameworks can be used in robust decision-making and optimizing resources. Also, I appreciated how data pipelines are developed, especially in the agriculture sector. I learned about requirement gathering, data collection, inspection and data storage, data processing, validation, and visualization in the agriculture sector. I also learned about how we can use block chains in supply chain management and smart farming systems. I appreciated how the authors highlighted the success of different projects after integrating big data with deep learning in agriculture space. For instance, I learned about precision pest management strategy
implemented in Punjab (India), the success of irrigation project in Israel and success of AI-based crop monitoring system in Germany. From these noble and scalable solutions, I learned how big data and deep learning algorithms can solve farmers' problems worldwide. Furthermore, I appreciated the focus and challenges, such as data privacy and security, lack of awareness, and lack of infrastructure and its cost, while using big data in agriculture. In this paper, I learned how big data and deep learning can solve the global problems of farmers. I understood that pests are the major players that damage crops worldwide. I understood that if more investment is done in agricultural related projects and in the infrastructure for making the usage of big data frameworks and deep learning frameworks scalable, secured, adaptable and fault-tolerant, many critical problems that arise in the field gets challenged. Finally, in the “A Stratified Seed Selection Algorithm for Kmeans Clustering on Big Data” (2024) paper [5], I understood how the authors address the limitations of k-means clustering by using probabilistic stochastic point process. I learned that this model uses entire datasets, unlike k-means, random sampling, and mini-batch methods. In this paper, I learned that this process always includes high-potential seeds. The concept of using k-means clustering to determine the quality of the seeds. I learned how probabilistic stochastic process point is used for determining the potential seed cluster, estimating the seed quality and selecting good quality seeds in a large scale. I learned how they are using the whole dataset to improve the quality of the cluster. In fact, I learned that this noble solution is computationally optimized and can cluster a dataset of 10.5 million samples with 28 features in just a couple of minutes. Furthermore, I appreciated how they are using clustering algorithms and big data frameworks for developing a unified solution that can be used in downstream applications. III. DISCUSSION (PROS) In the paper [1], I learned that PySpark and Hadoop handle and analyze large image datasets effectively. I noticed that the model achieves a higher accuracy of 95.76% for determining if the plant is diseased or healthy in the early stage, which improves the yield and reduces the crop losses. I learned how PySpark helps in distributed data processing. Likewise, I noticed how Hadoop’s HDFS confirms the fault tolerance by replicating the image data three times at each node. I understood how Hive helps in reading and managing massive datasets. Also, I understood how the combination of CNN, PySpark, Hadoop, and Tkinter creates a unified solution for plant disease detection and management when compared to traditional models. Also, I appreciated how they are using Tkinter-based web interface so that farmers can use this system with ease. In the paper [2], I learned that their automated system helps farmers detect diseases early. I appreciated how the authors explained the use of PySpark for processing large-scale image datasets efficiently. I understood how the use of deep learning (CNN and VGG-16 approach) in a system gives more accurate results with faster execution speed compared to traditional machine learning algorithms. I learned that PySpark processes large-scale datasets into multiple nodes, which makes the model scalable and efficient. I understood that the integration of PySpark and machine learning makes the worker nodes work much faster, increasing the execution speed to perform. Certain tasks. In this paper, I noticed that the faster execution process helps the system to perform real-time tip burn disease detection in strawberry leaves. In the paper [3], I learned that their research solves a farmer’s problem by using machine learning, Hadoop, and PySpark efficiently. Likewise, I admired how the integration of DenseNet-121 CNN, which is a powerful deep learning framework, with Apache Spark creates a scalable solution and can handle real-time processing of large datasets. I noticed that their system could detect four types of rice diseases with an accuracy of 99%. I learned that PySpark is used for distributed data processing that handles data efficiently. Also, I appreciated the explanation and importance of the underlying background of the Apache Spark, DenseNet-121 and Hadoop in the agricultural sector. In the paper [4], I learned about the advantages of integrating big data and deep learning in the agriculture sector, such as integrated pest management, identifying the damage in the seeds, and optimizing the resources. I appreciated how deep learning and big data can help in smart and sustainable farming. Besides that, I also learned about different phases of data pipelines starting from data collection to robust decision making. I understood how big data can be used to solve various problems, such as controlling the overdose of insecticides and other harmful chemicals in crops. For instance, I learned that Integrated Precision Pest Management Strategy controls the overuse of insecticides, while taking care of the quality of the crops and seeds. Likewise, I appreciated how we can use IPM for managing the pest population through regular monitoring and identifying the pests accurately. I learned that excessive use of insecticides may kill the insecticides, but it doesn’t improve crop productivity and result in higher yields compared to IPM. Thus, I understood how big data and deep learning can be used to monitor the pests and reduce the production cost of the crops. In the paper [5], I learned about how k-means clustering algorithm improved the clustering quality of the seeds. I admired how the entire dataset has been used for selecting high quality seeds. For instance, I learned about how a lot of work can be automated using big data and Artificial Intelligence (AI), which saves time, resources, seeds, and crops. I appreciated how the authors solved the initialization problem, which is encountered while using k-means. I also learned that probabilistic stochastic process points clusters datasets that contain 10.5 million records with 28 features in less than 2-3 minutes, showcasing how optimal, efficient and scalable this algorithm is. In this algorithm, I admired how they are using complete dataset aiming to get patterns from every data point. This is so useful for text data that can be represented in tabular form.
IV. DISCUSSION (CONS) In the paper [1], I learned that Hadoop and PySpark support distributed data processing. However, it is difficult to configure and manage a distributed environment. In addition, I learned that training machine learning models with distributed systems require extensive computing resources. Furthermore, I learned how challenging and costly it is to deploy the system in agriculture settings. I noticed they are automating the damage detection system. However, without monitoring and identifying the disease in real-time, the application cannot solve the farmers' problems. Additionally, from this paper, I learned that the model is trained only on the 2000 images (Kaggle dataset), which is small and may not work well in real-world scenarios. Likewise, I noticed that they are just using leaf images, which limits the application of the model and system in the agriculture setting. In the paper [2], I noticed that PySpark and Random Forest Classifier are used to determine the scalability and accuracy of the system by training the models on a few images. I noticed that the model needs to be trained on more diverse field datasets, considering different lighting conditions and backgrounds. Likewise, I learned how challenging it is to manage and handle deep learning based distributed systems when compared to single-node deep learning frameworks. In this paper, I learned that PySpark manages worker nodes and computers very fast. However, it still needs more memory and space, which is costly. Likewise, I learned that it might be challenging to optimize the performance when PySpark is used. Also, it adds an additional layer of complexity when the machine learning framework is integrated with the model and the system. In the paper [3], I noticed that a small dataset (3,472 images with only four categories of rice diseases) is used for training and determining the accuracy of the model. Likewise, the dataset excludes the images collected in different lighting and environmental conditions, which fails to generalize in agriculture settings. As fewer images are used for training, there is a high risk of overfitting and poor generalization on new, unseen images. I noticed that more diverse datasets are required for training the model. Likewise, I noticed that further analysis needs to be done to determine the measure of disease severity, which is critical and important information to farmers. In addition, although PySpark distributes the workload, the highperformance computing infrastructure is still required for training the model, which is costly and requires more resources. Also, I noticed that the performance of the model is not compared with other existing models, which challenges its adoption. Furthermore, I learned that the paper doesn’t focus much on real-world deployment challenges. In the paper [4], I learned about the challenges in adopting big data in agriculture, such as data privacy and security issues, data fragmentation issues, and infrastructure costs. In fact, I learned about the issues integrating data as it comes from different data sources and gathered under controlled lab and field conditions, making it difficult for further analysis and exploration. Likewise, I learned that it is difficult to train farmers and convince them to use digital tools, especially in rural areas. I understood that internet connectivity and electricity issues could be another major problem, especially in rural agricultural areas. For instance, in Nepal, there are remote agricultural areas where there is neither electricity nor people are literate enough to use digital tools for monitoring their crops. I noticed that it might be difficult to convince the farmers to use digital solutions and make them feel confident to rely on the result of the tool. This is a challenging part! In the paper [5], I noticed that the entire dataset is used for making the seed selection method more robust and scalable. However, it increases computational time and memory usage when compared to other sampling methods. In addition, I noticed that it could be challenging for organizations who have limited infrastructure and resources. I appreciated the work of using probabilistic stochastic point process for selecting the seeds. However, it is very difficult to tune and use this model compared to other heuristics methods. Likewise, I learned that although k-means algorithm is used for obtaining good clusters, this algorithm doesn’t guarantee that the obtained cluster is the optimal result across the entire dataset. Also, I noticed that the k-value in k-means needs to be defined by the users, which is directly used in determining the quality of the final clusters. This process is challenging when datasets are complex and large in scale. Also, I noticed that the algorithm required for handling outliers during the seed selection process is not mentioned in this paper. Likewise, I noticed that the metrics they have used are not listed in the paper, which is making it harder to evaluate and validate their results. V. CONCLUSION In this paper, I learned about the importance of big data, machine learning, deep learning, and computer vision in Agriculture, especially in automation. I admired how the integration of CNN and Apache Spark solved the critical farming problems, such as creating an automated system that identifies and detects the damage in plants, rice, and fruits. In paper [1], I noticed that after integrating CNN, PySpark and Tkinter, their plant disease detection system detected the damaged plants with accuracy of 95.76%. I noticed that more advanced deep learning models like Generative Adversarial Networks (GANs) and Recurrent Neural Networks (RNN) can be used along with the possibility of monitoring the damages in real-time. For instance, I noticed that it is also possible to monitor the pests through Integrated Pest Management (IPM), which improves the yield and preserves the standard quality of the plants while minimizing the excessive usage of the chemicals. I appreciate how researchers are working hard enough to solve the problems of farmers. Yet, more complex research utilizing high-performance computing platforms needs to develop and deploy that can work in an agriculture setting. Also, I learned how the Internet of Things, deep learning and AIpowered solutions are shaping the agricultural industry. I noticed that the implementation and impact of big data, data science, and AI in the agriculture sector is still booming. Thus, I realized that future research should focus more on using high-dimensional data captured in different biological and environmental conditions that can be used in agriculture related downstream applications. As technology is evolving day by day, agriculture companies can benefit and be profitable if those technologies are used strategically to get insights from their field data.
REFERENCES [1] Sharma, V., Kannan, S., Tanya, S., & Panda, N. (2024). Detecting plant diseases at scale: A distributed cnn approach with pyspark and hadoop. Procedia Computer Science, 235, 1044-1057. [2] Prathyuma, V., Hareesh Teja, S., Suganeshwari, G., & Divya, S. (2023, September). Assessing the Feasibility and Scalability of Using Spark for Identifying Tip Burn Diseases in Strawberry Leaves. In International Conference on Advances in Data-driven Computing and Intelligent Systems (pp. 343-354). Singapore: Springer Nature Singapore. [3] Prasad, M. G., Pratap, M. S., Jain, P., Gujjar, J. P., Kumar, M. A., & Kukreti, A. (2022, December). RDI-SD: an efficient rice disease identification based on apache spark and deep learning technique. In 2022 International Conference on Artificial Intelligence and Data Engineering (AIDE) (pp. 277-282). IEEE. [4] Shalik, S. S., & Meyyappan, M. (2025). An Overview of Data Science and Big Data Applications in Agriculture for Smart and Sustainable Farming. Journal of Experimental Agriculture International, 47(5), 759774. [5] Bajpai, N., Paik, J. H., & Sarkar, S. (2024). A Stratified Seed Selection Algorithm for K-means Clustering on Big Data. IEEE Transactions on Artificial Intelligence