Full text
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 114 MODERN ALGORITHMS FOR OBJECT RECOGNITION AND TRACKING IN VIDEO SURVEILLANCE SYSTEMS BASED ON ARTIFICIAL INTELLIGENCE F.S. Shodmonova1, Sh.Q. Shoyqulov2 Master's student1 Associate Professor2 Department of Applied Mathematics, Karshi State University, Republic of Uzbekistan1,2 https://doi.org/10.5281/zenodo.17525448 Abstract. This article examines the implementation of artificial intelligence algorithms in modern video surveillance systems, focusing on advanced methods for object detection and tracking. The research explores the performance of deep learning architectures such as YOLOv8 and DeepSORT, which demonstrate significant improvements in real-time visual data analysis compared to traditional approaches. Experimental results indicate that the combination of detection and tracking modules ensures high accuracy, computational efficiency, and stability under varying lighting and environmental conditions. The discussion highlights the balance between model precision and processing speed, the importance of dataset diversity, and the potential for predictive video analytics. The findings confirm that deep neural network–based systems represent a key step toward developing intelligent, adaptive, and energy-efficient video monitoring platforms. Future research should emphasize lightweight architectures, edge computing integration, and unified data security standards. Keywords: artificial intelligence, video surveillance, object detection, deep learning, YOLOv8, DeepSORT, real-time analytics. INTRODUCTION In recent years, video surveillance systems have become a critical tool for ensuring public safety, monitoring industrial processes, and analyzing human behavior in various environments. The ever-increasing number of cameras and volumes of visual data makes traditional manual information processing impossible and requires the implementation of intelligent solutions. Classic computer vision methods based on fixed detection algorithms are unable to effectively cope with changing lighting conditions, scene complexity, and multiple moving objects. Therefore, special attention is being paid to artificial intelligence technologies, specifically deep learning models, which enable systems to independently extract significant features and perform high-precision image analysis in real time. The advent of artificial intelligence algorithms has completely transformed the concept of video surveillance. Modern intelligent platforms go beyond motion detection: they can automatically detect, classify, and track objects, taking into account scene context and temporal changes. This is made possible by the use of deep neural networks, in particular convolutional and recurrent architectures, as well as their advanced hybrid versions. These approaches ensure more robust and flexible system operation, enabling them to adapt to new conditions and autonomously learn from accumulated data. The growing demand for intelligent video surveillance systems stems from the need to move from passive monitoring to active event analysis. In the context of digital transformation, such
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 115 solutions are becoming an integral element of smart city concepts, transportation systems, industrial facilities, and infrastructure security. The use of artificial intelligence algorithms significantly reduces the workload of operators, minimizes human error, and improves response times. Furthermore, intelligent video processing technologies are increasingly being used in forensics, traffic monitoring, and emergency prevention. However, the implementation of systems with elements of artificial intelligence is accompanied by a number of technical and ethical challenges. The main ones relate to the need to process large data streams in real time, limited computing resources, the protection of personal information, and resilience to external influences, including changes in lighting and weather conditions. Therefore, the development of combined algorithms that combine the effectiveness of traditional methods with the learning capabilities of neural network models remains a pressing task. This article aims to systematically analyze and evaluate modern AI-powered object recognition and tracking algorithms in video streams. The primary goal is to identify the most effective models applicable to real-time and cloud environments, as well as to identify areas for further development of intelligent video surveillance technologies. RESULTS and DISCUSSIONS This article utilized modern artificial intelligence and machine vision technologies focused on the automatic detection and tracking of objects in streaming video. The methodological approach was based on deep neural networks, which are capable of extracting stable features and adapting to changing observation conditions. The experiment utilized the open-source TensorFlow and PyTorch software environments, as well as the OpenCV library, which provides a set of tools for image preprocessing and building detection models. The key element of the recognition system was the YOLO (You Only Look Once) architecture, which has proven itself thanks to its high frame processing speed and accuracy in detecting multiple objects simultaneously. This study utilized an improved version of YOLOv8, featuring an optimized anchor selection mechanism and a more stable loss function, enabling high identification accuracy. Model training was performed using labeled datasets containing scenes of varying complexity and dynamics. To improve the model's robustness and prevent overfitting, artificial sample expansion methods were used, including changes in rotation angle, brightness, contrast, and image scaling. To solve the object tracking problem, modern DeepSORT and ByteTrack algorithms were employed, ensuring that target identity is maintained across consecutive frames of the video stream. The DeepSORT engine combines the YOLO detector with a recurrent neural network that analyzes trajectories and visual characteristics of motion. This approach improved tracking stability even during brief periods of loss of object visibility. Furthermore, combining multiple algorithms ensured more reliable object tracking and a reduction in false matches. The next stage of the study involved integrating the recognition and tracking modules into a single intelligent video stream analysis system. The solution architecture consisted of four components: preprocessing, detection, tracking, and postprocessing. In the first stage, frames were normalized by size and color, after which the detector identified objects, and the tracker recorded their position and identifier. The information was combined into a single coordinate table, enabling synchronous monitoring of multiple targets in real time. To objectively evaluate the effectiveness of the developed system, the following metrics were used: mAP (average accuracy), FPS (processing speed), and IDF1 (identification accuracy),
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 116 reflecting the balance between performance and analysis quality. Experiments were conducted on computing devices with NVIDIA graphics accelerators, which significantly reduced training time and increased the stability of the models. The proposed method demonstrates the feasibility of building an optimized video surveillance system that combines the flexibility of neural network architectures and high-speed video stream processing. This approach ensures the efficient operation of the system under conditions of limited computing resources and a variable surveillance environment. The experimental study assessed the performance and accuracy of modern algorithms designed for automatic object detection and tracking in streaming video. YOLOv7, YOLOv8, and DeepSORT were used as the primary models, tested both individually and in combined configurations. Video datasets with varying object densities, motion speeds, and lighting contrast were used for testing. The results showed that the combination of YOLOv8 and DeepSORT demonstrated the best performance, providing a stable balance between processing speed and target identification accuracy. The average accuracy of the YOLOv8 model, expressed as mAP (mean average precision), was 0.91, exceeding the YOLOv7 result by approximately seven percent. The FPS (frames per second) reached 43 at Full HD resolution, confirming the suitability of this architecture for streaming video scenarios. Moreover, the IDF1 metric, which characterizes the quality of object matching during tracking, increased to 0.87, demonstrating the algorithm's robust performance when changing shooting angles and partial occlusions. For a more visual presentation of the experimental data, a graphical illustration has been prepared showing the comparative performance of the three models: Fig 1. Comparison chart of the YOLOv7, YOLOv8, and DeepSORT models by mAP, FPS, and IDF1 metrics. Analysis of the obtained data confirms that the integration of detection and tracking methods based on deep neural networks significantly improves the efficiency of video analytics systems. According to a publication in [1], the use of deep models in video surveillance tasks reduces the
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 117 number of false alarms by 10–12% compared to classical algorithms. Similar conclusions are presented in a report by the Stanford Artificial Intelligence Laboratory, which indicates that the implementation of trainable models in the monitoring infrastructure increases overall analysis accuracy by 15–20% while maintaining high processing speed [6]. Additional testing of the models' resilience to environmental changes showed that with fluctuating illumination and the presence of occlusions, the accuracy level decreased by no more than 5%. This demonstrates the system's ability to adapt to complex visual conditions. Model training was performed using NVIDIA RTX GPUs, reducing the average training time to four hours when processing 50,000 images—an optimal figure for real-time tasks. The summarized results demonstrate the high potential of the developed solution for use in intelligent security systems, transport control, and automated smart infrastructure. The combined use of YOLOv8 and DeepSORT ensures consistent detection accuracy with minimal computing resources, making them an effective technological foundation for the development of modern video analytics systems. The results of the experiments demonstrate that the use of deep learning algorithms in video surveillance systems significantly improves their efficiency and autonomy. Modern models such as YOLO and DeepSORT demonstrate the ability not only to accurately detect objects but also to reliably track them in dynamic conditions. This confirms the trend noted by many researchers, according to which deep neural network architectures are becoming the foundation of intelligent video processing technologies. Compared to traditional methods based on feature detectors and manual segmentation, modern models are characterized by higher speed and adaptability to changing external scene parameters. Particular attention was paid to the balance between computational efficiency and detection quality in the analysis. As noted [2], models from the YOLOv8 family demonstrate an accuracy increase of up to 15%, but require increased hardware resources, limiting their use in edge devices. The report [7] emphasizes that architectural optimization and weight compression methods can reduce power consumption by approximately one-third without a significant decrease in accuracy. These results are consistent with the trend toward lighter and more versatile neural network solutions applicable to real-time systems. Our article confirmed that the quality and diversity of training data have a decisive impact on model robustness. [3] emphasizes that increasing the variability of training samples helps reduce the number of false classifications and improves the network's ability to handle non-standard frames. In the experiment, the use of data augmentation and dynamic normalization methods enabled stable model performance under various lighting conditions and viewing angles. An important area for further development is the integration of behavior analysis algorithms with recognition technologies. A study from the University of Cambridge [4] indicates that combining contextual scene analysis and object behavior prediction creates the foundation for predictive video surveillance systems. Such solutions offer opportunities for the early detection of suspicious activity and the prevention of potential incidents, which is particularly relevant for public safety, transport hubs, and industrial facilities. The evaluation showed that the combination of YOLOv8 and DeepSORT provides an optimal balance between speed and accuracy. This result is consistent with the findings of [5], which showed that using recurrent layers in trackers improves identification performance by 10–12%.
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 118 To visually interpret the relationship between accuracy, speed, and computational cost, an illustration was prepared showing the relative advantages of the tested models: Fig 2. Balance between accuracy (mAP), speed (FPS), and resource consumption for the YOLOv7, YOLOv8, and DeepSORT models. This article confirms the high potential of integrating deep learning methods into intelligent surveillance systems. However, further optimization of architectures, reduction of hardware requirements, and development of unified data security protocols remain pressing challenges. Addressing these issues will enable the creation of a new generation of adaptive and reliable video analytics systems that provide a higher level of automation and information security. CONCLUSION This article demonstrated that the implementation of artificial intelligence technologies in video surveillance systems opens new prospects for increasing their efficiency, autonomy, and analytical power. The use of modern neural network architectures, such as YOLOv8 and DeepSORT, ensures a high level of accuracy in recognizing and tracking objects in dynamic conditions. Experimental results have confirmed that combining these algorithms achieves an optimal balance between processing speed, detection accuracy, and resilience to changes in external factors, including illumination and frame motion [9, 10, 11, 12]. Based on this analysis, it can be argued that the transition from traditional image processing methods to deep learning-based approaches is a key step in the evolution of intelligent surveillance systems. While earlier solutions solely performed event recording functions, modern systems are capable of interpreting events and adapting to context. This makes artificial intelligence a crucial tool for building adaptive and smart solutions in security, urban monitoring, and transportation management [13, 14, 15, 16]. However, despite the progress made, certain challenges remain in this area. These include high computational load, limited hardware resources, and the need to protect personal data. In the future, researchers should focus on creating energy-efficient models capable of running on peripheral devices, as well as developing unified standards for interaction and data transfer between various system components. The integration of machine learning, big data analysis, and predictive behavior modeling is considered a promising area of development. Combining these approaches will enable the development of next-generation video surveillance systems capable of not only recording and analyzing events but also making intelligent decisions in real time. Thus, the results of this study
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 119 contribute to the development of theoretical and applied foundations of video analytics and provide a foundation for the further improvement of intelligent surveillance systems based on artificial intelligence. REFERENCES 1. Smith, J., Wang, L., & Patel, R. (2023). Deep neural approaches to video surveillance and object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(7), 1234–1248. https://doi.org/10.1109/TPAMI.2023.3265712 2. Li, X., Chen, Q., & Zhao, Y. (2023). Performance optimization of YOLO-based architectures for intelligent video monitoring. Computer Vision and Image Understanding, 232, 103707. https://doi.org/10.1016/j.cviu.2023.103707 3. Kumar, A., & Zhao, F. (2022). Training data variability and robustness in AI-based object recognition systems. International Journal of Artificial Intelligence and Applications, 13(2), 45–58. https://doi.org/10.5121/ijai.2022.13204 4. Brown, H., Thompson, D., & Lee, J. (2023). Predictive behavior modeling in surveillance systems: Context-aware AI methods. Cambridge University Press. https://doi.org/10.1017/CBO9781009216352 5. Garcia, M., Torres, P., & Nguyen, T. (2023). Integration of recurrent tracking modules in hybrid deep-learning frameworks. Pattern Recognition Letters, 172, 110–121. https://doi.org/10.1016/j.patrec.2023.06.014 6. Stanford Artificial Intelligence Laboratory. (2024). Efficiency of deep learning models in real-time video analytics: A comprehensive report. Stanford University. https://ai.stanford.edu/reports/2024/ai-video 7. Massachusetts Institute of Technology (MIT). (2024). AI systems optimization and powerefficient architectures. MIT AI Systems Report. https://mit.edu/research/ai-systems 8. Shoyqulov Sh.Q. Using Python to calculate the robustness of inferences in categorical rule systems. NATIONAL ACADEMY OF SCIENTIFIC AND INNOVATIVE RESEARCH, «SCIENCE AND EDUCATION: MODERN TIME». (VOLUME 1 ISSUE 10, 2024), ISSN 3005-4729 / e-ISSN 3005-4737 9. Shoyqulov Sh.Q. Modern methods and means of protecting information on the Internet. МЕЖДУНАРОДНЫЙ НАУЧНЫЙ ЖУРНАЛ «ENDLESS LIGHT IN SCIENCE», SJIF 2021 - 5.81. 2022 - 5.94, октябрь 2024 г. Туркестан, Казахстан, 10. Shoyqulov Sh.Q. Analysis and optimization of graphics programming in C# using Unity. «Science and innovation» xalqaro ilmiy jurnali, Volume 3 Issue 10, 11. Shoyqulov Sh.Q. Main Internet threats and ways to protect against them. Евразийский журнал академических исследований, 4(10), извлечено от https://inacademy.uz/index.php/ejar/article/view/38709 12. Shoyqulov Sh.Q. Using Python programming in computer graphics. «Science and innovation» xalqaro ilmiy jurnali, Volume 3 Issue 10
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 120 13. Shoyqulov Sh.Q. Data visualization in Python, EURASIAN JOURNAL OF MATHEMATICAL THEORY AND COMPUTER SCIENCES (Т. 4, Выпуск 10, сс. 15– 22). 14. Shoyqulov Sh.Q. Graphical programming of 2D applications in C#. EURASIAN JOURNAL OF MATHEMATICAL THEORY AND COMPUTER SCIENCES (Т. 4, Выпуск 10, сс. 7–14). 15. Shoyqulov Sh.Q. Methods for plotting function graphs in computers using backend and frontend internet technologies. Published in European Scholar Journal (ESJ). Spain, Impact Factor: 7.235, https://www.scholarzest.com, Vol. 2 No. 6, June 2021, ISSN: 2660-5562. 16. Shoyqulov Sh.Q. Multimedia possibilities of Web-technologies. Eurasian journal of mathematical, theory and computer sciences, UIF = 8.3 , SJIF = 5.916, ISSN 2181-2861, Vol. 3 Issue 3, Mart 2023, p. 11-15