Enhancing cricket performance analysis with human pose estimation and machine learning
Abstract
Producción Científica
Full text
Citation: Siddiqui, H.U.R.; Younas, F.; Rustam, F.; Flores, E.S.; Ballester, J.B.; Diez, I.d.l.T.; Dudley, S.; Ashraf, I. Enhancing Cricket Performance Analysis with Human Pose Estimation and Machine Learning. Sensors 2023,23, 6839. https:// doi.org/10.3390/s23156839 Academic Editor: Kah Phooi Seng Received: 13 June 2023 Revised: 21 July 2023 Accepted: 29 July 2023 Published: 1 August 2023 Copyright: © 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). sensors Article Enhancing Cricket Performance Analysis with Human Pose Estimation and Machine Learning Hafeez Ur Rehman Siddiqui 1, Faizan Younas 1, Furqan Rustam 2, Emmanuel Soriano Flores 3,4,5 , Julién Brito Ballester 3,6,7, Isabel de la Torre Diez 8,* , Sandra Dudley 9and Imran Ashraf 10,* 1Institute of Computer Science, Khwaja Fareed University of Engineering and Information Technology, Abu Dhabi Road, Rahim Yar Khan 64200, Punjab, Pakistan; [email protected] (H.U.R.S.); [email protected] (F.Y.) 2School of Computer Science, University College Dublin, D04 V1W8 Dublin, Ireland; furqan.r[email protected] 3Engineering Research & Innovation Group, Universidad Europea del Atlántico, Isabel Torres 21, 39011 Santander, Spain; [email protected] (E.S.F.); [email protected] (J.B.B.) 4Department of Project Management, Universidad Internacional Iberoamericana Campeche, Campeche 24560, Mexico 5Department of Projects, Universidad Internacional Iberoamericana Arecibo, Puerto Rico, PR 00613, USA 6Project Management, Universidade Internacional do Cuanza, Cuito EN250, Angola 7Fundación Universitaria Internacional de Colombia Bogotá, Bogotá 11001, Colombia 8Department of Signal Theory, Communications and Telematics Engineering, University of Valladolid, Paseo de Belén, 15, 47011 Valladolid, Spain 9Bioengineering Research Centre, School of Engineering, London South Bank University, 103 Borough Road, London SE1 0AA, UK; [email protected] 10 Department of Information and Communication Engineering, Yeungnam University, Gyongsan-si 38541, Republic of Korea *Correspondence: [email protected] (I.d.l.T.D.); [email protected] (I.A.) Abstract: Cricket has a massive global following and is ranked as the second most popular sport globally, with an estimated 2.5 billion fans. Batting requires quick decisions based on ball speed, trajectory, fielder positions, etc. Recently, computer vision and machine learning techniques have gained attention as potential tools to predict cricket strokes played by batters. This study presents a cutting-edge approach to predicting batsman strokes using computer vision and machine learning. The study analyzes eight strokes: pull, cut, cover drive, straight drive, backfoot punch, on drive, flick, and sweep. The study uses the MediaPipe library to extract features from videos and several machine learning and deep learning algorithms, including random forest (RF), support vector machine, knearest neighbors, decision tree, linear regression, and long short-term memory to predict the strokes. The study achieves an outstanding accuracy of 99.77% using the RF algorithm, outperforming the other algorithms used in the study. The k-fold validation of the RF model is 95.0% with a standard deviation of 0.07, highlighting the potential of computer vision and machine learning techniques for predicting batsman strokes in cricket. The study’s results could help improve coaching techniques and enhance batsmen’s performance in cricket, ultimately improving the game’s overall quality. Keywords: batsman stroke prediction; computer vision; machine learning; random forest 1. Introduction Human pose estimation (HPE) is a rapidly developing field of research that employs computer vision techniques to estimate the positions of various human body components in images or video footage. Despite recent advancements in computer vision, accurately understanding human actions from visual data is still challenging. Human body movements are often driven by unique activities, making identifying and categorizing them accurately difficult. Understanding a person’s body pose is crucial for identifying their actions, which is where HPE techniques come in handy. By recognizing and categorizing Sensors 2023,23, 6839. https://doi.org/10.3390/s23156839 https://www.mdpi.com/journal/sensors
Sensors 2023,23, 6839 2 of 16 human body joints, such as the head, arms, and torso, HPE can capture coordinates for each joint that define a person’s position [1]. In sports analytics, computer vision has become increasingly crucial for extracting valuable insights from various forms of visual data [ 2 ]. Coaches and athletes can use computer vision techniques to track and analyze movement patterns during games or practice sessions, providing valuable performance feedback, identifying areas for improvement, and making strategic decisions [ 3 ]. Additionally, computer vision can be used for activity recognition, outcome prediction, and injury prevention. Using computer vision in sports can revolutionize how we analyze and train athletes, improving their performance and reducing the risk of injury. Human pose estimation, in particular, is an exciting area of research within sports analytics. With advancements in camera technology and computer vision algorithms, tracking of athletes’ body movements during training and competition has become more accurate over time [4]. This technology has significant applications in sports performance analysis and injury prevention. Coaches and athletes can monitor progress, identify areas for improvement, and prevent potential injuries by tracking body movements. Human pose estimation can also provide insights into the biomechanics of athletic movements, helping coaches and trainers optimize training methods and improve performance. The application of human pose estimation in sports extends to various sports, including basketball, soccer, and volleyball, making it an area of growing interest among researchers exploring its potential for improving athletic performance and reducing the risk of injury. Human pose estimation through computer vision has revolutionized how cricket strokes are analyzed and predicted. By scrutinizing batsmen’s body posture and movements during a game, coaches and analysts gain detailed insights into their batting techniques and strategies [ 5 ]. Computer vision techniques are used to detect the orientation of the bat and the position of the batsman’s body, enabling the identification of different types of strokes played by the batsman. This data analysis helps recognize a batsman’s strengths and weaknesses, empowering coaches and players to optimize their training and gameplay. Furthermore, the integration of machine learning algorithms enables the system to forecast the type of shot the batsman is likely to play based on their previous performances [ 6 ]. Such predictions are instrumental in helping bowlers anticipate the shot and modify their strategy accordingly. For instance, if the system forecasts that the batsman is likely to play a cover drive, the bowler may adjust their line and length to make it more difficult to play that shot. In conclusion, human pose estimation using computer vision in cricket has exhibited enormous potential in enhancing performance analysis and improving training methods. It enables coaches and players to make data-driven decisions, ultimately improving their chances of winning. For accurate stroke prediction, the use of machine learning methods holds significant importance. In this regard, this study adopts a machine-learning approach for batsmen’s stroke prediction. This research makes several significant contributions to the field: • The study collects a comprehensive video dataset to classify different cricket strokes. In contrast to previous studies that only use image datasets and cover a maximum of five strokes, this study covers eight strokes, including ‘flick’, ‘back foot punch’, ‘pull’, ‘cut’, ‘cover drive’, ‘straight drive’, ‘on drive’, and ‘sweep’. • A novel technique is employed to extract features from the video dataset. The MediaPipe library extracts seventeen critical points of the human body. Based on these key points, the batsman’s stroke is accurately classified. • The study uses fine-tuned machine learning and deep learning models to classify the strokes based on the extracted feature dataset. Cross-validation is employed to validate the model’s performance, ensuring accurate results. • This research provides a more comprehensive and accurate approach to classifying cricket strokes. The novel technique that extracts features from video datasets and utilizes state-of-the-art machine learning and deep learning models helps improve classification accuracy.
Sensors 2023,23, 6839 3 of 16 The organization of the study is as follows: Section 2examines the relevant literature studies on pose estimation and stroke recognition. Section 3analyzes the workflow of the proposed methodology. The video stroke dataset and the technique used for feature extraction are also described. Results and discussions are presented in Section 4, and Section 5concludes this study. 2. Related Work Machine learning models have witnessed a wide adoption in various fields like image processing [ 7 – 9 ], text analysis [ 10 , 11 ], education [ 12 , 13 ], medical data analysis [ 14 ], etc., and sports is no exception. As a result, several studies have been presented involving the use of machine learning techniques in sports [15–17]. Human pose estimation for predicting players’ performance in sports has been investigated recently, leading to several techniques and approaches in this field. A recent study [ 18 ] proposed a batsman shorts estimation model to identify four different strokes in cricket: glance, drive, block, and cut. The study utilized an image dataset of cricket strokes and extracted feature vectors from head, feet, bat, and hand positions to train several models, including a k-nearest neighbor, support vector machine, and convolutional neural network (CNN)/AlexNet. The AlexNet model achieved the highest accuracy of 74.33%. Along the same directions, ref. [ 19 ] extracted 15 critical data points from an image dataset of different cricket strokes using MediaPipe. The dataset was used to develop a mobile application to help batsmen improve their accuracy. The random forest (RF) model achieved an F1 score of 87%. In another study [ 20 ], a dataset of 63 different backward and forward cricket strokes was collected and classified using a long short-memory (LSTM) network and bidirectional LSTM models. Both models achieved 100% accuracy. The authors used motion vectors and three-dimensional (3D) match recognition to classify eight angles of cricket strokes with high precision in [21]. Action recognition using deep learning was also applied to other sports like badminton, table tennis, and high jump. For instance, in a recent study ([ 22 ]), the ResNet-18, VGG-16, and GoogleNet models were used to classify badminton smashes. The ResNet-18 achieved a high accuracy of 97.51% and 98.66% on training and testing, respectively. On the Jeston Nano hardware, the GoogleNet model outperformed, achieving 83.04% and 97.0% accuracy on training and testing, respectively. In one study ([ 23 ]), a new approach was utilized to collect data on the footwork of badminton players. This study used a deep-learning method to extract two-dimensional (2D) and 3D coordinates of the players’ shoes. The model achieved an absolute positioning accuracy of 74%. These data provide valuable insights into the players’ movements, which can help improve their performance on the court. Study [ 24 ] employed a novel technique to gather data for the classification of different strokes played in table tennis. The authors collected a video dataset of the primary 11 strokes of 14 professional table tennis players and utilized CNN and other machine learning models to classify the strokes. The CNN model achieved an impressive accuracy of 99.37%. Similarly, ref. [ 25 ] studied classifying different human actions using a custom CNN model. The authors created two datasets, the first consisting of 10 actions obtained using the Kinect v2 sensor and the second comprising seven subjects performing 20 other actions. The model achieved 97.23% accuracy on the Kinect dataset and 87.1% on the MRS dataset. A 13-layered conventional neural network called ‘short net’ is presented in [ 26 ] to classify six different strokes. These strokes include ‘cut shot’, ‘straight drive’, ‘cover drive’, ‘pull shot’, ‘leg glance shot’, and ‘scoop shot’. The model achieved good accuracy with a minimum entropy score. All the previous work on cricket stroke recognization is summarized in Table 1. Table 1shows the dataset used in the previous studies, the stroke they classified, and the outperforming model with the reported accuracy.
Sensors 2023,23, 6839 4 of 16 Table 1. Summary of the literature review on cricket stroke prediction. Refs. Dataset Strokes Technique Accuracy [18] Images Glance, drive, block, and cut AlexNet 74.33% [19] Images Cut, cover drive, straight drive, pull, leg glance, scoop Random forest 87% [20] Videos Backward and forward LSTM 100% [18] Videos Strokes and gameplay AlexNet 96.66 3. Proposed Methodology The workflow of the proposed approach is presented in Figure 1. The study collected videos of eight different types of batsman strokes. The videos were preprocessed to remove any noise present to ensure accurate analysis. The MediaPipe library was used to extract human key points from the preprocessed videos, and a novel dataset was created based on these features. The dataset was preprocessed again to eliminate any remaining noise, and the analysis focused on 17 critical points of human movement. Before implementing machine learning and deep learning models on the dataset, it was split into test and train sets. The research dataset was used to train and test the models, and a performance evaluation was conducted to assess their effectiveness in real-time. Figure 1. Workflow of adopted methodology. 3.1. Video Cricket Strokes Dataset This study aims to create a comprehensive dataset of cricket stroke videos by collecting a diverse range of videos from various platforms. To ensure the dataset’s generalizability, the videos were collected from both YouTube channels and the Liaquat Pur cricket club, Pakistan, focusing on eight primary strokes: ‘pull’, ‘cut’, ‘cover drive’, ‘straight drive’, ‘backfoot punch’, ‘on drive’, ‘flick’, and ‘sweep’. Multiple videos of each stroke were collected to provide a diverse range of examples for analysis. The count plot in Figure 2 visually represents the number of videos collected for each stroke. The x -axis displays the number of videos, and the y -axis displays the type of strokes. This information provides an overview of the distribution of videos in the dataset.
Sensors 2023,23, 6839 5 of 16 Figure 2. Number of videos for different strokes. To ensure accurate analysis, the recorded video data were preprocessed, and any noise present was manually removed. Each stroke video has a length of approximately 1.5 to 2 s, providing a consistent length for analysis. A few sample frames from the recorded videos are shown in Figure 3, demonstrating the video quality. The resulting dataset provides a valuable resource for researchers to analyze and compare different cricket stroke techniques. The diverse range of videos ensures that the dataset is comprehensive and can be used to study the nuances of each stroke. Figure 3. Sample frames from different videos.
Sensors 2023,23, 6839 6 of 16 3.2. Feature Extraction from Videos Once the cricket stroke videos are preprocessed and the noise removed, the MediaPipe library extracts features from the videos. This library is a pre-built set of components that can be used to create complex machine-learning models for tasks such as pose estimation, facial recognition, hand tracking, and object detection. It can extract 33 landmarks from the human body pose estimation. The pose landmarks P can be used to represent the pose of a person in various ways. One common representation is the skeletal representation, where the pose landmarks are connected by lines to form a skeletal structure representing the person’s body. The skeletal representation can be represented as follows: S={li}(1) where li= (pi,pj),i,j∈ {1, 2, . . . , 17},i<j(2) The value S is the set of 16 lines that connect the 33 pose landmarks P to form the skeletal structure, and pi and pj are the two endpoints of the i th line. The pose estimation pipeline can be summarized as follows: I→f→P→S(3) where I is the input frame from the video, f is the deep neural network that performs the pose landmark estimation, P is the set of 33 pose landmarks, and S is the skeletal representation of the pose. For this study, only 17 landmarks were selected, as they are critical to detecting strokes, namely the nose, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle, right heel, left heel, left foot index, and right foot index. The MediaPipe library extracts the 17 landmark points and their x,y, and zcoordinate values from every video frame. The MediaPipe library also provides a visibility value that can be set to extract features from the videos. This study’s visibility value was set to extract 17 landmarks only. The OpenCV library passes every video to the MediaPipe library to extract the landmark points. Extracting these landmarks creates a new data frame containing 51 feature columns and one label column named as cricket stroke dataset. The working of the proposed approach is shown in Algorithm 1. Algorithm 1 Batsmen stroke prediction. Input: Video strokes dataset (VSD) Output: Stroke prediction {cover driver, pull, sweep, state drive, on drive, cut and back foot punch} 1: MPF←− MediaPipe(VSD) // VSD ∈ Video strokes dataset, MPF ∈ extracted features from the MediaPipe library. 2: TRF ←− RFtraining(MPFTe) // MPFTe∈MPF , here MPFTe is the training data of MPF. 3: RFPred ←− TRF(MPFTs) // MPFTs∈MP F, here MPFTs is the testing data of MPF, RFPred ∈{cover drive, pull, sweep, state drive, on drive, cut and back foot punch} 3.3. Cricket Stroke Exploratory Data Analysis This section deeply explores the cricket stroke dataset after extracting features from the videos. The new dataset contains 51 feature columns. The 51 feature columns correspond to the x , y , and z coordinates of each of the 17 selected landmark points. The label column contains the name of the stroke performed in the video. The column names for the features and labels are shown in Table 2.
Sensors 2023,23, 6839 7 of 16 Table 2. Features in the dataset. Attribute Dtype Attribute Dtype Attribute Dtype nosex float lshoulderx float64 nosey float64 lshouldery float64 nosez float64 lshoulderz float64 rshoulderx float64 lelbowx float64 rshouldery float64 lelbowy float64 rshoulderz float64 lelbowz float64 relbowx float64 rWristx float64 relbowy float64 rWristy float64 relbowz float64 rWristz float64 lWristx float64 rhipx float64 lWristy float64 rhipy float64 lWristz float64 rhipz float64 lhipx float64 rkneex float64 lhipy float64 rkneey float64 lhipz float64 rkneez float64 lkneex float64 rankelx float64 lkneey float64 rankely float64 lkneez float64 rankelz float64 lankelx float64 rheelx float64 lankely float64 rheely float64 lankelz float64 rheelz float64 lheelx float64 lfindexx float64 lheely float64 lfindexy float64 lheelz float64 lfindexz float64 rfindexx float64 rfindexy float64 rfindexz float64 The cricket strokes dataset (CSD) is a collection of numeric features extracted from videos using the MediaPipe library, resulting in 8998 records. However, the final dataset is not balanced, with different strokes having varying numbers of instances. Specifically, the dataset includes 1060 records for ‘straight drive’, 2276 instances for ‘on drive’, 1236 records for ‘cover drive’, 1011 rows for ‘cut’, 779 records for ‘pull’, 511 records for ‘sweep’, 908 instances for ‘flick’, and 1217 records for ‘back foot punch’. A summary of the dataset is presented in Table 3. The highest percentage belongs to ‘on drive’ with 25.29% instances, whereas ‘sweep’ has the lowest ratio to the total records at 5.68% of the total records. Table 3. Number of instances of every stroke. Strokes Records Percentage State Drive 1060 11.78% On Drive 2276 25.29% Cover Drive 1236 13.74% Cut 1011 11.24% Pull 779 8.66% Sweep 511 5.68% Flick 908 10.09% Backfoot Punch 1217 13.53% The cricket strokes dataset is analyzed in three-dimensional space. A Python library HyperTools created a cubic scatter plot. HyperTools uses dimensionality reduction to visualize high-dimensional data in a lower-dimensional space using the t-distributed
Sensors 2023,23, 6839 8 of 16 stochastic neighbor embedding (t-SNE) technique. The features extracted from the video are more detachable, and the machine learning model can easily classify these features, as shown in Figure 4. Figure 4. Feature space analysis. The pair plot is plotted on the dataset to check the correlation between different features. We extract the five most important features from the dataset using principal component analysis to plot the pair plot on these features. The pair plot shows that these points are more easily detachable, as shown in Figure 5. 3.4. Target Label Encoding Label encoding is a common technique used in machine learning to convert categorical variables into numerical representations. Label encoding is necessary because many machine learning algorithms require input data to be in numerical format. Label encoding assigns a unique numerical value to each category within a variable, allowing the algorithm to identify patterns and relationships within the data. The label column is encoded in this study, and every class is assigned a different serial number from 0 to 7. 3.5. Dataset Splitting Dataset splitting is a technique used in machine learning to partition a dataset into two subsets: a training dataset, and a testing dataset. The purpose of this is to assess the performance of a machine learning model on unseen data, which can help to identify whether the model is overfitting, underfitting, or generalized. This study splits the dataset into three different ratios, 70:30, 80:20, and 90:10, and gets the accuracy on all splits. On the 80:20 data split, the models give high accuracy.
Sensors 2023,23, 6839 9 of 16 0.5 0.0 0.5 1.0 1 0.5 0.0 0.5 1.0 1.5 2 0.5 0.0 0.5 1.0 3 0.75 0.50 0.25 0.00 0.25 0.50 0.75 4 1.0 0.5 0.0 0.5 1.0 1 1.0 0.5 0.0 0.5 1.0 5 1 0 1 2 2 1 0 1 3 1.0 0.5 0.0 0.5 1.0 4 1 0 1 5 Label Pull onDrive state drive Cover Drive back foot punch Flick Sweep Cut Figure 5. Pair plot on extracted features. 3.6. Model Training Various machine learning and deep learning models were applied to the cricket strokes dataset to classify the batsmen’s strokes. The machine learning models used in the study included LSTM, k-nearest neighbor (KNN), logistic regression (LR), decision tree (DT), support vector machine (SVM), and RF. A hyperparameter tuning process was applied to these models to obtain optimal results. The specific parameters used for the machine learning models are outlined in Table 4. 3.7. Performance Metrics Several performance matrices are used in this study to evaluate the performance of machine learning algorithms. Precision, recall, and F1 score are three standard evaluation metrics in machine learning classification tasks. With these, standard evaluation matrix geometric mean, Cohen’s kappa, and log loss are also measured to evaluate the performance of machine learning models.
Sensors 2023,23, 6839 16 of 16 18. Moodley, T.; van der Haar, D. Cricket stroke recognition using computer vision methods. In Information Science and Applications: ICISA 2019; Springer: Berlin/Heidelberg, Germany, 2019; pp. 171–181. 19. Devanandan, M.; Rasaratnam, V.; Anbalagan, M.K.; Asokan, N.; Panchendrarajan, R.; Tharmaseelan, J. Cricket Shot Image Classification Using Random Forest. In Proceedings of the 2021 3rd International Conference on Advancements in Computing (ICAC), Colombo, Sri Lanka, 9–11 December 2021; pp. 425–430. 20. Bandara, I.; Baˇci´c, B. Strokes classification in cricket batting videos. In Proceedings of the 2020 5th International Conference on Innovative Technologies in Intelligent Systems and Industrial Applications (CITISIA), Sydney, Australia, 25–27 November 2020; pp. 1–6. 21. Karmaker, D.; Chowdhury, A.; Miah, M.; Imran, M.; Rahman, M. Cricket shot classification using motion vector. In Proceedings of the 2015 Second International Conference on Computing Technology and Information Management (ICCTIM), Johor, Malaysia, 21–23 April 2015; pp. 125–129. 22. Yip, Z.Y.; Khairuddin, I.M.; Isa, W.H.M.; Majeed, A.P.A.; Abdullah, M.A.; Razman, M.A.M. Badminton Smashing Recognition through Video Performance by using Deep Learning. MEKATRONIKA 2022,4, 70–79. [CrossRef] 23. Luo, J.; Hu, Y.; Davids, K.; Zhang, D.; Gouin, C.; Li, X.; Xu, X. Vision-based movement recognition reveals badminton player footwork using deep learning and binocular positioning. Heliyon 2022,8, e10089. [CrossRef] 24. Kulkarni, K.M.; Shenoy, S. Table tennis stroke recognition using two-dimensional human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 25 June 2021; pp. 4576–4584. 25. Ahmad, Z.; Illanko, K.; Khan, N.; Androutsos, D. Human action recognition using convolutional neural network and depth sensor data. In Proceedings of the 2019 International Conference on Information Technology and Computer Communications, Paris, France, 29 April–2 May 2019; pp. 1–5. 26. Lazarescu, M.; Venkatesh, S.; West, G. Classifying and learning cricket shots using camera motion. In Proceedings of the Advanced Topics in Artificial Intelligence: 12th Australian Joint Conference on Artificial Intelligence, AI’99, Sydney, Australia, 6–10 December 1999; Proceedings 12; Springer: Berlin/Heidelberg, Germany, 1999; pp. 13–23. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.