Markerless 2D kinematic analysis of underwater running : A deep learning approach
Full text
This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY-NC-ND 4.0 https://creativecommons.org/licenses/by-nc-nd/4.0/ Markerless 2D kinematic analysis of underwater running : A deep learning approach © 2019 Elsevier Ltd. Accepted version (Final draft) Cronin, Neil; Rantalainen, Timo; Ahtiainen, Juha; Hynynen, Esa; Waller, Benjamin Cronin, N., Rantalainen, T., Ahtiainen, J., Hynynen, E., & Waller, B. (2019). Markerless 2D kinematic analysis of underwater running : A deep learning approach. Journal of Biomechanics, 87, 75-82. https://doi.org/10.1016/j.jbiomech.2019.02.021 2019
1 Markerless 2D Kinematic Analysis of Underwater Running: A Deep Learning Approach1 Neil Cronin1, Timo Rantalainen1, Juha Ahtiainen1, Esa Hynynen2, Ben Waller1,3 2 1 Faculty of Sport and Health Sciences, University of Jyväskylä, Finland3 2 KIHU- Research Institute for Olympic Sports, Jyväskylä, Finland4 3Physical Activity, Physical Education, Sport and Health Research Centre (PAPESH), Sports5 Science Department, School of Science and Engineering, Reykjavik University, Reykjavik, Iceland6 7 Original Article8 9 Correspondence:10 Neil Cronin11 University of Jyvaskyla, Neuromuscular Research Center, Faculty of Sport and Health Sciences, P.12 O. Box 35, FI-40014, University of Jyvaskyla, Finland13 tel: +358 40 805 373514 e-mail: [email protected] 16 Keywords: deepwater running, kinematics, deep learning, artificial intelligence, motion analysis17 18 Word count: 233919 20 21 22 23 24 25
2 Abstract26 27 Kinematic analysis is often performed with a camera system combined with reflective markers28 placed over bony landmarks. This method is restrictive (and often expensive), and limits the ability29 to perform analyses outside of the lab. In the present study, we used a markerless deep learning-30 based method to perform 2D kinematic analysis of deepwater running, a task that poses several31 challenges to image processing methods. A single GoPro camera recorded sagittal plane lower limb32 motion. A deep neural network was trained using data from 17 individuals, and then used to predict33 the locations of markers that approximated joint centres. We found that 300-400 labelled images34 were sufficient to train the network to be able to position joint markers with an accuracy similar to35 that of a human labeler (mean difference <3 pixels, around 1cm). This level of accuracy is sufficient36 for many 2D applications, such as sports biomechanics, coaching/training, and rehabilitation. The37 method was sensitive enough to differentiate between closely-spaced running cadences (45-8538 strides per minute in increments of 5). We also found high test-retest reliability of mean stride data,39 with between-session correlation coefficients of 0.90-0.97. Our approach represents a low-cost,40 adaptable solution for kinematic analysis, and could easily be modified for use in other movements41 and settings. Using additional cameras, this approach could also be used to perform 3D analyses.42 The method presented here may have broad applications in different fields, for example by enabling43 markerless motion analysis to be performed during rehabilitation, training or even competition44 environments.45 46 Introduction47 48 Kinematic analysis is used to characterise changes in joint angles during human movement. This49 information can be combined with other sources, e.g. force data, to build a more complete picture of50
3 how a movement is performed (Winter, 1991), and thus has important implications for various51 fields such as sports biomechanics, injury risk assessment and rehabilitation (see Colyer et al. 201852 for a review). Kinematic analysis is often performed with a camera system combined with a set of53 reflective markers placed over bony landmarks, allowing a digital model of the moving person to be54 reconstructed (van der Kruk and Reijne, 2018). However, the use of reflective markers can restrict55 the settings in which data can realistically be collected, and many existing camera-based methods56 still rely on expensive hardware and software. Moreover, in an aquatic environment, the use of57 markers is impractical because they impede normal movement and are prone to significant motion58 artifact.59 60 Recently, several attempts have been made to develop markerless methods, which in theory could61 be used outside of the laboratory and allow movement to be analysed in more natural, unconstrained62 conditions (see Drory, Li, and Hartley 2017 for a comprehensive overview). In particular, methods63 that rely on artificial intelligence have demonstrated promising results (see Colyer et al., 2018 for64 review), and have the potential to revolutionise the way movement analysis is performed due to65 their powerful ability to ‘learn’ patterns in data. In the present study, we used DeepLabCut66 (Insafutdinov et al., 2016; Mathis et al., 2018; Pishchulin et al., 2015) to track the locations of67 (approximated) lower limb joint centres and used this information to perform 2D kinematic analysis68 of deepwater running, a task that poses several challenges to image processing methods, such as69 poor contrast and changes in light intensity. DeepLabCut is an open-source method that combines a70 residual neural network (ResNet-50) pretrained on ImageNet with deep convolutional and71 deconvolutional neural network layers (Insafutdinov et al., 2016) to predict the ‘learned’ locations72 of individual points in an image using feature detectors (He et al. 2015). The network ‘learns’73 marker locations by being trained on labeled data, which consists of individual images accompanied74 by a human-defined label of the ‘correct’ marker location. During training, the weights are adjusted75
4 iteratively so that for each image, the network assigns high probabilities to target marker locations76 and low probabilities to all other regions. Training thus allows the network to ‘learn’ feature77 detectors for each user-defined marker, rather than relying on hard-coded, pre-defined features.78 79 In this study we demonstrate that a modified version of the DeepLabCut method can be used for80 accurate 2D kinematic analysis of deepwater running filmed using a single GoPro camera. We used81 this method to determine lower limb segment lengths and joint angles, and we present various other82 parameters that could be useful in motion analysis applications.83 84 Methods85 86 Participants. A total of 21 individuals (age: 24±4 years, height: 177±10cm, mass 67±9; 13 males87 and 8 females) volunteered to participate and provided written informed consent. The study was88 approved by the University’s ethics committee, and testing was conducted in accordance with the89 most recent Helsinki declaration.90 91 Experimental protocol. Participants performed bouts of deepwater running whilst immersed to92 shoulder level, and were tethered to the edge of the pool by a non-elastic cable attached to a93 buoyancy aid (Aquawallgym©, Hungary). A single GoPro camera (Hero 3 model) was enclosed in94 a waterproof case and positioned underwater in the sagittal plane to the participants’ left side at a95 distance of approximately 5 m. A custom-made calibration frame (2m x 2m) was used to calibrate96 the field of view for each participant and test. The camera was then set to record at 60Hz whilst97 participants ‘ran’ at different cadences controlled by a metronome (increased by 5 strides per98 minutes (spm) from 45 to 90 spm). A subset of participants were tested a second time99 approximately 1 week after the first test, to enable test-retest comparisons to be performed. A deep100
5 neural network was trained and then used to predict the locations of several markers that101 approximated joint centres. Predicted joint coordinates were used to determine lower limb segment102 lengths and joint angles from the left leg, which was closest to the camera.103 104 Deep Neural Network. The method used here largely followed the method described by Mathis et105 al. (Mathis et al., 2018; v1). We first trained the network using 500 images from 17 randomly106 chosen participants (i.e. 28-30 images per participant), leaving aside data from the remaining 4107 participants (see below). The training images were randomly selected using a custom-written script108 in Matlab (Mathworks, v2016b). These images were cropped (dimensions: 580 x 480 pixels) and109 then manually labelled, with markers placed on the lateral side of the trunk (approximately mid-way110 between the shoulder and hip), greater trochanter, lateral femoral condyle, lateral malleolus, and 5th 111 meta-tarsal head. The labelled images were used to train a deep neural network with a 90% training,112 10% test split. The ResNet model was initialised with weights trained on ImageNet (He et al.,113 2015), and the cross-entropy loss between the predicted score-map and the ground truth score-map114 was minimised using stochastic gradient descent (Insafutdinov et al., 2016). The network was115 trained for 200,000 iterations using a single Tesla K80 GPU via Microsoft Azure’s cloud platform116 running Python (Python Software Foundation; v.3.5) and Tensorflow (Abadi et al. 2018; v.1.2.1).117 The training process was repeated with smaller training sets (400, 300, 200 and 100 images118 respectively), to determine the minimum number of images required to reach satisfactory predictive119 performance for this task. The number of frames used for training was selected based on previous120 work using a similar method (Mathis et al., 2018), and for each trained model, frames were121 randomly assigned to the test or training set.122 123 Evaluation of deep neural network performance. To compare between joint coordinates labelled124 by a human and those labelled by the network, pairwise Euclidean distances were computed for125
6 each marker location (root mean square error: RMSE). In the results section the RMSE values are126 shown for individual joints or as the average across all joints, as appropriate. To quantify the127 evolution of the training error, and to enable training to be resumed later if needed, the Tensorflow128 weights were stored every 10,000 iterations. As noted above, data from 4 randomly chosen129 participants were excluded completely from the training set. After training of the neural networks130 was complete, videos from these 4 participants were evaluated by each neural network model,131 thereby serving as additional test data. This approach was chosen to enable out of sample132 predictions that were completely independent of the training process, thus giving some indication of133 the generalisability of our trained models.134 135 Determining joint angles and segment lengths. Segment lengths were initially computed in terms136 of pixels, using the coordinate data of each point exported during the analysis of each video.137 Segment lengths were calculated based on the distance formula: =(− )+ (− ),138 where d = segment length in pixels, and x and y values denote the coordinates of the two points that139 make up a segment. A scaling factor was calculated for each participant and trial based on the140 corresponding calibration frame video, and used to scale segment lengths. Joint angles at the hip,141 knee and ankle were determined using the atan2 function in Matlab. In several cases, the neural142 network did not attempt to place a marker because the target joint location was blocked by the hand143 or moved beyond the image field of view. To overcome the effect of missed (and misplaced)144 markers on the resulting kinematic and segment length data, raw data were first filtered with a145 median filter (10-20 data points generally yielded good results) followed by a Butterworth 4th order146 low-pass filter (Figure 3). In some cases, e.g. when a marker was missing for several consecutive147 frames, it was necessary to experiment with different filtering procedures.148 149 Results150
7 151 Deep Neural Network Performance152 153 Using the full training set of 500 labelled images, the mean training error across all images was 1.4154 pixels. The mean test error was 2.92 pixels (approximately 1cm). This model represents the best155 performance achieved out of all of the tested models. As seen in Figure 1A, training performance156 was similar between all of the tested models after 200,000 iterations. Test performance, i.e. how157 well the network predicts marker locations on images it has not ‘seen’ during training, was similar158 between models trained on 300-500 images, suggesting that 300 training images was sufficient for159 this task. However, test performance clearly decreased with training datasets of 100-200 images,160 indicating overfitting of these models during the training stage. For all models, training time varied161 between approximately 9-12 hours.162 163 *** FIGURE 1 HERE ***164 165 For the better performing models, both training and test errors were largely independent of which166 marker was being tracked, whereas for the poorer performing models, the disparity between167 different markers was much larger (Figure 1B and C). It should be noted that the trunk marker was168 not placed by the network in around 20% of frames due to the hand blocking the target area. In169 some images, the 5th metatarsal marker also was not placed because the foot moved beyond the170 camera field of view. However, these issues did not substantially affect kinematic tracking (Figures171 3 and 4).172 173 *** FIGURE 2 HERE ***174 175
8 Figure 2 shows some examples of the same images labelled by the 100 and 500 models. In some176 cases, the models make similar predictions, compared to each other and to a human labeller (e.g.177 middle image in Figure 2). In other images, the 100 model consistently makes larger errors, and in178 the most extreme cases, identifies an ostensibly correct location but on the wrong limb (left image179 in Figure 2), which largely accounts for the bigger test errors of the poorer performing models (for180 further model comparisons see Supplementary Video 1). Segment length calculations yielded181 consistent traces across consecutive stride cycles, particularly for models trained on more images.182 Segment lengths varied somewhat throughout a stride (see Figure S1), due to minor fluctuations in183 marker locations, as well as the inevitable 3D rotation of the lower limb that cannot be quantified184 with this method.185 186 Kinematics of deepwater running187 188 Using this method we obtained consistent joint angle traces over several consecutive stride cycles.189 Figure 3 shows examples of data computed from three different 10s videos obtained from different190 individuals whose data were not seen by the neural network during training. These videos were191 processed entirely by a trained neural network, and did not require a human labeler at any stage (see192 also Supplementary Videos 2-4).193 194 *** FIGURE 3 HERE ***195 196 Based on visual identification from the videos, it was possible to approximate the start of individual197 stride cycles. Figure 4 shows the results of this segmentation for a single 20s trial from a participant198 whose data were not seen by the neural network during training.199 200 *** FIGURE 4 HERE ***201