scieee AI-readable full text Open interactive document viewer

CGP Box: An effective direction representation strategy for oriented object detection in remote sensing images

Guan, Qiuyu,Qu, Zhenshen,Zeng, Ming,Shen, Jianxiong,Du, Jingda

Abstract

In recent years, the emergence of convolutional neural networks (CNN) has greatly promoted the development of the object detection field, and many CNN-based detectors have achieved excellent performance on object detection in remote sensing images. To accurately locate the target, oriented bounding box (OBB) is usually used in remote sensing objects, such as the angle-based OBB, to represent the target. Nevertheless, the critical loss instability caused by the periodicity of the angle is always difficult to solve. In this paper, we propose a novel strategy called the Center-Guide points (CGP) box method that uses the guide points to locate the target, which breaks the limit of the angle-based thinking pattern to solve the critical loss instability problem. To be specific, we define a new guide-points selection rule and prediction structure, which replaces the traditional method of using angle values to indicate the direction. Furthermore, we propose the matching method of centre points and guide points, which is a box decoding method that matches the object and the corresponding guide points. Finally, an attention learning module called the Gaussian Center-Line (GC-L) Attention module based on the Gaussian centre-line is proposed to improve the accuracy of guide points. These strategies are applied to the key point detection framework and tested on three classical-oriented object remote sensing datasets. The results show that our method is effective and competitive.

Full text

Full Terms & Conditions of access and use can be found at https://www.tandfonline.com/action/journalInformation?journalCode=tres20 International Journal of Remote Sensing ISSN: (Print) (Online) Journal homepage: https://www.tandfonline.com/loi/tres20 CGP Box: An effective direction representation strategy for oriented object detection in remote sensing images Qiuyu Guan, Zhenshen Qu, Ming Zeng, Jianxiong Shen & Jingda Du To cite this article: Qiuyu Guan, Zhenshen Qu, Ming Zeng, Jianxiong Shen & Jingda Du (2021) CGP Box: An effective direction representation strategy for oriented object detection in remote sensing images, International Journal of Remote Sensing, 42:17, 6666-6687, DOI: 10.1080/01431161.2021.1941389 To link to this article: https://doi.org/10.1080/01431161.2021.1941389 © 2021 The Author(s). Published by Informa UK Limited, trading as Taylor & Francis Group. Published online: 27 Jul 2021. Submit your article to this journal Article views: 492 View related articles View Crossmark data CGP Box: An effective direction representation strategy for oriented object detection in remote sensing images Qiuyu Guan a , Zhenshen Qu a , Ming Zeng a , Jianxiong Shen b and Jingda Du a a Space Control and Inertial Technology Research Center, Harbin Institute of Technology, Harbin, China; b Institut de Robòtica i Informàtica Industrial, CSIC-UPC, Barcelona, Spain ABSTRACT In recent years, the emergence of convolutional neural networks (CNN) has greatly promoted the development of the object detection field, and many CNN-based detectors have achieved excellent performance on object detection in remote sensing images. To accurately locate the target, oriented bounding box (OBB) is usually used in remote sensing objects, such as the angle-based OBB, to represent the target. Nevertheless, the critical loss instability caused by the periodicity of the angle is always difficult to solve. In this paper, we propose a novel strategy called the Center-Guide points (CGP) box method that uses the guide points to locate the target, which breaks the limit of the angle-based thinking pattern to solve the critical loss instability problem. To be specific, we define a new guide-points selection rule and prediction structure, which replaces the traditional method of using angle values to indicate the direction. Furthermore, we propose the matching method of centre points and guide points, which is a box decoding method that matches the object and the corresponding guide points. Finally, an attention learning module called the Gaussian Center-Line (GCL) Attention module based on the Gaussian centre-line is proposed to improve the accuracy of guide points. These strategies are applied to the key point detection framework and tested on three classical-oriented object remote sensing datasets. The results show that our method is effective and competitive. ARTICLE HISTORY Received 31 January 2021 Accepted 5 June 2021 1. Introduction Object detection is a classical branch in the field of computer vision (Lin et al. 2017a; Ren et al. 2015; Dai et al. 2016; Liu et al. 2016a; Zhou, Wang, and Philipp 2019). Its goal is to analyse the input image and output the location information and classification information of the specified object category. In recent years, due to the wide application of deep learning technology, this field has made significant progress. However, compared with natural scenes, object detection in remote sensing images is still a challenging work. Objects in remote sensing images tend to be variable in orientation, have a large aspect ratio, and sometimes are closely arranged (vehicles in the parking lot, ships in the harbour). Therefore, using Horizontal Bounding Box (HBB) as labels of objects has great disadvantages. As shown in Figure 1(a), a HBB that accurately surrounds the object will CONTACT Zhenshen Qu [email protected] Harbin Institute of Technology, Harbin, China INTERNATIONAL JOURNAL OF REMOTE SENSING 2021, VOL. 42, NO. 17, 6666–6687 https://doi.org/10.1080/01431161.2021.1941389 © 2021 The Author(s). Published by Informa UK Limited, trading as Taylor & Francis Group. This is an Open Access article distributed under the terms of the Creative Commons Attribution-NonCommercial-NoDerivatives License (http://creativecommons.org/licenses/by-nc-nd/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited, and is not altered, transformed, or built upon in any way. cover multiple tightly arranged ships. Such object positioning ambiguity will cause the network to be unable to accurately distinguish the boundaries of adjacent objects in the learning process. Usually, the Intersection over Union (IoU) is used in the field of object detection to evaluate the degree of overlap of two bounding boxes, that is, the ratio of the area of intersection to the area of union. Moreover, the detection boxes with large IoU in the same category will be suppressed by the Non-Maximum Suppression (Neubeck and Luc 2006) algorithm, and only the HBB with the highest score will be left in the end. Oriented Bounding Box (OBB) (Figure 1(b)) eliminates the above defects and is more in line with the requirements of remote sensing object detection. Therefore, accurate prediction of OBB direction is an important part of this field. A branch of (Cheng, Zhou, and Han 2016; Deng et al. 2017; Long et al. 2017) directly predicts the angle value to represent OBB, which is usually expressed as ðx;y;w;h;θÞ, we call it θ-based OBB, where ðx;yÞis the centre point coordinates of the bounding box, w and h are its width and height, respectively, θ is the inclination angle. This type of detector performs extremely well in OBB detection and achieves top-notch results. However, the angle value depicted by the radian system will cause ambiguity at the critical point of the period due to the periodicity, resulting in severe fluctuations of loss value and affecting convergence. For example, as shown in Figure 2, if the angle period is ½ π=2;π=2�, the loss of the two prediction boxes (Predicted Box 1 and 2) near the critical point (π=2) of the period is very different from the truth boxes ((Δθ1;Δθ2) in the figure). Therefore, in order to eliminate the ambiguity of object box representation, we propose a novel OBB representation method called Center-Guide points Box (CGP Box). Unlike other OBB representations, CGP Box uses guide points to determine the direction of the object, which avoids periodic ambiguity and can converge to the optimal solution more stably. Inspired by the general object detector, most of the work(Xia et al. 2018; Lin et al. 2017; Ding et al. 2019) in the field of remote sensing object detection is based on the R-CNN framework, which is a kind of classic anchor-based network. The anchor-based detector uses the anchor to obtain the region proposal, and then makes the secondary correction on the preselected region to obtain accurate results. In general, the acquisition of anchor requires sufficient prior knowledge. Many works have proposed oriented anchor to apply Figure 1. Examples of different label method in HRSC2016 dataset. INTERNATIONAL JOURNAL OF REMOTE SENSING 6667 R-CNN framework to the field of oriented object detection in remote sensing images. However, the introduction of direction information also leads to a larger number of required anchors and more prior knowledge, which ultimately makes network prediction inefficient. Besides, when the network predicts multiple corner points, the output is usually one-dimensional, which will lead to severe fluctuations of loss. For example, the network outputs three bounding boxes ðE1;E2;E3;E4Þ, ðE1;E3;E2;E4Þ, and ðE4;E3;E2;E1Þ, where Ei¼ ðxi;yiÞji¼1;2;3;4f gis the coordinate of the corner points. Although they all refer to the same bounding box, the calculated losses are completely different due to different orders. Therefore, we adopt the anchor-free structure as the basic framework of CGP Box, which is a full convolutional network structure with multiple probabilistic heatmaps as the output. Due to the uniqueness of the guide points and centre points in the heatmap, loss fluctuation is avoided. In this paper, we proposed an anchor-free network that integrates the Center-Guide points box regression strategy (mentioned in Section 3.2). Besides, to more accurately locate the position of the guide points, we also proposed a supervised central attention learning module called the Gaussian Center-Line (GC-L) Attention Module (mentioned in Section 3.4). Due to the lack of computing resources, we chose to verify the validity of each part on the HRSC2016 (Liu et al. 2016b) dataset and finally compared the test set results on the DOTA (Xia et al. 2018) dataset and UCAS-AOD (Zhu et al. 2015) dataset. In summary, the contributions of this paper are as follows: Figure 2. Due to the periodicity of angle, the loss values calculated by the prediction box with the same estimation error are greatly different. 6668 Q. GUAN ET AL. (1) We proposed an OBB representation strategy based on Center-Guide points called CGP Box and merged it with the anchor-free network to solve the hidden ambiguities in other representation methods. Experiments on HRSC2016 proved that this method is superior to the θ-based OBB. (2) We designed the Gaussian Center-Line Attention learning module to further improve the positioning accuracy of the guide points. (3) We designed the decoding method of CGP Box and verified the experimental results under different parameters. We tested our approach on three public datasets (DOTA, HRSC2016, and UCAS-AOD) and compared them with other the Stateof-the-Art (SOTA) algorithms. 2. Related work Although remote sensing object detection has its particularity, it still belongs to a branch of general object detection. Therefore, at present, most remote sensing object detectors are based on general detectors, such as R-CNN series (Girshick et al. 2014; Girshick 2015; Ren et al. 2015), SSD (Liu et al. 2016a), YOLO series (Van Etten 2018; Redmon et al. 2016; Redmon and Farhadi 2017), Retinanet (Lin et al. 2017b), FCN (Long, Shelhamer, and Darrell 2015), etc 2.1. Angle-based Oriented Object Detectors The method of learning the angle value of the object was called the angle-based OBB detector. RRPN (Ma et al. 2018) was the first to propose R-RPN structure based on the rotated anchor, which enabled the network to learn direction information. RRCNN (Lin et al. 2017) proposed RRoI pooling to obtain the features of directed regions based on the R-RPN structure and developed the Non-Maximum Suppression (Neubeck and Luc 2006) algorithm to improve the suppression rate of overlapping targets. RDFPN (Yang et al. 2018) employed the feature pyramid network structure to improve the detection performance on ship targets. RoI-Transformer (Ding et al. 2019) designed RoI Learner to convert horizontal anchor to rotated anchor, effectively reducing the computing pressure brought by the rotated anchor. SCRDet (Yang et al. 2019b) used supervised attention modules to improve detection performance by fusing channel information. R3Det (Yang et al. 2019a) and RSDet (Qian et al. 2019) put forward their solutions to the unstable problem caused by angle value regression. 2.2. Point-based Oriented Object Detectors The method of finding four-point coordinates of OBB was called the point-based detector. Based on Faster-RCNN, R2CNN (Jiang et al. 2017) located the target by detecting two corner points and the width of the bounding box. Besides, the strategy of predicting OBB corner offset based on HBB was also a classic method, such as TextBoxes++(Liao, Shi, and Bai 2018) based on single-stage SSD detector, FR-O (Xia et al. 2018) and ICN (Azimi et al. 2018) based on RCNN framework. O2-DNET (Wei et al. 2019) treated the oriented object as pairs of middle lines and predicted the intersection point, which was an emerging anchorINTERNATIONAL JOURNAL OF REMOTE SENSING 6669 free keypoint method, such as CornerNet (Law and Deng 2018). Our method was also implemented and tested under this framework. 2.3. Segmentation-based Oriented Object Detectors The method to obtain the OBB based on the instance segmentation image of the target was called the segmentation-based detector. There was not a lot of work on this field, Mou and Zhu (2018) proposed a semantic bounding-aware unified multitask ResFCN to handle vehicle instance segmentation in the aerial images. STN (Jaderberg et al. 2015) and SENET (Hu, Shen, and Sun 2018) introduced spatial and channel attention mechanism to extract key feature respectively. Wang et al. (2019) proposed Mask OBB and used Inception Lateral Connection Network to enhance FPN. Our Gaussian Center-Line Attention module was similar to these methods but very different. 3. Proposed method 3.1. Framework The basic framework of the method used in this paper is shown in Figure 3, which is an anchor-free key point detection network based on the CenterNet (Zhou, Wang, and Philipp 2019) backbone. We choose the subsampling part of ResNet (He et al. 2016) as the feature extractor of the network. When the size of the feature map is 1 32 times the input size, the transpose convolution with deformable convolution (Dai et al. 2017) module is used to carry out the upsampling of the feature map, and finally, the feature map is 1 4 times the input size. GC-L Attention Module is a supervised Gaussian Center-Line feature map learning module. The output and input feature maps of the module are combined by element-wise addition to highlight the central feature of the object. The regression branch is composed of a CenterPoint Map, a GuidePoint Map, a Radius Map, Figure 3. The basic framework of the Center-Guide points OBB regression network. 6670 Q. GUAN ET AL. a RatioMap, and an Offset Map. Channels in the figure represent the number of channels in the heatmap. The details of each part are explained one by one in the following sections. 3.2. Centre-Guide points OBB regression strategy 3.2.1. The composition of Center-Guide points OBB The Center-Guide points OBB is a direction-representation strategy with no angle value, as shown in Figure 4(a). The blue dotted line box is the ground truth of the object, w and h are the long and short sides of the rectangle box, respectively. The red point Q is the centre point, and the direction guideline (orange dotted line) of the object is drawn through the red centre-point Q. Draw a circle with Q as the centre and r as the radius (dotted green circle), intersecting with the orange direction line at point G. The direction of the vector ~ QG is the direction of the object. As shown in Figure 3, the positions of Point Q and Point G are obtained from CenterPoint Map Pq and GuidePoint Map Pg. Let Figure 4(a) serves as the input image for the network in Figure 3, and let q;g2 R2 be the origin coordinates of point Figure 4. Example of the composition of Center-Guide points OBB regression strategy. The truth values of key points are composed by Gaussian kernel. INTERNATIONAL JOURNAL OF REMOTE SENSING 6671 Q and point G in the input image, respectively. After network down-sampling, the projection coordinates of Q and G on Pq and Pg in Figure 3 are q0and g0, which are expressed as: q0¼q T jk g0¼g T jk (1) where bc is the symbol of round down. Since the features on the images near the centre point or the guide point are very similar, setting the ground-truth value at only the corresponding position in the heatmap to 1 will cause foreground and background confusion and affect convergence, which is unreasonable. So we make the region near the point as ground truth, as shown in Figure 4(c), the label value at the coordinates of the centre and guide points on the heatmap is set to 1 (the point where the centre of the heatmap is 1), and the Gaussian kernel is used to generate non-zero negative samples around it (the non-1 area of the heatmap). This operation produces a locally high energy region near the positive sample as shown in Figure 4(c), and the rest region of the heatmap is set to 0. The Gaussian kernel is applied as follows: Pq xyc ¼exp  ðxqx0Þ2þðyqy0Þ 2σ ! Pg xyc ¼exp ð ðxgx0Þ2þðygy0Þ 2σ ! (2) Among them, σ is a hyperparameter that changes adaptively with the scale of the object, following CornerNet (Law and Deng 2018). The radius r and the ratio coefficient of the long and short sides α are obtained by regression of Radius Map Pr2 RW T�H T�1 and Ratio Map Pα2 ½0;1�W T�H T�1. The long and short sides of the object box ðw;hÞare expressed as: h¼2�r w¼α�h(3) The position of the centre point and direction point will incur quantization loss during downsampling. The loss can be expressed as Oq¼ jq0qjand Og¼ jg0gj, and the deviations of these two parts share a heatmap Offset Map Po2 RW T�H T�2 to correct. 3.2.2. Decoding of guide point OBB In the inference stage, we first get the centre point by extracting m local maximum point integer position coordinates Dq¼ fð~ xq i;~ yq iÞgm i¼1 on the heatmap Pq. The corresponding value of each centre point ð~ xq i;~ yq iÞon the three regression heatmaps of Pr, Pα and Po are ri, αi and δxq i. The long side wi and short side hi of bounding box can be solved by Eq (3) according to ri and αi. For the guide point, we take out 2 �m local maximum points coordinates Dg¼ fð~ xg j;~ yg jÞg2m j¼1 on the heatmap Pg, similar to the extraction method of the centre points. Depending on the position of the guide points, the corresponding offset δxg j will also be obtained on the heatmap Pg. The centre point and guide point obtained 6672 Q. GUAN ET AL. here are integer coordinates that need to be corrected with the offset coordinates. The corrected coordinate values can be calculated according to the following formula: ðxi;yiÞ ¼ ð~ xiþδxi;~ yiþδyiÞ(4) Therefore, the corrected points ðxq i;yq iÞand ðxg j;yg jÞcan be obtained. As shown in Figure 5, each centre point generates an vaild area, and the guide points that fall within the area (vaild points) will match the centre point, and the points outside the area (invalid points) will be discarded, where η2 ð0;1Þis a super hyperparameter that controls the size of the vaild area. The best match guide point ðxg i;yg iÞis: Disqg ¼ j ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi ðxq ixg jÞ2þðyq iyg jÞ2 qrij;j2 ½0;2mÞ;i2 ½0;mÞ;i;j2N(5) ðxg i;yg iÞ ¼ arg minðDisqgÞ;if Disqg �η�ri None ;Otherwise �(6) Finally, the direction angle of the object θi is determined by the vector formed by the centre point and the guide point. θi¼arccosðyg iyq i xg ixq iÞ(7) where θi2 ½0;πÞ. According to the geometric relationship, the bounding box of the object is calculated by ðxq i;yq i;wi;hi;θiÞ. Figure 5. Matching method of centre point and guide point. INTERNATIONAL JOURNAL OF REMOTE SENSING 6673 10.2. Comparison with the State-of-the-Art methods We compared our method with the State-of-the-Art methods on three datasets. Our method had excellent performance on HRSC2016, and had achieved competitive results on other datasets. 11. Results on HRSC2016 As shown in Table 3, we conducted experiments based on ResNet18 and ResNet50 on the HRSC2016 test dataset. Compared with other algorithms, our method had the SOTA performance (90.46%mAP) only using ResNet18. For a fair comparison, we chose the same input size (800 �800) and backbone (ResNet50) as other methods and achieved 90.43%mAP. In the case of real-time demand, the combination of ResNet18 and (800 �800) could be used, which could achieve 90.2%mAP at 58FPS. For accuracy, a combination of ResNet50 and (1024 �1024) could be selected, which had the best detection performance (90.52%mAP) and competitive performance on inference speed (23FPS). The Visualization results on the test dataset are shown in Figure 11. Figure 10. Visualization of the results of different strategies and components on HRSC2016 test dataset. 6680 Q. GUAN ET AL. 12. Results on DOTA Due to the limitation of computing power, it was difficult for our equipment to support more complex backbone on our method for sufficient training and testing on the DOTA dataset. Therefore, we only tested the network performance which ResNet18 as the backbone. Part of the results are shown in Figure 12 and Table 4. Under the condition of the ResNet18 backbone, our method not only achieved excellent performance in the categories of BD, LV, SP and RA, but also achieved competitive results on the whole (71.35%mAP). At the same time, we welcomed other researchers to use more complex backbone to test the performance of our detectors, such as ResNet50, ResNet101, ResNet152, etc. Figure 11. Visualization of detection results on HRSC2016 dataset. Table 3. Comparison with state-of-the-art methods on the HRSC2016 dataset. method Backbone Input Size mAP FPS R2CNN (Jiang et al. 2017) VGG16 800 �800 73.07 2 RRPN (Ma et al. 2018) ResNet101 800 �800 79.08 3.5 R2PN(Zhang et al. 2018)VGG16 – 79.6 – RetinaNet-H(Yang et al. 2019a) ResNet101 800 �800 82.89 14 RRD(Liao et al. 2018) VGG16 384 �384 84.3 – RoITransformer(Ding et al. 2019) ResNet101 512 �800 86.2 6 Gliding Vertex(Chen, G.S. Xia, and X.Bai 2020) ResNet101 – 88.20 – RetinaNet-R(Yang et al. 2019a) ResNet101 800 �800 89.18 10 R3Det(Yang et al. 2019a)ResNet101 800 �800 89.26 12 R3Det-DCL(Yang et al. 2020)ResNet101 – 89.46 – PolarDet(Zhao et al. 2020) ResNet50 800 �800 90.13 32 Ours ResNet18 800 �800 90.20 58 ResNet50 800 �800 90.43 25 ResNet18 1024 �1024 90.46 50 ResNet50 1024 �1024 90.52 23 INTERNATIONAL JOURNAL OF REMOTE SENSING 6681 Table 4. Comparison with state-of-the-art detectors on DOTA dataset. method Backbone Plane BD Bridge GTF SV LV Ship TC BC ST SBF RA Harbour SP HC mAP FR-O ResNet50 79.42 77.13 17.7 64.05 35.3 38.02 37.16 89.41 69.64 59.28 50.3 52.91 47.89 47.4 46.3 54.13 RRPN ResNet101 80.94 65.75 35.34 67.44 59.92 50.91 55.81 90.67 66.92 72.39 55.06 52.23 55.14 53.35 48.22 61.01 R2CNN ResNet101 88.52 71.2 31.66 59.3 51.85 56.19 57.25 90.81 72.84 67.38 56.69 52.84 53.08 51.94 53.58 60.67 R-DFPN ResNet101 80.92 65.82 33.77 58.94 55.77 50.94 54.78 90.33 66.34 68.66 48.73 51.76 55.1 51.32 35.88 57.94 ICN ResNet101 81.36 74.3 47.7 70.32 64.89 67.82 69.98 90.76 79.06 78.2 53.64 62.9 67.02 64.17 50.23 68.16 RoITransformer ResNet101 88.64 78.52 43.44 75.92 68.81 73.68 83.59 90.74 77.27 81.46 58.39 53.54 62.83 58.93 47.67 69.56 O2-DNet Hourglass104 89.31 82.14 47.33 61.21 71.32 74.03 78.62 90.76 82.23 81.36 60.93 60.17 58.21 66.98 61.03 71.04 R3Det ResNet101 89.54 81.99 48.46 62.52 70.48 74.29 77.54 90.80 81.39 83.54 61.97 59.82 65.44 67.46 60.05 71.69 SCRDet ResNet101 89.98 80.65 52.09 68.36 68.36 60.32 72.41 90.85 87.94 86.86 65.02 66.68 66.25 68.24 65.21 72.61 Ours ResNet18 89.00 82.83 34.12 63.91 74.04 74.70 79.35 90.81 80.19 84.74 55.77 67.51 64.72 73.15 55.47 71.35 6682 Q. GUAN ET AL. 13. Results on UCAS-AOD The detection results of UCAS-AOD are shown in Table 5. Here, we used the ResNet50 backbone for training and testing, and achieved 95.08%mAP. There was not much difference between this result and the optimal detector, indicating that our method has good universality in remote sensing images. storage-tank small-vehicle ground-track-field Plane Bridge Baseball-diamond basketball-court tennis-court Ship large-vehicle soccer-ball-field harbor swimming-pool helicopter roundabout Figure 12. Visualization of detection results on DOTA dataset. Table 5. Comparison with state-of-the-art detectors on UCAS-AOD dataset. Method Plane Car mAP YOLOv2(Redmon and Farhadi 2017) 96.60 79.20 87.90 R-DFPN (Yang et al. 2018) 98.90 82.50 89.20 DRBox (Liu, Pan, and Lei 2017) 94.90 85.00 89.95 O2-DNet(Wei et al. 2019)93.21 86.72 89.96 S2ARN (Bao et al. 2019) 97.60 92.20 94.90 RetinaNet-H (Yang et al. 2019a) 97.34 93.60 95.47 ICN (Azimi et al. 2018) – – 95.67 FADet(Li et al. 2019) 98.69 92.72 95.71 Ours 97.880 92.282 95.08 INTERNATIONAL JOURNAL OF REMOTE SENSING 6683 14. Discussion of Limitations In the experiments, we found that the setting of batch size on the DOTA dataset had a great influence on the result, and the best result could be achieved when the batch size was greater than 32. Due to device limitations, it was difficult for us to use more complex base networks under the condition of batch size = 32, such as ResNet50 and ResNet101. However, no such phenomenon was found in the HRSC2016 or the USAC-AOD dataset. We believed that the reason was that DOTA dataset has more abundant features and target categories, and training under a small batch was easy to overfit some features and categories. Therefore, we will open-source the code, and welcome interested workers to continue our work. 15. Conclusion Oriented object detection in remote sensing image is always an important research field. In the process of direction prediction, there is ambiguity in many methods to represent OBB. In this paper, we propose a novel OBB representation called CGP Box. We apply CGP Box to an anchorfree key point detection framework and compare it with the traditional representation method based on angle value. The experiment proves that the CGP Box-based method is better than the angle-based representation method. To further improve the accuracy of Guide Point positioning, we propose Gaussian Center-Line Attention Learning Module. Our experiments on three classic remote sensing public datasets (DOTA, HRSC2016, UCAS-AOD) show that CGP Box performs competitively even with a low-complexity backbone, and we would like to welcome other workers optimize CGP performance with a more complex backbone in the future. List of acronyms CNN Oriented Bounding Box OBB Oriented Bounding Box HBB Horizontal Bounding Box CGP Center-Guide Points GC-L Gaussian Center-Line IoU Intersection over Union BD baseball-diamond GTF ground-track-_eld SV small-vehicle LV large-vehicle TC tennis-court BC basketball-court ST storage-tank SBF soccer-ball-_eld RA roundabout SP swimming-pool HC helicopter SOTA the State-of-the-Art mAP mean Average Precision GSD Ground Sample Distance 6684 Q. GUAN ET AL. Notes 1. https://captain-whu.github.io/DOTA/index.html 2. https://www.kaggle.com/guofeng/hrsc2016 3. https://www.ucassdl.cn/resource.asp Disclosure statement No potential conflict of interest was reported by the author(s). ORCID Qiuyu Guan http://orcid.org/0000-0003-0263-5600 References Azimi, S. M., E. Vig, R. Bahmanyar, M. Körner, and P. Reinartz. 2018. “Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery.” In Asian Conference on Computer Vision, Cham, Springer, 150–165.. Bao, S., X. Zhong, R. Zhu, X. Zhang, L. Zhuqiang, and L. Mengyang. 2019. “Single Shot Anchor Refinement Network for Oriented Object Detection in Optical Remote Sensing Imagery.” Ieee Access 7: 87150–87161. doi:10.1109/ACCESS.2019.2924643. Chen, G. S. Xia, and X. Bai. 2020. “Gliding Vertex on the Horizontal Bounding Box for Multi-oriented Object Detection.” IEEE Transactions on Pattern Analysis and Machine Intelligence, 1452 - 1459. Cheng, G., P. Zhou, and J. Han. 2016. “Learning Rotation-invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images.” IEEE Transactions on Geoscience and Remote Sensing 54 (12): 7405–7415. doi:10.1109/TGRS.2016.2601622. Dai, J., Q. Haozhi, Y. Xiong, L. Yi, G. Zhang, H. Han, and Y. Wei. 2017. “Deformable Convolutional Networks.” In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 764–773. Dai, J., L. Yi, H. Kaiming, and J. Sun. 2016. “R-fcn: Object Detection via Region-based Fully Convolutional Networks.” In Advances in Neural Information Processing Systems, NY, United States, 379–387. Deng, Z., H. Sun, S. Zhou, J. Zhao, and H. Zou. 2017. “Toward Fast and Accurate Vehicle Detection in Aerial Images Using Coupled Region-based Convolutional Neural Networks.” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 10 (8): 3652–3664. doi:10.1109/JSTARS.2017.2694890. Ding, J., N. Xue, Y. Long, G.-S. Xia, and L. Qikai. 2019. “Learning Roi Transformer for Oriented Object Detection in Aerial Images.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 2849–2858. Everingham, M., L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. 2010. “The Pascal Visual Object Classes (Voc) Challenge.” International Journal of Computer Vision 88 (2): 303–338. doi:10.1007/s11263-009-0275-4. Girshick,R. 2015. “Fast R-cnn.” In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 1440–1448. Girshick, R., J. Donahue, T. Darrell, and J. Malik. 2014. “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 580–587. He, K., X. Zhang, S. Ren, and J. Sun. 2016. “Deep Residual Learning for Image Recognition.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 770–778. INTERNATIONAL JOURNAL OF REMOTE SENSING 6685 Hu, J., L. Shen, and G. Sun. 2018. “Squeeze-and-excitation Networks.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, , Salt Lake City, UT, USA, 7132–7141. Jaderberg, M., K. Simonyan, A. Zisserman, and K. Kavukcuoglu. 2015. “Spatial Transformer Networks.” arXiv Preprint arXiv:1506.02025Spatial Transformer Networks 2: 2017–2025 . Jiang, Y., X. Zhu, X. Wang, S. Yang, L. Wei, H. Wang, F. Pei, and Z. Luo. 2017. “R2cnn: Rotational Region Cnn for Orientation Robust Scene Text Detection.” arXiv Preprint arXiv:1706 09579. Kingma, D. P., and B. Jimmy. 2014. “Adam: A Method for Stochastic Optimization.” arXiv Preprint arXiv:1412.6980. Law, H., and J. Deng. 2018. “Cornernet: Detecting Objects as Paired Keypoints.” In Proceedings of the European Conference on Computer Vision (ECCV), Springer, Cham, 734–750. Li, C., X. Chunyan, Z. Cui, D. Wang, T. Zhang, and J. Yang. 2019. “Feature-attentioned Object Detection in Remote Sensing Imagery.” In 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan, China, 3886–3890. IEEE. Liao, M., B. Shi, and X. Bai. 2018. “Textboxes++: A Single-shot Oriented Scene Text Detector.” IEEE Transactions on Image Processing 27 (8): 3676–3690. doi:10.1109/TIP.2018.2825107. Liao, M., Z. Zhu, B. Shi, G.-S. Xia, and X. Bai. 2018. “Rotation-sensitive Regression for Oriented Scene Text Detection.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , Salt Lake City, UT, USA, 5909–5918. Lin, T.-Y., P. Dollár, R. Girshick, H. Kaiming, B. Hariharan, and S. Belongie. 2017a. “Feature Pyramid Networks for Object Detection.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 2117–2125. Liu, L., Z. Pan, and B. Lei. 2017. “Learning a Rotation Invariant Detector with Rotatable Bounding Box.” arXiv Preprint arXiv:1711 09405. Liu, W., D. Anguelov, D. Erhan, C. Szegedy, S. Reed, F. Cheng-Yang, and A. C. Berg. 2016a. “Ssd: Single Shot Multibox Detector.” In European Conference on Computer Vision, Springer, Cham, 21–37. Liu, Z., H. Jingao, L. Weng, and Y. Yang. 2017. “Rotated Region Based CNN for Ship Detection.” In 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China, 900–904. Liu, Z., H. Wang, L. Weng, and Y. Yang. 2016b. “Ship Rotated Bounding Box Space for Ship Extraction from High-resolution Optical Satellite Images with Complex Backgrounds.” IEEE Geoscience and Remote Sensing Letters 13 (8): 1074–1078. doi:10.1109/LGRS.2016.2565705. Long, J., E. Shelhamer, and T. Darrell. 2015. “Fully Convolutional Networks for Semantic Segmentation.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 3431–3440. Long, Y., Y. Gong, Z. Xiao, and Q. Liu. 2017. “Accurate Object Localization in Remote Sensing Images Based on Convolutional Neural Networks.” IEEE Transactions on Geoscience and Remote Sensing 55 (5): 2486–2498. doi:10.1109/TGRS.2016.2645610. Ma, J., W. Shao, Y. Hao, L. Wang, H. Wang, Y. Zheng, and X. Xue. 2018. “Arbitrary-oriented Scene Text Detection via Rotation Proposals.” IEEE Transactions on Multimedia 20 (11): 3111–3122. doi:10.1109/TMM.2018.2818020. Mou, L., and X. X. Zhu. 2018. “Vehicle Instance Segmentation from Aerial Image and Video Using a Multitask Learning Residual Fully Convolutional Network.” IEEE Transactions on Geoscience and Remote Sensing 56 (11): 6699–6711. doi:10.1109/TGRS.2018.2841808. Neubeck, A., and V. G. Luc. 2006. “Efficient Non-maximum Suppression.” 18th International Conference on Pattern Recognition (ICPR’06), Hong Kong, China, Vol.3 850–855. Qian, W., X. Yang, S. Peng, Y. Guo, and J. Yan. 2019. “Learning Modulated Loss for Rotated Object Detection.” arXiv Preprint arXiv:1911 08299. Redmon, J., S. Divvala, R. Girshick, and A. Farhadi. 2016. “You Only Look Once: Unified, Real-time Object Detection.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 779–788. Redmon, J., and A. Farhadi. 2017. “YOLO9000: Better, Faster, Stronger.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 7263–7271. Ren, S., H. Kaiming, R. Girshick, and J. Sun. 2015. “Faster R-cnn: Towards Real-time Object Detection with Region Proposal Networks.” In Advances in Neural Information Processing Systems, 91–99. 6686 Q. GUAN ET AL. Van Etten, A. 2018. “You Only Look Twice: Rapid Multi-scale Object Detection in Satellite Imagery.” arXiv Preprint arXiv:1805 09512. Wang, J., J. Ding, H. Guo, W. Cheng, T. Pan, and W. Yang. 2019. “Mask Obb: A Semantic Attention-based Mask Oriented Bounding Box Representation for Multi-category Object Detection in Aerial Images.” Remote Sensing 11 (24): 2930. doi:10.3390/rs11242930. Wei, H., L. Zhou, Y. Zhang, L. Hao, R. Guo, and H. Wang. 2019. “Oriented Objects as Pairs of Middle Lines.” arXiv Preprint arXiv:1912 10694. Xia, G.-S., X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang. 2018. “DOTA: A Large-scale Dataset for Object Detection in Aerial Images.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, edited by Y. Xu, F. Mingtao, Q. Wang, and Y. Wang, Kai, 3974–3983. Yang, X., L. Hou, Y. Zhou, W. Wang, and J. Yan. 2020. “Dense Label Encoding for Boundary Discontinuity Free Rotation Detection.” arXiv Preprint arXiv:2011 09670. Yang, X., Q. Liu, J. Yan, and L. Ang. 2019a. “R3DET: Refined Single-stage Detector with Feature Refinement for Rotating Object.” arXiv Preprint arXiv:1908 05612. Yang, X., H. Sun, F. Kun, J. Yang, X. Sun, M. Yan, and Z. Guo. 2018. “Automatic Ship Detection in Remote Sensing Images from Google Earth of Complex Scenes Based on Multiscale Rotation Dense Feature Pyramid Networks.” Remote Sensing 10 (1): 132. doi:10.3390/rs10010132. Yang, X., J. Yang, J. Yan, Y. Zhang, T. Zhang, Z. Guo, X. Sun, and F. Kun. 2019b. “Scrdet: Towards More Robust Detection for Small, Cluttered and Rotated Objects.” In Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea (South), 8232–8241. Zhang, Z., W. Guo, S. Zhu, and Y. Wenxian. 2018. “Toward Arbitrary-oriented Ship Detection with Rotated Region Proposal and Discrimination Networks.” IEEE Geoscience and Remote Sensing Letters 15 (11): 1745–1749. doi:10.1109/LGRS.2018.2856921. Zhao, P., Q. Zhenshen, B. Yingjia, W. Tan, Y. Ren, and P. Shiliang. 2020. “PolarDet: A Fast, More Precise Detector for Rotated Target in Aerial Images.” arXiv Preprint arXiv:2010 08720. Zhou, X., D. Wang, and K. Philipp. 2019. “Objects as Points.” arXiv Preprint arXiv:1904 07850. Zhu, H., X. Chen, W. Dai, F. Kun, Y. Qixiang, and J. Jiao. 2015. “Orientation Robust Object Detection in Aerial Images Using Deep Convolutional Neural Network.” In IEEE International Conference on Image Processing., Quebec City, QC, Canada. INTERNATIONAL JOURNAL OF REMOTE SENSING 6687