scieee AI-readable full text Open interactive document viewer

UltraPoint: improving SuperPoint for faster image keypoint extraction

Linder, Yaroslav

Abstract

We present UltraPoint, a lightweight neural network model for keypoint detection and description that improves upon the SuperPoint architecture. UltraPoint uses a streamlined encoder based on depthwise separable convolutions to drastically reduce computation—achieving about a 6× reduction in FLOPs and a 4× reduction in parameters compared to SuperPoint—while preserving the quality of the local features. On the HPatches benchmark, UltraPoint attains repeatability and localization accuracy on par with the original SuperPoint. The model is significantly smaller than current state-of-the-art learned feature extractors, enabling real-time performance (60+ FPS on a mid-range GPU) even on devices with limited computational resources. UltraPoint is trained via self-supervised learning (homographic adaptation on real images) and can also leverage knowledge distillation from a pre-trained SuperPoint to retain high descriptor discriminability. The resulting model provides an excellent trade-off between speed and accuracy for tasks like image matching, visual odometry, SLAM, and augmented reality. We also outline future extensions such as incorporating rotation invariance and applying model quantization to further improve efficiency.

Full text

UltraPoint: improving SuperPoint for faster image keypoint extraction Ihor Olkhovatyi* Yaroslav Linder [email protected] [email protected] Taras Shevchenko National University of Kyiv Abstract The object of the research is the set of algorithms and models for automatic keypoint detection and description on images. The aim of this work is to improve the SuperPoint neural network architecture by lightweighting its encoder to speed up the model and increase efficiency (reducing required computations) while preserving the quality of the obtained features. Methods: The modified model was developed and implemented in Python Results: An improved architecture named UltraPoint was developed for keypoint detection and description. UltraPoint uses depthwise separable convolutions in the encoder, reducing the model’s FLOPs by 6× and the number of parameters to 0.3 M (more than 4× smaller) while maintaining feature quality. The proposed modification achieved competitive mean localization error and repeatability on the HPatches dataset compared to the original SuperPoint, while being significantly smaller in model size and parameters than all current state-of-theart (SOTA) solutions. Practical application: The developed model can be applied to a wide range of computer vision tasks that require fast and reliable keypoint detection and description. These include image matching, visual odometry, SLAM (simultaneous localization and mapping) systems, object recognition, and augmented reality. Owing to its efficiency, the model is advantageous in scenarios with strict real-time processing requirements and deployment on platforms with limited computational resources. Significance: This work’s significance lies in creating a lightweight and efficient model for keypoint detection and description that reduces computational cost without sacrificing result quality. The proposed UltraPoint architecture aligns with modern computer vision trends focused on model optimization for mobile and embedded devices with constrained resources. In particular, the use of depthwise separable convolutions greatly reduces the model’s parameter count while preserving its representational power. Experiments confirm the advantages of the proposed approach over existing solutions in terms of the speed vs. detection quality trade-off. Recommendations for further development: Future directions include making the descriptor extraction invariant to large rotations and other homographic transformations; integrating the proposed model with other components of visual systems to optimize overall performance; evaluating SLAM system accuracy when using UltraPoint for visual odometry; and exploring model quantization and pruning to further reduce computational costs. Keywords: lightweight neural networks, selfsupervised learning, keypoint detection, descriptor learning, convolutional neural network, depthwise separable convolution, homographic adaptation, synthetic data, image matching, visual feature representation, local feature description, real-time feature detection, knowledge distillation. 1. Introduction 1.1. Overview The first step in solving many geometric computer vision problems – such as simultaneous localization and mapping (SLAM), Structure-from-Motion (SfM), camera calibration, and image matching – is detecting keypoints in an image. Keypoints are 2D image coordinates that are stable and repeatable under varying lighting and viewpoint conditions. The subfield of mathematics and computer vision known as multiple view geometry Convolutional neural networks (CNNs) have demonstrated superiority over hand-crafted methods in nearly all image processing tasks. In particular, fully convolutional neural networks that predict 2D “keypoint” locations have been well-studied for various tasks, such as human pose estimation Instead of using manual annotations to define keypoints on real images, the authors of SuperPoint MagicPoint works surprisingly well on real data despite domain adaptation issues Homographic adaptation warps the input image multiple times to help the keypoint detector see the scene from different perspectives and at different scales. The authors apply homographic adaptation in combination with the MagicPoint detector to improve its quality and generate pseudo-labels for keypoints on real images (see Figure 11 in the SuperPoint paper). As a result, detection becomes more repeatable and works on a broader set of data from any domain; the resulting improved detector is named SuperPoint. In this work, we go further: using the trained SuperPoint model, we can re-label any dataset and then use those labeled images to train a lightweight version of the model – dubbed UltraPoint – that operates much more efficiently by leveraging depthwise separable convolutions in the style of MobileNetV2 1.2. Current State of the Research Problem For a considerable time, SuperPoint has remained one of the primary solutions for image keypoint detection. Naturally, over the years there have been attempts to replace it with faster models (e.g. ALIKE by Zhao et al. 1.3. Comparison with Other Solutions Competitors to SuperPoint for keypoint detection on mobile or low-computation scenarios are still the classical methods, since they are faster than most neural network solutions. Traditional keypoint detectors have been thoroughly evaluated in prior work The SuperPoint architecture is inspired by the latest advances in applying deep learning to keypoint detection and descriptor learning. Our research builds on the SuperPoint idea and inherits all of its advantages. On the other hand, LIFT Table 1 (below) compares key hardware-related metrics – number of operations (FLOPs), number of parameters, and achievable frame processing rate (FPS) – for several current detector-descriptor systems. D2-Net The ALIKE family (ALIKE-N and ALIKED-N) strikes a balance between speed and accuracy via a compact U-Net-style backbone and a multi-layer LoFTRtype head SuperPoint remains a “golden mean.” Despite its 1.30 M parameters, the model can run in real time (∼53 FPS) and is easily adaptable to new domains. The proposed UltraPoint inherits SuperPoint’s repeatability but replaces standard convolutions with depthwise separable convolutions, reducing parameters 4× (to 0.30 M) and FLOPs over 6× (to 1.02 GFLOPs) Speed: UltraPoint exceeds SuperPoint’s speed (∼60 FPS vs ∼53 FPS), while remaining sufficiently accurate for SLAM and visual localization tasks on mobile platforms without a GPU. Table 1. Comparison of model size and speed on RTX 2060. Method GFLOPs (↓) Params (M, ↓) FPS (↑) ALIKE-N 7.91 0.318 84.96 ALIKED-N 4.05 0.677 77.40 SuperPoint 6.54 1.301 52.63 UltraPoint 1.02 0.301 60.52 Computational Efficiency: UltraPoint shows the lowest computational cost in the group, yielding FPS second only to ALIKE-N, but with lower memory usage. For systems with strict energy limits (drones, IoT robots), this FLOPs/parameters ratio is critical. Modularity and Adaptation: Unlike ALIKED-N which “hard-codes” rotational invariance, UltraPoint can be fine-tuned to arbitrary domains, retaining the ability to extend to specific transformations without significant parameter growth. Overall, UltraPoint demonstrates that through depthwise separable convolutions and distillation from SuperPoint, it is possible to dramatically cut both parameters and FLOPs with no noticeable loss in detection quality 2. Theory 2.1. The Keypoint Detection Task In computer vision, a keypoint (also interest point) is an image pixel (or small region) that is characterized by local uniqueness and stability under viewing condition changes. Keypoints typically correspond to geometrically significant image structures – such as corners, junctions of edges, blobs, or areas of sharp local contrast. The main requirement for keypoints is repeatability: they should be consistently detectable under various scene and imaging transformations (scaling, rotation, illumination changes, noise, deformation, etc.). Formally, let an image be given as I: Ω →R, where Ωis the image domain (e.g. a rectangular pixel grid), and I(p)is the intensity at point p. The keypoint detection task is to find a set of points Psuch that each point p∈Pcorresponds to a local maximum of some salience function f: Ω →R. In other words, for each candidate point p∗,f(p∗)has a high value relative to its neighbors (indicating an “informative” region like a corner or blob) and low values on flat or textureless regions. Formally, each p∗should satisfy: f(p∗)≥f(q),∀q∈N(p∗), meaning p∗is a local maximum of f. The specific form of the salience function fdepends on the chosen algorithm: for example, in the Harris corner detector, fis based on the gradient covariance matrix; in SIFT, fis the Difference-of-Gaussians; in SURF In addition to coordinates, each keypoint is often assigned extra attributes – such as scale, orientation, or a descriptor – i.e. a numerical vector characterizing the local image patch around the point. Thus, a full keypoint description can be given as a tuple (x, y, s, θ, d)where (x, y)are coordinates, sis scale, θis orientation, and d is the descriptor vector. (Often, only the descriptor and location are retained.) The keypoint pipeline is typically divided into two sub-tasks: •Detection – finding keypoint locations pwith the desired invariance properties. •Description – computing a descriptor dfor each keypoint that allows reliable matching of points between different images. After detection and description, keypoints and their descriptors are used in downstream computer vision tasks: establishing correspondences between images, motion tracking, image mosaicking, SLAM, object recognition, etc. The effectiveness of such systems largely depends on the quality of the detected keypoints and the discriminative power of their descriptors. 2.2. Classical Keypoint Detection Algorithms We next review three classic algorithms: SIFT, SURF, and ORB 2.2.1. SIFT SIFT (Scale-Invariant Feature Transform) – proposed by David Lowe in 1999–2004 – is an algorithm for detecting and describing local features invariant to image scale and rotation. SIFT features exhibit a high repeatability and strong discriminative power: even a single keypoint can often be correctly matched between images with high probability. The algorithm consists of the following main steps: Scale-space extrema detection: Construct a scale space by repeatedly convolving the image with Gaussians of increasing size. Formally, define: L(x, y, σ)=G(x, y, σ)∗I(x, y), where I(x, y)is the original image, G(x, y, σ)is a Gaussian kernel with scale parameter σ, and ∗denotes convolution. This produces a set of smoothed images L(x, y, σ)at various scales. To discretely approximate the continuous scale space, an image pyramid is used. At each octave of the pyramid, the image is progressively blurred. The Difference of Gaussians (DoG) between adjacent scales is computed to detect potential keypoints: D(x, y, σ) = L(x, y, kσ)−L(x, y, σ), where kis a fixed multiplier between scales. The resulting 3D volume (position and scale) is searched for local extrema: each pixel is compared to its 8 neighbors in the current scale and 9 neighbors in each of the neighboring scales (26 neighbors in total in the 3×3×3neighborhood). Any pixel that is a maxima or minima relative to all those neighbors is considered a keypoint candidate. Keypoint localization and filtering: Each candidate’s location and scale are refined (using interpolation to find a more precise extremum position), and unstable points are discarded. Points with low contrast (i.e. |D(x, y, σ)| value below a threshold) or those on edges are eliminated. The edge response criterion uses the Hessian matrix eigenvalues: if the ratio of eigenvalues ris above a threshold, the point lies on an elongated edge and is rejected. These steps improve the stability and precision of the keypoint set. Orientation assignment: Each detected keypoint is assigned a consistent orientation (angle) to achieve rotation invariance. This is done by considering the gradient magnitudes and orientations in the neighborhood of the keypoint at the appropriate scale. For the image at the keypoint’s scale σ, compute gradients ∇xand ∇yfor pixels around the keypoint. An orientation histogram (e.g. 36 bins covering 360°) is formed. The highest peak in the histogram gives the keypoint’s dominant orientation. If there are other significant peaks (within 80% of the maximum), additional keypoints are created at the same location and scale but with those alternate orientations. Descriptor computation: A descriptor is computed by sampling the gradient orientations around the keypoint (after rotating to the keypoint’s dominant orientation) and forming a robust representation (e.g. a 128dimensional vector). SIFT’s descriptor is constructed by dividing the neighborhood into 4×4subregions and accumulating a gradient orientation histogram (8 bins) in each, yielding a 128-length feature vector. The descriptor is typically normalized to enhance invariance to illumination changes. 2.2.2. SURF SURF (Speeded-Up Robust Features) – proposed by Bay et al. in 2006 SURF’s detector is based on the Hessian matrix determinant. Keypoints are found as scale-space extrema of the Hessian determinant det(H), which can be computed quickly via box filters (approximating secondorder Gaussian derivatives) and an integral image. This yields approximate LoG (Laplacian of Gaussian) extrema faster than difference-of-Gaussians. SURF uses a distribution of Haar-wavelet responses within a keypoint neighborhood to form the descriptor (called SURF-64 or SURF-128 depending on length). Like SIFT, it is invariant to rotation (by aligning to the dominant orientation) and robust to scale. SURF achieves comparable results to SIFT on many tasks but with improved speed, especially when using hardware acceleration due to its reliance on convolutionlike operations that integrate well with integral images. 2.2.3. ORB ORB (Oriented FAST and Rotated BRIEF) – by Rublee et al. in 2011 FAST detector: ORB uses FAST Orientation: ORB computes an orientation for each FAST keypoint using the intensity centroid method (essentially, moments within a patch to find a dominant direction). BRIEF descriptor: ORB employs the BRIEF ORB was designed as an open-source, license-free alternative to SIFT/SURF. It is very fast to compute and match, though its descriptors (being binary) can sometimes be less discriminative than SIFT’s in difficult conditions. ORB is less robust to significant viewpoint or affine changes than SIFT (it uses a discrete scale pyramid and thus has more limited scale invariance, and no explicit affine invariance). However, ORB’s speed and low resource usage make it an attractive choice for realtime and mobile applications. ORB and other classical methods are sensitive to certain factors: dramatic viewpoint or affine transformations (ORB’s discrete scale pyramid makes it even less robust to these), changes in lighting or image noise can also affect detectors and descriptors. Nonetheless, classical feature pipelines remain competitive in scenarios where computational budget is extremely tight or where training data for learning-based methods is insufficient. 2.3. Neural Networks 2.3.1. Why Use Neural Networks for Keypoints? Modern deep neural networks offer learnable, datadriven approaches to detection and description tasks that can surpass classical hand-engineered methods. The key reasons to apply neural networks for keypoint detection include: •Ability to learn invariance and features directly from data rather than relying on fixed image filters. •Joint detection and description: networks can be trained to optimize the end-to-end matching performance. •Adaptability to new domains through transfer learning or fine-tuning, which is difficult for fixed algorithms. However, training requires sufficient data and careful consideration of what constitutes a “ground truth” keypoint in an unsupervised or self-supervised setting (since manual labeling of generic keypoints is impractical). Figure 1. Graphs of common activation functions: Sigmoid, Tanh, ReLU, and Softmax. The sigmoid activation (top left) maps any input to the range (0,1) in an S-shaped curve. The hyperbolic tangent (tanh) (top right) squashes inputs to (−1,1). The ReLU (Rectified Linear Unit, bottom left) outputs max(0, x), effectively zeroing out negative inputs. It is computationally efficient and helps mitigate vanishing gradients. The softmax (bottom right) is used in output layers for classification; it normalizes a vector of values into a probability distribution. 2.3.2. Basic Concepts and Working Principles of Neural Networks A neural network consists of layers of interconnected units (neurons) with learnable weights. Each neuron computes a weighted sum of its inputs and passes it through a non-linear activation function. Networks can approximate complex mappings by stacking many layers (hence “deep” learning). Key concepts include: •Feed-forward networks (MLP): Each layer fully connects to the next. These were among the first networks used (e.g. perceptron, multi-layer perceptron). •Convolutional Neural Networks (CNNs): Specialized for images Neural networks learn by adjusting weights to minimize a loss function on training data, typically through backpropagation and gradient descent (or variants like SGD, RMSprop, Adam optimizer, etc.). These fundamental neural network architectures provide building blocks that can be leveraged for keypoint detection tasks. In particular, CNNs are the workhorse for image feature learning, and modern keypoint detectors like SuperPoint are fully convolutional networks specialized to predict per-pixel keypoint probabilities and descriptors. 2.4. Metrics for Keypoint Detection and Description To evaluate keypoint detectors and descriptors, several metrics are commonly used: •Precision: The fraction of detected keypoints (or matches) that are relevant (true positives). It measures accuracy of detection. If we denote true positives (correct detections) as TP and false positives as FP, then Precision =T P T P +F P . •Recall: The fraction of true keypoints that were detected. It measures completeness. Using false negatives FN (missed detections), Recall =T P TP+F N . •Repeatability: A keypoint is considered repeatable if the same physical point (after a known homography or geometric transformation) is within a certain distance (e.g. εpixels) of a detected point in the corresponding image. Repeatability is the percentage of keypoints that meet this criterion across image pairs – effectively, the probability that a keypoint detected in one image will be found at the corresponding location in another view 3. Proposed Solution 3.1. SuperPoint Model Details To understand our proposed changes, we first review the SuperPoint model, as UltraPoint follows a similar training process and overall architecture. SuperPoint Shared Encoder: The encoder is a CNN (akin to a VGG-style feature extractor) that processes the input image and produces a dense feature map. In the original SuperPoint, the encoder has 8 convolutional layers with 3×3 kernels, periodically downsampled to produce a lower-resolution feature map (typically 1/8th of the input resolution) Keypoint Decoder: Takes the shared feature map and upsamples it (using bilinear or nearest-neighbor upsampling with no trainable parameters) back to full resolution. It applies a final 3×3 convolution to produce a keypoint heatmap – essentially a probability map where each pixel’s value indicates the confidence of a keypoint at that location. Non-maximum suppression (NMS) is then applied to pick out discrete keypoint locations from this heatmap. Descriptor Decoder: Also operates on the shared feature map (potentially at a lower spatial resolution). It applies a set of convolutions to produce a dense grid of descriptor vectors (e.g. 256-dimensional descriptor at each cell of an 1/8-scale grid). These descriptors are usually ℓ2-normalized to unit length During training, SuperPoint uses a softmax across heatmap pixels (with a background class) and a crossentropy loss to learn keypoint detection, combined Figure 2. SuperPoint architecture (diagram reference Figure 3. Homographic adaptation process (as used in SuperPoint’s self-supervised training with a descriptor loss (e.g. a hinge-based triplet loss or a pairing that encourages descriptors of corresponding points to be similar and non-corresponding to be different). 3.2. UltraPoint 3.2.1. Designing a Lightweight Architecture Our UltraPoint design goal was to significantly reduce model size and computation while maintaining performance. To achieve this, we drew inspiration from MobileNetV2 –We replace the 3×3 full convolutions in SuperPoint’s encoder with depthwise separable convolutional blocks. A depthwise separable convolution factorizes a standard convolution into a depthwise (per-channel) convolution followed by a pointwise (1×1) convolution. This drastically reduces the number of multiplications. In MobileNetV2, such blocks are often combined with expansion (to a higher channel count) and projection (to a lower channel count) with residual connections. We incorporate similar inverted residual bottleneck blocks to preserve representational power. –We keep the dual-head structure (detector and descriptor remain separate decoders). This not only simplifies training (each head focuses on its task) but also makes it easier to integrate UltraPoint into existing pipelines (you can use the detector or descriptor independently if needed). Maintaining two heads did not significantly increase parameters compared to a single-head approach, since the majority of parameters lie in the shared encoder. By adopting the above architectural changes, we achieved a ∼6× reduction in FLOPs (to ≈1.02 GFLOPs) and ∼4.3× reduction in parameters (to ≈0.30 M) without significant loss of keypoint detection or descriptor quality Figure 4. UltraPoint architecture overview. 3.2.2. Training Process We explore two training approaches for UltraPoint: training from scratch using the self-supervised pipeline, and knowledge distillation from SuperPoint: –Self-supervised training from scratch: We mimic the original SuperPoint training procedure In practice, we found that combining both strategies yields the best results: we initialize UltraPoint’s training with knowledge distillation (quickly giving it a good descriptor quality baseline), then finetune with the self-supervised homographic adaptation and pseudo-label scheme to adapt to any domain specifics and possibly improve repeatability. The result is a model that preserves SuperPoint’s strengths (repeatability and adaptability) but is lighter and faster by a large margin. 4. Experimental Results We evaluated UltraPoint against SuperPoint and other methods on standard benchmarks, focusing on keypoint repeatability and matching performance. Key results: Synthetic shapes & HPatches: On the Synthetic Shapes dataset and the HPatches benchmark, UltraPoint achieved mean localization error (MLE) and repeatability on par with SuperPoint Speed and Efficiency: On an NVIDIA RTX 2060 GPU, UltraPoint processed ∼60.5 FPS vs ∼52.6 FPS for SuperPoint, and ∼85 FPS for ALIKE-N (the fastest in our comparison) Table 2. Comparison of model size and speed on RTX 2060. Best three FPS are marked in red, green, and blue. Model GFLOPs Params (M) FPS D2-Net 889.40 7.635 7.63 R2D2 464.55 0.484 4.10 DISK 98.97 1.092 11.81 ALIKE-L 19.68 0.653 56.66 ALIKE-N 7.91 0.318 84.96 ALIKED-N 4.05 0.677 77.40 SuperPoint 6.54 1.301 52.63 UltraPoint 1.02 0.301 60.52 Matching and Homography Estimation: We computed homography estimation accuracy on HPatches with varying correctness thresholds. UltraPoint’s detector repeatability and localization error were similar to SuperPoint (so geometric accuracy of detected points was maintained) Table 3. Homography evaluation on HPatches with different correctness thresholds ε. Model Homography Detector Desc. Overall ε=1 ε=3 ε=5 Rep. MLE mAP M UltraPoint 0.227 0.351 0.370 0.451 1.616 0.339 0.185 SuperPoint 0.306 0.468 0.486 0.459 1.472 0.444 0.420 Table 4 illustrates detector repeatability on HPatches (split by illumination and viewpoint change categories, with different NMS radius settings), showing UltraPoint vs SuperPoint: As seen, UltraPoint can slightly exceed SuperPoint’s repeatability in some cases (e.g. viewpoint Figure 5. UltraPoint training process using distillation learning. Figure 6. Example of pseudo-labels on real images (MSCOCO dataset Figure 7. Examples of visualization of UltraPoint model results on images from the validation part of the MS-COCO dataset with a lower detection threshold yields higher repeatability), likely because the distilled model foFigure 8. Examples of UltraPoint predictions (red) and annotations (green) on synthetic images from the Synthetic Shapes dataset (using augmentations). Table 4. Detector repeatability on HPatches (higher is better). Method Illumination Viewpoint NMS=4 NMS=8 NMS=4 NMS=8 UltraPoint (τ=0.007) 0.608 0.451 0.282 0.138 UltraPoint (τ=0.0025) 0.645 0.459 0.464 0.326 SuperPoint 0.670 0.532 0.237 0.132 Random 0.101 0.103 0.100 0.104 cuses on the most reliable points. In terms of descriptor matching, we evaluated the nearest-neighbor mean Average Precision (NN mAP) and matching score on HPatches sequences. UltraPoint achieved an NN mAP within a few points of SuperPoint (and higher than other lightweight methods like ORB). Its matching score (after RANSAC refinement) was also comparable, albeit slightly lower than SuperPoint due to a few missed matches under very large viewpoint changes. Overall, UltraPoint provides a very good compromise: nearly the same matching performance at a fraction of the cost. Figure 9. Validation recall curve during training on Synthetic Shapes. This graph tracks the keypoint detector’s recall on a validation set as training progresses (higher is better). UltraPoint (blue curve) quickly reaches over 80% recall. The sharp drops correspond to learning rate schedule changes or model re-initializations (for example, when switching from synthetic pre-training to real image training). After retraining on real data with pseudo-labels, the recall stabilizes at a high value (∼0.8), indicating that UltraPoint detects a large fraction of the true (pseudo-labeled) keypoints. The final small dip and rise are due to fine-tuning steps. 5. Conclusions In this thesis, we conducted a comprehensive study of methods for automatic keypoint detection and description in images, and proposed an improved, lightweight architecture UltraPoint that builds upon the ideas of SuperPoint. The main results and conclusions can be summarized as follows: *Theoretical and Methodological Foundations: We performed a critical analysis of classical and neural approaches. We showed that while SIFT, SURF, and ORB provide invariance to scale and rotation, they remain sensitive to complex affine distortions, abrupt illumination changes, and noise. Their fixed detector-descriptor pipeline and hardwired parameters limit flexibility and adaptability to new domains and hardware constraints. *Role of Self-Supervised Learning: We demonstrated the advantages of the MagicPoint →SuperPoint strategy with three-phase training (Synthetic Shapes →homographic adaptation → training on pseudo-labels). This approach eliminates the need for human annotation and greatly extends the applicability of the model to unlabeled data and new domains. *Architecture and Implementation Aspects of UltraPoint: ·Lightweight Encoder: We proposed replacing VGG-style 3×3 convolutions with depthwise separable convolution blocks and inverted bottlenecks (as in MobileNetV2). This reduced FLOPs by 6× (to ≈1.02 GFLOPs) and parameters by 4.3× (to ≈0.3 M) without significant loss of detection or descriptor performance References [1] Programming language Python. Van Rossum, G., & Drake, F. L. Python 3 Reference Manual. Scotts Valley, CA: CreateSpace, 2009. [2] Machine learning framework PyTorch. Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library, 2019. [3] Hartley, R., & Zisserman, A. Multiple View Geometry in Computer Vision. 2nd Ed. Cambridge University Press, 2003. [4] Wei, S.-E., Ramakrishna, V., Kanade, T., & Sheikh, Y. Convolutional pose machines.InCVPR, 2016. [5] Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., & Berg, A. C. SSD: Single Shot Multibox Detector.InECCV, 2016. [6] Lee, C.-Y., Badrinarayanan, V., Malisiewicz, T., & Rabinovich, A. RoomNet: End-to-End Room Layout Estimation. In ICCV, 2017. [7] DeTone, D., Malisiewicz, T., & Rabinovich, A. SuperPoint: Self-Supervised Interest Point Detection and Description. In CVPR Workshop on Deep Learning for Visual SLAM, 2018. [8] Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., & Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv:1704.04861, 2017. [9] Zhao, X., Wu, X., Miao, J., Chen, W., & Chen, P. C. Y., Li, Z. ALIKE: Accurate and Lightweight Keypoint Detection and Descriptor Extraction.IEEE Transactions on Multimedia, 25(2023): 3101–3112. DOI: 10.1109/TMM.2022.3155592. [10] Zhao, X., Wu, X., Chen, W., Chen, P. C. Y., Xu, Q., & Li, Z. ALIKED: A Lighter Keypoint and Descriptor Extraction Network via Deformable Transformation.IEEE Transactions on Instrumentation & Measurement, 72(2023): 1–16. DOI: 10.1109/TIM.2023.3271000. [11] Schmid, C., Mohr, R., & Bauckhage, C. Evaluation of interest point detectors.International Journal of Computer Vision, 37(2): 151–172 (2000). [12] Mikolajczyk, K., & Schmid, C. A performance evaluation of local descriptors. IEEE PAMI, 27(10): 1615–1630 (2005). [13] Rosten, E., & Drummond, T. Machine Learning for High-Speed Corner Detection. In ECCV 2006, pp. 430–443. Lecture Notes in Computer Science 3951, 2006. DOI: 10.1007/11744023 34. (Introduction of FAST) [14] Lowe, D. G. Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision, 60(2): 91–110 (2004). DOI: 10.1023/B:VISI.0000029664.99615.94. (SIFT) [15] Yi, K. M., Trulls, E., Lepetit, V., & Fua, P. LIFT: Learned Invariant Feature Transform.InECCV 2016, Lecture Notes in Computer Science, vol. 9910, pp. 467–483. Springer, 2016. DOI: 10.1007/978-3-319-46466-4 28. [16] Rublee, E., Rabaud, V., Konolige, K., & Bradski, G. ORB: An Efficient Alternative to SIFT or SURF. In ICCV, pp. 2564–2571, 2011. DOI: 10.1109/ICCV.2011.6126544. [17] Dusmanu, M., Rocco, I., Pajdla, T., Pollefeys, M., Sivic, J., Torii, A., & Sattler, T. D2-Net: A Trainable CNN for Joint Description and Detection of Local Features.InCVPR, 2019, pp. 8084–8093. [18] Revaud, J., Weinzaepfel, P., de Souza, C., Pion, N., Csurka, G., Cabon, Y., & Humenberger, M. R2D2: Repeatable and Reliable Detector and Descriptor. In NeurIPS, 2019. [19] Tyszkiewicz, M. J., Fua, P., & Trulls, E. DISK: Learning local features with policy gradient.InNeurIPS, 2020. [20] Simonyan, K., & Zisserman, A. Very Deep Convolutional Networks for LargeScale Image Recognition. Proc. of ICLR, 2015. Pre-print arXiv:1409.1556. DOI: 10.48550/arXiv.1409.1556. (VGG) [21] Sun, J., Shen, Z., Wang, Y., Bao, H., & Zhou, X. LoFTR: Detector-Free Local Feature Matching with Transformers. In CVPR, 2021, pp. 8922–8931. DOI: 10.1109/CVPR46437.2021.00881. [22] Bay, H., Tuytelaars, T., & Van Gool, L. SURF: Speeded Up Robust Features. In