Features extraction for image identification using computer vision
Abstract
This study examines various feature extraction techniques in computer vision, the primary focus of which is on Vision Transformers (ViTs) and other approaches such as Generative Adversarial Networks (GANs), deep feature models, traditional approaches (SIFT, SURF, ORB), and non-contrastive and contrastive feature models. Emphasizing ViTs, the report summarizes their architecture, including patch embedding, positional encoding, and multi-head self-attention mechanisms with which they overperform conventional convolutional neural networks (CNNs). Experimental results determine the merits and limitations of both methods and their utilitarian applications in advancing computer vision.
Full text
Corresponding author: Venant Niyonkuru Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution License 4.0. Features extraction for image identification using computer vision Venant Niyonkuru 1, *, Sylla Sekou 2 and Jimmy Jackson Sinzinkayo 3 1 Department of Computing and Information System, Kenyatta University, Kenya. 2 Department of Mathematics, Institute for Basic Science, Technology and Innovation, Pan-African University, Kenya. 3 Department of Software Engineering, College of Software, Nankai University, China. World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 Publication history: Received on 01 June 2025; revised on 12 July 2025; accepted on 14 July 2025 Article DOI: https://doi.org/10.30574/wjarr.2025.27.1.2647 Abstract This study examines various feature extraction techniques in computer vision, the primary focus of which is on Vision Transformers (ViTs) and other approaches such as Generative Adversarial Networks (GANs), deep feature models, traditional approaches (SIFT, SURF, ORB), and non-contrastive and contrastive feature models. Emphasizing ViTs, the report summarizes their architecture, including patch embedding, positional encoding, and multi-head self-attention mechanisms with which they overperform conventional convolutional neural networks (CNNs). Experimental results determine the merits and limitations of both methods and their utilitarian applications in advancing computer vision. Keywords: Feature Extraction; Positional Embeddings; Self-Attention; Vision Transformers (ViTs) 1. Introduction Feature extraction is a critical stage in the computer vision domain that is the backbone of transforming raw image data with high amounts into compact, descriptive representations that enable object detection, image categorization, segmentation, and scene interpretation. Traditionally, feature extraction methods have developed over time based on the need for creating descriptors that are invariant to scaling, rotation, lighting, and perspective, but computationally effective (Dosovitskiy et al, 2020; Jiang, 2009; Lowe,2004; Grill et al, 2020). Traditional feature extraction techniques like Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Oriented FAST and Rotated BRIEF (ORB) have been instrumental for initial computer vision systems (Lowe,2004; Rublee et al, 2011, Morrow, 2000) . These algorithms engineer features from local image properties, finding keypoints and constructing descriptors to facilitate matching among different images. While resistant in the majority of scenarios, these hand-crafted features are often prone to difficulty with complexity, scalability, and sometimes devoid of semantic context. Deep learning transformed feature extraction by the power to learn hierarchical representations directly from data without needing hand-designed features. Convolutional Neural Networks (CNNs) emerged as the standard by leveraging local spatial correlation and shared weights but with the expense of local receptive fields, which limit their capacity to learn long-range dependencies in images (Dosovitskiy et al, 2020; Ali, et al, 2023; Krizhevsky et al, 2012,Morrow, 2000). Here, Vision Transformers (ViTs) have emerged as a highly promising substitute that brings the self-attention mechanism of NLP into computer vision (Dosovitskiy et al, 2020). ViTs work by dividing images into fixedsize patches, flattening them, and linearly embedding them. Positional embeddings help to maintain the spatial information, and the patch embedding is fed into multi-head self-attention to capture global context (Dosovitskiy et al, 2020, Patwardhan et al, 2023, Montrezol, 2024). This paradigm change helps ViTs capture the relationships of the entire
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1342 image and surpass limitations intrinsic to CNNs and delivering superior performance across a wide array of vision benchmarks. Also, newer architectures such as Generative Adversarial Network (GAN)-based models and contrastive learning techniques have added to the list of tools used to learn semantic features from images (Ali et al, 2024, Cao et al, 2018; Kovács et al, 2023). These are aimed at learning discriminative and generative representations that are useful over a broad range of tasks ranging from image generation to self-supervised learning (Grill et al., 2020; Ansar et al., 2024). This study comprehensively examines these varied feature extraction approaches, demystifying the principle behind Vision Transformers and their position within the wider computer vision context. Experimental results clarify their individual strengths, compromises, and practical usability, sketching the outline for the best feature extraction approaches to use in real-world applications(Purchase, 2012). 2. Related work Traditional approaches are SIFT (Lowe, 2004), SURF (Bay et al., 2008), and ORB (Rublee et al., 2011), which have served as standard baseline approaches to image matching and recognition. These approaches rely on handcrafted descriptors to obtain local features. They operate well in structured or low-variation visual scenes. However, they cannot deal with scale variation, illumination variation, and occlusion. One of the major breakthroughs as exemplified by the emergence of deep learning models, particularly Convolutional Neural Networks (CNNs), was when these models learned to learn end-to-end discriminative hierarchical features from raw images (Krizhevsky et al., 2012). CNNs were more generalizable on a wide variety of vision tasks and thus remained the standard for a number of years. In recent times, Vision Transformers (ViTs) have been strong competitors that are based on self-attention mechanisms for obtaining long-range relations in images (Dosovitskiy et al., 2020). ViTs outperformed CNNs on big-benchmark benchmarks, particularly when they were trained on very large datasets. Simultaneously, Generative Adversarial Networks (GANs) have not only been utilized for image synthesis but also for feature extraction, depending on discriminators for obtaining detailed, high-level features. Furthermore, contrastive learning techniques such as BYOL (Grill et al., 2020) and SimCLR have enhanced self-supervised feature learning by optimizing the agreement between multiple copies of an image that are transformed differently. Recent large-scale surveys (Ali et al., 2023; Patwardhan et al., 2023) cover developments in these architectures, presenting trends and open questions. However, there are fewer papers providing an explicit comparison of these different approaches under the same experimental setting. This paper fills this gap by comparing classical descriptors, CNNs, ViTs, and GAN-based models on an identical setup of popular benchmarks and measures. 3. Methodology 3.1. Vision Transformer (Vits) 3.1.1. Definition and functionality Vision Transformers (ViTs) are deep learning models that leverage self-attention mechanisms to process image data, offering improved performance over traditional convolutional neural networks (CNNs) (Dosovitskiy et al, 2020) . 3.1.2. Architecture ViTs divide an image into fixed-size patches, linearly embed them, and feed them into a transformer encoder. The key components include: • Patch Embedding Layer: Converts image patches into token embeddings. • Positional Encoding: Adds spatial information to tokens. • Multi-Head Self-Attention: Captures long-range dependencies in an image. • Feed-Forward Network (FFN): Processes token representations for classification tasks.
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1343 Figure 1 Transformer Encorder How ViTs Work • The image is broken into non-overlapping patches • Each patch is flattened and subsequently passed through a linear projection. • The transformer encoder converts the patch embeddings through self-attention. • classification head produces predictions from the last encoded representation. 3.1.3. Image Patches The process starts with dividing an image into small, fixed-size patches, and that is a simple transformation step. This process has a direct analogy in natural language processing (NLP) where a sentence is segmented into individual units such as words or subword tokens. Just like how every token within a sentence carries contextual meaning, every patch within an image captures localized visual context. In this analogy, the entire image is taken as a sentence, and its patches are akin to tokens, which enable transformer-based models originally designed for text to be used on visual data. Figure 2 Image to Image Patches Both vision transformers (ViT) and natural language processing (NLP) partition large inputs (i.e., sentences in text or entire images into smaller ones, e.g., tokens in text or image patches). For instance, processing an entire 224×224 pixel image directly would entail an impossibly large number of calculations, approximately 2.5 billion comparisons. But by dividing the very same image into 256 patches, each 14×14 pixels, the computation load of one attention layer becomes incredibly smaller approximately 9.8 million comparisons.
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1344 Figure 3 Vision transformers 3.1.4. Linear Projection Following patch division of the image, each patch is then converted from a 2D array to a 1D vector using a linear projection, effectively projecting raw pixel information into a set of patch embeddings. Figure 4 Linear Projection The role of the linear projection layer is to transform each image patch into a fixed-size vector representation, the aim being to maintain meaningful relations so visually similar patches produce similar embeddings. This transformation brings the data into a form compatible with the input format needed by the transformer model. Two further processing steps remain before these embeddings can be used. 3.1.5. Learnable Embeddings One of the important features added in widely used transformer models such as BERT is the inclusion of a special classification token, also known as [CLS]. This token is placed at the beginning of every input sequence and is meant to capture the sentence-level representation for classification tasks.
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1345 Figure 5 Bert Tokenizer There is a unique token, [CLS], in BERT that is added to the beginning of all input sequences. This token is embedded like any other and passed through the encoder layers of the model. The [CLS] token is special in that it doesn't represent any specific word of the input it begins as a neutral or uninitialized vector. In addition, during pretraining, this final output at the [CLS] position is fed as input to a classification layer. This encourages the model to encode information from the entire sentence into this single vector, learning an effective representation of the input. Vision Transformers (ViT) do exactly the same thing with a learnable embedding that serves the same purpose as the [CLS] token in BERT, providing a summary representation for image-level classification tasks. Figure 6 Transformer Encoder to Linear Projection 3.1.6. Positioning Embedding Transformers do not have an inherent perception of sequence or spatial arrangement of input tokens or patches. However, preserving order is important in language, where word reordering can dramatically alter meaning. The same is true for visual information: when the components of a picture are mixed up, as in a jigsaw puzzle, identification of the whole picture becomes extremely challenging. This is also true for transformer models, which require an additional mechanism to understand the relative position of these parts. To address this, positional embeddings are added. In Vision Transformers (ViT), these are learned and of the same dimension as the patch embeddings. Following the division of the image into patches and adding the special classification token, each element is added to its respective positional embedding. These position vectors are also trained along with the model and can further be fine-tuned later. They gradually come to denote spatial relationships, usually identical to proximate locations in the grid particularly in the same column or row such as:
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1346 Figure 7 Embeddings Position Once positional embeddings are added, the patch embeddings are complete. These enhanced embeddings are then passed into the Vision Transformer (ViT), and they are processed in the same way as regular tokens in a standard transformer model Imprementation
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1347 The training dataset consists of 60,000 images across 11 unique classes. In order to obtain the equivalent humanreadable labels for these classes, the following steps may be used: ClassLabel has 11 classes: ['airplane', 'automobile', 'bird', 'cat', 'deer', 'dog',.]. Each entry in the dataset contains two features: `img` and `label`. The `img` feature contains a 32x32 pixel image which is of type PIL and with three color channels of RGB (red, green, blue). . 3.1.7. Feature extraction Before sending images to the Vision Transformer (ViT) model, a feature extractor is used to handle preprocessing. This involves resizing and normalizing images, converting them into tensors referred to as "pixel_values."
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1348 The feature extractor may be initialized with the Transformers library of Hugging Face, as shown below: The feature extractor configuration shows that normalization and resizing are set to true. Normalization is performed across the three color channels using the mean and standard deviation values stored in "image_mean" and "image_std" respectively. Therefore, it is optimal to use an image that is slightly larger than needed, since reducing by a small amount usually preserves visual quality and avoids introducing visible degradation in image quality.
World Journal of Advanced Research and Reviews, 2025, 27(01), 1341-1351 1349 Evaluation and Prediction The Trainer evaluates during training but we can also quickly do a more qualitative verification (or estimation) by passing through a single image with the model and feature_extractor. We will pass the following image: The picture is of poor visual quality and does not have distinguishing features, so visual categorization based on the picture is difficult. However, the label given classifies the subject as a cat. We will now go ahead and test the model's prediction for this picture.