Full text
Image Feature Extraction Acceleration Jorge Fernández-Berni, Manuel Suárez, Ricardo Carmona-Galán, Víctor M. Brea, Rocío del Río, Diego Cabello and Ángel Rodríguez-Vázquez Abstract Image feature extraction is instrumental for most of the best-performing algorithms in computer vision. However, it is also expensive in terms of computational and memory resources for embedded systems due to the need of dealing with individual pixels at the earliest processing levels. In this regard, conventional system architectures do not take advantage of potential exploitation of parallelism and distributed memory from the very beginning of the processing chain. Raw pixel values provided by the front-end image sensor are squeezed into a high-speed interface with the rest of system components. Only then, after deserializing this massive dataflow, parallelism, if any, is exploited. This chapter introduces a rather different approach from an architectural point of view. We present two Application-Specific Integrated Circuits (ASICs) where the 2-D array of photo-sensitive devices featured by regular imagers is combined with distributed memory supporting concurrent processing. Custom circuitry is added per pixel in order to accelerate image feature extraction right at the focal plane. Specifically, the proposed sensing-processing chips aim at the acceleration of two flagships algorithms within the computer vision community: J. Fernández-Berni (B )·R. Carmona-Galán ·R. del Río ·Á. Rodríguez-Vázquez Institute of Microelectronics of Seville (CSIC - Universidad de Sevilla), C/ Américo Vespucio s/n, 41092 Seville, Spain e-mail: [email protected] R. Carmona-Galán e-mail: [email protected] R. del Río e-mail: [email protected] Á. Rodríguez-Vázquez e-mail: [email protected] V.M. Brea ·M. Suárez ·D. Cabello Centro de Investigación en Tecnoloxías da Información (CITIUS), University of Santiago de Compostela, Santiago de Compostela, Spain e-mail: victor[email protected] D. Cabello e-mail: [email protected] © Springer International Publishing Switzerland 2016 A.I. Awad and M. Hassaballah (eds.), Image Feature Detectors and Descriptors, Studies in Computational Intelligence 630, DOI 10.1007/978-3-319-28854-3_5 109
110 J. Fernández-Berni et al. the Viola-Jones face detection algorithm and the Scale Invariant Feature Transform (SIFT). Experimental results prove the feasibility and benefits of this architectural solution. Keywords Image feature extraction ·Focal-plane acceleration ·Distributed memory ·Parallel processing ·Viola-Jones ·SIFT ·Vision chip 1 Introduction 1.1 Embedded Vision Embedded vision market is forecast to experience a notable and sustained growth during the next few years [1]. The integration of hardware and software technologies is reaching the required maturity to support this growth. At hardware level, the ever-increasing computational power of Digital Signal Processors (DSPs), Field Programmable Gate Arrays (FPGAs), General-Purpose Graphics Processing Units (GP-GPUs) and vision-specific co-processors permit to address the challenging processing requirements usually demanded by embedded vision applications [2]. At software level, the development of standards like OpenCL [3] or OpenVX [4]as well as tools like OpenCV [5]orCUDA[6] allow for rapid prototyping and shorter time to market. A noticeable trend within this ecosystem of technologies is hardware parallelization, commonly in terms of processing operations [7,8]. However, improving performanceisnotonlyamatterofparallelizingcomputationaltasks.Memorymanagement and dataflow organization are crucial aspects to take into account [9,10]. In the case of memory management, the limitation arises from the so-called memory gap [11], leading to a substantial amount of idle time for processing resources due to slow memory access. The influence of a well-designed dataflow organization on the system performance is intimately related to this limitation. The overall objective must be to avoid moving large amounts of information pieces back and forth between system components via intermediate memory modules [2]. Optimization on this point must be planned after a comprehensive analysis of the processing flow featured by the targeted algorithm [9]. Particularly, early vision involving pixel-level operations must be carefully considered as it normally constitutes the most demanding stage in terms of processing and memory resources. 1.2 Focal-Plane Sensing-Processing Architecture When all these key factors shaping performance are closely examined from an architectural point of view, a major disadvantage of conventional system architectures becomes evident. As can be observed in Fig.1, vision systems typically consist of a
Image Feature Extraction Acceleration 111 Fig. 1 Conventional architecture of embedded vision systems: image sensor, high-speed Analogto-Digital Conversion (ADC), memory and processing resources (DSP, GPU etc.) front-end imager delivering high-quality images at high speed to the rest of system components. This arrangement, by itself, generates a critical bottleneck associated with the huge amount of raw data rendered by the imager that must be subsequently stored and processed from scratch. But even more importantly, it precludes a first stage of processing acceleration from taking place just at the focal plane in a distributed and parallel way. Notice that the imager inevitably requires the physical realization of a 2-D array of photo-sensitive devices topographically assigned to their corresponding pixel values. This array can be exploited as distributed memory where the data are directly accessible for concurrent processing by including suitable circuitry at pixel level. As a result, the imager will be delivering pre-processed images, possibly in addition to the original raw information in case the algorithm needs it to superpose the processing outcome—e.g. highlighting the location of a tracked object. This architectural approach, referred in the literature as focal-plane sensing-processing [12] and represented in Fig.2, presents two fundamental advantages when compared to that of Fig.1. First of all, it enables a drastic reduction of memory accesses during low-level processing stages, where pixel-wise operations are common. Secondly, it permits to design ad-hoc circuitry to accelerate a vision Fig. 2 Proposed architecture for focal-plane acceleration of image feature extraction. The pixel array is exploited as distributed memory including per-pixel circuitry for parallel processing
112 J. Fernández-Berni et al. algorithm according to its specific characteristics. This circuitry can even be implemented in the analog domain for the sake of power and area efficiency since the pixel values at the focal plane have not been converted to digital yet. On the flip side, the incorporation of processing circuitry at pixel level reduces, for a prescribed pixel area, the sensitivity of the imager as less area is devoted to capture light. This drawback could be overcome by means of the so-called 3-D integration technologies [13,14]. In this case, a sensor layer devoting most of its silicon area to capture light would be stacked and vertically interconnected onto one or more layers exclusively dedicated to processing. While not mature enough yet for reliable implementation of sensing-processing stacks, 3-D manufacturing processes will most surely boost the application frameworks of the research hereby presented. All in all, this chapter introduces two full-custom focal-plane accelerator sensingprocessing chips. They are our first prototypes aiming respectively at speeding up the image feature extraction of two flagships algorithms within the embedded vision field: the Viola-Jones face detection algorithm [15] and the Scale Invariant Feature Transform (SIFT) [16]. To the best of our knowledge, no prior attempts pointing to these algorithms have been reported for the proposed sensing-processing architectural solution. The chapter is organized as follows. After briefly describing both algorithms, we justify the operations targeted for implementation at the focal plane. We demonstrate that these operations feature a common underlying processing primitive, the Gaussian filtering, convenient for pixel-level circuitry. We then explain how this processing primitive has been implemented on both chips. Finally, we provide experimental results and discuss the guidelines of our future work on this subject matter. 2 Vision Algorithms 2.1 Viola-Jones Face Detection Algorithm The Viola-Jones sliding window face detector [15] is considered a milestone in real-time generic object recognition. It requires a cumbersome previous training, demanding a large number of cropped frontal face samples. But once trained, the detectionstageisfastthanks to the computation of the integralimage, an intermediate image representation speeding up feature extraction, and to a cascade of classifiers of progressive complexity. A basic scheme of the Viola-Jones processing flow is depicted in Fig.3. Despite its simplicity and detection effectiveness, the algorithm still requires a considerable amount of computational and memory resources in terms of embedded system affordability. Different approaches have been proposed in the literature in order to increase the implementation performance: by exploiting the highly parallel computation structure of GPUs [17,18]; by making the most of the logic and memory capabilities of FPGAs [19,20]; by custom design of specialized digital hardware [21]etc.
Image Feature Extraction Acceleration 113 Fig. 3 Simplified scheme of the Viola-Jones processing flow In order to evaluate the possibilities for focal-plane acceleration, our interest focuses on pixel-level operations. For the Viola-Jones algorithm, these operations take place during the computation of the integral image, defined as: II(x,y)= x x=1 y y=1 I(x,y)(1) where I(x,y)represents the input image. That is, each pixel composing II(x,y)is equal to the sum of all the pixels above and to the left of the corresponding pixel at the input image. The first advantage of the integral image is that its calculation permits to compute the sum of any rectangular region of the input image by accessing only four pixels of the matrix II(x,y). This is critical for real-time operation, given the potential large number of Haar-like features to be extracted—2135 in total for the OpenCV baseline implementation. The second advantage is that the computation of the integral image fits very well into a pipeline architecture—typically implemented in DSPs—by making use of the following pair of recurrences: r(x,y)=r(x,y−1)+I(x,y) II(x,y)=II(x−1,y)+r(x,y)(2) with r(x,0)=0 and II(0,y)=0. The matrix II(x,y)can thus be obtained in one pass over the input image. Despite these advantages, the purely sequential approach defined by Eq.(2)is still computationally expensive and memory access intensive [20,22]. It usually
114 J. Fernández-Berni et al. accounts for a large fraction of the total execution time due to its linear dependence on the number of pixels of the input image [23]. Thus, its parallelization would boost the performance of the whole algorithm. In the next sections, we will propose an acceleration scheme that can clearly benefit from the concurrent operation and distributed memory provided by focal-plane architectures. 2.2 Scale Invariant Feature Transform (SIFT) The SIFT algorithm constitutes a combination of keypoint detector and corresponding feature descriptor encoding [16]. It can be broken up into four main steps: 1. Scale-space extrema detection: generation of the Gaussian and subsequent Difference-of-Gaussian (DoG) pyramids, searching for the extrema points in the DoG pyramid. 2. Accurate keypoint location in the scale space. 3. Orientation assignment to the corresponding keypoint, searching for the main orientation or main component from the gradient in its neighborhood. 4. Keypoint descriptor: construction of a vector representative of the local characteristics of the keypoint in a wider neighborhood with orientation correction. Numerous examples of SIFT implementations on different platforms have been reported: general-purpose CPU [24], GPU [25,26], FPGA [27,28], FPGA +DSP [29], specific digital co-processors [30] etc. As for the Viola-Jones, the lowest-level operation of the SIFT, namely the generation of the Gaussian pyramid, dominates the workload of the algorithm, reaching up to 90% of the whole process [31]. Figure4 Fig. 4 Gaussian pyramid with its associated DoGs
Image Feature Extraction Acceleration 115 shows anexampleof Gaussian pyramidwithits associated DoGs. It is made upof sets of filtered images (scales). Every octave starts with a half-sized downscaling of the previous octave. The filter bandwidth, σ, applied for a new scale within each octave is the one applied in the previous scale multiplied by a constant factor k:σn=kσn−1. Each octave is originally divided into an integer number of scales, s,sok=21/s. A total of s+3[16] images must be produced in the stack of blurred images for extrema detection to cover a complete octave. Once the Gaussian pyramid is built, the scales are subtracted from each other, obtaining the DoGs as an approximation to the Laplacian operator. Our objective is therefore to accelerate the SIFT Gaussian pyramid generation by means of in-pixel circuitry performing concurrent processing. For the sake of relaxation on the hardware requirements, we carried out a preliminary study to determine the number of octaves and scales to be provided by our focal-plane sensor-processor. For this study, we used a publicly available version of SIFT in MATLAB [32]. Every octave is generated from a scale of the previous octave downsized by a 1/4 factor (1/2×1/2), decreasing the pixels per octave. Therefore, the maximum potential keypoints decrease rapidly with the octaves o=0,1,2...as M×N/2o×2, with M×N being the size of the input image. Assuming a resolution of 320×240 pixels (QVGA), we obtained the keypoints for two images under many scales and rotation transformations. The reason of this moderate resolution is that the area to be allocated for in-pixel processing circuitry makes it difficult to reach larger resolutions in standard CMOS technologies with a reasonable chip size. The results for two of the applied transformations together with the test images are represented in Fig.5. Clearly, the 3 first octaves render almost all the keypoints. Concerning scales, we have two opposite contributions. On the one hand, less scales per octave means more distance between scales, causing more pixels to exceed the threshold to be sorted out as keypoints. On the other hand, reducing scales also means to diminish the total number of potential keypoints. Both combined effects make it difficult to choose a specific value for scales as in the case of the octaves. The result of the scale analysis for the same respective test images and transformations as in Fig.5is depicted Fig. 5 The number of keypoints hardly increases from the 3 first octaves. This will be the reference value for our implementation
116 J. Fernández-Berni et al. Fig. 6 The number of keypoints increases with the number of scales per octave. Trading this result for computational demand and hardware complexity leads to a targeted number of 6 scales. The same test images and transformations as in Fig.5are respectively used in Fig.6. It shows that the amount of keypoints increases monotonically with the scales. Nevertheless, increasing the scales per octave is not an option because of its corresponding computational demand and hardware complexity. Trading all these aspects, we conclude that 6 scales suffice for Gaussian pyramid generation at the focal plane. This figure coincides with the number of scales proposed in [16]. 3 Gaussian Filtering We demonstrate in this section that Gaussian filtering is the common underlying processing primitive for both the Viola-Jones and SIFT algorithms. While the role of Gaussian filtering is well defined for the latter, it is not obvious at all for the former. In order to understand the relation, we first need to establish a formal mathematical framework. Gaussian filtering is best illustrated in terms of a diffusion process. The concept of diffusion is widely applied in physics. It explains the equalization process undergone by an initially uneven concentration of a certain magnitude. A typical example is heat diffusion. Mathematically, a diffusion process can be defined by considering a function V(x,t)defined over a continuous space, in this case a plane, for every time instant. At each point x=(x1,x2), the linear diffusion of the function V(.) is described by the following well-known partial differential equation [33]: ∂V ∂t=∇·(D∇V)(3) where Dis referred to as the diffusion coefficient. If Ddoes not depend on the position: ∂V ∂t=D∇2V(4)
Image Feature Extraction Acceleration 117 and realizing the spatial Fourier transform of this equation, we obtain: ∂ˆ V(k) ∂t=−4π2D|k|2ˆ V(k)(5) wherekrepresentsthewavenumbervectorinthecontinuousFourierdomain.Finally, by solving this equation we have: ˆ V(k,t)=ˆ V(k,0)e−4π2Dt|k|2(6) where ˆ V(k,t)is the spatial Fourier transform of the function V(.) at time instant t and ˆ V(k,0)is the spatial Fourier transform of the function V(.) at time t=0, that is, just before starting the diffusion. Equation (6) can be written as a transfer function: ˆ G(k,t)=ˆ V(k,t) ˆ V(k,0)=e−4π2Dt|k|2(7) which, by defining σ=√2Dt, is transformed into: ˆ G(k,σ)=e−2π2σ2|k|2(8) This transfer function corresponds to the Fourier transform of a spatial Gaussian filter of the form: G(x,σ)=1 2πσ2e−|x|2 2σ2(9) and therefore the diffusion process is equivalent to the convolution expressed by the following equation: V(x,t)=1 2πσ2e−|x|2 2σ2∗V(x,0)(10) We can see that a diffusion process intrinsically entails a spatial Gaussian filtering which takes place along time. The width of the filter is determined by the time the diffusion is permitted to evolve: the longer the diffusion time, t, the larger the width of the corresponding filter, σ. This means that, ideally, any width is possible provided that a sufficiently fine temporal control is available. From the point of view of the Fourier domain, we can define the diffusion as an isotropic lowpass filter whose bandwidth is controlled by t. The longer t, the narrower the bandwidth of the filter around the dc component (Fig.7). Eventually, for t→∞, all the spatial frequencies but the dc component are removed. Furthermore, this dc component is completely unaffected by the diffusion, that is, ˆ G(0,t)=1∀t. It is just this characteristic of the Gaussian filtering what constitutes the missing link with the computation of the integral image. When discretized and applied to a set of pixels, this property says that a progressive Gaussian filtering eventually leads to the average of the values the
124 J. Fernández-Berni et al. Fig. 13 Elementary sensing-processing pixel of the Viola-Jones focal-plane accelerator chip Fig. 14 Photograph of the Viola-Jones focal-plane accelerator chip and the FPGA-based system where it has been integrated A direct transformation of the simplified scheme of Fig.9into a stacked structure is possible, as shown in Fig.16. The top tier would exclusively include photo-diodes and some readout circuitry whereas the bottom tier would implement the reconfigurable diffusion network. The interconnection between both tiers would be carried out by the so-called Through-Silicon Vias (TSVs). This structure keeps maximum parallelism at processing while drastically increasing resolution and sensitivity.
Image Feature Extraction Acceleration 125 Fig. 15 Example of on-chip integral image computation and comparison with the ideal case Fig. 16 Transformation of the simplified scheme of Fig.9into a stacked structure 5.2 SIFT Focal-Plane Accelerator Chip The SIFT accelerator chip presents a similar floorplan to that of the Viola-Jones prototype. However, its elementary sensing-processing cell significantly differs. A simplified scheme is depicted in Fig.17. The constituent blocks are mainly four photo-diodes,thelocalanalogmemories(LAMs), thecomparatorforA/Dconversion and the switched capacitor network. During the acquisition stage, the photo-diodes, the capacitor C and the LAMs work together to implement a technique known as correlated double sampling [38] that improves the image quality. The LAMs jointly with the diffussion network carry out the Gaussian filtering. The capacitor C and the inverter make up the A/D comparator that would drive a register in the bottom
126 J. Fernández-Berni et al. Fig. 17 Simplified scheme of the elementary sensing-processing cell designed for the SIFT focal-plane accelerator chip tier by a TSV on CMOS-3D technologies, or peripheral circuits on conventional CMOS. Every cell is 4-connected to its closest neighbors in the North, South, East and West directions. Given that every cell includes four photo-diodes, 4 internal and 8 peripheral interconnections are required. Two microphotographs of the chip together with the different components of the camera module built for test purposes are reproduced in Fig.18. This prototype, also manufactured in a standard 0.18µm CMOS process, features a resolution of 176×120 pixels and can generate 120 Gaussian pyramids per second with a power consumption of 70mW. One of the operations required for Gaussian pyramid generation is downscaling. As previously commented, the 3 first octaves are the most important ones in the performance of SIFT. This corresponds with downscaling at ratios 4:1 and 16:1 for octaves 2 and 3, respectively. The chip includes the hardware required to implement this spatial resolution reduction. An example is shown in Fig.19. The images to the left are represented with the same sizes in order to visually highlight the effects of downscaling. Another example, in this case of on-chip Gaussian filtering, is shown in Fig.20. The upper left image constitutes the input whereas the three remaining images, from left to right and top to down, correspond to σ=1.77, (clock cyles n=19), σ=2.17 (n=29), and σ=2.51 (n=39). More details about the performance of this chip can be found in [39]. This chip was conceived, from the very beginning, for implementation in 3-D integration technologies [40]. Unfortunately, these technologies are not mature enough yet for reliable fabrication. Manufacturing costs of prototypes are also extremely high for the time being, with long turnarounds, exceeding 1 year. In these circumstances, we were forced to redistribute the original two-tier circuit layout devised for a CMOS 3-D stack in order to fit it into a conventional planar CMOS technology. The result is depicted in Fig.21.
Image Feature Extraction Acceleration 127 Fig. 18 Photograph of the SIFT focal-plane accelerator chip together with the camera module where it has been integrated O1 O2 O3 Fig.19 On-chipimage resolution reductionby 4:1 and16:1 as part of the calculation ofthe pyramid octaves
128 J. Fernández-Berni et al. Fig. 20 Different snapshots of on-chip Gaussian pyramid Fig. 21 Redistribution of circuits for Gaussian pyramid generation when mapping the original CMOS 3D-based architecture onto a conventional planar CMOS technology 5.3 Performance Comparison Comparing the performance of the implemented prototypes with state-of-the-art focal-plane accelerator chips is not straightforward since every realization addresses a different functionality. As an example, we have included the most significant characteristics of our prototypes together with two recently reported focal-plane sensorprocessor chips in Table1. The Viola-Jones chip embeds extra functionalities in addition to the computation of the integral image [41] while featuring the largest resolution and the smallest pixel pitch, with a cost in terms of a reduced fill factor and increased energy consumption. Concerning the SIFT chip, one of the reasons of the energy overhead is the inherent high number of A/D conversions of the whole Gaussian pyramid plus the input scene, which amounts to 40 A/D conversions of the entire pixel array. Still, the acceleration at the focal plane provided by this chip
Image Feature Extraction Acceleration 129 Table 1 Comparison of the implemented prototypes with state-of-the-art focal-plane sensorprocessor chips Reference Ref. [42] Ref. [43]Viola-Jones chip SIFT chip Function Edge filtering, tracking, HDR 2-D optic flow estimation HDR, integral image, Gaussian filtering, programmable pixelation Gaussian pyramid Tech. (µm) 0.18 0.18 0.18 0.18 Supply (V) 0.5 3.3 1.8 1.8 Resolution 64×64 64×64 320×240 176×120 Pixel pitch (µm) 20 28.8 19.6 44 Fill factor (%) 32.4 18.32 5.4 10.25 Dyn. range (dB) 105 −102 − Power consumption 1.25 0.89 23.9 26.5 (nW/px·frame) pays off when comparing with more conventional solutions, as shown in Table2. The power consumption of conventional CMOS imagers from Omnivision [44] featuring the image resolution tackled by the corresponding processor is incorporated in each of the entries related to conventional solutions. We have not accounted for accesses to external memories first because such costs would also be present if our chip were part of a complete hardware platform for a particular application; and Table 2 Comparison of the SIFT focal-plane accelerator chip with conventional solutions Hardware solution Functionality Energy/frame Energy/pixel Mpx/s SIFT chip Gaussian 176×120 resol. 26.5nJ/px 2.64 180nm CMOS pyramid 70mW @ 8ms 0.56mJ/frame Ref. [45] Gaussian VGA resol. 15.5µJ/px 2.26 OV9655 + Core-i7 pyramid 90mW @ 30fps +35W @ 136ms 4.8J/frame Ref. [46] Gaussian VGA resolution 240µJ/px 0.15 OV9655 + Core-2-Duo pyramid 90mW +35W @2.1s 73.7J/frame Ref. [47] Gaussian 350×256 resol. 4.4µJ/px 0.91 OV6922 +pyramid 30mW +4W Qualcomm @ 98.5ms Snapdragon S4 0.4J/frame
130 J. Fernández-Berni et al. second because they are hardly predictable even with memory models. The energy cost of our chip outperforms that of an imager +conventional processor unit—even a low-power unit—in three orders of magnitude with similar processing speed. This leads to a combined speed-power figure of merit which makes our chip outperform conventional solutions in the range of three to six orders of magnitude. 6 Conclusions and Future Work Focal-plane sensing-processing constitutes an architectural approach that can boost the performance of vision algorithms running on embedded systems. Specifically, early vision stages can greatly benefit from focal-plane acceleration by exploiting the distributed memory and concurrent processing in 2-D arrays of sensing-processing pixels. This chapter provides an overview of the fundamental concepts driving the designandimplementationof twofocal-planeacceleratorchipstailored,respectively, for the Viola-Jones and the SIFT algorithms. These are the first steps within a longterm research framework aiming at achieving image sensors capable of simultaneouslyrendering high-resolution high-quality rawimagesandvaluablepre-processing at ultra-low energy cost. The future work will be singularly biased by the availability of monolithic sensing-processing stacks. 3-D technologies will remove the tradeoff arising when it comes to allocating silicon area for sensors and processors on the same plane. High sensitivity and high processing parallelization will be compatible on the same chip. 3-D stacks will also foster alternative ways of making the most of vertical across-chip interconnections, from transistor level up to system architecture. In summary, 3-D integration technologies are the natural solution to develop feature extractors with low power budget without degrading image quality. Our prototypes on planar processes already consider future migration to these technologies, and this will continue to be a compulsory requirement of forthcoming designs. Acknowledgments This work has been funded by: Spanish Government through projects TEC2012-38921-C02MINECO (European Region DevelopmentFund, ERDF/FEDER), IPT-20111625-430000 MINECO and IPC-20111009 CDTI (ERDF/FEDER); Junta de Andalucía through project TIC 2338-2013 CEICE; Xunta de Galicia through projects EM2013/038, AE CITIUS (CN2012/151,ERDF/FEDER),and GPC2013/040 ERDF/FEDER;Officeof Naval Research(USA) through grant N000141410355. References 1. Market analysis, embedded vision alliance. http://www.embedded-vision.com/industryanalysis/market-analysis 2. Kolsch, M., Butner, S.: Hardware considerations for embedded vision systems. In: Kisacanin, B., Bhattacharyya, S.S., Chai, S. (eds.) Embedded Computer Vision, Advances in Pattern Recognition Series, pp. 3–26. Springer, London (2009) 3. Open computing language. https://www.khronos.org/opencl/
Image Feature Extraction Acceleration 131 4. OpenVX: portable, power-efficient vision processing. https://www.khronos.org/openvx/ 5. Open source computer vision. http://opencv.org/ 6. Compute unified device architecture. http://www.nvidia.com/object/cuda_home_new.html 7. Bailey, D.: Design for Embedded Image Processing on FPGAs. Wiley, Singapore (2011) 8. Kim, J., Rajkumar, R., Kato, S.: Towards adaptive GPU resource management for embedded real-time systems. ACM SIGBED Rev. 10, 14–17 (2013) 9. Tusch, M.: Harnessing hardware accelerators to move from algorithms to embedded vision. In: Embedded Vision Summit. Embedded Vision Alliance, Boston (2012) 10. Horowitz, M.: Computing’s energy problem (and what we can do about it). In: International Solid-State Circuits Conference (ISSCC), pp. 10–14. San Francisco (2014) 11. Wilkes, M.V.: The memory gap and the future of high performance memories. SIGARCH Comput. Archit. News 29, 2–7 (2001) 12. Zarándy, A. (ed.): Focal-Plane Sensor-Processor Chips. Springer, New York (2011) 13. Campardo, G., Ripamonti, G., Micheloni, R.: Scanning the issue: 3-D integration technologies. Proc. IEEE 97, 5–8 (2009) 14. Courtland, R.: ICs grow up. IEEE Spectr. 49, 33–35 (2012) 15. Viola, P., Jones, M.: Robust real-time face detection. Int. J. Comput. Vis. 57, 137–154 (2004) 16. Lowe, D.: Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 60, 91–110 (2004) 17. Jia, H., Zhang, Y., Wang, W., Xu, J.: Accelerating Viola-Jones face detection algorithm on GPUs. In: IEEE International Conference on Embedded Software and Systems, pp. 396–403. Liverpool (2012) 18. Masek, J., Burget, R., Uher, V., Guney, S.: Speeding up Viola-Jones algorithm using multi-core GPU implementation. In: IEEE International Conference on Telecommunications and Signal Processing (TSP), pp. 808–812. Rome (2013) 19. Acasandrei, L., Barriga A.: FPGA implementation of an embedded face detection system based on LEON3. In: International Conference on Image Processing, Computer Vision, and Pattern Recognition. Las Vegas (2012) 20. Ouyang, P., Yin, S., Zhang, Y., Liu, L., Wei, S.: A fast integral image computing hardware architecture with high power and area efficiency. IEEE Trans. Circuits Syst. II(62), 75–79 (2015) 21. Kyrkou, C., Theocharides, T.: A flexible parallel hardware architecture for adaboost-based real-time object detection. IEEE Trans. Very Large Scale Integr. VLSI Syst. 19, 1034–1047 (2011) 22. Gschwandtner, M., Uhl, A., Unterweger, A.: Speeding up object detection fast resizing in the integral image domain. Technical Report, University of Salzburg (2014) 23. de la Cruz, J.A.: Field-programmable gate array implementation of a scalable integral image architecture based on systolic arrays. Master Thesis, Utah State University (2011) 24. Kumar, G., Prasad, G., Mamatha, G.: Automatic object searching system based on real time SIFT algorithm. In: IEEE International Conference on Communication Control and Computing Technologies, pp. 617–622. Ramanathapuram (2010) 25. Cornelis,N.,VanGool,L.:Fastscale invariantfeaturedetectionandmatchingonprogrammable graphics hardware. In: IEEE Computer Vision and Pattern Recognition Workshops, pp. 1–8. Anchorage (2008) 26. Cohen, B., Byrne, J.: Inertial aided SIFT for time to collision estimation. In: IEEE International Conference on Robotics and Automation, pp. 1613–1614. Kobe (2009) 27. Cabani,C., MacLean, W.J.:A proposed pipelined-architecture for FPGA-basedaffine-invariant feature detectors. In: IEEE Computer Vision and Pattern Recognition Workshops, pp. 121. New York (2006) 28. Nobre, H., Kim, H.Y.: Automatic VHDL generation for solving rotation and scale-invariant template matching in FPGA. In: IEEE Southern Conference on Programmable Logic, pp. 21–26. Sao Carlos (2009) 29. Song, H., Xiao, H., He, W., Wen, F., Yuan, K.: A fast stereovision measurement algorithm based on SIFT keypoints for mobile robot. In: IEEE International Conference on Mechatronics and Automation (ICMA), pp. 1743–1748. Takamatsu (2013)
132 J. Fernández-Berni et al. 30. Gao, H., Yin, S., Ouyang, P., Liu, L., Wei, S.: Scale invariant feature transform algorithm based on a reconfigurable architecture system. In: 8th IEEE International Conference on Computing Technology and Information Management (ICCM), pp. 759–762. Seoul (2012) 31. Noguchi, H., Guangji H., Terachi, Y., Kamino, T., Kawaguchi, H., Yoshimoto, M.: Fast and low-memory-bandwidth architecture of SIFT descriptor generation with scalability on speed and accuracy for VGA video. In: IEEE International Conference on Field Programmable Logic and Applications (FPL), pp. 608–611. Milano (2010) 32. Andrea Vedaldi’s implementation of the SIFT detector and descriptor. http://www.robots.ox. ac.uk/vedaldi/code/sift.html 33. Jahne, B.: Multiresolution signal representation. In: Jahne, B., Haubecker, H., Geibler, P. (eds.) HandbookofComputerVisionandApplications(volume2).AcademicPress,SanDiego(1999) 34. Fernández-Berni, J., Carmona-Galán, R., del Río, R., Rodríguez-Vázquez, A.: Bottom-up performance analysis of focal-plane mixed-signal hardware for Viola-Jones early vision tasks. Int. J. Circuit Theory Appl. (2014). doi:10.1002/cta.1996 35. Fernández-Berni, J., Carmona-Galán, R., Carranza-González, L.: FLIP-Q: a QCIF resolution focal-plane array for low-power image processing. IEEE J. Solid-State Circuits 46, 669–680 (2011) 36. Allen, P.E.: Switched Capacitor Circuits. Springer, New York (1984) 37. Suárez,M.,Brea, V.M., Cabello,D., Pozas-Flores, F.,Carmona-Galán,R., Rodríguez-Vázquez, A.: Switched-capacitor networks for scale-space generation. In: IEEE European Conference on Circuit Theory and Design (ECCTD), pp. 190–193. Linkoping (2011) 38. Enz, C.C., Temes, G.C.: Circuit techniques for reducing the effects of op-amp imperfections: autozeroing, correlated double sampling, and chopper stabilization. Proc. IEEE 84, 1584–1614 (1996) 39. Suárez, M., Brea, V.M., Fernández-Berni, J., Carmona-Galán, R., Cabello, D., RodríguezVázquez, A.: A 26.5 nJ/px 2.64 Mpx/s CMOS vision sensor for gaussian pyramid extraction. In: IEEE European Solid-State Circuits Conference (ESSCIRC), pp. 311–314. Venice (2014) 40. Suárez, M., Brea, V.M., Fernández-Berni, J., Carmona-Galán, R., Liñán, G., Cabello, D., Rodríguez-Vázquez, A.: CMOS-3-D smart imager architectures for feature detection. IEEE J. Emerg. Sel. Top. Circuits Syst. 2, 723–736 (2012) 41. Fernández-Berni, J., Carmona-Galán, R., del Río, R., Kleihorst, R., Philips, W., R., RodríguezVázquez, A.: Focal-plane sensing-processing: a power-efficient approach for the implementation of privacy-aware networked visual sensors. Sensors 14, 15203–15226 (2014) 42. Yin, C., Hsieh, C.: A 0.5V 34.4μW 14.28kfps 105dB smart image sensor with array-level analog signal processing. In: IEEE Asian Solid-State Circuits Conference (ASSCC), pp. 97– 100. Singapore (2013) 43. Park, S., Cho, J., Lee, K., Yoon, E.: 243.3pJ/pixel bio-inspired time-stamp-based 2D optic flow sensor for artificial compound eyes. In: IEEE International Solid-State Circuits Conference (ISSCC), pp. 126–127. San Francisco (2014) 44. Omnivision image sensors. http://www.ovt.com/products/ 45. Murphy, M., Keutzer, K., Wang, P.: Image feature extraction for mobile processors. In: IEEE International Symposium on Workload Characterization (IISWC), pp. 138–147. Austin (2009) 46. Huang, F., Huang, S., Ker, J., Chen, Y.: High-performance SIFT hardware accelerator for real-time image feature extraction. IEEE Trans. Circuits Syst. Video Technol. 22, 340–351 (2012) 47. Wang, G., Rister, B., Cavallaro, J.: Workload analysis and efficient openCL-based implementationof SIFTalgorithm ona smartphone.In: IEEEGlobal Conferenceon Signaland Information Processing, pp. 759–762. Austin (2013)