Design of a programmable mixed-signal CMOS image-processing chip in 0.8 /spl mu/m CMOS
Abstract
An operational vision-chip prototype with a wide-range of potential applications in artificial-vision systems is presented. Its functionality includes concurrent image-transduction, programmable image-processing, image-storage, and algorithmic control over a network of 20/spl times/22 identical cells. The prototype has been designed and manufactured in 0.8 /spl mu/m CMOS standard technology, and has a total area of 30 mm/sup 2/. Experimental results are reported.
Full text
I997 IEEE International Symposium on Circuits and Systems, June 9-12,1997, Hong Kong Design of A Programmable Mixed-Signal CMOS Image-Processing Chip in 0.8~” CMOS A. RodriguezVa‘zquez, S. Espejo, R. Dominguez-Castro and R. Carmona Instituto de Microelectr6nica de Sevilla-Universidad de Sevilla Edificio CICA, CRarfia s/n, 41012-Sevilla, SPAIN FAX:: 34 5 4231832; Phone:: 34 5 4239923 email: angel @cnm.us.es ABSTRACT An operational vision-chip prototype with a wide-range of potential applications in artificial-vision systems is presented. Its functionality includes concurrent image-transduction, programmable image-processing, image-storage, and algorithmic control over a network of 20 x 22 identical cells. The prototype has been designed and manufactured in 0.8pm CMOS standard technology, and has a total area of 3b2. Experimental results are reported. 1. INTRODUCTION Conventional image-processing systems use a CCD camera for parallel acquisition of the input image, and serial transmission of the digitalized image to a separate processing element. It results in huge data rates which conventional computers are not capable to analyze in real-time. For instance, a 3-colour@512 x 512 pixel camera delivers about F million bytehecond, where F is the frame rate. Such a huge rate may be managed by conventional computers for operations such as auto-focus, image stabilization, control of the luminance/chrominance, etc. However, real-time completion of the spatio-temporal operations required for understanding images requires bulky and sophisticated processors. In contrast to this, the smallest insects, albeit equipped with really tiny brains, are capable to analyze complex time-varying scenes in real-time [ 11. This contrast between artificial and “natural” vision systems is due to the inherent parallelism of the latter. Particularly, the cells of the natural retina combine photo-transduction and collective parallel processing for the realization of low-level image processing operations (light adaptation, feature extraction, motion analysis, etc.) concurrently to the acquisition of the image [ 11. Inspired on this, new generations of image processing systems have addressed the incorporation of distributed parallel processing already at the plane of the image sensor. One possible strategy uses flip-chip bonding of separated imagers and 2-D processors. Other possibility is to incorporate the sensory and the processing circuitry on the same semiconductor substrate [2]. CMOS technologies offer unique features for this latter type of chips due to the availability of good CMOS photo-transduction devices and the possibility to realize linear and nonlinear processing functions with simple CMOS circuitry [3]. 0-7803-3583-X/97 $10.00 01997 IEEE 725 A number of CMOS retinas have been previously reported in literature [2]. In many cases their development have emphasized light adaptation, i.e. the capability to adapt the response of the 2-D optical sensor to the lighting conditions of the incoming image, while image processing has remained secondary. Some circuits which incorporate processing capabilities are intended for fixed processing function [2][4]. On the other hand, the programmable circuits found in literature have neither accurate standard control interfaces nor the capability of flexible operation. Besides, their development lack systematics because there is not a clearly defined design path from the top processing algorithm down to the circuit design [5]. To establish such a path requires first to capture the retina processing function into signal processing algorithms. Different remarkable efforts have beer1 undertaken in this direction. In particular, researchers of the University of California at Berkeley and the Hungarian Academy o Sciences have recently set-up a powerful methodological framework for the systematic formulation of the low-level image processing function of the retina [6][7][8]. It is based on the observation that most image processing operations can be formulated as well-defined tasks on signal values placed over regular 2-D spatial distributions, and with direct interactions among signals limited to local receptive fields. Consequently, they are directly mappable onto Cellular Nonlinear (or Neural) Networks (CNNs), which are arrays of nonlinear dynamic analog processing units (cells), arranged on regular grids where direct interactions among cells are limited to finite local neighborhoods. By enabling these interactions to be programmable, and incorporating the possibility to store a program, and to sequentially realize it over an 2 x N - D image memory, the CNN have evolved into the CNN Universal Machine (CNN-UM) [SI. Since CNNs realize the vast majority of image processing tasks through proper selection of the interaction strengths and/or task sequencing [9], the CNN-UM can be considered as a general-purpose image processing computer on a chip. This paper presents a realization of the CNN-UM in 0.8pm CMOS single-poly double-metal technology. It is intended to the processing of binary images and incorporate the features of 2-D signal acquisition, light adaptation and programmable focal plane array processing. It is also capable to internally store intermediate images and processing coefficients, and using them in any order and any number of times. Consequently, the chip is capable to operate as
powerful front-end for the realization of simple and medium-complexity artificial vision tasks, including sequential and bifurcated-flow algorithms. Also, although the fundamental processing function is analog and continuous-time, as in other CNN chips [lO][ll], the interface of our prototype is completely digital, making it extremely easy to control with conventional computing systems. 2. SYSTEM ARCHITECTURE As shown in Fig.1, the chip contains an array of 20 x 22 identical cells and some peripheral U0 and control circuitry. Although its main internal processing-functions are analog, the external interface of the chip is completely digital, allowing a straightforward integration in digitally controlled, higher-level computing or control systems. Each cell in the array performs image acquisition, storage and processing functions (Fig.1). Most of the cell area (80%) is dedicated to the programmable analog processing circuitry, based on the Cellular Neural Network (CNN) computing paradigm [7][8]. Each cell has an input variable (U&, an output variable 0,) and a state variable (x,). The output is a saturation-type nonlinear version of the state variable. Every cell performs a nonlinear transient evolution driven by weighted summations zacdyd+ zbcpd of the output and input levels of its 9 nearest neighbors (including itself and those in vertical, horizontal, and both diagonal directions). The 18 weighting parameters aCd and b,, plus an additional offset term d,, which are invariant from cell to cell, determine the processing function of the array. All of these coefficients can be programmed in this chip with a dynamic range of 7 bits plus sign (8 bits plus sign for the offset term). A digital RAM memory located in the periphery of the array provides on-chip storage for eight complete sets of coefficients, which can later on be selected through an external control. Four pixel-memories in each cell provide on-chip storage for four complete images, allowing the realization of complex, sequential andor bifurcated image-processing algorithms without the need of intermediate image loading/downloading processes. These pixel-memories centralize the information flow within the cells (Fig. 1). A programmable two-input boolean operator is also included in every cell, providing the additional capability of parallel logical operations among images. Since spatial boundary conditions are important for processing, the cell array is surrounded by a ring of border cells with programmable output variable. Fig.2 illustrates the analog processing circuitry of each cell. The 18 weighting contributions are implemented using only 9 multipliers. For this purpose, one of the weighted summations is previously computed on-chip by programming the 9 multipliers of every cell with coefficients b,, and introducing the input value U, of each cell as the initial condition of its (disabled) integrator. The resulting incoming contribution bcdud for each cell is then stored in an analog memory (included in its circuitry), and thereafter added to the incoming signals. After this step, the 9 multipliers are reprogrammed with coefficients aCd , and the integrators are enabled after setting their initial conditions x(0). Apart from halving the number of multipliers, this strategy can be shown to provide a functional cancellation of the output-referred offset of the multipliers. The programmable offset-term contribution is generated by an identical multiplier driven by a reference signal. 3. CIRCUIT DESIGN The design of the circuitry has involved extensive structural and parametric optimization to reduce the silicon area while keeping the accuracy in the analog operations. This has affected the choice of the processing algorithm itself [10][12], the choice of the circuit structures for the interconnection synapses [3], and the optimization of the sizes on the basis of the statistical modeling of mismatch [13]. Fig.3 shows the circuit implementation of the most relevant blocks of the analog processing circuitry and the optical interface. The multiplier [ 141 employs four transistors operating in their ohmic region, and two source followers as weight-signal buffers, in a fully differential architecture. The integrator (Fig.3b) consists of two simple current conveyors (common-mode feedback circuitry not shown) and two capacitors. These capacitors are in fact obtained from the gate capacitance of the multipliers connected to the integrator output, avoiding the use of specific silicon area for this purpose. Remind that multiplier transistors operate in strong inversion and ohmic region, thus their gate capacitance is fairly linear. Also, the capacitance area-density is maximal. The nonlinear resistor is shown in Fig.3~. Finally, Fig.3d and e show the fundamental blocks of the optical interface: the sensor and the interface circuitry respectively. The photosensor employs a floating-base vertical PNP bipolar device as sensitive device, and an additional vertical PNP in a Darlington configuration as current amplifier. The interface circuitry replicates the photogenerated current levels to obtain the mean value of the distribution TPh and substracts it from local values. Node SUM in Fig.3e is common to all cells. Thus, the interface circuitry constitutes a global adaptive scheme which ensures appropriate contrast levels by shifting the observed scene to obtain a zero-mean distribution of pixel values. Optional extemal circuitry can be employed to control the mean of the distribution when needed, for instance for highly regular images with dominant background. Images can also be electrically loaded and downloaded through a 22 lines bus, on a row by row basis. The inherent non-linearity of the (resistively loaded) source follower employed in the multipliers is not relevant because the analog weight signals are generated from their digitally coded values using adaptive loops comprising the same multiplier architecture and a linear D/A converter, thus ensuring linearity between the effective weight value and its digital codification. The 10 weight-signal adaptation-stages (9 programmed coefficients plus the offset term) are located in the periphery of the cell array, adjacent to their associated digital memory blocks (Fig.1). The use of internal analog 726
weight signals, and external digital codification of their values simplifies the external interface of the chip (making it digital) and facilitates the on-chip storage of weight values while maintaining the number of weight-signals (metal lines) to be transmitted to the cells in the array within reasonable limits. 4. RESULTS The chip has been fabricated in a standard 0.8pm n-well CMOS technology (1-poly, two metals). The second metal layer is completely dedicated to global (common to all cells) signals (analog weight voltages, power supplies, control and I/O signals, etc.). These metal lines run over the cell circuitry for higher area efficiency. For this purpose, local (cell-level) routing avoids the use of the metal-2 layer. The possible effects of the metal-2 lines running over analog-operating transistors have been avoided trough extensive use of symmetry in the layout and the differential architecture of the analog circuitry. Figures 4a,b and c contain static measurements of the fundamental analog blocks: the multiplier (Fig.4a), the V-I Characteristic at the iow-impedance input of the current conveyors (Fig.4b), and the I-V characteristic of the nonlinear resistor (Fig.4~). Fig.4d summarizes measurements taken from 10 multiplier samples, showing maximum deviation from linearity (below 0.4%) versus output-current offset (below 1.6%), both relative to full scale. At a system level, the chip has been successfully applied to many basic image processing functions (filtering, borders, peeling, etc.) as well as to more complex tasks based on concatenated image processing, like motion detection, motion estimation (speed and direction), and texture recognition [ 151. Fig.5 illustrates the image processing algorithm employed for motion detection in a particular direction and speed range [ 151. Table I summarizes the most relevant characteristics of the chip, and Fig.6 contains its microphotograph. Image-processing time varies with the specific application (i.e., the coefficient values), and ranges from about three analog time-constants (see Table I) to about 20. 5. CONCLUSIONS We have briefly described the design of a fully programmable vision-chip with a wide range of potential applications. The processing function is based on the Cellular Neural Network paradigm, while the phototransduction relies on vertical BJTs available on standard CMOS technologies and includes an automatic contrast-centering circuitry. Additional features like internal image memories, algorithmic control, and programmable logic operators provide a high versatility for simple and medium complexity artificial-vision applications. Although the internal operation is fundamentally analog, the interface of the prototype is completely digital, making it directly controllable by conventional computing devices. The achieved cell density is of 27.5cell&nm2 -- much larger than achieved in other programmable CNN implementations [ 111 with similar analog accuracy levels. 6. REFERENCES M.M. Gupta and G.K. Knopf “Neuro-Vision Systems: A Tutorial”. Neuro-Vision Systems: Principles and Applications, New York, EEE Press 1994. C.Koch and H.Li (editors): Vision Chips: Implementing Vision Algorithms using Analog VLSI Circuits. New York, IEEE Press 1994. S. Espejo: VLSI Design and Modeling of Cellular Neural Networks. Ph.D. dissertation, University of Seville, 1994. S. Espejo, A. Ilodriguez-Vizquez, R. Dom’nguez-Castro, J.L. Huertas, and E. Sinchez-Sinencio, “Smart-Pixel Cellular Neural Networks in Analog Current-Mode CMOS Technology”, IEEE Journal ofSolid-State Circuits, Vol. 29, pp. 895-905, August 1994. K. Kyuma, E. Lange, J. Ohta, A. Hermanns, B. Banish and M. Oita, “Artificial Retinas - Fast, Versatile Image Processors”, Nature, Vol. 312, pp. 197-198, November 1994. F. Werblin, A. Jacobs and J. Teeters, “The Computational Eye”, IEEE Spectrum, pp. 30-31, May 1996. L.O. Chua and T. Roska, “The CNN Paradigm”. IEEE Trans. Circuits and Systemd, Vol. CAS-40, pp 147-156, March 1993. T. Roska and L.O. Chua, “The CNN Universal Machine: An Analogic Array Computer”, IEEE Transactions on Circuits and Systems-11, Vol., 40, No.-3, March 1993. T. Roska and L. Kkk, Analogic CNN Program Libraiy, Analogical and Neural Computing Laboratory Memo. DNS-5-1994, Budapest, 1994. [lo] A. Rodriguez-Vizquez, S. Espejo, R. Dom’nguez-Castro, J.L. Huertas, and E. Sinchez-Sinencio, “Current-Mode Techniques for the Implementation of Continuousand Discrete-Time Cellular Neural Networks”, IEEE Trans. on Circuits and Systems II, Vol40, No. 3, pp 132-146, March 1993. [ 111 P. Kinget and M. Steyaert: “An Analog Parallel array Processor for Real-Time Sensor Signal Processing”, 1996 Int. Solid State Circuits Conference, paper 6.1. [ 121 S. Espejo, R. Carmona, R. Dom’nguez-Castro and A. Rodriguez-Vizquez, “A VLSI-oriented Continuous-Time CNN Model”, Int. J. Circuit Theory and Applicationsl, Vol. 24, pp. 341-356, May 1996. [13] M.J.M Pelgram, A.C.J. Duinmaijer and A.P.G. Webers: “Matching Properties of MOS Transistors”. IEEE J. Solid-state Circuits, Vo1.24, pp 1433-1440, October 1989. [14] B. Song: “CMOS RF Circuits for Data Communication Applications”. IEEE J. Solid-state Circuits, Vo1.21, pp 310-317, April 1986. [15] P. Foldesy, A. Zarandy, P. Szolgay and T. Roska: Measurement Results of the 20x 22 CNNUM Chip, Anlogic and Neural Computing Laboratory, Hungarian Academy of Sciences, February 1996. 727
1 U0 cells Power supply ...................... 4.5 to 5.5 V Power dissipation ............... 1.1 W @ 5V Number of cells Cell size ....... Photosensor current (Darl.). - 0.8 FA @ 100 Package .............................. PGA-120 Weights range ..................... 7 bits + sign Offset term range ................ 8 bits + sign ' Weights accuracy ................ -1% Programmable analog coeff. 19 Programmable logic values. 4 Photosensor & Threshold I I I I I I1 I I Ill I I I I I I I I1 1111 I I 4f . Pixel Memories CIRCUTI'RY Programmable Logic Gate Y ' lux Fig.1: System block-diagram and elementary cell architecture. Table I: System specifications and features. To Neighbors From Neighbors - ..-, ,.. , ,______., .. ............. I p; 'd ,----- :e j L---J -- I&] I* i ........... .. Fig5 One example of system level experimental Fig.2: Analog Processing Circuitry Diagram. detection and estimation [E]. result: motion Fig.3: Basic circuit blocks: a) multiplier, b) integrator, c) nonlinear resistor, d) photosensor and e) adaptive threshold circuitry. Zfl ,-,/ c 4 b) I, (A) C) vxp-vxIy (V) 4 output ogset (RFS) a) vxp-vxIy (V) Fig.4: Electrical measurements of basic circuit blocks: a) multiplier, b) CCIP x-input V-I curve, c) non-linear resistor V-I curve, d) measured errors from 10 multiplier samples, relative to full scale (RFS). 728