scieee AI-readable full text Open interactive document viewer

Hardware Demonstrators of Oscillatory Neural Networks: A Versatile Approach to Brain-Inspired Computing, from Embedded Systems to Phase-Transition Materials

Jiménez Través, Manuel

Abstract

The ever-increasing amount of information that computing platforms are required to handle presents a tremendous challenge, particularly concerning power consumption, data rate, and processing speed. The separation of processing and memory constitutes a significant efficiency bottleneck in well-established systems based on the von Neumann architecture. For example, the rise of Edge Computing and the Internet of Things underscores the need for more efficient and decentralized computing architectures. As the demand for real-time processing, reduced latency, and low power consumption intensifies, traditional CMOS-based architectures are approaching their inherent limits, and innovative perspectives at different levels are gaining considerable interest. Research covering from materials and devices to architectures and systems, as well as novel computing paradigms, aims to provide medium-term solutions. Specifically, computation that resorts to dynamical systems, such as coupled electrical oscillators, has emerged as a fast and energy-efficient computing paradigm. In this approach, the problem solution is encoded within the system energy function and achieved through its inherent dynamics in a collective-intelligent, parallel manner. Following a brain-inspired approach to this computing paradigm, the Oscillatory Neural Networks (ONNs) consist of many oscillator (neuron) circuits interconnected by electrical elements (synapses). Recent advances in materials technology and the emergence of novel devices have enabled the implementation of compact oscillators with very low power consumption. In particular, oscillators built with the phase-transition material vanadium dioxide (VO2) have received significant interest, and their exploitation in ONNs has become an active area of research. This Doctoral Thesis presents the development of different hardware demonstrators of ONNs and their application as Associative Memories for pattern recognition tasks and as Ising Machines to solve combinatorial optimization problems. First, a fully digital ONN was designed as a proof of concept of the oscillatory-based computing paradigm. It was implemented in FPGA technology and two prototypes were experimentally validated: a digit-recognition application and a responsive obstacle-avoidance system in a mobile robot. Later, an analog CMOS-ONN ASIC was designed and fabricated in a commercial TSMC 65nm technology, integrating a CMOS emulator of the electrical behavior of the VO2 material. This ASIC enabled an early evaluation of the VO2-based ONN response. Finally, demonstrators built with fabricated VO2 oscillators and discrete components were evaluated, showcasing the utility of VO2-based ONNs to solve computationally hard problems in very few cycles. Alongside the development of these demonstrators, a comprehensive review of suitable learning rules was conducted to enhance the capabilities of the ONN as Associative Memory. This led to the conception of a novel iterative algorithm, IRPUSH, which outperforms all the existing learning rules compatible with ONN. It ensures the learning of highly correlated patterns and achieved high retrieval accuracy in pattern recognition tasks, even with reduced weight precision, showing a minimal loss of accuracy. These efforts have significantly contributed to validating and deepening our understanding of the oscillatory-computing paradigm using VO2-based ONNs. The insights derived from the development of the demonstrators, regarding the fast and parallel convergence of the ONNs to the problem solution, have proven the potential for energy-efficient computation. Furthermore, a solid foundation has been established on how electrical ONNs should be operated and configured considering the specific application while exploring methods to overcome the intrinsic constraints and limitations of such a complex dynamical system.

Full text

IMS -cnm Instituto de Microelectrónica de Sevilla Hardware Demonstrators of Oscillatory Neural Networks A Versatile Approach to Brain-Inspired Computing, from Embedded Systems to Phase-Transition Materials Manuel Jiménez Través 2024 Ph.D. Dissertation Hardware Demonstrators of Oscillatory Neural Networks: A Versatile Approach to Brain-Inspired Computing, from Embedded Systems to Phase-Transition Materials Ph.D. Dissertation by Manuel Jiménez Través for the degree of Doctor of Philosophy by the Universidad de Sevilla within the Doctorate Program of Sciences and Physical Technologies Supervisors: Dr. María José Avedillo de Juan Dr. Juan Núñez Martínez Tutor: Dr. Bernabé Linares Barranco Sevilla, 2024 Instituto de Microelectrónica de Sevilla, IMSE-CNM (CSIC, Universidad de Sevilla) C\ Américo Vespucio 28, 41092, Sevilla, España Departamento de Electrónica y Electromagnetismo Demostradores en Hardware de Redes Neuronales Oscilatorias: Un Enfoque Versátil a la Computación Neuro-Inspirada, desde Sistemas Embebidos a Materiales de Transición de Fase Tesis Doctoral por Manuel Jiménez Través para optar al título de Doctor en Filosofía por la Universidad de Sevilla dentro del Programa de Doctorado de Ciencias y Tecnologías Físicas Directores: Dr. María José Avedillo de Juan Dr. Juan Núñez Martínez Tutor: Dr. Bernabé Linares Barranco Sevilla, 2024 A mi familia, tanto la de sangre como la forjada en el camino. A mi madre y, en especial, a la memoria de mi padre. vii Acknowledgments This Ph.D. Dissertation has been developed at Instituto de Microelectrónica de Sevilla, Centro Nacional de Microelectrónica (IMSE-CNM, CSIC, Universidad de Sevilla). It has been supported under the framework of the European H2020 Project "NeurONN: Two-Dimensional Oscillatory Neural Networks for Energy Efficient Neuromorphic Computing," (Grant Number: 871501). Also, it has been partially supported by PEROSPIKER (Horizon Europe ERC-2022-ADG - Grant ID: 101097688), Consejería de Universidad, Investigación e Innovación de la Junta de Andalucía (BioVEO - Junta de Andalucía PAIDI 2020 - ProyExcel_00060), and by Consejería de Transformación Económica, Industria, Conocimiento y Universidades de la Junta de Andalucía, dentro del Programa Operativo FEDER 2014-2020 (Project US-1380876). A 3-month secondment, funded by NeurONN and supervised by Dr. Siegfried Karg, has been carried out at IBM Research - Zurich, in collaboration with the Science of Quantum and Information Technology department. First of all, I would like to express my immense gratitude to my thesis supervisors, María José Avedillo and Juan Núñez, for their constant guidance, patience and encouragement. There is no doubt that they have contributed significantly to my growth both professionally, scientifically and personally throughout all these years. Likewise, I extend my thanks to Bernabé Linares for his willingness, support and work, making my participation in NeurONN and my secondment in Switzerland possible. I would also like to thank the staff of the Technical Unit at IMSE for their indispensable work, which often goes unnoticed, as well as their camaraderie and good spirit. Throughout the development of NeurONN, I had the privilege of learning from highly professional colleagues who created a very stimulating environment. I would like to thank all the members of the consortium during these years, in particular Aida Todri for her excellent work as coordinator, as well as my fellow PhD students. I am especially grateful to Sigi Karg, Olivier Maher, and Bernd Gotsmann for their warm welcome during my stay at IBM Research. They, together with the other group members and friends living in Switzerland, made this experience truly unique and enriching. I am grateful for the learning, kindness and support I received from my colleagues and friends both in difficult times and moments of laughter, whether inside or outside IMSE. There are too many unforgettable people to name in these lines. It is impossible not to consider family to those who have been by my side for so many years. I dedicate a special mention to my mechatronics friends (including Mario), to Juan Normando, and to Jesús Parra, for being valuable sources of guidance and support. A special thanks to Rafael Vallejo Jurado for his contribution to the cover. I would like to end by dedicating these lines to my family, especially my grandparents and my mother, Juani, as well as my brother from another mother, Manuel. I will also fondly carry with me all that I have learned from my uncle Pepe and my aunt Chari, and their dedication. Thank you all for being there. xv Figure 3.4: Noisy A-J patterns from the noise test set with (a) 10, (b) 20, (c) 30 and (d) 40 corrupted pixels. .......................................................................................................................................... 43 Figure 3.5: Quantization impact on accuracy for 60-neuron networks trained to store 5 digit patterns (‘0’-‘4’). ....................................................................................................................................... 45 Figure 3.6: Proposed IRPUSH learning rule. .............................................................................. 48 Figure 3.7: Retrieval accuracy versus noise level for networks trained with IRPUSH with different T values. .......................................................................................................................................... 50 Figure 3.8: Accuracy versus PF for IRPUSH with T=95. ........................................................... 51 Figure 3.9: (a) Retrieval accuracy versus noise level using IRPUSH with partial factors of 100% and 33% and T=150. (b) Maximum noise level tolerated keeping retrieval accuracy above 95% using IH and IRPUSH with partial factors of 100% and 33%. Considered T value was the one providing the best result for each case. ......................................................................................... 53 Figure 3.10: Retrieval accuracy comparison with other learning rules. ...................................... 54 Figure 4.1: (a) Schematic of a differential oscillator. (b) Schematic of the emulator circuit of the VO2 device. .................................................................................................................................. 59 Figure 4.2: Physical layout of the differential oscillator included in the ASIC. ......................... 59 Figure 4.3: (a) Schematic of oscillator calibration voltage selection. (b) Scheme for a differential oscillator. ...................................................................................................................................... 60 Figure 4.4: (a) Schematic of the synapsis. (b) DONN implementation from [124]. ................... 60 Figure 4.5: Physical layout of the synaptic circuit. ..................................................................... 60 Figure 4.6: Layout of the fabricated circuit. ................................................................................ 61 Figure 4.7: Footprint of the QFN64 package and names of the signals. ..................................... 63 Figure 4.8: Schematic of the designed Schmitt-Trigger buffer. .................................................. 64 Figure 4.9: (a) Block diagram and (b) photograph of the experimental setup ............................ 66 Figure 4.10: (a) Diagram of instruments involved in the set-up. (b) Front panel of the GUI-app for controlling the ONN experiments. ......................................................................................... 67 Figure 4.11: Experimental waveforms for a differential oscillator with analog outputs. ............ 68 Figure 4.12: Experimental waveforms for differential oscillator outputs after digitalization. .... 68 Figure 4.13: Impact of the configuration of the output buffer stage on the duty cycle of the output voltage. ........................................................................................................................................ 69 Figure 4.14: Experimental waveforms corresponding to the outputs of one of the oscillators after digitalization applying two calibration voltages: (a) 1.2V and (b) 0.9V. .................................... 69 Figure 4.15: Outputs of two uncoupled oscillators with SHIL: (a) out-of-phase, (b) in-phase. (c) Without SHIL. .............................................................................................................................. 70 Figure 4.16: Experimental waveforms for positive output of oscillators 1, 4 and 9. .................. 71 Figure 4.17: Two coupled oscillators in which the synapse is programmed with (a) Negative coupling: VP=0V and VN=0.95V and (b) Positive coupling: VP=0.95V and VN=0V. .................. 73 Figure 4.18: (a) Three oscillators with all-to all connections. Outputs (b) without SHIL and (c) with SHIL. ........................................................................................................................................... 74 Figure 4.19: (a) The learned patterns and (b) the test patterns selected for the 3x3 DONN AM. ..................................................................................................................................................... 76 Figure 4.20: Training patterns for the three-patterns AM evaluation. ........................................ 78 Figure 4.21: Graphs with 9 nodes proposed for the Max-Cut hard-problem: (a) G9A, (b) G9B, (c) G9C. ............................................................................................................................................ 79 Figure 4.22: Temporal evolution of the number of trials that achieved the optimum cut value for each graph: (a) G9A, (b) G9B, (c) G9C. ..................................................................................... 80 Figure 4.23: Number of measured optimum states upon each trial. (a) G9A, (b) G9B, (c) G9C. Inset tables depict the percentage of trials with the same number of measured optimum. ......... 81 Figure 5.1: Illustration of fabricated die structure. ...................................................................... 84 Figure 5.2: Pictures of (a) the IBM laboratory setup and (b) the interface PCB. ........................ 85 xvi Figure 5.3: (a) Scanning electron microscope image of crossbar VO2 device. (b) Schematic of the nano-oscillator implemented with the VO2 device. (c) Electrical measurement on the crossbar VO2 device. (d) Oscillation voltage measurements of a single VO2 crossbar oscillator. Material adapted with permission from [62]. ............................................................................................ 86 Figure 5.4: (a) Schematic of the VO2-based oscillator with SHIL injected through capacitance at the output node. (b) Schematic of four coupled oscillators with square configuration. Average phase and measured waveforms for: (c) case without SHIL, and (d) with SHIL applied. .......... 88 Figure 5.5: Experimental results of coupled relaxation VO2 oscillators, involving 4 to 9 nodes to solve the Max-Cut problem: a) Input problem. b) Solution found, and measured. c) Phase relationships of the ONN outputs. d) Number of cycles required to get to the stable state. Material reproduced with permission from [62]. ................................................................................................................... 90 Figure 5.6: Distribution of the cuts attained for graphs involving 6 to 9 coupled VO2 oscillators to solve the Max-cut problem. The optimal solution is found in most instances for graphs C, E, and G. Material reproduced with permission from [62]. ............................................................. 91 Figure 5.7: Distribution of the number of satisfying assignments 𝐾 attained for graphs involving 6 to 9 coupled VO2 oscillators to solve the Max-3SAT problem. Material from [62]. ............... 93 Figure 5.8: Experimental results involving 3 to 6 nodes to solve the Graph coloring problem: a) Input problem, b) Oscillator graph, measured c) Waveforms and d) Phase relationships of the ONN outputs. e) Number of cycles required to get to the stable state. ....................................... 95 Figure 5.9: Schematic and measured waveforms of DONN. (a) Positive coupling. (b) Negative coupling. ...................................................................................................................................... 97 Figure 5.10: Waveforms of the 4-neuron DONN showing a stable solution. Stored patterns encoded in the negative weight matrix (Equation 5.7). ............................................................... 98 Figure 5.11: Schematic of the VO2-based minority gate with the capacitively-coupled inputs and SHIL. ........................................................................................................................................... 99 Figure 5.12: Minority gate operation. Voltage waveforms of the SHIL and the inputs signals. ................................................................................................................................................... 100 Figure 5.13: Minority gate operation. (a) Waveforms for ‘000’ input, evaluated two times in the same experiment. (b) inset A and (c) inset B respectively show the times where inputs are activated, the previous and posterior output logic-value encoded in the phase difference with respect to a reference signal. ..................................................................................................... 101 Figure 5.14: Minority gate operation. (a) Waveforms for ‘100’ input, evaluated two times in the same experiment. (b) Inset A and (c) inset B respectively show the times where inputs are activated, the previous and posterior output logic-value encoded in the phase difference with respect to a reference signal. ..................................................................................................... 102 Figure 5.15: Schematic and measured waveforms of four uncoupled oscillators with SHIL signal added to VDD. .............................................................................................................................. 103 Figure 5.16: Measured waveforms of four oscillators configured as the square graph injecting FHIL signal added to VDD. Different FHIL amplitudes are shown: (a) 125mV, (b) 250mV, (c) 0.5V, and (d) 1.0V. .......................................................................................................................................... 104 Figure 5.17: Measured waveforms of oscillators configured as (a) the four-node square graph and (b) the six-node graph, injecting a SHIL signal with 250mV amplitude added to VDD. ....................... 104 Figure 5.18: Measured waveforms of five isolated oscillators (a) without SHIL injection and (b) with SHIL injection generated with an additional double-frequency oscillator. .................................... 105 Figure 5.19: A 4-node Max-Cut problem. (a) Optimal solution (4 cuts). (b) Energy landscape. ....... 107 Figure 5.20: Evolution of the number of cuts with time for (a) two different noise levels and (b) two different SHIL amplitudes. ................................................................................................. 108 Figure 5.21: Waveforms for two SHIL schedules, a) constant amplitude and b) increasing amplitude. .................................................................................................................................. 109 Figure 5.22: The proposed SHIL schedule and the parameters that describe its configuration. 109 Figure 5.23: 6-node Max-Cut problem with two optimum solutions with 8 cuts. .................... 110 xvii Figure 5.24: SHIL schedules used for the evaluation of the oscillatory IM. ............................. 110 Figure 5.25: Success ratio versus HD between the initial state and the closest optimum solution for the three SHIL schedules. Simulation parameters: VMAX = 0.4V, σNOISE = 1mV, TCOMP3. ... 112 Figure 5.26: Average and standard deviation of the success rate for the three SHIL schemes evaluated at five different computation times. .......................................................................... 112 Figure 5.27: a) 8-node Max-Cut problem reported in [86] with 12 cuts (b) Histogram representing the probability of success of obtaining at least 10, 11 or 12 cuts. ............................................. 113 xviii List of Tables Table 2.1: Distances, nearest pattern to retrieve and ONN results for 5x3 experiment. ............. 22 Table 2.2: Distances, nearest pattern to retrieve and ONN results for 10x6 experiment. ........... 24 Table 2.3: Resources and maximum system clock frequency on Zybo-Z20 FPGA for ONN designs. ........................................................................................................................................ 25 Table 2.4: Time measurements from post-implementation simulations for ONN designs. ........ 26 Table 2.5: Resources on Zybo-Z20 for ONN designs with the programmable synapses. ............. 26 Table 2.6. Successful inferences and average time for ONN designs and the HNN model. ....... 27 Table 2.7: Alternative 40 neuron ONN training configuration evaluation.................................. 32 Table 2.8: 8-ONN training configuration evaluation. ................................................................. 33 Table 3.1: Hamming Distances – 10x6 Digits. ........................................................................... 40 Table 3.2: Capacity results on digit set 10x6. ............................................................................. 41 Table 3.3: Hamming Distances – 12x8 Characters. .................................................................... 41 Table 3.4: Capacity results on character set 12x8. ...................................................................... 42 Table 3.5: Accuracy results for digit set 10x6............................................................................. 44 Table 3.6: Accuracy results for character set 12x8. .................................................................... 44 Table 3.7: Quantization impact on digit set 10x6 – 5 patterns. ................................................... 45 Table 3.8: Quantization impact on character set 12x8 – 4 patterns............................................. 46 Table 3.9: Quantization impact on character set 12x8 – 10 patterns. .......................................... 46 Table 3.10: Capacity with different T values and weight precision. ........................................... 47 Table 3.11: Minimum number of bits and number of learning iterations with different T values for IH and IRPUSH without partial update for the 26 patterns example. ................................... 49 Table 3.12: Minimum number of bits and number of learning iterations with different T values for IRPUSH using a partial factor of 33% and 100% (no partial update). .................................. 49 Table 3.13: Retrieval accuracy with different T values and weight precision. ........................... 50 Table 3.14: Accuracy of IRPUSH vs T with partial update. ....................................................... 52 Table 3.15: Comparison of minimum number of bits required to store the 26 patterns. ............ 55 Table 4.1: Singled-ended oscillators design parameters. ............................................................ 59 Table 4.2: ASIC external signals. ................................................................................................ 62 Table 4.3: Configuration of the DONN output multiplexers according to the control bits. ........ 64 Table 4.4: Configuration of the output multiplexer for the set of simple circuits according to control bits. .................................................................................................................................. 65 Table 4.5: Frequency characterization of uncoupled oscillators. ................................................ 71 Table 4.6: Oscillator calibration results. ..................................................................................... 72 Table 4.7: Frequency characterization of uncoupled oscillators applying calibration. ............... 72 Table 4.8: DONN experiment results corresponding to the AM................................................. 76 Table 4.9: Summary of experimental results corresponding to the AM. .................................... 77 Table 4.10: ASIC results for solving Max-Cut problem associated to the 9-node graphs. ......... 79 Table 5.1: Values for the circuit parameters used for the Max-Cut problems in Figure 5.5. Material from [62]...................................................................................................................................... 89 Table 5.2. Circuit parameters values used to solve the Max-3SAT problems in Figure 5.7. ...... 93 Table 5.3: Circuit parameters values used for the Graph Coloring problems in Figure 5.8. ....... 94 Table 5.4: Truth table of 3-input majority and minority gates. ................................................... 98 Table 5.5: Initial states that converged to an optimum solution with SHIL1, SHIL2, and SHIL3, respectively. ............................................................................................................................... 111 xix List of Acronyms AM: Associative Memory ASIC: Application-Specific Integrated Circuit CMOS: Complementary Metal-Oxide-Semiconductor COP: Combinatorial Optimization Problem CSV: Comma-Separated Value DEB: Descent Exponential Barrier DONN: Differential Oscillatory Neural Network D&O: Diederich and Opper FPGA: Field Programmable Gate Array FSM: Finite-State Machine FP: Floating Point GKM: Gardner-Krauth-Mezard HD: Hamming Distance HNN: Hopfield Neural Network IH: Iterative Hebbian IM: Ising Machine IMT / MIT: PTM’s state Insulating to Metallic / Metallic to Insulating Transition INS / MET: PTM’s state: Insulating / Metallic IRPUSH: Iterative Random Partial Update Symmetric Hebbian MC: Monte Carlo MSB: Most Significant Bit MOSFET: Metal-Oxide-Semiconductor Field-Effect Transistor OBC: Oscillatory-Based Computing ONN: Oscillatory Neural Network PCB: Printed-Circuit Board PTM: Phase-Transition Material QFN: Quad-Flat No-Leads SHIL: Second-Harmonic Injection Locking SMA: SubMiniature version A Chapter 1 1.Introduction Motivation Today's computing platforms face the dual challenge of handling vast amounts of information and executing complex operations. This significantly impacts power consumption, and as environmental sustainability becomes an increasing concern, minimizing the carbon footprint of computing processes has become even more crucial. Additionally, these systems deal with the 'bottleneck' limit associated with von Neumann-type architectures. Meanwhile, conventional complementary metaloxide-semiconductor (CMOS) technology is physically constrained in terms of energy efficiency, as its continued miniaturization leads to increased energy losses from leakage currents [1], [2]. The rise of Edge Computing and the Internet of Things further underscores the need for more efficient and decentralized computing architectures [3], [4], [5]. As the demand for real-time processing, reduced latency, and low power consumption intensifies, traditional CMOS-based architectures approach their inherent limits, necessitating a paradigm shift towards innovative solutions. Diverse perspectives, spanning the physical, circuit, architectural, and system levels, are currently being investigated to tackle these challenges. Innovative approaches, aimed at providing mediumterm solutions, explore alternative devices to either replace or complement the MOS transistor, as well as alternative computing paradigms. One such non-von Neumann approach resorts to exploiting dynamical systems [6], [7], [8], [9]. The fundamental concept involves mapping the solution of a problem to the energy function of an appropriately designed dynamical system [9], [10], [11], [12]. As the system evolves to minimize its energy, it inherently computes the solution to the problem. This approach (“let’s physics compute”) draws inspiration from the natural world, where real-time information processing can be observed in dynamical systems such as the activation patterns of neural circuits [13], [14], cellular signaling mechanisms [15], and information flow in social networks [16]. The physical system performs ‘computation’ through its inherent dynamics in a collective, parallel manner, resulting in competitive computation times and reduced energy consumption [7], [17], [18]. Oscillatory-based computing (OBC) [17], [19] belongs to this category of novel computing paradigms. It leverages the rich, non-linear dynamics of a system of coupled oscillators to implement mathematical functions that link up a system output and input states. The idea of OBC comes from the 1950s in the field of logic [20], [21], where, instead of voltage levels, logical values are encoded in the phase differences between an oscillatory signal and a reference. This is of significant interest for several reasons, notably that the switching between logical values does not require, in principle, energy expenditure, and that noise immunity is also improved compared to the encoding with logical levels [22]. OBC encompasses a wide variety of operating principles and architectures. First, there is a distinction between systems that work with oscillators of ideally identical frequency, where processing corresponds to obtaining a stable phase pattern or phase synchronization (PSK, Phase Shift Key), and systems that work with oscillators of different frequencies, where processing corresponds to frequency synchronization. (FSK, Frequency Shift Key) [23]. 2 Chapter 1: Introduction The connection of a multitude of oscillator circuits, via electrical elements that act as synapses, creates an intelligent collective system also referred to in the literature as an oscillatory neural network (ONN) [8], [24], [25], [26]. ONNs mimic the functioning of certain neurons in specific areas of the brain, which, thanks to their synchrony, can process and transmit information quickly and efficiently. This processing by synchronism is not exclusive to the human body and can be found elsewhere in nature: for example, in the central pattern generators and synchronized locomotion of animals [27], [13], [14], [28], [29], in the flashing rhythms of fireflies [30], [31], or in the interaction between coupled pendulum clocks [30]. In recent years, due to advances in material technology, ONNs have received significant interest and have become an active area of research thanks to the emergence of devices that operate on the basis of very different physical phenomena, with the ability to implement compact oscillators with very low power consumption [17], [32], [33]. Vanadium dioxide (VO2) devices, in particular, stand out for the hysteresis in their characteristic I-V curve, which enables the easy implementation of relaxation oscillators [34], [35], [36]. This Thesis is framed in the context of the European project "NeurONN: Two-Dimensional Oscillatory Neural Networks for Energy Efficient Neuromorphic Computing" [37], [38], [39], whose objectives aimed to develop ONN platforms using VO2 metal-insulator transition (MIT) devices and 2D memristors, as well as, to deliver an ONN design methodology toolbox covering aspects from ONN architecture design to algorithms, facilitating adoption by users to unleash the potential of ONN technology. Our ONNs operate resorting to the PSK operating principle, where the information is encoded in the phase of the oscillations, and computations are carried out through the dynamic interactions and synchronization of these oscillators. The Thesis explores different hardware implementations of ONNs applied to pattern recognition tasks and to solving combinatorial optimization problems (COPs). The final demonstrators of this Thesis are built using VO2 oscillators. The rest of the Chapter is structured as follows: Section 2 describes the implementation of the oscillator using the VO2 device. Section 3 introduces the system of coupled oscillators, the ONNs. Section 4 and Section 5 are dedicated to the applications of pattern recognition and COP solving, respectively. The objectives of the Thesis are outlined in Section 6, and finally, the structure of the Thesis is highlighted in Section 7. VO2-based oscillator VO2 stands out as a fascinating phase-transition material (PTM), exhibiting remarkable properties that enable the implementation of highly efficient and compact oscillators with minimal energy consumption based on physical phenomena. Its distinctiveness arises from a temperature-driven phase transition, a phenomenon in which VO2 undergoes a rapid change in its crystal structure. Figure 1.1a illustrates this phase-transition behavior. At temperatures below a critical point of around 68ºC (340 K), VO2 exists in a semiconducting state, demonstrating insulating properties. However, as the temperature surpasses this critical point, a swift transition occurs, transforming VO2 into a metallic state with significantly altered electronic and optical behavior. This transition is accompanied by dramatic alterations in conductivity (by 3 to 5 orders of magnitude), allowing VO2 to seamlessly shift from being an electrical insulator to a conductor. This process is accompanied by a structural transition from a monoclinic to a rutile crystal structure. 1.2. VO2-based oscillator 3 These insulator-metal transitions can be driven by electrical stimuli. Thus, it can also be described as a two-terminal device with the I-V characteristic shown in Figure 1.1b [40]. Under no electrical stimuli, VO2 tends to stabilize in the insulating phase. When the applied voltage increases and the current density flowing through it reaches JC-IMT, insulator-to-metal transition (IMT) occurs. Once in the metallic state, when the voltage decreases and the current density drops below JC-MIT, metalto-insulator transition (MIT) takes place. VIMT (VMIT) is the voltage at which the insulator-to-metal transition, (IMT) (metal-to-insulator transition, MIT) transition occurs. RINS (RMET) is the resistance in the insulating (metallic) state. MIT and IMT transitions are abrupt but not instantaneous. The hysteresis in the I-V characteristic allows the implementation of a compact oscillator. Figure 1.2a shows a VO2-based oscillator [36], [42], [43]. Note that the resistor and the VO2 can be exchanged. To achieve consecutive, self-sustained IMT and MIT in the device, it is necessary to set its working point in the negative-differential regime of the VO2 I-V curve (Figure 1.2b). By carefully selecting an appropriate resistance value, such that the load line intersects the unstable region defined by the transition points, the system can switch periodically from a high (insulator) to a low (metallic) resistive state [36], [44]. This behavior can also be achieved by connecting the VO2 device in series with a transistor [42]. Figure 1.1: (a) Material resistivity of VO2 against temperature. The phase change presents a structural transition: monoclinic (rutile) crystal when the material is in insulating (metallic) state. Reproduced from [41] . (b) I - V electrical characteristic of the two - terminal VO 2 device. a) b) 4 Chapter 1: Introduction Figure 1.2c depicts waveforms for the oscillator output. The state of the VO2 is also shown to better illustrate the circuit behavior. VO2,STATE=‘INS’ means the device is in the insulating state, and VO2,STATE=‘MET’ corresponds to the device in the metallic state. Assuming the VO2 is in an insulating state (point “A” in Figure 1.2c), the oscillator output is discharged through the transistor. This increases the voltage drop across the VO2 (VDD – VOUT), causing the current through it to increase. Once enough current density circulates (JC-IMT), it switches to the metallic state. Equivalently, once the VO2 voltage reaches VIMT, the switching to the metallic state occurs. The output is then charged through the VO2. This charging is very fast because of the low RMET value. The voltage seen by the VO2 decreases until it reaches VMIT, at which point the MIT occurs. In terms of energy, VO2-based relaxation oscillators have shown good performance. Projections in several papers reported their potential to reduce energy per cycle and to be competitive with ring oscillators and other non-conventional implementations [17], [34], [41], [44]. Consequently, different groups have been exploring the capabilities of VO2-based oscillators for ONN implementations [34], [45], [46], [47], [48], [49], [50], [51], [52], [53]. Within the NeurONN project, fabrication techniques were improved to enhance the performance of VO2 material, and crossbar configurations were explored for their greater advantages over other fabrication methods, such as planar, in terms of power consumption and scalability [41], [54]. V V DD MIT - V OUT time MET INS VO 2,STATE V V DD IMT - Figure 1.2: (a) Schematic of the VO2-based relaxation oscillator. (b) I-V and oscillationenabling negative resistance region determined by the load line. (c) Output voltage waveform of the VO2based oscillator and internal state of the VO 2 device. b) a) c) 1.5. ONNs for solving COPs 11 This is a simplified version of the Hamiltonian, where interaction with an external field is omitted. Solving an Ising model corresponds to determining the spin configuration which minimizes H (ground state) [70], [71]. Figure 1.8 depicts a generic Ising model together with a typical Hamiltonian shape. Figure 1.8: (a) Ising model representation. (b) Hamiltonian energy landscape. A special type of computer, called Ising machine (IM) or Ising solver, solves Ising models. Thus, it has been put forward as an efficient way of performing optimization. An IM solves a class of optimization problems by mapping them to the Ising Hamiltonian of a spin glass system and finding its ground state solution. There are many efforts to develop efficient Ising solvers or IMs. In [70], IMs are classified into three main categories according to their operating principle: classical annealing, quantum annealing, and dynamical system evolution, although different operating principles can be combined. Examples of digital dedicated annealers are the Fujitsu digital annealer [72], [73], the Toshiba bifurcation machine [74] or the FPGA-based solution reported in [75]. In the analog domain, it has been proposed to exploit the intrinsic randomness of nanodevices to avoid the area and energy cost associated with the pseudo-random number generators required in their digital counterparts, also improving the quality of randomness. For example, using magnetic tunnel junctions [61], [76]. In [77] a quantum annealer built by D-Wave Systems is reported. However, it suffers from various problems. These included the need to operate at a temperature close to absolute zero as well as the difficulty of scaling up the machine size, a consequence of its use of quantum techniques. The third type of IM is based on the evolution of a dynamical system in which the state of the system is naturally driven towards the lowest energy state of the Ising model. In particular, coupled oscillators exhibit such behavior and several oscillator-based IMs have been proposed using very different types of oscillators, including optical parametric oscillators [78], [79], [80], electronic conventional ones, such as discrete LC oscillators [81], [82], [83], ring oscillators [59], [60], differential oscillators [84] or Schmitt Trigger oscillators [85], as well as other built based on non-conventional devices like PTMs [46], [86], ferroelectric transistors [87], magnetic-tunnel junction or spin-torque devices [61], [88]. It is convenient to explain that there is a significant difference between the two explored applications for ONNs. Both AMs and IMs rely on the system dynamics naturally evolving to states with less energy, but the system can be trapped at a local minimum. From the point of view of optimization problems, this impedes obtaining the optimum solution. Unlike this, in the AM this local minimum can correspond to the stored pattern closest to the applied test pattern. In the IM case, being able to escape from it is positive, as it could improve the quality of the solution, but it is negative in the AM since a wrong pattern could be retrieved, reducing accuracy. a) b) 12 Chapter 1: Introduction Max-Cut The well-known unweighted Max-Cut problem can be mapped to an Ising model. A maximum cut of a graph is a partition of the vertices of the graph into two complementary sets such that the number of edges between them is the largest possible. The problem of finding a maximum cut in a graph is known as the Max-Cut problem. Formally, it can be formulated as follows. Given an unweighted graph G = (V, E), where V are the vertices and E are the edges between them, the solution of the Max-Cut problem provides a disjoint partition V1 ∪ V2 = V such that the cardinal of E’ ={(u,v) ∈ E, u ∈ V1, v ∈ V2} is maximized. Assuming: 𝑥  =  1 𝑖𝑓 𝑖 ∈ 𝑉  − 1 𝑖𝑓 𝑖 ∈ 𝑉  (1.5) the value of the cut can be written as: 𝐶 =  𝑤  ,  (       )     =    𝑤  ,  −    𝑤  ,            (1.6) where wi,j = 1 if (i,j) ∈ E and 0 otherwise. Thus, the relationship with the Ising model is evident. C can be written in terms of the Hamiltonian in (1): 𝐶 = 12  𝑤  ,  − 12 𝐻 ( 𝑥 )    (1.7) with Ji,j = -wi,j. Note that the first term is constant and that by minimizing H, the value of cut, C, is maximized. Figure 1.9 shows a four-node graph as an example of the Max-Cut problem and the electrical simulation of four-oscillator IM solving it. In this case, the maximum cut value is 4 and it can be obtained with whichever combination of subgroups containing two vertices. The waveforms show that the oscillator phases evolve so that OUT1 (OUT3) is in phase with OUT2 (OUT4). Thus, the problem solution can be obtained from the stable phase pattern of the coupled oscillators. 1 4 3 2 OUT1 OUT2 OUT3 OUT4 Figure 1.9: Example of (a) four-node graph showing one Max-Cut optimum solution. (b) Simulated waveforms of a 4-oscillator oscillatory IM solving the problem. a) b) 1.6. Objectives 13 Objectives The purpose of this Thesis is to contribute to the advancement of the state of the art of ONN using phase-transition VO2 nano-oscillators and the development of different demonstrators that support the potential of this novel paradigm, especially in resource-constrained scenarios that demand fast and energy-efficient solutions. The specific objectives are: 1) To develop ONN demonstrators including a first digital proof-of-concept demonstrator, a CMOS demonstrator in which VO2 behavior is emulated, and demonstrators that use manufactured VO2 devices. 2) To develop appropriate learning algorithms for ONNs functioning as AMs to apply them to pattern recognition tasks. 3) To explore the use of ONNs as IMs for COP resolution. Structure of the contents The rest of the Thesis is organized as follows: In Chapter 2 the design of the first fully digital ONN is presented and characterized on an FPGA, representing a proof of concept of phase-based OBC hardware. Functional validation was performed on pattern recognition tasks. Furthermore, the ONN-FPGA was experimentally validated on a mobile robot that performs obstacle avoidance, being the ONN responsible for the integration and processing of the information from the distance sensors and, after, commanding the robot movement direction. Chapter 3 covers the exploration of HNN literature on learning rules for AMs, with the aim of translating compatible algorithms to ONN training, enhancing its storage and noise robustness capabilities while respecting the hardware implementation constraints. A novel iterative algorithm, with outstanding AM performance and adequate for ONN synapse limitations, is proposed. Chapter 4 contains the design of an analog CMOS-ONN application-specific integrated circuit (ASIC) demonstrator fabricated in TSMC 65nm technology. The demonstrator integrates a CMOS emulator of the VO2 material to allow the early evaluation of the response of the projected VO2-based ONN. The ASIC evaluation board and the whole setup are also described and experimental measurements and characterization of the analog ONN dynamics are reported, along with the demonstration of the ONN problem-solving capabilities as AM and IM. Chapter 5 collects different ONN experiments using fabricated VO2 devices at IBM-Research Zurich facilities. The fabricated device and the experimental setup are described, as well as the implementation of oscillatory IMs, which are evaluated on well-known COPs. Other experiments are also reported as the solving of Graph Coloring problems, the implementation of differential ONNs, the phase-encoded minority logic gate and the evaluation of alternative synchronization injection techniques. Finally, the factors impacting the ONN operation as IM are analyzed and an operation protocol is proposed to achieve the best results. Last, Chapter 6 gathers a general discussion of the developed contents and the relevant conclusions and highlights the future research direction paved by the presented work. 14 Chapter 1: Introduction Chapter 2 2.Digital ONN on FPGA Introduction This Chapter covers the development of a fully digital ONN: design, validation, implementation, and characterization. Inside the digital ONN, neurons are built with digital oscillators, and information is encoded on the phase differences between them. The aim of this design is the demonstration of an ONN system able to tackle interesting AI tasks resorting to phase-encoded computation in the digital domain. It is easily reconfigurable and serves to prove the ONN computing paradigm and to demonstrate different applications. For example, it has been used for pattern recognition and a robotic navigation task on an FPGA evaluation board. That is, the primary goal of this system is to showcase the computing paradigm and anticipate the behavior of the analog ONN based on PTM, introduced in Chapter 1, which holds the potential to establish an energy-efficient computing platform. Therefore, an FPGA implementation was chosen, as it offers a fast and cost-effective prototype while maintaining the intended functionality. This Chapter is structured as follows. Initially, it provides a detailed description of the digital ONN design, offering readers a comprehensive insight into its internal structure and operational mechanism. Subsequently, the design undergoes a thorough validation and characterization process, involving diverse case studies. Additionally, significant improvements in the design are highlighted, such as enabling programmable weights and the integration of an online learning operation mode. Practical demonstrators are also presented, showcasing the ONN versatility in real-world scenarios. The Chapter concludes with a synthesis of key findings, offering insightful conclusions drawn from the exploration of the digital ONN paradigm. Description of the digital ONN design As a proof of concept of the ONN paradigm, a fully digital version has been designed and evaluated. Following the typical FPGA design flow, a top-down approach was employed. The design, described in the hardware description language (HDL) Verilog, ensures versatility with various parameters. It draws strong inspiration from an existing design proposed by Thomas Jackson in [89] but avoiding its analog components. In Jackson's design, the synapses are implemented by a resistor network, and a critical analog comparator is required at the input of each digital neuron to connect to it. It is reported that the comparator demands careful design. In our implementation, the synapses block is an arithmetic circuit, and there are no analog comparators in the neurons. Several ONN versions have been implemented, showing progressive improvement in terms of FPGA resources, accuracy, frequency, and reconfigurability. 16 Chapter 2: Digital ONN on FPGA System description The digital system is composed of the following blocks: ‘neuron’, ‘synapses’ and the control modules. Figure 2.1 depicts the architecture. External ports are indicated with a pad symbol and internal signals are denoted with arrows. As inputs, the system takes an external clock signal (clk), an asynchronous reset, and two dedicated ports to read the initialization signals: load and data_input. As outputs, the neuron outputs are provided along with another two output signals, steady_check and inconsistent_check, indicating whether the ONN state was able to converge or not to a stable state. Additionally, the ONN state is copied into a serial accessible register. Figure 2.1: Fully digital ONN architecture. Neuron unit Figure 2.2 depicts the block diagram of the neuron. It has one 1-bit input, Ninput, and one 1-bit output, Noutput, along with clocks and control signals to be introduced later. The output signal represents the oscillation of the neuron and the input signal, generated by the ‘synapses’ block, represents the weighted sum of the output of the other neurons and determines the evolution of the neuron state. The neuron works such that Noutput is aligned in phase with Ninput. For that, the phase required for Noutput is computed by the ‘phase calculator’ sub-block, updating the neuron state which is provided as input to a ‘phase-controlled oscillator,’ as shown in Figure 2.2. The state value determines the phase of Noutput, which is the output of the latter sub-block. 2.2. Description of the digital ONN design 17 Figure 2.2: Entity description and operating principle of single oscillatory neuron. The implementation of the phase-controlled oscillator is similar to Jackson’s design [89]. A circular shift register and a multiplexer are used. Through the selection bits of the multiplexer, corresponding to the neuron state value, any of the shift register stages can be selected as the neuron output. The shift register has 16 stages. The 16-bit pattern {1111111100000000} continuously cycles through it, so that 16 square signals with distinct phases are available and can be selected to become the neuron output. Figure 2.3a depicts the logic diagram of the ‘phasecontrolled oscillator’ and Figure 2.3b, the waveforms corresponding to stage 0 and stage 8. Note they are 180° apart from each other. The register controlling the multiplexer stores the neuron state or equivalently the phase of the neuron output. Note different clocks are used to control the state register (driven by the system clock) and the shift register. The latter is driven by a slow clock generated from the system clock so that the multiplexer output will cycle, as long as its control register remains immutable, with a period, Tosc = 16 · Tslowclk. Initially, the relation between the slow clock and the system clock given by a frequency divider was Tslowclk = 32 · Tclk, corresponding to a 5-bit division. Latest explorations on the impact of this factor showed that the design can be sped up using an equivalent frequency divider of 2 bits, thus, the resulting slow clock period is Tslowclk = 4 · Tclk. The ‘phase calculator’ block is composed of two edge detectors and a Finite-State Machine (FSM). The associated state diagram is shown in Figure 2.4. One edge detector is used for Ninput, generated by the synapsis block, and the other for the neuron output, Noutput. A conventional implementation with three serially connected flip-flops and a simple logic operation on the flipflops outputs is used. The FSM measures the time difference between the rising edges of the neuron input (output of synapses block) and the neuron output by using a counter which stores the number of clock cycles. The order of the edges is detected and registered using two different states (firstNinp and firstNout) and the variable sig. When the second rising edge is detected (secondEdge), the phase difference is obtained by scaling the counter value with the same factor from the slow clock, which drives the circular shift-register in the phase-controlled oscillator. The neuron state is updated with the contribution of the computed phase difference, so that the output signal phase is aligned with the input signal phase. Note the input signal phase is determined by the phase of the remaining neurons and the weights associated to the corresponding interaction between them. 18 Chapter 2: Digital ONN on FPGA The state registers of all the neurons are serially connected, forming a scan path that allows the state of the ONN to be loaded and read in series. This is done with the signals ser_state_in and ser_state_out. Also, it has an output signal, state_changed, that goes high when the neuron state is modified with respect to its value during the previous clk cycle. Figure 2.3: (a) Logic diagram of implemented phase-controlled oscillator. (b) Output waveform of internal shift-register stages 0 and 8. Figure 2.4: Logic diagram of FSM for the phase calculator. -0 dif <= counter >> clk_div_factor+1 01 sig <= 1 11 01 11 000, 11 counter <= 0 secondEdge update idle firstNout counter <= +1 10 sig <= 0 10 firstNinp counter <= +1 state <= sig ? state + diff: state -dif 2.2. Description of the digital ONN design 19 Synapses block The ‘synapses’ block is in charge of generating the neuron input signals, as illustrated in Figure 2.5. As it has been above mentioned, unlike Jackson’s design, an arithmetic logic circuit, instead of a resistor network, is used in our fully digital implementation of the ONN. Input signal to i-th neuron is generated as: 𝑁𝑖𝑛𝑝𝑢𝑡 [ 𝑖 ] = 𝑠𝑖𝑔𝑛 (  𝑤   −  𝑤   ) (2.1) where j extends to those neurons with Noutput[j]=1 and k to those with Noutput[k]=0. Oscillatory outputs from the neurons determine the signs of the weights when they are summed in the ‘arithmetic circuits.’ The sign of the summed weights in each circuit provides the 1-bit input signal used by ‘neuron’ block. This is the costliest component in terms of resources. Depending on the size of the ONN, the sparsity of the weight matrix, and the resolution of the weights, it will be possible to fit a combinational implementation of this block within the FPGA or a sequential implementation will be required. The latest version adapted the synapses block using a synthesis-friendly approach that allows a reduction in the total add operations achieving both a reduction of required resources and a higher maximum frequency operation. Figure 2.5: Block diagram of digital ONN showing internal sub-blocks. Control block An FSM in the ‘serial2states’ module manages the initialization of the ONN state. Upon detecting a rising edge of the load signal, the data_input bit is transmitted serially, through the scan path configured with the state registers of the neurons, loading the initial state of each neuron in the ONN. Remember that the input to be processed is encoded in the initial state of the ONN. Once the end of the datastream is reached, determined by monitoring an internal counter that increments at the positive edge of clk, the internal full_tick signal is raised. This action triggers the oscillations of the neuronal network, allowing the states to evolve. 20 Chapter 2: Digital ONN on FPGA The ‘freq_divider’ module scales the system clock frequency to provide the signal slowclk that controls the neurons´ oscillation period. The system determines the convergence of the ONN resorting to the N_state_changed signal, that indicates if any neuron state has changed, and a timer which is reset each time the previous signal arises. If the states don´t change in a determined time, steady_check is activated while inconsistent_check remains at low value, indicating stability. Meanwhile, an additional watchdog timer is constantly counting at each clock cycle, unless a reset or load is asserted, in such a way that it produces the activation of both system output signals when the count surpasses an established limit. Design validation and characterization A parameterized Verilog register-transfer-level description of the ONN described in the previous section has been validated, synthesized, and evaluated using Xilinx’s Vivado Design Suite 2018.2 [90] and the XC7Z020-1CLG400C as the target device [91]. Design parameters include the number of neurons, synapses resolution (number of bits for the weights), phase resolution (number of neuron states), and oscillation period. Functional validation As part of the digital design flow, testbenches are used to simulate the operation of the digital ONN. Experiments with a 5x3 ONN and a 10x6 ONN configured as AMs have been carried out. It means that after the training stage, the ONN stores patterns as stable attractors. This is done offline, computing the synaptic weights through the application of a learning rule. Then, during the inference stage, the ONN state evolves from a typically corrupted initial state to one of the stored patterns. Figure 2.6 illustrates the operation of the 5x3 digital ONN. The weights are selected to store the representation of the digit ‘0’ (‘image0’) as one of its stable patterns. It is associated with a 5x3 matrix of pixels representing grey-coded images. Each grey value pixel is encoded into the phase (or state) of the neuron associated to it. White is encoded as state 0 and black as state 8. Operation starts with the img_load pulse. The input pattern is loaded into the ONN with the signal serial_out connected to port data_input. The ONN state is serially initialized through this port (the region in Figure 2.6 where serial_out exhibits activity). Note that once this process finishes (serial_out stops changing), the state of the ONN (signal state, concatenation of the states of the neurons) stores the input image (noisy ‘image0’ à888808800800088). Then, the phase of each neuron is updated according to the output of the other ones and the weights. The ONN state evolves until a stable one is reached. Training patterns are stable, and ideally, the one that more closely resembles the input image is retrieved. As previously said, the 5x3 representation of digit ‘0’ (‘image0’ à 888808808808888) is one of the stored patterns. When the ONN is initialized with an incomplete version of it, the digit ‘0’ is recovered. Transitions between states are marked in the figure before reaching the correct value. 2.3. Design validation and characterization 27 ONN internal state. The state of each neuron is represented with 4 bits. It can be observed that both the “0” and the “1” patterns are recognized. Note the activation of steady signal. However, when pattern “2” is applied, the closest training pattern is retrieved (“0”). Finally, Figure 2.11c depicts the addition of pattern “2” as additional training pattern. Afterwards, when pattern “2” is applied as input for inference, it is recognized and it does not converge to “0” as it happened before being learned by the ONN. Digital ONN and HNN As the ONN is a single-layer design, it is unfair to directly compare it to standard benchmark sets based on other Deep Neural Network implementations. The closest artificial neural network to compare with the ONN is the HNN. It is crucial to clarify to the reader that comparing the digital ONN with a digital implementation of the HNN [95], [96], [97], in terms of power consumption, resources, or speed is not meaningful. The primary goal of the presented digital ONN is to serve as a proof-of-concept for the phase-based oscillatory computing paradigm. This concept will later be translated into highly efficient analog computing platforms built with PTMs. However, it is interesting to assess the behavior of the HNN model in the target applications with which the digital ONN is evaluated. Considering that both networks are energy-based fullyconnected neural networks of a single layer, the training methods and network-state dynamics align easily. In particular, Table 2.6 summarizes the comparison of results obtained for the 5x3 and 10x6 experiments with the digital ONN designs and the HNN mathematical model. The table highlights a high affinity in successful inferences. Moreover, four out of the five incorrectly retrieved patterns in the 10x6 case study trained with Hebb are common to both the ONN and the HNN. Notably, the required cycles for the ONN represent the inference time divided by the oscillation period, while the iterations for HNN indicate the number of required synchronous updates of the states. The ONN architecture allows states to continuously evolve and update when required, achieving correct inferences in fewer cycles than the number of iterations required by the HNN. Table 2.6. Successful inferences and average time for ONN designs and the HNN model. 5x3-Hebb 5x3-Storkey 10x6-Hebb 10x6-Storkey ONN ( #success / # test) 11 / 12 11 / 12 15 / 20 20 / 20 HNN ( #success / #test) 12 / 12 12 / 12 15 / 20 20 / 20 ONN (aver. cycles) 1.15 2.00 1.20 1.12 HNN (aver. iterations) 2.08 2.00 2.00 2.05 28 Chapter 2: Digital ONN on FPGA Figure 2.11: Illustration of learning operation mode: (a) Training the network with patterns ‘1’ and ‘2’. (b) Recognition of three test patterns. (c) Second training step with additional pattern and recognition test. Weight matrix for “0” W. matrix for “0” and “1” Apply “0” Apply “1” Apply “2” Retrieve “0” Retrieve “0” Retrieve “1” Learn “2” Retrieve “2” a) b) c) Apply “2” 2.4. Demonstrators and applications 29 Demonstrators and applications Different demonstrators of ONN applications were developed in cooperation with other partners in the NeurONN project. Particularly, they were implemented on FPGA at the Laboratory of Computer Science, Robotics and Microelectronics of Montpellier (LIRMM-CNRS), using our digital ONN. Also, we were in charge of training the ONNs for the different applications. We briefly describe two of them. Digits recognition The first demonstrator does digits recognition from a camera stream and prints the result on an external screen. Input images are presented by a phone to the camera. The camera is connected to the development board and sends images to the ONN. When the computation time is over, the output pattern is printed on an external screen. The camera is a Pcam 5c camera and communicates via MIPI CSI-2 (MIPI Camera Serial Interface 2) with the Zybo-Z7 development board. The external screen is connected via HDMI (High-Definition Multimedia Interface). The image streaming from the camera to the HDMI screen comes from a Digilent Github project named Zybo-Z7-Pcam-5c, compatible with Vivado software. The digital ONN implementation is integrated inside the processing flow. The image from the camera is converted in greyscale, binarized in black/white pixels, and scaled down to 10x6 pixels image taking the major color between black and white pixels. The ONN output is restructured into a 1280x720 pixels image to be printed on the screen as shown in Figure 2.12. Figure 2.12: Design of the digit recognition application. For this application, the ONN is trained with a set of 5 patterns representing digits from 0 to 4 following the Hebbian rule. The test set is composed of the 5 trained patterns and 20 corrupted images like for design validation. As expected, results are equal to the simulated tests, with 5 images not correctly recognized. Few test patterns along with the corresponding inferred pattern are shown in Figure 2.13. This first demonstrator shows the usability of the ONN inside a complete image recognition application. 30 Chapter 2: Digital ONN on FPGA Figure 2.13: Noisy test patterns and ONN inferred patterns of digits recognition application. Obstacle avoidance: 3 distance sensors The second demonstrator deals with a robotic application. It uses the ONN as a robot brain to avoid obstacles. The 10x6 ONN is embedded inside a small mobile robot from Arduino equipped with 3 proximity sensors, see Figure 2.14. To integrate the ONN inside a robot, it has been implemented into a CMOD-A7 development board from Digilent which is smaller than the Zybo-Z7. The Arduino robot uses an off-the-shelf Arduino board to control the three proximity sensors and motors of the robot, see Figure 2.15. Each sensor can measure a distance between 0 cm and 20 cm. Measurements are discretized to be encoded into an image of 10x6 pixels. A thermometer encoding is used with a 6-value discretization (from 0 to 5), as exemplified in Figure 2.16. For each sensor, the higher the white level is, the further away the obstacle is. Eight reference patterns have been selected to train the ONN, following the Hebbian rule, and a driving command is associated for each, as shown in the Figure 2.17. A video available on [98] shows the capability of the robot to move in a close area without touching any wall or obstacle. Figure 2.14: Schematic of the robot obstacle avoidance application. 2.4. Demonstrators and applications 31 Figure 2.15: Picture of the Arduino robot equipped for obstacle avoidance application. Figure 2.16: Entry image example of the ONN for obstacle avoidance application. Figure 2.17: Stored patterns and associated commands for obstacle avoidance. CMOD - A7 Arduino Sensors 32 Chapter 2: Digital ONN on FPGA Obstacle avoidance: 8 distance sensors On the basis of the results obtained in the 3-sensor demonstrator, a more complex scenario was addressed with 8 proximity sensors. The first step was designing the ONN itself. That is, to determine the size of the ONN that we are going to use, as well as its weight or training matrix. The latter requires selecting a set of patterns to store in the ONN (training patterns) that are relevant to our application. The ONN will map the patterns or images generated by the sensors in this set of patterns which will be translated into control commands for the robot movement. Extending our experience from the three-sensor system, a 40-neuron ONN was selected. Each sensor information is encoded using five neurons (or pixels). So, we visualize patterns as 5x8 images. The weights matrixes were generated applying the Hebbian rule to the set of training patterns. The role of the ONN is retrieving one of those patterns from any possible input image. That is, information from position sensors is integrated, abstracting details in order to simplify taking decisions. Different sets of training patterns have been derived and evaluated by simulation in order to identify those with higher potential. Table 2.7 summarizes some explored solutions. Relevant criteria to assess potential include:  Number of output patterns. The lower this number is, the more abstraction of information is carried out by the ONN. At the limit, it might work with a very simple and reduced set of control commands.  Convergence rate. It is defined as the fraction of input images which converges to a stable output pattern. Ideally it should be 100%. If this is not the case, the control algorithm must monitor the signal generated by the digital design indicating this non-convergence after a reasonable amount of time and must take a suitable decision.  Accuracy. Fraction of input images which converges to a stable output pattern at minimum distance. Ideally, a reduced number of output patterns with 100% convergence rate and 100% accuracy would be desirable. From Table 2.7 it can be observed that W_1 is very competitive in terms of convergence rate and accuracy but generates a larger number of outputs patterns than the other solutions. Thus, we considered using a second ONN to further abstract information and eventually simplifying the control algorithms. For that, an ONN with eight neurons (one per sensor or per column in the 40neuron ONN) was trained to recognize the set of patterns in Figure 2.18 plus both white and black backgrounds. Table 2.8 summarizes relevant information for this training. Table 2.7: Alternative 40 neuron ONN training configuration evaluation. Training Set Description # output patterns Convergence rate Accuracy 40_W_1 Extending the 3-sensor solution. All possible combinations of black and white columns in the 5x8 image. 256 100% 100% 40_W_2 Instead of considering single columns, sets of two columns are used like in W_1. 16 48% 48% 40_W_2_nbg Same as W_2 but without training B/W backgrounds. 16 94% 79% 40_W_4 Sets with four columns group being sliced by 1 column and backgrounds. 42 98% 52% 40_W_4_nbg Same as W_4 but without training backgrounds. 8 93% 69% 2.4. Demonstrators and applications 33 Figure 2.18: Trained patterns for the second 8-neuron ONN stage. Table 2.8: 8-ONN training configuration evaluation. Description # patterns # training patterns 16 # input patterns 256 # output patterns 80 # input patterns that returned a training pattern 192 Convergence rate 100% Accuracy 74% From Table 2.8, it can be observed that there are many more output patterns than training ones. That is, there are many stable output patterns which are not training ones. They are called spurious patterns. However, it can be also seen that most input patterns converge to training patterns. Because of this we also consider this to be a solution with higher potential for our application. Note the actual suitability of a training can depend on many different factors such as the robot itself and the application in which the OAV system is to be integrated. Figure 2.19 illustrates the two-cascaded ONNs scheme. It was incorporated into a unicycle mobile robot with two wheel-drive connected to two direct current motors and eight infra-red proximity sensors, as shown in Figure 2.20. The output pattern obtained with the second ONN was postprocessed on Arduino to find a free direction which is commanded to the robot to follow. A video is available on [99]. Figure 2.19: Two-cascaded ONNs for obstacle-avoidance system improvement. 34 Chapter 2: Digital ONN on FPGA Figure 2.20: Final robot demonstrator doing obstacle avoidance with an embedded digital ONN. Conclusions As a preliminary step towards the analog ONN using PTMs, a fully digital ONN has been developed. The design has been synthesized into an FPGA technology for fast prototyping and validation of suitable applications. This FPGA implementation has been comprehensively characterized and evaluated as AM for different ONN sizes and learning rules. It demonstrates that the system is able to converge to a stable phase pattern within very few oscillation cycles, supporting the potential of the computing paradigm in terms of energy efficiency. Also, significant improvements in the design have been developed, such as the enabling of programmable weights and the integration of an online learning operation mode. The experiments have also been evaluated using a mathematical model of the HNN to assess its suitability as an analytical tool for predicting the potential results obtained with the ONN and the quality of the computed weights. The high degree of similarity in the results suggests that leveraging the extensive literature on learning rules with the Hopfield model can provide valuable insights into developing algorithms compatible with the hardware constraints of the ONN. Finally, demonstrators have been developed to showcase the capability of using the designed digital ONN for complete applications. Specifically, implementations for a digit recognition application and a responsive obstacle avoidance system in a mobile robot have been completed. Chapter 3 3.Learning for Oscillatory AM Introduction As it was already shown in previous Chapters, a single fully connected ONN layer can implement an AM comparable to that of a HNN. It possesses the ability to store and retrieve patterns, thus distinguishing between two different stages in its operation: the learning of the patterns, and the inference phase, where the state of the system evolves until it matches one of the stored patterns [10]. Learning is the process of obtaining appropriate values for weights, thereby encoding the information to be stored in the relationships between the weight values. The weights, which can be grouped into a matrix, determine the synaptic strengths, therefore, determining the interactions, between neurons. This process is often referred to as training of the neural network [100], [101], [102]. Training algorithms found in Machine Learning are typically classified as supervised or unsupervised, on the basis of whether or not knowledge of the target output is used during learning. Depending on the conditions in which the learning is executed, distinctions can also be made between one-shot or iterative, incremental or non-incremental, and local or non-local learning. A learning method is incremental when it only uses information from the new pattern to learn and the current weight matrix to compute the new one, whereas it is iterative if it presents several repetitions of the same information to the training algorithm until it is learned. Locality occurs if the increment in the weight between two neurons only depends on the desired state for those two neurons or their hidden potential. This is an attractive property because it makes learning biologically plausible [64], [103], [104]. Learning can also be considered offline or online depending on whether it requires knowledge of the network response before updating the weights. As suggested in Chapter 2, Learning for ONNs can exploit their similarity with HNNs. An extensive amount of literature is available about learning in HNNs. However, not all of these algorithms are useful for ONN training due to the constraints imposed by their physical implementation. Given that ONN synapses comprise electrical couplings, the computed weights cannot be directly translated to electrical conductance parameters: a mapping step is required [105]. Furthermore, physical implementation imposes restrictions on the weight matrix: bidirectional coupling in oscillator connections requires the matrix to be symmetric; null diagonal elements representing no self-interactions; weight precision is limited by how many electrical strengths the system can distinguish; and layout parasites can impact the performance by modifying the effective synaptic strength and operating frequency. The original HNN model also includes a bias parameter representing equivalently the contribution of one fixed-state neuron. Since the ONN architecture does not include such a contribution, it cannot appear in the learning algorithm. 36 Chapter 3: Learning for Oscillatory AM Weight matrices derived by the application of the well-known Hebb’s rule are symmetric, with null diagonal and without bias parameter, and hence Hebbian learning rule is the most widely adopted method for configuring ONN-based AMs in the very little literature addressing this topic, despite its well-known limitations [19], [26], [50], [51], [89], [105], [106], [107], [108]. Also, weight matrices derived by the Storkey learning rule satisfy those conditions. In fact, it was applied in some applications described in Chapter 2 showing its superiority with respect to Hebbian rule. There are many other learning algorithms with far better performance than the two rules mentioned above. This Chapter explores the HNN literature on learning rules with the aim of developing algorithms suitable for ONN training, enhancing its performance beyond what can be achieved with the two simple rules applied thus far. The rest of the Chapter is structured as follows: Section 2 reviews various HNN learning rules from the literature. In Section 3, we outline our approach to enhance the performance of the Storkey rule. Section 4 introduces a novel iterative rule tailored for ONNs. A comparison of different learning rules is presented in Section 5. Finally, Section 6 summarizes the conclusions. Learning rules for AMs Chapter 1 introduced two well-known learning rules (Hebb and Storkey, [63], [64]), which are reviewed here to facilitate the understanding of this Chapter. As introduced in Chapter 1, for a set of target P patterns of length N, 𝜉𝜖{−1,+1}, to be stored, the coupling strength between neurons corresponds to the correlation of their state values and can be determined as: 𝑤  =  𝜉   𝜉       (3.1) where component 𝜉 represents the desired state of neuron i when pattern k shall be retrieved. Memory capacity limitations [109] storing highly correlated patterns stimulated further research into learning mechanisms. The work in [64] proposed an alternative rule that considers current knowledge of the network during the training steps. That is to say, the weight updates are obtained considering the previous weights. 𝑤   = 𝑤     + 1 𝑁 𝜉   𝜉   − 1 𝑁  𝜉   ℎ   + ℎ   𝜉    ℎ   =  𝑤   𝜉       ,    ,  (3.2 ) This algorithm serves to remove lower-order noise associated with the interaction of the different attractors, where subtraction is applied by computing pre-synaptic and post-synaptic forms of the local field: ℎ  and ℎ . It has been demonstrated that while the absolute capacity of the Hebbian algorithm (no retrieval errors allowed) is given by C=  (), Storkey’s rule achieves C=  () [64], [109], [110]. Both learning rules feature a one-shot, local, incremental, unsupervised training method. 3.3. Improving AM capabilities with Storkey learning rule 43 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 123 45678 1 2 3 4 5 6 7 8 9 10 11 12 12345678 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 1234 5678 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 1234 5678 1 2 3 4 5 6 7 8 9 10 11 12 1234567 8 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 Noise Level: 30 pixels 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 1234 5678 1 2 3 4 5 6 7 8 9 10 11 12 1234567 8 1 2 3 4 5 6 7 8 9 10 11 12 1 2 3 4 5 6 7 8 1 2 3 4 5 6 7 8 9 10 11 12 Noise Level: 40 pixels 123456 78 1 2 3 4 5 6 7 8 9 10 11 12 1234567 8 1 2 3 4 5 6 7 8 9 10 11 12 12345 678 1 2 3 4 5 6 7 8 9 10 11 12 Noise Level: 0 pixels2 Noise Level: 0 pixels1 a) b) c) d) Figure 3.4: Noisy A-J patterns from the noise test set with (a) 10, (b) 20, (c) 30 and (d) 40 corrupted pixels. The accuracy significantly improves with the application of repetition in all cases. Notably, the advantages are more pronounced when transitioning from one repetition to two, compared to the progression from two to three. It is noteworthy that repeating the natural order twice yields accuracy levels not much poorer than those achieved with the best-performing larger randomorder trials. For instance, in the digits example, the accuracy was 88% with repetition compared to 90.1% with the best random-order trial, and in the characters case study, it was even better, 78.5% versus 60.4%, respectively. Additionally, it is essential to highlight the computational efficiency of repetition compared to exploring numerous random orders. Given this efficiency, repetition emerges as the preferred approach for enhancing Storkey rule performance. 44 Chapter 3: Learning for Oscillatory AM Table 3.5: Accuracy results for digit set 10x6. Learning Rule Success Count (#) Max. Accuracy (%) Aver. Accuracy (%) Random Order (100) 1 80.7 80.7 Random Order (200) 39 87.6 80.9 Random Order (300) 53 90.1 81.7 Repeated x2 TrainSet 1 88.0 88.0 Repeated x3 TrainSet 1 90.4 90.4 Random & Rep. x2 (100) 96 91.9 87.2 Random & Rep. x2 (200) 193 92.3 87.4 Random & Rep. x2 (300) 285 92.3 87.2 Random & Rep. x3 (100) 100 94.3 91.0 Random & Rep. x3 (200) 200 94.6 91.1 Random & Rep. x3 (300) 300 94.1 91.0 Table 3.6: Accuracy results for character set 12x8. Learning Rule Success Count (#) Max. Accuracy (%) Aver. Accuracy (%) Random Order (500) 2 56.2 56.1 Random Order (1000) 3 60.4 59.6 Repeated x2 TrainSet 1 78.5 78.5 Repeated x3 TrainSet 1 85.1 85.1 Random & Rep. x2 (100) 85 82.4 78.2. Random & Rep. x2 (500) 398 82.9 77.7 Random & Rep. x2 (1000) 800 83.9 77.6 Random & Rep. x3 (100) 100 88.0 86.2 Random & Rep. x3 (500) 500 87.7 86.0 Random & Rep. x3 (1000) 1000 88.1 86.0 Impact of weight precision Given the restrictions on the weights to be compatible for analog ONNs, an evaluation of our approach concerning the impact of quantization was pursued. It has been explored storing from ‘0’ to ‘4’ digit-patterns for the digits experiment, since this is the largest set successfully handle by the simple Storkey solution. Storkey solution as well as those obtained with the repetitive learning approaches using the natural order of the patterns have been evaluated with weights quantized from 5 to 2 bits. The applied quantization method is based on: 1) Normalization with respect to the maximum absolute weight value to obtain 𝑤≤1. 2) Scaling by the largest absolute value which can be represented with the target number of bits (nbit). 3) Rounding. Note that the MSB bit is dedicated to the sign, while the remaining bits represent the integer value, up to a maximum: 2−1. That is, the magnitude value for nbit precision, 𝑤 ∗, is obtained as: 𝑤  ∗ = 𝑟𝑜𝑢𝑛𝑑 󰇧  2    − 1  𝑤  max  abs ( 𝑊 )  󰇨 (3.11) 3.3. Improving AM capabilities with Storkey learning rule 45 Table 3.7 gathers obtained accuracy using the same set-up as the previous experiments. These data are also graphically displayed is Figure 3.5. Accuracy is not reported in those cases for which the five patterns are not stored. It is worth to point out that although the cases with 2-bit weights are able to keep the target patterns as fixed points, their retrieval capability as attractor is really poor. The very low accuracy values achieved indicates their uselessness as AM. Results evinces that repeated training approach allows to retain the target number of patterns while keeping satisfactory noise robustness in spite of the loss of precision due to quantization. These results provide the insight that the obtained solutions are competitive when few bits precision is required, showing scores not far from the FP precision even quantizing up to three bits. Moreover, accuracy obtained by repeating four times the training and quantizing to four bits (96.8%) is larger than the accuracy exhibited by the single Storkey training with full precision (90%). Table 3.7: Quantization impact on digit set 10x6 – 5 patterns. Accuracy (%) Precision Storkey Repeated x2 Repeated x3 Repeated x4 FP 90.0 96.2 99.2 99.8 5-bit 89.6 94.4 99.6 99.8 4-bit - - 96.2 96.8 3-bit - 87.8 93.2 91.6 2-bit 9.4 7.2 6.8 6.8 Figure 3.5: Quantization impact on accuracy for 60-neuron networks trained to store 5 digit patterns (‘0’-‘4’). 46 Chapter 3: Learning for Oscillatory AM The performance of the simple Storkey rule on the character set experiment was so poor (only the first four pattern where stored when using the natural order), that we consider also ten patterns for the weight quantization experiment. Table 3.8 and Table 3.9 depict obtained results for four and ten patterns respectively. For the ten patterns experiment, we have been able to improve the accuracy achieved with full precision weights and twice training iterations with natural orders (78%) using just 4 bits by increasing the number of repetitions (83.1%). Table 3.8: Quantization impact on character set 12x8 – 4 patterns. Accuracy (%) Precision Storkey Repeated x3 Repeated x5 FP 83.8 92.7 92.5 5-bit 85.9 91.9 92.2 4-bit - 84.8 84.9 3-bit - 55.3 80.9 2-bit - - - Table 3.9: Quantization impact on character set 12x8 – 10 patterns. Accuracy (%) Precision Storkey Repeated x3 Repeated x5 FP - 85.1 86.5 5-bit - 84.4 86.2 4-bit - 80.0 83.1 3-bit - - - These results show that the advantages of repetitive Storkey approach are kept after quantization. So, it is considered as a practical solution to enhance the performance of analog ONNs which have a severe limited weight precision. Iterative Random Partial Update Symmetric Hebbian (IRPUSH) We have also proposed an iterative rule that is based on D&O’s Rule I [111], described in Section 3.2, and henceforth referred to as the Iterative Hebbian (IH) rule. Figure 3.6 shows the flow chart of the novel learning algorithm denoted Iterative Random Partial Update Symmetric Hebbian (IRPUSH). It consists of two nested loops, which are repeated until the algorithm converges to a solution that satisfies condition (Equation 3.3) (upd_flag=0) or until a limit number of iterations is reached. The external loop is for patterns and the inner loop for neurons. This is exactly the same as in the IH rule. The key difference is the weight updating procedure. First, in order to force symmetry in D&O’s algorithm, it is imposed that the new algorithm must have 𝑤=𝑤. In the IH rule, weight updates caused by neuron 𝑖 not satisfying (Equation 3.3) have no impact on the compliance or non-compliance status of the remaining neurons because only 𝑤 ∀ 𝑗∈{1,…,𝑁 / 𝑗≠𝑖} are modified. Due to the symmetry constraint, however, modification of 𝑤 implies modification of 𝑤, and so other neurons are also affected. From a different point of view, IH does not modify weights 𝑤 if neuron 𝑗 satisfies (Equation 3.3) (for a given pattern), but this is not the case when symmetry is forced. This could negatively impact the performance of the learning rule since some neurons which already satisfied (Equation 3.3) might not fulfill it after the weight update forced by the symmetry constraint. In order to somehow counteract this effect, we 3.4. Iterative Random Partial Update Symmetric Hebbian (IRPUSH) 47 propose that only a randomly selected fraction of the weights be updated at each step. This is motivated in [119], which it states that a reduced number of connections (weights) while applying the iterative learning leads to a redistribution of the lost information to the remnant ones. Possibly, it can contribute to the resulting weight matrix being more suitable for the quantization. Thus, the main differences, pointed out with letters inside a circle box in Figure 3.6, are: A. Symmetry is imposed. B. The partial factor determines the number of connections (PF) that are randomly selected from the current update set. After the update step, the chosen ones (candidatePF) are removed from the update set. All the weights are restored back in the update set once its cardinality is under PF, imposed by the partial factor. C. The algorithm terminates when the NP inequalities are satisfied or a maximum number of training iterations is reached. It is interesting to point out that the obtained matrix solution is always the same, in case of null partiality during the update step. With the partiality approach, a random fraction of the weights associated to the current evaluated neuron is modified with Hebbian learning. The randomness in this procedure allows to explore further different solutions in the weights assignment that totally complies with the embedding conditions of the training set when the algorithm converges. The limit on training iterations is implemented to prevent prolonged execution, which can occur when learning a challenging training set. To evaluate the performance of the AM trained with IRPUSH in the pattern recognition task, the experiment with the previous 12x8 character dataset was carried out. Different 96-neuron networks trained with our IRPUSH algorithm and using different values for parameter 𝑇 and partial factor were evaluated for storing the proposed set of 10 characters. The partial factor, which determines the number of weights that are modified when a neuron is updated, was first constrained to 100%: i.e., all 95 weights associated with one neuron being changed. Table 3.10 summarizes the capacity results obtained in this analysis and assesses the use of weights with reduced precision. The quantization method is the same as described in Equation 3.11. For purposes of comparison with a competitive candidate, the IH algorithm was also evaluated under similar conditions. Since update increments are not scaled by N, large T values were chosen. Thus, case 𝑇=1 in the original work corresponds to 𝑇=N in our implementation. Table 3.10: Capacity with different T values and weight precision. IRPUSH, Capacity (#) IH, Capacity (#) T FP 5-bit 4-bit 3-bit FP 5-bit 4-bit 3-bit 0 10 10 9 6 10 8 7 0 10 10 10 10 5 10 10 9 7 50 10 10 10 8 10 10 10 8 110 10 10 10 9 10 10 10 9 150 10 10 10 10 10 10 10 9 200 10 10 10 10 10 10 10 9 400 10 10 10 10 10 10 10 8 As can be seen in Table 3.10, every network retained the 10 patterns using FP weights. However, IRPUSH outperformed the IH capacity with 3-bit precision, being able to store the 10 patterns for the largest values of 𝑇. IRPUSH also found solutions at lowest values of 𝑇 with 5-bit and 4-bit precision. 48 Chapter 3: Learning for Oscillatory AM Figure 3.6: Proposed IRPUSH learning rule. 3.4. Iterative Random Partial Update Symmetric Hebbian (IRPUSH) 49 To further explore the capabilities of the IRPUSH algorithm to store patterns, a larger case study has been evaluated. A training set of P=26 patterns (Figure 3.7) with 32x32 (N=1024) pixels have been used. Figure 3.7: Characters A-Z, binary patterns with dimension 32x32 used as training patterns. Table 3.11 compares IRPUSH without partial update and IH for different T values. Two different parameters are shown. The minimum weight precision able to store the 26 characters and the number of learning iterations required by each algorithm. Table 3.11: Minimum number of bits and number of learning iterations with different T values for IH and IRPUSH without partial update for the 26 patterns example. # bits required # iterations T IH IRPUSH (100%) IH IRPUSH (100%) 0 13 10 9,844 8,093 10 10 8 10,137 8,380 50 8 7 11,186 9,123 110 7 7 13,025 10,513 150 7 7 14,121 11,356 200 7 7 15,188 12,594 400 7 6 20,856 17,181 As in the previous simpler example, it can be observed that increasing the threshold value used to derive the weight matrix allows reducing the weight precision. Also, the superiority of IRPUSH is clear. A solution with just 6 bits was found for IRPUSH but not for IH in our experiment. IH requires a larger number of iterations than IRPUSH. Table 3.12 compares IRPUSH with no partial update to IRPUSH with 33% partial factor. Advantages of IRPUSH (33%) in terms of number of bits required is clearly observed. It uses a higher number of iterations than IRPUSH with 100% partial factor for the larger T values when the random update is applied. This is due to a reduced increment of the neuron potential, ℎ, per each update since a smaller number of weights are modified, therefore requiring more iterations to reach the threshold T. Table 3.12: Minimum number of bits and number of learning iterations with different T values for IRPUSH using a partial factor of 33% and 100% (no partial update). # bits required # iterations T IRPUSH (33%) IRPUSH (100%) IRPUSH (33%) IRPUSH (100%) 0 7 10 7,007 8,093 10 5 8 7,467 8,380 50 5 7 9,293 9,123 110 5 7 12,171 10,513 200 5 7 16,450 12,594 400 5 6 26,107 17,181 50 Chapter 3: Learning for Oscillatory AM Noise robustness results Noise robustness was evaluated for those weight matrixes that stored the 10 training patterns from the 96-neuron character dataset (shaded cells in Table 3.10). As previously described, the retrieval accuracy for this experiment was measured as the percentage of test patterns which were correctly inferred. That is, the percentage of cases where the closest training pattern, indicated by the minimum HD, is retrieved. Figure 3.7 shows the retrieval accuracy of networks trained with IRPUSH versus noise level. It is evident that increasing the value of 𝑇 produced larger basins of attraction, pushing forward the number of noisy pixels that the network could tolerate before the number of erroneous inferences began to rise. However, the benefits of increasing 𝑇 were quickly reduced for 𝑇>50, as intuitively observed in Figure 3.7. Figure 3.7: Retrieval accuracy versus noise level for networks trained with IRPUSH with different T values. The accuracy results obtained with IRPUSH and IH for different 𝑇 values are reported in Table 3.13, together with their performance with limited precision. The advantages of IRPUSH were greatly reduced as 𝑇 increased, displaying very similar performance to IH. IRPUSH accuracy loss at 4-bit precision remained under 3% with respect to FP, while accuracy in 3-bit cases was degraded by around 20%. On the other hand, the 4-bit IH presented accuracy degradation of up to 6% with respect to FP. Table 3.13: Retrieval accuracy with different T values and weight precision. IRPUSH, Accuracy (%) IH, Accuracy (%) T FP 5-bit 4-bit 3-bit FP 5-bit 4-bit 3-bit 0 63.4 57.7 - - 37.1 - - - 10 76.2 70.5 77.1 - 57.8 53.4 - - 50 87.8 86.9 85.6 - 83.6 83.0 80.5 - 110 89.7 89.6 87.9 - 89.5 87.7 86.7 - 150 90.6 90.2 88.3 67.0 90.1 89.0 84.6 - 200 91.2 90.4 89.2 71.3 91.0 90.9 86.3 - 400 91.5 90.3 89.5 72.0 91.6 91.1 89.3 - 3.4. Iterative Random Partial Update Symmetric Hebbian (IRPUSH) 51 Keeping in mind that IRPUSH was used with a partial factor equal to 100%, the only difference with respect to IH was that IRPUSH constrained the weight matrix to be symmetric. The results show that imposing symmetry did not penalize the accuracy, as indicated in Section 3.4, where the enforcement of symmetry was observed to potentially hinder the performance of IRPUSH. In fact, our approach obtained better results for lower T values. The better accuracy results described above were probably a side effect of the larger number of update operations resulting from having imposed symmetry. Unlike the original asymmetric update in IH, 𝑤  (and 𝑤  ) could both be modified when processing neuron i or neuron j. Of even greater interest, however, was the success of our approach in storing the ten patterns with 3-bit weights, as pointed out in the capacity evaluation, for T ≥150. The next step was to evaluate IRPUSH for partial factors below 100% so that the proposed random partial update mechanism was actually applied. That is, the networks were also evaluated by applying the philosophy of updating a reduced number of weights, each time an update step took place. They had a common T value of 95 (corresponding to set the Rule I threshold to unity). The partial factor was reduced from 100% to 1% in steps of 25%, plus the selected factors of 33% and 10%. All the networks stored the training set perfectly. It is interesting to note that the algorithm returns a different weight matrix solution each time it is run, because the connections to be updated are randomly selected. A battery of 100 experiments for each case was therefore carried out. Figure 3.8 reports their respective maximum, average, and standard deviation values. Note that the case in which all neurons were updated (100% partial factor or PF equal to 95), an identical solution was therefore returned every time. It was found that reducing the number of updated connections resulted in an increase in accuracy with regard to IRPUSH with a partial factor of 100%. Figure 3.8: Accuracy versus PF for IRPUSH with T=95. The impact of 𝑇 was also explored for two particular cases: partial factors of 33% and 25%. Table 3.14 shows the accuracy results for different 𝑇 values and weight precisions. Note that accuracy increases with T for FP and 5-bit weights; however, for 3-bit precision, both increments and decrements are observed. This is explained because of the larger impact of rounding when precision is significantly reduced. 52 Chapter 3: Learning for Oscillatory AM Table 3.14: Accuracy of IRPUSH vs T with partial update. Accuracy, 33% factor (%) Accuracy, 25% factor (%) T FP 5-bit 4-bit 3-bit FP 5-bit 4-bit 3-bit 0 39.0 40.4 35.0 - 30.5 30.5 - - 10 80.7 82.0 80.9 - 82.4 81.0 78.2 - 50 91.7 91.4 90.8 84.4 90.8 90.9 89.5 84.0 110 92.6 92.4 91.9 90.2 93.1 92.7 91.9 87.5 150 92.7 93.0 92.5 90.9 92.8 92.4 93.5 75.0 200 93.2 93.0 92.2 83.3 92.9 92.7 92.2 84.6 400 93.4 93.2 93.3 - 92.3 92.5 93.3 82.6 It is worth noting that IRPUSH with the random partial update mechanism was quite satisfactory not only in the number of cases storing the 10 patterns. Comparing the results obtained for the most reduced numbers of bits with partial update (Table 3.14) and without (Table 3.13), it can be observed that much better results are achieved with partial update. For example, accuracy up to 90.9 with 33% of partial factor and three bits is obtained while the best results without partial update and that precision is only 72%. That is, the results support our hypothesis concerning the benefits of the partial update procedure. To further compare the inference robustness of the IRPUSH algorithm against noise, Figure 3.9a depicts the retrieval performance of IRPUSH with partial factors of 100% and 33% against different noise levels for the same T value of 150. Similar results were obtained with FP and a partial factor of 100% and with 3-bit precision and a partial factor of 33%. Both outperform 3-bit precision with a partial factor of 100%. This clearly illustrates the advantage of the random partial update mechanism using limited precision and demonstrates its ability to avoid serious degradation of the retrieval capabilities under low weight precision, when compared with the FP precision case. Figure 3.9b compares the maximum number of noisy pixels for which the evaluated networks achieved an accuracy greater than 95%. It shows the best result obtained among different T values in each case. The partial update strategy for 3-bit weights had a major impact, raising this figure of merit from 7 to 27. 4.2. Analog CMOS-ONN demonstrator description 59 Figure 4.1: (a) Schematic of a differential oscillator. (b) Schematic of the emulator circuit of the VO2 device. Table 4.1: Singled-ended oscillators design parameters. Parameter Value C 50fF R1 35kΩ R2 40kΩ WSHIL, LSHIL 120nm, 480nm WEMU, LEMU 1µm, 500nm Figure 4.2: Physical layout of the differential oscillator included in the ASIC. Synapse As an analogy of the Wheatstone bridge, Figure 4.4a shows the schematic of the synaptic circuit, which is capable of providing positive, negative, and zero weights. Being a four-terminal circuit makes it appropriate for differential structures. Depending on the gate voltages, it is possible to have positive (VP>VN), negative (VN>VP), or zero weight (VP=VN). The two PMOS transistors are used for controlling the current between the neurons. Transistors between the positive branches can transfer current when (VP-V1+)< VTH, because VP is applied to their gate. In addition, transistors between the negative branches can transfer current when (VN-V2+)<VTH, because VN is applied to their gate. Therefore, current between the neurons is controlled using these PMOS transistors. Figure 4.4b depicts the topology of the fabricated differential ONN (DONN). All transistors are equally sized with W=360nm and L=120nm. The physical layout of the synaptic circuit is shown in Figure 4.5. The area of the synaptic unity cell is 63.2m2. Each synapse control voltage can be connected to six different voltages by means of programmable switches controlled by programming registers, similarly to the oscillator calibration shown in Figure 4.3a. 60 Chapter 4: ASIC CMOS-ONN Demonstrator Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Q Q G RB CLR D Serial Parallel DataReady Serial dataSerial data CTRLP CTRLN Vosc1 Vosc2 Vosc3 Vcal Oscillator Oscillator OSCi Vosc4 Vosc5 Vosc6 CLK Figure 4.3: (a) Schematic of oscillator calibration voltage selection. (b) Scheme for a differential oscillator. b)a) Figure 4.4: (a) Schematic of the synapsis. (b) DONN implementation from [124]. Figure 4.5: Physical layout of the synaptic circuit. a) b) 4.2. Analog CMOS-ONN demonstrator description 61 ASIC description The ASIC consists of the following blocks: - ONN of 9 differential oscillators (9-neuron DONN) fully interconnected with each other through 36 synapses. The DONN includes a control system from which the voltages defining the synapse weights can be selected, as well as the oscillator calibration mechanism to improve network synchronization. The oscillator outputs are connected to digital pads through Schmitt-Trigger buffers with configurable hysteresis to regenerate rail-to-rail signal swing and to cope with the waveform shape of the relaxation oscillators. - Basic circuits. Specifically, a single-ended oscillator, a differential oscillator and two differential oscillators connected through a synapse similar to the one used in the DONN have been included. Their outputs are digital as in the DONN circuit. In addition, a differential oscillator has been included whose outputs are connected directly to analog pads in order to be able to observe the waveforms without digitizing. Figure 4.6 shows the layout of the fabricated circuit, showing the three types of circuits included (3x3 ONN, simple circuits, differential oscillator with analog outputs) and their connections to the pad ring. The 3x3 DONN occupies a rectangle of 776µm · 747µm with some empty area inside and the complete chip area including pad ring is 1710µm · 1710µm. Basic circuits 3x3 ONN Differential oscillator with analog outputs Figure 4.6: Layout of the fabricated circuit. Description of the signals/pads Table 4.2 summarizes the main characteristics of the ASIC signals, divided into the following categories: - Polarization and ground signals. Supply and ground voltages for the core (VDD, VSS) and the pad-ring power supply and ground (VDDPST, 3.3V, and VSSPST, 0V, respectively) are included. - Oscillator supply voltages. Since differential oscillators are being used and there are 9 neurons in the DONN, 18 signals are required. These signals are generated externally and applied to digital pads that generate step signals between 0V and 1.2V. Controlling the relative timing on initial phase selection step is essential: a delay between both signals corresponding to half a period involves applying input stimuli with opposite phases. The initialization signals of the i-th oscillator are denoted as VDDOSC<i> and VDDOSCx<i>. 62 Chapter 4: ASIC CMOS-ONN Demonstrator - Control system signals. Includes the clock signal (CLK), the signal that codifies the information to be loaded into the calibration/programming voltage selection registers (DATA_CTRL), and the signal that indicates that the information has been loaded into the serial registers and serial-toparallel conversion can be done (DATA_READY). - Oscillator calibration signals. These six signals can take values between 0V and 1.2V and allow the oscillator frequency to be tuned. They are referred to as IN_OSC<1> to IN_OSC<6>. - Synapse programming signals. These twelve signals are used to set the weights to the synapses. They take values between 0V and 1.2V and are denoted as IN_SYN<1> to IN_SYN<12>. The IN_SYN<11> and IN_SYN<12> signals are also used to set the thresholds of the DONN output Schmitt-Trigger buffers. - Synchronization signal. It is the digital signal VSYNC, with a frequency double that of the DONN, used to enable the SHIL mechanism. This external input ranges from 0V to 3.3V. - Output multiplexer control signals. Named OUT1_CTRL, OUT2_CTRL and OUT3_CTRL. Used to configure the multiplexer-output stage to allow access to the DONN outputs and the basic circuits ones. - Digital output signals. These are the signals OUT<1> to OUT<10> for the DONN outputs and OUT_SIMPLE for the output of the basic circuits. They show a voltage range between 0V and 3.3V. - Analog output signals. From the isolated differential oscillator connected to analog pads, referred to as OUT_OSC and OUT_OSCx. The output voltage range is between 0V and 1.2V. Table 4.2: ASIC external signals. Name Number of pins Purpose Type Pad type Voltage range VDD 2 Core supply voltage I/O Supply 1.2V VSS 2 Core ground I/O Ground 0V VDDPST 1 Pad-ring supply voltage I/O Supply 3.3V VSSPST 1 Pad-ring ground I/O Ground 0V VSSA 1 Analog ground I/O Ground 0V VDDOSC 9 Initialization Input Digital 0V-3.3V VDDOSCx 9 CLK 1 Control logic Input Digital 0V-3.3V DATA_CTRL 1 DATA_READY 1 IN_OSC 6 Oscillator calibration Input Analog 0V-1.2V IN_SYN 12 Synapses programming Input Analog 0V-1.2V VSYNC 1 SHIL signal Input Digital 0V-3.3V OUT1_CTRL 1 MUX selection Input Digital 0V-3.3V OUT2_CTRL 1 OUT3_CTRL 1 OUT 10 Outputs DONN Output Digital 0V-3.3V OUT_SIMPLE 1 Output simple circuits Output Digital 0V-3.3V OUT_OSC 1 Analog outputs Output Analog 0V-1.2V OUT_OSCx 1 4.2. Analog CMOS-ONN demonstrator description 63 Figure 4.7 describes the footprint of the 64-pin quad flat no-lead (QFN) package chosen for the ASIC and the correspondence between its pins and the input/output signals of the circuit. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 64 63 62 61 60 59 58 57 56 55 54 53 52 51 50 49 48 47 46 45 44 43 42 41 40 39 38 37 36 35 34 33 OUT_SIMPLE OUT<1> OUT<2> OUT<3> OUT<4> OUT<5> OUT<6> OUT<7> OUT<8> OUT<9> OUT<10> VDD VDDPST VSS VSSPST VDDOSC<9> VDDOSCx<8> VDDOSC<8> VDDOSCx<7> VDDOSC<7> VDDOSCx<6> VDDOSC<6> VDDOSCx<5> VDDOSC<5> VDDOSCx<4> VDDOSC<4> VDDOSCx<3> VDDOSC<3> VDDOSCx<2> VDDOSC<2> VDDOSCx<1> VDDOSC<1> VDDOSCx<9> VSYNC DATA_CTRL DATA_READY CLOCK VDD VSS CTRL1_OUT CTRL2_OUT CTRL3_OUT VSSA OSC_ANA OSCx_ANA IN_OSC<1> IN_OSC<2> IN_OSC<3> IN_OSC<4> IN_OSC<5> IN_OSC<6> IN_SYN<1> IN_SYN<2> IN_SYN<3> IN_SYN<4> IN_SYN<5> IN_SYN<6> IN_SYN<7> IN_SYN<8> IN_SYN<9> IN_SYN<10> IN_SYN<11> IN_SYN<12> Figure 4.7: Footprint of the QFN64 package and names of the signals. Control logic for calibration and programming This mechanism is described as follows: - Each oscillator and each synapse have as many flip-flops as switches to be controlled. That is, 12 for each oscillator (see Figure 4.5b) and 12 for each synapse. They all are configured in a single shift register, generating, therefore, a connection of 12·9+12·36=540 memory elements. These registers are controlled by the clock signal. A control word is serially loaded in the shift register. It contains the calibrating and programming bits indicating which switches are closed and which are not. Obviously, for each oscillator or synapse only one of its switches should be closed. - Once the control word is fully loaded, a signal that indicates that the data is ready to be loaded is activated and the information contained in the shift registers is loaded in parallel to the flipflops directly controlling the switches. 64 Chapter 4: ASIC CMOS-ONN Demonstrator Digital Output Stage Since the outputs of the 9 differential oscillators (18 in total) range between values away from ground and VDD, it is necessary to carefully design an output stage that delivers perfectly reconstructed signals to the digital pads. The output stage is further divided into two sub-stages: 1. A digitization of the analog signal from the oscillator must be performed. A frequent problem in the output of these oscillators is that a rise and fall can appear around the middle of VDD which can cause glitches if a conventional buffer is placed at the output. To solve this problem, a programmable Schmitt-Trigger buffer is used, as shown in Figure 4.8. By means of the VCTRLP and VCTRLN voltages, the transition thresholds of the Schmitt-Trigger can be controlled, thus suppressing the glitch and keeping the transitions within the range of the oscillator output variation voltages. VCTRLP and VCTRLN signals are internally connected to signals IN_SYN<11> and IN_SYN<12>, respectively. V DD VCTRLN VCTRLP V DD OUT IN Figure 4.8: Schematic of the designed Schmitt-Trigger buffer. 2. Once all the outputs have been properly digitized, 10 4-input multiplexers (control signals CTRL1_OUT and CTRL2_OUT) have been used to select different combinations of outputs, as shown in Table 4.3. In addition, one of the combinations monitors the DATA_CTRL signal at the output of the last shift register. Table 4.3: Configuration of the DONN output multiplexers according to the control bits. OUT (CTRL1_OUT, CTRL2_OUT) (0,0) (1,0) (0,1) (1,1) OUTPAD,1 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇  ,  𝑂𝑈𝑇  ,  OUTPAD,2 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  OUTPAD,3 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇  ,  𝑂𝑈𝑇  ,  OUTPAD,4 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  OUTPAD,5 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇  ,  𝑂𝑈𝑇  ,  OUTPAD,6 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  OUTPAD,7 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇  ,  𝑂𝑈𝑇  ,  OUTPAD,8 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇        ,  OUTPAD,9 𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝑂𝑈𝑇  ,  𝑂𝑈𝑇  ,  OUTPAD,10 𝑂𝑈𝑇        ,  𝑂𝑈𝑇  ,  𝑂𝑈𝑇        ,  𝐷𝐴𝑇𝐴 _ 𝐶𝑇𝑅𝐿 4.3. Experimental validation and characterization 65 Unlike the DONN oscillator outputs, in the basic circuits it is not necessary to use the SchmittTrigger buffer. Therefore, a conventional buffer is used to digitize the signal. The observable signals at the output are 7, so it is necessary to use a multiplexer controlled by 3 signals (CTRL1_OUT, CTRL2_OUT and CTRL3_OUT). The remaining output has been used to display the synchronization signal. The correspondence between the configuration of the control bits and the output is summarized Table 4.4. Table 4.4: Configuration of the output multiplexer for the set of simple circuits according to control bits. CTRL1_OUT CTRL2_OUT CTRL3_OUT OUT_SIMPLE 0 0 0 Single-ended oscillator 1 0 0 Differential oscillator (pos) 0 1 0 Differential oscillator (neg) 1 1 0 Coupled oscillators (OSC1 pos) 0 0 1 Coupled oscillators (OSC1 neg) 1 0 1 Coupled oscillators (OSC2 pos) 0 1 1 Coupled oscillators (OSC2 neg) 1 1 1 Synchronization signal Experimental validation and characterization The experimental verification of the ASIC has been performed using a custom designed test printed-circuit board (PCB) for this purpose. The block diagram of the experimental set-up is shown in Figure 4.9a, together with a photograph (Figure 4.9b). Additionally, used test equipment is also depicted. Note that an FPGA is used for programming the DONN (selecting the assignment of synapses and calibration voltages) and for controlling the time-delayed power-on of the differential oscillators. A physical micro-switch on-board allows to configure the initial state of the system, on which basis the input patterns are applied to the DONN. The main features of the experimental set-up are summarized below: - The initialization of the oscillators is performed from two digital signals. They step in from low to high level with a delay equivalent to half a period of the oscillators, and are applied to the supply voltages of the differential oscillators. The order in which they are applied is given by the position of the switch associated with each oscillator (white box with colored switches located at the top of the PCB picture). These signals are also generated by the custom digital design in the FPGA after that the Digital Discovery module triggers a START signal. Additionally, the supply voltages can be immediately turned off with a RESET signal from the Digital Discovery module. -The synchronization signal is obtained using the Tektronix AFG3102 function generator, whose output is connected to the SMA connector on the PCB. - The generation of the calibration and programming voltages (pins 8 to 25 in Figure 4.7) is carried out in the PCB with a simple circuit consisting of an operational amplifier, a potentiometer, resistors and capacitors. - Digital outputs can be observed using oscilloscope probes or the logic analyzer included in the Digilent Digital Discovery module. - Analog outputs can be observed using oscilloscopes probes. A Keysight DSOX4104A oscilloscope has been used. 66 Chapter 4: ASIC CMOS-ONN Demonstrator Figure 4.9: (a) Block diagram and (b) photograph of the experimental setup b) a) 4.3. Experimental validation and characterization 67 Figure 4.10a depicts an alternative high-level view of this experimental set-up. The arrows with different colors indicate the type of interaction and the arrow tip, the destiny of the interaction. A MATLAB script allows the automated configuration and initialization of every instrument. A GUI-app, shown in Figure 4.10b, was developed on the Digital Discovery vendor software [126] to have easy and simultaneous control over the programming and the turn-on of the DONN and the logic analyzer measurements. The digital design in the FPGA reads the control signals (VDD_ON, LOAD_CFG and RESET) and generates the programming and the operating signals for the DONN: clock and bitstream for the programming, the SHIL trigger and the oscillator supply voltages. Depending on the position of the FPGA switches, it selects the network configuration to be programmed and, also, if the experiment will be executed single or multiple times. Once the experiment is started, the logic analyzer acquires the measurement of the nine digital outputs as time-stamped events that are saved in a comma-separated value (CSV) file. The CSV files are later post-processed with MATLAB scripts to obtain the relevant information: the oscillators phase difference, oscillating periods, stability, time to achieve a stable phase relation (DONN state), etc. Figure 4.10: (a) Diagram of instruments involved in the set-up. (b) Front panel of the GUI-app for controlling the ONN experiments. a) b) 68 Chapter 4: ASIC CMOS-ONN Demonstrator Evaluating oscillator behavior Although the DONN system has only digital outputs, an isolated differential oscillator, identical to the ones in the DONN, was also included in the chip and connected to analog pads in order to be able to observe its behavior. Figure 4.11 depicts experimental waveforms for that analog oscillator. Expected behavior is observed with the two outputs being out of phase. As previously pointed out, oscillations are not full swing. Output average voltage ranges from 379mV to 763mV. Note the small oscillation amplitude justifies the required carefully designed output stage included for digitalization. Figure 4.11: Experimental waveforms for a differential oscillator with analog outputs. Figure 4.12 depicts the two outputs of a differential oscillators after digitalization with the output stage described in previous Section. It can be also observed that they are out of phase showing correct operation. Figure 4.12: Experimental waveforms for differential oscillator outputs after digitalization. Figure 4.13 illustrates the critical role of the output stage. It depicts obtained waveforms for one of the differential oscillators for different configurations of the output buffer stage. It can be observed that the duty cycle of the digital oscillator outputs changes. In the last configuration, an undesired glitch is observed. Avoiding this behavior was the reason to include the Schmitt-Trigger buffer instead of a simpler one for converting the oscillator output to full swing. 4.4. Operating the CMOS-ONN as AM 75 - The pattern inferred in post layout simulation. This post-layout simulations have been carried out without mismatching. SHIL signal is applied. - Experimental retrieved pattern. Different experiments are reported. They differ in the applied synaptic voltages, as well as in the operation conditions, as can be seen in the table. Five trials have been carried out for each input pattern. The number of times P1 is retrieved followed by the number of times P2 is retrieved is depicted. Reading operation is carried out 100µs after the application of the test patterns. It can be observed that simulation results match perfectly those obtained with the HNN model. For the ten patterns that are closer to one of the stored patterns than to the others, that one is retrieved, as expected. For the six test patterns which are at equal distance from both stored patterns, a stable output pattern which is neither of the stored ones is obtained. Experiment EXP1 corresponds to reproducing the conditions of the simulation. Since mismatch was not considered in the simulations, and so, the frequency of all the oscillators was identical, calibration has been applied. SHIL was applied as in simulations. SHIL frequency was different because fabricated oscillators were faster than predicted by post-layout analysis. Also note that a slightly higher synapse voltage was required in the experiment. EXP1 results exhibit differences with respect to simulation results. First, no spurious patterns were retrieved. Instead, for those test patterns at equal distance from P1 and P2, one of them was retrieved. However, not always the same one. That is, a non-deterministic behaviour was observed. Secondly, for two of the test patterns closer to P2 than to P1 (T3 and T8), one (T3) or two trials (T8) failed to retrieve the expected pattern. However, for both test patterns, the correct one was inferred in the majority of the trials. EXP2 differed from EXP1 only in the synapse voltage applied, which have been increased from 0.85V to 0.95V. Note, that then both T3 and T8 correctly retrieved P2 in the five trials. However, T7 did not. We applied ten times T7 and retrieved eight times the expected pattern. Although the number of trials was too low to allow to extract final conclusions, obtained results seemed to indicate that increasing the strength of the coupling, the DONN performance was improved. Additionally, two cases where test evaluation returned a spurious pattern with both the model and post-layout simulation, resulted into unstable phases´ relationships (T11, denoted as ‘x#’ in Table 4.8) when evaluated with the real ASIC. In EXP3, oscillators were not calibrated. Results were worse in terms of tolerance to noise. Note that only in three of the test patterns, the correct expected pattern was retrieved in all the trials. Finally, in EXP4, the oscillators were calibrated but SHIL is not applied. Better results than in EXP3 were obtained. No results were reported for the DONN operating without calibration and without SHIL because this experiment was not successful. 76 Chapter 4: ASIC CMOS-ONN Demonstrator Figure 4.19: (a) The learned patterns and (b) the test patterns selected for the 3x3 DONN AM. Table 4.8: DONN experiment results corresponding to the AM. Expected HNN ASIC Post-Layout Simulation ASIC EXP1 ASIC EXP2 ASIC EXP3 ASIC EXP4 SHIL - - 8MHz 14.3MHz 14.3MHz 14.3MHz No Positive weights - - Vp = 0.8V Vn = 0V Vp = 0.85V Vn = 0V Vp = 0.95V Vn = 0V Vp = 0.95V Vn = 0V Vp = 0.95V Vn = 0V Negative weights - - Vp = 0V Vn = 0.8V Vp = 0V Vn = 0.85V Vp = 0V Vn = 0.95V Vp = 0V Vn = 0.95V Vp = 0V Vn = 0.95V Zero weights - - Vp = 0V Vn = 0V Vp = 0V Vn = 0V Vp = 0V Vn = 0V Vp = 0V Vn = 0V Vp = 0V Vn = 0V Calibration - - - Yes Yes No Yes P1 P1 P1 P1 5/0 5/0 5/0 5/0 P2 P2 P2 P2 0/5 0/5 0/5 0/5 T1 (4) P1, P2 Spurious Spurious 1/4 2/3 3/2 2/1/x2 T2 (3) P1 P1 P1 5/0 5/0 4/1 5/0 T3 (3) P2 P2 P2 1/4 0/5 1/4 0/5 T4 (4) P1, P2 Spurious Spurious 2/3 5/0 2/3 0/1/x4 T5 (3) P2 P2 P2 0/5 0/5 0/5 0/5 T6 (4) P1, P2 Spurious Spurious 1/4 4/1 1/4 1/2/x2 T7 (2) P2 P2 P2 0/5 2/3 1/4 5/0 T8 (3) P2 P2 P2 2/3 0/5 2/3 0/5 T9 (3) P1 P1 P1 5/0 5/0 5/0 5/0 T10 (2) P1 P1 P1 5/0 5/0 3/2 5/0 T11 (4) P1, P2 Spurious Spurious 2/3 3/0/x2 0/5 2/0/x3 T12 (3) P1 P1 P1 5/0 5/0 5/0 5/0 T13 (4) P1, P2 Spurious Spurious 3/2 5/0 2/3 3/0/x2 T14 (3) P1 P1 P1 5/0 5/0 4/1 5/0 T15 (3) P2 P2 P2 0/5 0/5 1/4 0/5 T16 (4) P1, P2 Spurious Spurious 4/1 5/0 1/4 0/1/x4 4.4. Operating the CMOS-ONN as AM 77 In order to measure the repeatability and the stability of the DONN operating as AM, another set of experiments with a large number of trials was carried out. Additionally, the external SHIL generator was synchronized with the power-on of the oscillators to ensure that the situation of the SHIL signal and the start of the oscillations was identical for each trial. In previous described experiments we were using a SHIL signal that was completely independent from the DONN control signals. Table 4.9 depicts the obtained results for 100 trials. Synapses control voltages correspond to those of EXP2. For each trial, the reading operation was carried out at different time instants after the application of the test patterns, concretely at 3, 10 and 720 oscillation cycles from the beginning. Considering that the oscillation period was approximately 140ns, these measurements corresponded to 42ns, 1.4µs and 100.5µs. Particularly, T7 was the only test pattern closest to a single store pattern that did not converge to the expected pattern in 90 out of 100 at the first reading at 3 cycles, but the DONN quickly inferred and stabilized in the correct pattern from the second reading at 10 cycles, 1 us later. Additionally, it can be observed from the third read at 720 cycles that the retrieved pattern is kept. Even for the test pattern at equal distance from both stores patterns one of them is retrieved in the third reading without spurious. So, it is concluded that the DONN successfully stores the two patterns. Table 4.9: Summary of experimental results corresponding to the AM. Pattern Expected Read (@ 3 cycles) Read (@ 10 cycles) Read (@ 720 cycles) P1 P1 100/0 100/0 100/0 P2 P2 0/100 0/100 0/100 T1 (4) P1, P2 99/1 99/1 99/1 T2 (3) P1 100/0 100/0 100/0 T3 (3) P2 0/100 0/100 0/100 T4 (4) P1, P2 53/25/22sp 54/43/3sp 54/46 T5 (3) P2 0/100 0/100 0/100 T6 (4) P1, P2 97/1/2sp 97/2/1sp 98/2 T7 (2) P2 0/10 0/100 0/100 T8 (3) P2 0/100 0/100 0/100 T9 (3) P1 100/0 100/0 100/0 T10 (2) P1 100/0 100/0 100/0 T11 (4) P1, P2 0/75/25sp 0/96/4sp 0/100 T12 (3) P1 100/0 100/0 100/0 T13 (4) P1, P2 0/69/31sp 0/95/5sp 0/100 T14 (3) P1 100/0 100/0 100/0 T15 (3) P2 0/100 0/100 0/100 T16 (4) P1, P2 1/99 0/100 0/100 The exercise of storing three patterns was also addressed. Figure 4.20 illustrates them. Note that the coupling matrix associated with this configuration required a greater weight precision since it used more than one unique value for positive and negative couplings. Strong positive (negative) couplings were coded with VP=1.2V and VN=0V (VP=0V and VN=1.2V), while the weak positive (negative) couplings with VP=0.48V and VN=0V (VP=0V and VN=0.48V) and uncoupled with VP=VN=0V. Several experiments with 100 trials for each run were performed to prove the storage capabilities of the network. The injected SHIL signal was operating at 14.2MHz. 78 Chapter 4: ASIC CMOS-ONN Demonstrator Figure 4.20: Training patterns for the three-patterns AM evaluation. The average of the successful trials for the first and the third pattern were 96.8% and 92%. However, the second pattern tended to converge to a spurious state. It was only retrieved in 13.6% on average while the spurious was achieved in the 84% of the trials. Although post-layout simulation results showed the DONN to store these three patterns, simulation set up was too much ideal. Thus, it is not surprising that, at least in our restricted experiments, we were not able to achieve it. These results suggested that further experiments with the calibration and programming voltages could be required. Operating the CMOS-ONN as IM It was described in the introductory Chapter that an ONN can be used to implement an IM able to obtain the ground state of an Ising model. It was also introduced that the Max-Cut optimization problem can be formulated as an Ising model. In this section we described the results of using our DONN for obtaining the Max-Cut of different 9-node graphs. Figure 4.21 depicts the three 9node graphs that have been explored. According to Section 1.5, in order to solve Max-Cut problems, the ONN must be programmed such that edges are translated to negative couplings. The corresponding synapses were coded with VP=0V and VN=0.85V while the uncoupled had VP=VN=0V. It was required to apply SHIL at 14MHz to implement the binary spins of the IM. The optimum cut for the different graphs is depicted in Table 4.10 along with the results of the experiments carried out with the ASIC. As explained, the cut is the number of edges shared between two nodes that belong to the two different subsets (determined by the assigned spin). Particularly, in this IM implementation, the Ising spin was represented with the binarized oscillator phase, with respect to a selected phase reference. The state of the system corresponded to the whole collection of spins. Each graph instance was evaluated gathering 100 trials each time. The duration of each trial was 100μs. During the post-processing with the developed MATLAB scripts, the oscillator states for each trial were evaluated in segments of 5μs, which led to 20 measured states that were evaluated in terms of cuts. That is, the binary states were introduced as input in one of the scripts to compute the obtained cut value. So, for each run, 2000 states were read (20 measurements · 100 trials). The number of the states providing an optimum cut value is reported in Table 4.10 (‘Optimum Cases’ column). Additionally, the different numbers of cut values obtained are also reported. Note that in most cases, the experiments were conducted with all the switches in an identical position (OFF). This corresponds, as was already explained, to all the oscillators being initialized in the same state or phase. For graph G9A, a different initial state with only one switch (‘Switch 1’) ON was also tried. It can be clearly observed that the obtained results are significantly different. The optimum solution is more frequently read from the DONN state with the second initial state. The dependence of the obtained results on the initial DONN state is not surprising. In fact, in the previous section, this was exploited to implement the AM functionality. Different test patterns converged to different stable states. The network evolves to minimize energy but 4.5. Operating the CMOS-ONN as IM 79 distinct minimum energy states are reached depending on the starting point. In the IM application, a non-optimum solution to the associated Max-Cut problem corresponds to a local minimum of energy, while a global one is the optimum solution. Figure 4.21: Graphs with 9 nodes proposed for the Max-Cut hard-problem: (a) G9A, (b) G9B, (c) G9C. Table 4.10: ASIC results for solving Max-Cut problem associated to the 9-node graphs. Graph Instance Optimum Cut Initial state Optimum Cases (#) Obtained Cut values G9A 8 All switches OFF 1619 2, 4, 6, 8 G9A 8 Switch 1 ON 1905 6, 8 G9B 9 All switches OFF 1262 2 to 9 G9C 11 All switches OFF 1640 5 to 11 Figure 4.22 shows the temporal evolution of the number of trials that achieved the optimum cut in each experiment. That is, for each of the 20 measured points, the number of trials that obtained the optimum cut value is depicted. Figure 4.22 shows a quick evolution of a large fraction of trials for G9A and G9B graphs to the optimum solution in the first 5μs. The graph G9C has a higher number of edges. The impact of this more complex configuration was that the growth of the number of solutions that found an optimal solution is slower in comparison with the other graphs. It can also be observed that the DONN behaved differently for G9B compared to the other two. A high number of trials found an optimum solution quickly, but they were not stable through time. Figure 4.23 depicts the number of trials versus the number of optimum readings and the number of optimum readings for each trial is collected in the inset tables. For example, for G9A, the 64% of the trials achieved an optimum solution in 19 of the 20 readings and 16% of the trials failed to obtain an optimum solution. The later number reduces to 0 and 3 for G9A with switch on and G9C respectively. The first point to note is that, despite having the SHIL signal synchronized, not all the trials are identical. This also occurred for some test patterns at equal distance from the stored ones in the AM application. It is due to the impact of distinct sources of noise. However, there is much more variability among trials in the IM experiments than in the AM. The level of that noise sensitivity depends on many factors including the shape of the energy landscape of the oscillator system. In this sense, it is interesting to note the differences between the two-pattern AM example and the Max-Cuts experiments. For that, we look the systems from the point of view of the constraints associated to their weight matrixes. The AM example has weights +1, -1 and 0. A +1 (-1) constraint is satisfied if the two associated oscillators are in phase (out of phase). There is no constraint associated to null weight. The systems implementing the Max-Cut problem have -1 and 0 weight values. The two stored patterns in the AM application satisfy all the constraints. However, for the three Max-Cut experimented systems, none network state satisfies all the constraints. It is said that there is contention in the system. O1 O2 O3 O4 O5 O6 O7 O8 O9 1 1 1 1 1 1 1 a) b) c) 80 Chapter 4: ASIC CMOS-ONN Demonstrator Figure 4.22: Temporal evolution of the number of trials that achieved the optimum cut value for each graph: (a) G9A, (b) G9B, (c) G9C. 0 10 20 30 40 50 60 70 80 90 100 110 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 Optimum Cases (#) Measurement time (μs) G9A (sw1=0) G9A (sw1=1) 0 10 20 30 40 50 60 70 80 90 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 Optimum Cases (#) Measurement time (μs) G9B 0 10 20 30 40 50 60 70 80 90 100 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 Optimum Cases (#) Measurement time (μs) G9C a) b) c) 4.5. Operating the CMOS-ONN as IM 81 Figure 4.23: Number of measured optimum states upon each trial. (a) G9A, (b) G9B, (c) G9C. Inset tables depict the percentage of trials with the same number of measured optimum. 0 2 4 6 8 10 12 14 16 18 20 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100 Optimum Cases (#) Trial (#) G9A (sw1=0) G9A (sw1=1) 0 2 4 6 8 10 12 14 16 18 20 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100 Optimum Cases (#) Trial (#) G9B 0 2 4 6 8 10 12 14 16 18 20 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100 Optimum Cases (#) Trial (#) G9C a) b) c) 82 Chapter 4: ASIC CMOS-ONN Demonstrator Conclusions The analog CMOS-ONN ASIC demonstrator in the NeurONN project and its set-up description have been presented in this Chapter, reporting the oscillatory dynamics and the evaluation on two target applications. It consists of a 9-neuron CMOS-ONN resembling a VO2-based ONN. It uses a CMOS sub-circuit emulating the I-V characteristic of VO2 devices to build an architecture based on differential oscillators. The synapse is implemented with a 6-transistor bridge topology enabling resistive coupling among oscillators. Both positive and negative weights can be realized. The fabricated ONN-ASIC is programmable, including all-to-all connections, with a large degree of controllability and observability to be able to deep into the synchronization dynamics of coupled non-linear oscillators, on which basis the ONN computation is carried out. Several experiments have been carried out with different coupling and calibration configurations, along with a frequency characterization of the oscillators. The impact of SHIL and the coupling strength has been explored with respect to the synchronization of the oscillators. Different experiments operating the CMOS-ONN as AM and IM have also been reported, showing the utility of the demonstrator to showcase real scenarios of analog computation with oscillators. Both applications required SHIL to enhance the stability of the measurements since it has been observed that, despite a good synchronization could be achieved with the voltage calibration, some sudden phase shifts occurred due to non-controllable noise sources, alternating the ONN state between different stable situations. In contrast with the digital ONN described in Chapter 2, the insights collected during the development and test of this demonstrator have a meaningful added value, helping with the planning of the tasks and goals of the final VO2 demonstrators reported in Chapter 5. Chapter 5 5.Experiments with Coupled VO2-based Oscillators Introduction Previous Chapters describes two ONN demonstrators built with commercial technologies. Chapter 2 presents a digital implementation prototyped in an FPGA, while Chapter 4 focuses on an analog ONN integrated in a 65nm CMOS technology. Both AM and IM functionalities were explored experimentally using the second demonstrator. The acquired experience was applied to build the final demonstrators using fabricated VO2 devices for the oscillators. These demonstrators were carried out as a joint collaboration in the context of the NeurONN project at the IBM-Research Zurich facilities, where they have the capability to fabricate the VO2 material. As mentioned in Chapter 1, IMs with coupled oscillators that exploit their dynamics of evolution towards a minimal energy state have been proposed and experimentally validated in different technologies. Notably, experimental implementations of an Ising solver using coupled phasetransition nano-oscillators with VO2 devices have been reported [46], [86], [128]. Experimentally calibrated numerical simulations show they exhibit high energy efficiency and potential for improving by several orders of magnitude over CPU, GPU, quantum, and photonic approaches [86]. Furthermore, in [86] the performance of the VO2-based IM is improved by emulating classical annealing by applying a SHIL signal with an increasing amplitude. Various experimental tests were conducted there to explore the operation of the VO2-based ONNs. Emphasis was placed on the VO2-based IMs for solving Max-Cut optimization problems. Also, the solving of Max-3Sat problems were addressed. Interestingly, the Graph Coloring problem associated to the same graphs can also be solved with the ONNs when operating them without forcing the binarization of the phases. Note that these applications require implementing an ONN with only negative weights, which can be successfully realized through capacitive coupling of single-ended VO2-based oscillators. Efforts were also made to evaluate its application as AM for pattern recognition. However, this typically requires both positive and negative weights. This translates into both resistive and capacitive coupling in the ONN using single-ended oscillators, but resistive coupling proved to be quite challenging because of the parasitic effects of the setup. This limitation can be overcome using differential oscillators. Nevertheless, due to the reduced number of VO2 devices available to probe simultaneously, and since the differential oscillator uses two of them, only a small-scale DONN could be experimented with. Other experiments were also explored as the phase-encoded minority gate or different synchronization injection techniques. In addition, motivated by the annealing reported in [86], a set of simulation experiments has been designed to gain insight into how to operate ONNs to improve their performance as IMs. 84 Chapter 5: Experiments with Coupled VO2-based Oscillators The remainder of the Chapter is organized as follows. Section 5.2 contains the description of the CMOS-compatible VO2-based oscillators and the experimental setup in which different experiments exploring VO2-based ONNs were carried out. Section 5.3 describes the results of applying the ONNs to solve the Max-Cut and the Max-3SAT optimization problems. Section 5.4 reports on other experiments carried out with the ONNs. Section 5.5 covers the simulation experiments conducted to derive a suitable operation protocol for IMs. Finally, Section 5.6 is devoted to conclusions. Description of fabricated VO2-based oscillators and experimental setup The VO2 devices were fabricated on a silicon platform with a hafnium oxide interlayer to create a VO2 granular film comprised between two metallic (platinum) electrodes whose cross-section define an area where current can flow and generate thermal filaments through Joule heating. This manufacturing process followed semiconductor industry standards, facilitating the integration of the devices at the Back-End-of-Line while ensuring their compatibility with CMOS technology. The detailed fabrication process and the basics on fundamental operation can be found in [54]. Each fabricated die held 4 arrays with 13 fields, where each field was composed of 9 two-terminal VO2 crossbar devices, as Figure 5.1 illustrates. There were devices with different cross-section sizes. The VO2 thickness was kept constant at 60nm. The number of devices simultaneously accessible was delimited by the 18 electrical probes available, which corresponded to a single field. The measured VO2 samples were contained in a vacuum chamber, shown inside the oval mark in Figure 5.2a, at controlled temperature (typically 220K). The probes were wired to an external custom PCB (square mark in Figure 5.2a and Figure 5.2b) where the discrete components were mounted and the instruments for the generation and the measurement of the signals were connected. The interface PCB had a stacked design, so the upper PCB included the holes to place the discrete coupling elements in the so-called coupling matrix. It contained two triangular-matrix configurations that corresponded with 36 possible interconnections for each configuration. Also, a programmable resistor module (circle mark in Figure 5.2a) was incorporated in the instruments for the series resistance required to implement the oscillators. The configuration of the instruments was carried out with a custom application built with LabView, allowing to define all the required parameters and arbitrary waveforms. Figure 5.1: Illustration of fabricated die structure. 5.3. Experimental IM with VO2-based oscillators 91 Figure 5.6 presents the results of repeated Max-Cut experiments on graphs C, D, E, and G, using six to nine interconnected VO2 oscillators. The fraction of trials per number of obtained cuts is shown. In the case of graph C, which has the highest connection density (η > 0.7), the network attain stability in the ground state in over 62% of the trial runs. For all graphs except D, the best solution is found in the majority of cases, exceeding 42% for the largest graph employing nine oscillators. Similar to the findings demonstrated in Figure 5.5, the ONNs shown in Figure 5.6 consistently reach their stable states within 15 oscillation cycles. This highlights the potential for faster time execution of massive parallel computations through coupled oscillators [18]. It should be noted that most graphs in Figure 5.5 and Figure 5.6 are either planar, i.e. they can be drawn on a plane without edges intersecting, or nearly planar [132]. Although algorithms exist to solve the MaxCut problem in polynomial time for planar graphs [132], our results in Figure 5.5 and Figure 5.6 demonstrate notable efficiency by attaining solutions within 15 oscillation cycles. It will be essential to reevaluate this level of performance when dealing with denser non-planar graphs involving a greater number of VO2 oscillators. Figure 5.6: Distribution of the cuts attained for graphs involving 6 to 9 coupled VO2 oscillators to solve the Max-cut problem. The optimal solution is found in most instances for graphs C, E, and G. Material reproduced with permission from [62]. 92 Chapter 5: Experiments with Coupled VO2-based Oscillators Max-3SAT problem The 3SAT problem consists in finding a Boolean combination of variables in a way that satisfies a given formula ℱ [133]. This formula is made up of smaller parts called clauses 𝐶, and each one contains three specific pieces of information called literals [133]. ℱ = 𝐶  ∧ 𝐶  ∧ … ∧ 𝐶    ∧ 𝐶  ( 5.2 ) Where 𝐶 is the inclusive disjunction of three literals 𝑥, such that: 𝐶  =  𝑥   ∨ 𝑥   ∨ 𝑥    (5.3) Having a dedicated hardware solver for 3SAT presents a significant potential to accelerate the computation of NP-complete problems, as these problems can all be reduced to 3SAT [133]. The NP-hard Max-3SAT problem seeks the maximum number of satisfied clauses 𝐾. If the solution of Max-3SAT says that all the clauses are satisfiable, then the corresponding 3SAT problem is satisfiable [69], [134]. There are different formulations to encode this NP-hard problem into the Ising paradigm [69]. Typically, the Max-3SAT problem addressed with IM requires the addition of an external field that interacts with all the spins in one direction. Thus, the Ising Hamiltonian associated to the energy-state landscape is formulated as: 𝐻 = −  𝐽  𝑠  𝑠         −  ℎ  𝑠      (5.1) Where 𝐽 is the coupling coefficient or the interaction between units 𝑖 and 𝑗, 𝑠 is the spin (up ↑ or down ↓) of unit 𝑖. Also, ℎ is the external bias, that represents the interaction of unit 𝑖 with an external unit with fixed spin value. The external bias is capacitively injected as an additional signal with the same frequency as the coupled oscillators. This is similar to applying SHIL but the injected signal is the first harmonic and the resulting impact is that all the oscillators´ phases are pushed to be in-phase together. Both SHIL and the first-harmonic injection locking (FHIL) were provided through their corresponding capacitors connected to the outputs. Developing the Maximum Independent Set formulation [134], [135] to construct the equivalent graph, each literal is mapped to a node (VO2 oscillator), while each clause corresponds to an interconnection (capacitor), forming the triangles such as the ones shown in the three examples in Figure 5.7. Furthermore, all complementary literals are interconnected (blue connections in Figure 5.7), as they cannot simultaneously satisfy the two corresponding clauses. Figure 5.7 also presents the results of repeated Max-3SAT experiments on graphs involving six (𝓕𝟏) and nine (𝓕𝟐 and 𝓕𝟑) oscillators. Note that the maximum value of 𝑲 is determined by the number of clauses. The values of the electrical circuit parameters chosen for each graph are reported in Table 5.2. ℱ  = ( 𝑥 1 ∨ 𝑥 2  ∨ 𝑥 3 ) ∧ ( 𝑥 1 ∨ 𝑥 2 ∨ 𝑥 3  ) (5.4) ℱ  = ( 𝑥 1 ∨ 𝑥 2  ∨ 𝑥 3 ) ∧ ( 𝑥 1  ∨ 𝑥 2  ∨ 𝑥 3  ) ∧ ( 𝑥 1 ∨ 𝑥 2 ∨ 𝑥 3  ) (5.5) ℱ  = ( 𝑥 1 ∨ 𝑥 2  ∨ 𝑥 3 ) ∧ ( 𝑥 1  ∨ 𝑥 2 ∨ 𝑥 3  ) ∧ ( 𝑥 1 ∨ 𝑥 2  ∨ 𝑥 3  ) (5.6) 5.4. Additional experiments 93 Figure 5.7: Distribution of the number of satisfying assignments 𝐾 attained for graphs involving 6 to 9 coupled VO2 oscillators to solve the Max-3SAT problem. Material from [62]. Table 5.2. Circuit parameters values used to solve the Max-3SAT problems in Figure 5.7. Graph VDD (V) RS (kΩ) CL (nF) Cc (nF) C2-HIL (pF) C1-HIL (pF) V2-HIL (V) 𝑽 𝒉 (V) 𝒇 (kHz) Active area (nm3) 𝐹  9.0 45 10.0 0.68 470 270 6.0 4.0 3.10 300 × 300 × 60 𝐹  6.75 48 10.0 1.0 470 270 9.0 4.0 2.47 150 × 150 × 60 𝐹  6.75 46 10.0 1.0 270 470 6.0 8.0 2.33 150 × 150 × 60 Figure 5.7 shows the system capacity to identify the Max-3SAT in all instances 𝓕𝟏, 𝓕𝟐, and 𝓕𝟑, with a success rate reaching up to 75% in the case of six coupled oscillators. Across all experimental trials, the network consistently achieved stability within 23 oscillation cycles, a significant improvement compared to digital methods. To potentially enhance the system ability to solve Max-3SAT, an approach could involve balancing and adapting better the parameters for each specific graph. For example, ensuring the coupling coefficients 𝐽 are higher than the bias ℎ, which translates into their respective capacitance values, would prevent getting trapped in local energy minima and ensure oscillators representing complement variables stabilize with opposite spins. The chosen amplitudes (Vh and V2-HIL) and capacitances (C1-HIL and C2-HIL) of the injected signals also impact significantly the eventual convergence state of the network. Further exploration of the influence of these values is required to maximize the likelihood of successfully solving the Max-3SAT problem. Nonetheless, the optimal solution is found in all instances and the results presented in Figure 5.7 serve as a proof of concept, proving 3SAT solvers can be implemented with a network of capacitively-coupled VO2 oscillators. Additional experiments This section covers a broad range of experiments carried out with the coupled VO2-based oscillators in addition to the IM application described in previous section: graph coloring problems; the differential oscillator configuration to emulate a differential neuron; the demonstration of a minority gate using discrete components and VO2 oscillators; and alternative methods for SHIL injection. 94 Chapter 5: Experiments with Coupled VO2-based Oscillators Graph Coloring problem The Graph Coloring problem consists in assigning a color to the nodes in a graph using the minimum number of colors, such that no connected nodes share the same color [53], [136]. Within the oscillatory paradigm, the oscillators represent the nodes and the capacitive couplings matches the edges. The color to each node is assigned from the phase ordering once a steady state is reached. Starting with one reference oscillator, the closest oscillator phase is checked if it is connected with the reference, and it is assigned a different color in the positive case. This procedure is iteratively done until all the oscillators have been assigned to a color. In Figure 5.8, the experimental results of four graph coloring problems involving three to six nodes are shown. The input geographical problem is mapped onto a network of VO2 oscillators, where the coupling capacitors represent borders between individual countries. The circuit rapidly converges to a steady state within 10 oscillation cycles, a significantly faster process compared to testing all potential combinations. For the most complex graph, the system converges to the solution within 10 oscillation cycles, wherein the phase order defines a color assignment for each node. Table 5.3 provides details regarding the different parameter values used in each experimental set, along with corresponding waveforms showcasing the stable state in Figure 5.8. The parameter sets should be considered as possible examples. Other combinations where the supply voltage (VDD) and the series resistance (RS) are adjusted with the devices active area could work as well. An adaption of the parameters for graphs with different number of nodes and edges is required as the capacitances vary. This makes it evident that there is a considerable challenge in establishing any form of metric for selecting values to color our map and converge rapidly towards a solution [53]. In our case, the careful choice of these values resulted in a cluster diameter, that is relatively small [53]. For the Central European and South American graphs, the cluster diameter averaged 33.5° and 30.0°, respectively. The inherent sparsity of these graphs without all-to-all connectivity makes coloring more challenging [53], resulting in larger cluster diameters compared to the Northern Europe and East Asia graphs. In the South America graph, characterized by nonuniform connectivity with varying degrees of connections on each node, the combined and unbalanced repelling effect of the coupling capacitances establishes a phase ordering among the oscillators [53]. This phase ordering can only approximate the minimum vertex coloring, causing uneven cluster spacing in the solution [53], [137]. More effective mapping techniques employing a circular ordering in color assignment could be employed to omit the impact of the cluster diameter in the reading of the solution, particularly for larger-scale graphs [53]. Table 5.3: Circuit parameters values used for the Graph Coloring problems in Figure 5.8. Graph VDD (V) RS (kΩ) CL (nF) Cc (nF) Active area (nm3) Northern Europe 5.5 40 10.0 2.2 100 × 50 × 60 Central Europe 9.0 40 10.0 0.68 300 × 300 × 60 East Asia 7.0 30 10.0 2.2 300 × 80 × 60 South America 5.5 42 11.5 1.0 100 × 50 × 60 5.4. Additional experiments 95 Figure 5.8: Experimental results involving 3 to 6 nodes to solve the Graph coloring problem: a) Input problem, b) Oscillator graph, measured c) Waveforms and d) Phase relationships of the ONN outputs. e) Number of cycles required to get to the stable state. 96 Chapter 5: Experiments with Coupled VO2-based Oscillators Differential Oscillatory Neural Network The CMOS ONN described in Chapter 4 used differential oscillators instead of single-ended ones in order to enable implementation of both positive and negative coupling using only resistive coupling elements. The implementation and coupling of VO2-based differential oscillators were also experimentally evaluated. This subsection contains electrical measurements of capacitively coupled differential oscillators. First, positive and negative couplings were demonstrated using only capacitors. Given that the capacitor between two oscillators imposes an anti-phase behavior, the positive coupling was implemented through the interconnection of the differential outputs in a crossed manner. That is, the positive output of one oscillator was connected to the negative output of the other oscillator and vice versa. The negative coupling was done connecting the corresponding positive outputs together (and the same with the negatives). Both couplings are illustrated in the schematics shown in Figure 5.9 along with the measured waveforms. The parameters for these experiments are: active area: 200×200×60nm3, VDD = 8V, Rs = 50kΩ, CL = 10nF, CD = 2.2nF, CC = 150pF. The positive output of the oscillators has been marked with cross symbols. It can be observed that in both cases, the outputs of each oscillator are out-of-phase as it is expected. Also, positive outputs are in-phase when the oscillators are positively coupled (Figure 5.9a) and out-of-phase when negatively coupled (Figure 5.9b). Following the described approach, a DONN with 4 differential oscillators was evaluated. Note it implied the use of 8 single-ended oscillators. The coupling elements were chosen to encode the following: 𝑊 = 󰇯 0 − 1 − 1 0 − 1 − 1 − 1 − 1 − 1 − 1 − 1 − 1 0 − 1 − 1 0 󰇰 (5.7) This matrix configures an IM for the Max-Cut problem of the fully-coupled four-node graph. An optimum solution for this graph is whatever division in groups of two outputs. It is worth to mention that this matrix, when interpreted as an AM, stores the patterns shown in Figure 5.10. Figure 5.10 shows the measured waveforms of the described DONN. Again, the positive output of the oscillators has been marked with cross symbols. Figure 5.10 shows that in this particular case, the obtained relationship is oscillator 1 and 3 in one phase and 2 and 4 in the other. Note that each one of the complemented outputs are found in a different group with respect to its corresponding output. 5.4. Additional experiments 97 Figure 5.9: Schematic and measured waveforms of DONN. (a) Positive coupling. (b) Negative coupling. b) a) 98 Chapter 5: Experiments with Coupled VO2-based Oscillators Figure 5.10: Waveforms of the 4-neuron DONN showing a stable solution. Stored patterns encoded in the negative weight matrix (Equation 5.7). Minority gate operation A majority (minority gate) gate is a logic gate where the output replicates the logic value (the complement of the logic value) in the majority of inputs. Table 5.4 depicts the logic functions implemented by majority and minority 3-inputs gates. Table 5.4: Truth table of 3-input majority and minority gates. X1 X2 X3 Major. Minor. 0 0 0 0 1 0 0 1 0 1 0 1 0 0 1 0 1 1 1 0 1 0 0 0 1 1 0 1 1 0 1 1 0 1 0 1 1 1 1 0 5.4. Additional experiments 99 The discretized behavior of the oscillating phase of the VO2 oscillator can be exploited to implement logic functions. Particularly, authors in [67] proposed the phase-encoded logic using VO2-based oscillators to implement the majority functionality. In the phase-encoded paradigm, logic input values are translated into binary phases in the oscillatory signals: 0º for low-value and 180º for high-value. The phase output signal of the oscillator replicates the most common value among the input phases. Logically, the oscillating frequency is the same for every input signal and the VO2 oscillator output signal. As a proof-of-concept of the phase-encoded logic gate implemented with VO2-based oscillator, an experimental validation of a 3-input minority gate was conducted with the fabricated VO2 devices and discrete electronic components. The schematic of the evaluated circuit is shown in Figure 5.11. Given that our particular implementation employs a capacitor as the coupling element, the correct operation takes place when the output phase is the complementary to the most repeated input phase. In fact, the capacitor constitutes a negative interaction between the phases of the coupled signals, pushing them apart in anti-phase (complemented) behavior. The particular parameters for this experiment are: active area: 150×150×60nm3, VDD = 7V, Rs = 40kΩ, CL = 10nF, CC = 150pF, VINPUTp,p = 6V, CSHIL = 270pF, VSHILp,p = 4V. Figure 5.12 illustrates the operation protocol. A synchronization signal (SHIL) forces the discretization of the output phase while a reference oscillatory signal is considered to have a common measurement framework. The input signals are kept inactive before the evaluation phase. Once they are activated, they oscillate for few cycles (10 cycles, 2ms), encoding in their phases the input to be evaluated by the minority gate. SHIL is turned off while the activation of the inputs to ease the interaction between the oscillator and the inputs. Once SHIL is activated again, the output phase is frozen and the response of the logic gate is read. VO 2 R VO 2 OUT R V SHIL C V DD V DD V INPUT-1 V INPUT-2 V INPUT-3 Figure 5.11: Schematic of the VO2-based minority gate with the capacitively-coupled inputs and SHIL. 100 Chapter 5: Experiments with Coupled VO2-based Oscillators Figure 5.12: Minority gate operation. Voltage waveforms of the SHIL and the inputs signals. The behavior of the implemented minority gate has been exhaustively validated. Figure 5.13 and Figure 5.14 depict waveforms associated to two different representative input combinations: the three input signals being in-phase and one input signal being out-of-phase with the other two. Note that in order to validate the circuit operation, each input combination is evaluated twice. Each time with a different initial output oscillation phase. Evaluations are marked as A and B in Figure 5.13a and Figure 5.14a. There are insets for each of them to better show the operation details. In Figure 5.13a, it can be observed that the gate output (blue waveform) is initially in-phase with the reference signal (green waveform). All inputs in-phase with the reference signal (red waveform) are activated (with SHIL off). When SHIL is activated (read operation mode) the output frozen in out of phase with the reference signal as expected according to the minority functionality. Second evaluation (Figure 5.13c) starts with the output out of phase. It can be observed that it remains out of phase after evaluation. Similarly, Figure 5.14b depicts the same initial situation being the gate output and the reference signal in-phase. For the evaluated input, one of the input phases is out-of-phase with the other two and with the reference signal. As in previous example, the obtained response is out-of-phase with the reference signal, as long as the evaluated input shared the same expected logic output value. Figure 5.14c shows the waveforms of the second evaluation with the similar expected behavior. 0 2 4 SHIL 0 5 Input 2 5 5.5 6 6.5 7 7.5 8 8.5 9 9.5 Time (ms) 0 5 Input 3 0 5 Input 1 5.5. Operation of ONNs as IMs 107 Figure 5.19: A 4-node Max-Cut problem. (a) Optimal solution (4 cuts). (b) Energy landscape. A key component of hardware IMs that can be fundamental in efficiently solving COPs is the internal noise of the circuit, which helps escape from suboptimal solutions (local energy minimum). In fact, in digital IMs, random number generators are used for this [138], [139]. Moreover, different authors have stated the critical role of inherent stochasticity in the implementation of IMs, which is identified with a probabilistic exploration [59], [86]. Because of this, they report that many trials are conducted so that the chance of obtaining the optimum solution increases. It is well-known and has been reported in the literature that actual VO2 devices exhibit variability in the state transition voltages. In [140], it is reported that VIMT exhibits variability more than seven times larger than for VMIT. It is also reported that it varies stochastically for all cycles without a reduction in the mean or range and can be approximated by a normal distribution. Additionally, the impact of the insulating-to-metal variability in the performance of coupled oscillator systems is much greater than the one associated with the reverse phase transition [141]. Thus, a cycle-to-cycle variability of its insulating to metal voltage (VIMT) has been included in the VO2 model. A normal distribution is associated with it. That is, VIMT changes in time following a normal distribution around its nominal value, with a given standard deviation denoted NOISE or noise level. It translates in cycle-to-cycle variations in the oscillation period (jitter). The chosen name refers to the fact that it, not only allows us to model the experimental variability observed in said parameter, but also allows to capture the effect of other noise sources that manifest as jitter in the oscillation period. Note that there is also mathematical noise associated with the integration algorithms and their tolerances. Figure 5.20a shows simulation results for the cut value associated with the evolution of the 4oscillator system representing the graph of Figure 5.19a with two different noise levels. Identical initial state and synchronization signals have been used in both simulations. It can be observed that the system with a very small noise level becomes trapped in a local minimum with a cutvalue equal to 2, not a maximum. The system with more noise can reach the global optimum (minimum energy spin configuration). Figure 5.20b depicts the result of simulations with a very small noise level and two different amplitudes for the synchronization signal. This experiment shows how SHIL counteracts noise. Reducing the amplitude of the synchronization signal allows escaping from the local minimum. 108 Chapter 5: Experiments with Coupled VO2-based Oscillators That is, as already mentioned, there is a relationship between the different factors involved in the dynamic evolution of the coupled oscillator system. Another point to consider is the SHIL signal scheduling. Different schedulings have been proposed in the literature. In particular, an improvement in the success probability of reaching the optimum solution has been reported by linearly increasing SHIL amplitude compared to a constant SHIL [128]. This has been compared to a thermal annealing process in which the reduction of the temperature is associated with the reduction of the stochastic noise in the system as a consequence of the increase in the amplitude of SHIL signal. Figure 5.21 illustrates the potential of using a SHIL with smoothly increases its amplitude and its impact on the system. It can be observed that the application of the constant amplitude SHIL signal freezes the system in a non-optimum phase relationship with three oscillators in one phase (cut value equal to 2). This is in agreement with [79] in which it is stated that the freeze-out effects associated with the discretization of the oscillator phases are a limitation for the synchronization dynamics. In [142], we showed that the success probability can also be increased by delaying the application of the synchronization signal. In both cases, we are giving time for the system to evolve towards a minimal energy state before forcing a binary discretization of the oscillator phases, which occurs for an enough large SHIL amplitude. Contrary, if a strong enough synchronization signal is applied from the very beginning or too early, the system gets stuck in a stable state in the neighborhood of the initial one which, in general, can be suboptimal. That is a local minimum. Figure 5.20: Evolution of the number of cuts with time for (a) two different noise levels and (b) two different SHIL amplitudes. a) b) 5.5. Operation of ONNs as IMs 109 Figure 5.21: Waveforms for two SHIL schedules, a) constant amplitude and b) increasing amplitude. IM operation protocol On the basis of the analysis of the experiments carried out, we propose the following oscillatory IM operating protocol.  SHIL scheduling according to Figure 5.22. The VMAX parameter is the signal amplitude, and KPER controls the linear increasing/decreasing rate of its amplitude. It is measured in cycles of period, TOSC. The amplitude rises from 0 to VMAX in KPER cycles. The repetitive scheme of sweeping SHIL amplitude allows further exploration of the associated energy landscape of the Ising model, such as the use of many trials reported in the literature, and enables an important escape mechanism from local minimum solutions when SHIL strength is reduced.  The solution is selected as the best one from multiple readings. This aims to counteract the possibility of escaping a global minimum or worsening the quality of the solution. This could occur because of excessive noise in the system or as a consequence of SHIL strength reduction. Readings must be carried out with enough SHIL signal strength so that phases are binarized and so that the reading process is simplified. Readings are carried out around the points at which the SHIL signal reaches its maximum amplitude. Thus, how many readings take place depends on how long the IM is operated (TCOMP) and KPER. Figure 5.22: The proposed SHIL schedule and the parameters that describe its configuration. a) b ) 110 Chapter 5: Experiments with Coupled VO2-based Oscillators To evaluate the proposed operation protocol, a second Max-Cut instance, depicted in Figure 5.23 has been used in the simulation experiments. Note that it exhibits critical difference with respect to that previously analyzed (graph in Figure 5.19). The associated oscillator system presents phase contention. This means that it is not possible to fulfil all phase constraints imposed by the capacitors that connect pairs of oscillators in the associated IM. In addition, it is a non-balanced graph in the sense that nodes exhibit a different number of connections. The IM associated with the graph in Figure 5.23 has been evaluated with three different SHIL schedules illustrated in Figure 5.24: an already introduced increasing amplitude SHIL signal (SHIL1), a SHIL signal periodically switching between two different amplitudes (SHIL2) [143] and the proposed one (SHIL3). The multiple reading approach has been used in all the cases. For each of them, 20 Monte Carlo (MC) runs have been performed for each of the 64 possible initial binary phase combinations. MC analysis is used because of the random distribution included in the VO2 model. This experiment was carried out for three different values (no noise, 1mV and 10mV). Three TCOMP values were analyzed. TCOMP1 corresponds to 1.5 periods of the SHIL3 signal (2 readings), TCOMP2 corresponds to 9.5 periods and 10 readings, and TCOMP3 to 17.5 periods and 18 readings. For selected KPER, TCOMP1 is 12µs. VMAX = 0.4V has been chosen so that it is strong enough to produce the desired discretization of the phases and its amplitude is similar to that of the oscillators themselves. KPER = 20 has been selected. Figure 5.23: 6-node Max-Cut problem with two optimum solutions with 8 cuts. Figure 5.24: SHIL schedules used for the evaluation of the oscillatory IM. 1 6 5 2 3 4 1 6 5 2 3 4 5.5. Operation of ONNs as IMs 111 Table 5.5 shows the number of initial system states from which an optimum solution is obtained for different computation times. That is, at least one of their associated 20 MC simulations obtained an optimum solution. Note that a single simulation is carried out when no noise is modeled. It is clear that noise is beneficial in increasing the number of successful initial states, although this figure of merit alone does not completely support the above statement, but the improvement observed (especially for the SHIL1 and TCOMP1) between 1mV and 10mV of noise also suggests this. It can also be seen that the number of SHIL1 results is much worse than the other two for the shortest computation time. This can be explained since the capability of exploring the energy landscape and so escaping from a local minimum reduces once a strong enough SHIL is applied. Both SHIL2 and SHIL3 benefit from the reduction of the SHIL strength. When the computation time is increased, the chance of escaping due to noise is increased, and so the number of successful initial states for the SHIL1 experiment also increases significantly. Also, it can be observed that the noise level has little impact but for the SHIL1 and for the shortest computation time. This agrees with our above explanation. The higher the noise level increases the probability of escaping from local minima. Table 5.5: Initial states that converged to an optimum solution with SHIL1, SHIL2, and SHIL3, respectively. NOISE TCOMP1 TCOMP2 TCOMP3 0mV (no noise) 9, 18, 26 18, 16, 48 21, 49, 58 1mV 31, 61, 64 58, 64, 64 62, 64, 64 10mV 41, 62, 64 55, 64, 64 62, 64, 64 Much more insight into system performance can be obtained if the fraction of MC runs achieving the optimum solution is also analyzed. It can be considered as the success probability. A value of 1 for this figure of merit means that the 20 runs produce an optimum solution. Figure 5.25 shows such a fraction for each of the 64 initial states versus its normalized HD to the closest optimum solution. A HD equal to 0 means that the initial state is an optimum solution to the associated optimization problem. The differences are notorious among the three SHIL schedules. On the one hand, a fraction of successful runs is significantly higher for the proposed SHIL (black points are over 0.8 for all the initial states) than for the other two. The worst results are obtained with SHIL1. The second interesting aspect to be pointed out is that the results are almost independent of the initial state (the 64 points have very little dispersion) for the proposed SHIL, while this is not so for the other SHILs, especially for SHIL1. In fact, very few black points appear since many of them occupy the same position. This reduced dependence is also a consequence of reducing the SHIL strength during the operation, allowing a better exploration of the solution space. 112 Chapter 5: Experiments with Coupled VO2-based Oscillators Figure 5.25: Success ratio versus HD between the initial state and the closest optimum solution for the three SHIL schedules. Simulation parameters: VMAX = 0.4V, σNOISE = 1mV, TCOMP3. It is also interesting to analyze the fraction of successful runs for different computation times. In this analysis, the average and standard deviation over the 64 initial states have been used. A value of 1 for the average means that 1280 (64x20) runs produced an optimum result. Figure 5.26 depicts the average of the success rate including bars for the standard deviation. Note that there are two additional computation times with respect to previous experiments (TCOMP12 and TCOMP23). The limitations of SHIL1 are also evident from this figure. Both SHIL2 and SHIL3 achieved better results in TCOMP1 than SHIL1 in TCOMP3, although TCOMP1 is around ten times shorter than TCOMP3. The smaller standard deviation exhibited by SHIL3 also indicates that the initial state is less relevant for SHIL3, as already pointed out when describing Figure 5.25. In [86], the Max-Cut problem depicted in Figure 5.27a was experimentally solved using VO2 oscillators and the SHIL schedule identified in this work as SHIL1. They report a Max-Cut value equal to 10 with a 96% success rate with KPER over 600 oscillation cycles. We have simulated the same graph with SHIL3. Computation time have been fixed to 600 cycles for a fair comparison. In particular, three SHIL3 cycles with KPER = 100 have been applied. 20 Monte Carlo simulations for each of 50 different initial states have been carried out. Figure 5.27b depicts obtained results. Note that Max-Cut values higher than 10 have been obtained. In particular, the optimum solution to the problem (maximum-cut equals to 12) have been obtained in 46.4% of the trials. Moreover, maximum-cut values equal or greater than 10 have been achieved in 96.1% of all the simulations. The experiment in [86] has been also simulated with our VO2 models. The optimum Max-Cut value is obtained in 14.8% of the trials. Figure 5.26: Average and standard deviation of the success rate for the three SHIL schemes evaluated at five different computation times. 5.5. Operation of ONNs as IMs 113 Additionally, these results indicate that the proposed SHIL scheme could also translate into advantages in terms of energy efficiency. In [86], as already mentioned, analysis results that show the potential of VO2 based oscillatory IMs to drastically increase operations per second and per watt with respect to platforms conventionally used to solve optimization problems are reported. The ability of SHIL3 to achieve higher success rates in shorter computation times could only contribute to increasing this figure of merit with respect to [86] that applied SHIL1. Figure 5.27: a) 8-node Max-Cut problem reported in [86] with 12 cuts (b) Histogram representing the probability of success of obtaining at least 10, 11 or 12 cuts. 114 Chapter 5: Experiments with Coupled VO2-based Oscillators Conclusions This Chapter collects the experiments carried out with the fabricated VO2 devices during the joint collaboration at the IBM-Research Zurich facilities. The implemented demonstrators aimed to prove the capability of the VO2-based ONNs to implement IMs to solve COPs. It has been benchmarked different Max-Cut and Max-3-SAT problems with graphs up to 9 nodes, exploring the use of capacitively injected SHIL to ensure the synchronization of the fabricated oscillators. The intrinsic noise and variability of the VO2 and the set-up helps the ONN to further explore the Hamiltonian energy landscape, converging within very few cycles. The dependence of the supply voltage with the size of the fabricated material is exposed, being necessary a fine selection of applied voltage and the value of the passive discrete components to ease the synchrony of the oscillators. Moreover, additional experiments have been reported showcasing the graph-coloring solver and other circuits implementation as the VO2-based differential oscillators or the phase-based minority gate, as well as, different techniques of injecting the synchronization signal over the supply voltage terminal or being provided by an additional and buffered twice-frequency oscillator. Finally, a comprehensive study of the SHIL application, or protocol, to enhance the efficacy of the IM solver has been provided. The interplay between noise variability and the tuning of the strength of SHIL is assessed to increase the accuracy of the COPs solving. The simulation work in Section 5.5 shows the potential of increasing and decreasing the SHIL amplitude to escape from local minimum energy states in oscillatory IMs and the role of noise in this energy landscape exploration. Chapter 6 6.Conclusions This Ph.D. Dissertation has contributed to the validation and understanding of the oscillatory computing paradigm based on VO2-based ONNs, through several proof-of-concept demonstrators and practical applications. It has established a solid foundation for how electrical ONNs should be operated and configured while exploring methods to overcome the intrinsic constraints and limitations of such a complex non-linear dynamical system. The main contributions are presented following the order of presentation in this work:  The first fully digital ONN hardware has been designed as a proof-of-concept demonstrator of the oscillatory computing paradigm. It demonstrated the important feature of converging to a stable phase pattern within very few oscillation cycles, supporting the potential of the computing paradigm in terms of energy efficiency. Additionally, the design has been provided with fully programmable weights and an online learning operation mode.  The digital ONN has been implemented on an FPGA, constituting a fast prototype to validate suitable real applications. Specifically, prototypes for a digit recognition application and a responsive obstacle avoidance system in a mobile robot have been experimentally validated.  The extensive HNN literature on learning rules for AMs has been explored with the aim of developing learning algorithms compatible with the constraints imposed by analog ONNs.  Different training approaches with the Storkey learning rule have been reported. The limits in storage capacity were extended with any of the proposed approaches, being noteworthy that the rate of successful solutions when exploring the random order in the training set, was boosted to 100% when three repetitive iterations were applied. Specifically, the repetitive training showed a highly satisfactory impact on obtaining a successful weight matrix using reduced weight precision. For example, training with three iterations of the Storkey rule, the 60-neuron network storing 5 patterns obtained a success of 93.2% at retrieval accuracy using three bits. The same case with FP precision achieved 99.2% while the trained network with a single iteration of the Storkey rule only got 90%, not being able to succeed using a weight precision lower than 5 bits (89.6%).  A novel iterative algorithm (IRPUSH) has been proposed. It features configurable margin stability for the stored patterns, allowing the obtention of multiple successful matrix solutions, even with reduced weight precision without sacrificing AM capabilities. It has shown to be more competitive than any other evaluated algorithm, even when compared with the other nonconstrained iterative algorithm. In particular, IRPUSH was the unique rule that achieved a 96neuron network successfully storing all the 10-character patterns using a 3-bit weight precision and remarkably without a significant loss of retrieval accuracy (90.9% against the 92.7% obtained with the FP precision). The best score achieved by another algorithm, 88.8% applying DEB with FP precision, did not overpassed 3-bit IRPUSH performance. 116 Chapter 6: Conclusions  An analog 9-neuron CMOS-ONN ASIC has been designed and fabricated in TSMC 65nm technology. This demonstrator integrated a CMOS emulator of the VO2 material. Replicating the VO2 voltage-current characteristic enabled an early evaluation of the VO2-based ONN response. It has been designed with a large degree of controllability and observability to be able to deep into the synchronization dynamics of ONNs. As well, the ASIC evaluation board and the whole experimental set-up have been developed.  Measurements and characterization of the analog CMOS-ONN and the oscillator dynamics have been carried out. The operation of the ONN as AM and IM have been demonstrated, showing the utility of the ASIC to showcase real scenarios of analog computation with oscillators. Both applications proved the important role of SHIL.  The AM demonstrator storing two patterns proved the capability to recognize all the test patterns in the first three cycles. The results matched well with the obtained with the HNN model and the post-layout simulations.  The IM demonstrator was applied to three different Max-Cut problems. Concretely, the evaluated graphs showed an averaged success of 67%, 82%, and 89% finding the optimum case when measured after 35, 70, and 245 oscillating cycles, respectively.  Experiments with fabricated VO2 devices have been carried out during the secondment at IBM-Research Zürich facilities. The VO2-based ONN operation has been explored with an emphasis on solving COPs with the Ising formulation. ONNs up to 9 oscillators encoding Max-Cut and Max-3-SAT COPs have been implemented and evaluated. The more complex 9-node Max-Cut problem attains stability in the ground state (the optimal solution) in over 42% of the trial runs. All the IM demonstrators converged within very few cycles. Additionally, other applications have been explored resorting to the VO2-based oscillators.  A detailed analysis of the factor impacting the ONN performance as IM using electrical simulation has been carried out. It was evinced that the intrinsic noise and variability of the VO2 helped the ONN to further explore the Hamiltonian energy landscape. On this basis, an operation protocol for IMs has been developed which has been shown to improve the quality in the search for optimal solution for the COPs resolution. A SHIL signal is repeatedly and gradually turned on and off as a mechanism to escape from the local minimum. While considering that the specific objectives of the Ph.D. Dissertation have been satisfactorily accomplished, the developed research with our ONN demonstrators leaves open questions that can lead to future research efforts. As an immediate example, a second CMOS-ONN ASIC could be designed adding all the gained knowledge in different aspects. For example, from the learning aspect, the scalability challenge could be overcome: the IRPUSH algorithm can be adapted to limit the number of connections, therefore, reducing the quadratically-increasing number of connections in the all-to-all architecture. Also, design techniques to mitigate the impact of noise sources or layout parasites can be applied. It has been proved that the analog ONN benefits from the calibration voltage for tuning the frequency or the compensation provided by capacitor banks. In a similar way, subsequent ASIC designs could target a more modular approach enabling the electrical interfacing and the heterogeneous integration with other emerging devices, as could be memristive synapses or different oscillator implementations, advancing separately the design and test of alternative architectures. 123 [76] K. Y. Camsari, S. Salahuddin, and S. Datta, “Implementing p-bits With Embedded MTJ,” IEEE Electron Device Lett., vol. 38, no. 12, pp. 1767–1770, Dec. 2017, doi: 10.1109/LED.2017.2768321. [77] R. Hamerly et al., “Experimental investigation of performance differences between coherent Ising machines and a quantum annealer,” Sci. Adv., vol. 5, no. 5, p. eaau0823, May 2019, doi: 10.1126/sciadv.aau0823. [78] T. Inagaki et al., “A coherent Ising machine for 2000-node optimization problems,” Science, vol. 354, no. 6312, pp. 603–606, Nov. 2016, doi: 10.1126/science.aah4243. [79] Y. Haribara, S. Utsunomiya, and Y. Yamamoto, “A Coherent Ising Machine for MAX-CUT Problems: Performance Evaluation against Semidefinite Programming and Simulated Annealing,” in Principles and Methods of Quantum Information Technologies, Y. Yamamoto and K. Semba, Eds., Tokyo: Springer Japan, 2016, pp. 251–262. doi: 10.1007/978-4-43155756-2_12. [80] P. L. McMahon et al., “A fully programmable 100-spin coherent Ising machine with all-toall connections,” Science, vol. 354, no. 6312, pp. 614–617, Nov. 2016, doi: 10.1126/science.aah5178. [81] T. Wang, L. Wu, P. Nobel, and J. Roychowdhury, “Solving combinatorial optimisation problems using oscillator based Ising machines,” Nat. Comput., vol. 20, no. 2, pp. 287–306, Jun. 2021, doi: 10.1007/s11047-021-09845-3. [82] T. Wang, L. Wu, and J. Roychowdhury, “Late Breaking Results: New Computational Results and Hardware Prototypes for Oscillator-based Ising Machines,” Apr. 23, 2019, arXiv: arXiv:1904.10211. doi: 10.48550/arXiv.1904.10211. [83] T. Wang, L. Wu, and J. Roychowdhury, “New Computational Results and Hardware Prototypes for Oscillator-based Ising Machines,” in Proceedings of the 56th Annual Design Automation Conference 2019, Las Vegas NV USA: ACM, Jun. 2019, pp. 1–2. doi: 10.1145/3316781.3322473. [84] M. Graber and K. Hofmann, “A Versatile & Adjustable 400 Node CMOS Oscillator Based Ising Machine to Investigate and Optimize the Internal Computing Principle,” in 2022 IEEE 35th International System-on-Chip Conference (SOCC), Sep. 2022, pp. 1–6. doi: 10.1109/SOCC56010.2022.9908118. [85] M. K. Bashar, A. Mallick, D. S. Truesdell, B. H. Calhoun, S. Joshi, and N. Shukla, “Experimental Demonstration of a Reconfigurable Coupled Oscillator Platform to Solve the Max-Cut Problem,” IEEE J. Explor. Solid-State Comput. Devices Circuits, vol. 6, no. 2, pp. 116–121, Dec. 2020, doi: 10.1109/JXCDC.2020.3025994. [86] S. Dutta et al., “An Ising Hamiltonian solver based on coupled stochastic phase-transition nano-oscillators,” Nat. Electron., vol. 4, no. 7, Art. no. 7, Jul. 2021, doi: 10.1038/s41928021-00616-7. [87] J. Núñez, S. Thomann, H. Amrouch, and M. J. Avedillo, “Mitigating the Impact of Variability in NCFET-based Coupled-Oscillator Networks Applications,” in 2022 29th IEEE International Conference on Electronics, Circuits and Systems (ICECS), Oct. 2022, pp. 1–4. doi: 10.1109/ICECS202256217.2022.9970771. [88] D. I. Albertsson, M. Zahedinejad, A. Houshang, R. Khymyn, J. Åkerman, and A. Rusu, “Ultrafast Ising Machines using spin torque nano-oscillators,” Appl. Phys. Lett., vol. 118, no. 11, p. 112404, Mar. 2021, doi: 10.1063/5.0041575. [89] T. Jackson, S. Pagliarini, and L. Pileggi, “An Oscillatory Neural Network with Programmable Resistive Synapses in 28 Nm CMOS,” in 2018 IEEE International Conference on Rebooting Computing (ICRC), McLean, VA, USA: IEEE, Nov. 2018, pp. 1–7. doi: 10.1109/ICRC.2018.8638600. [90] “AMD VivadoTM Design Suite,” AMD. Accessed: Oct. 26, 2024. [Online]. Available: https://www.amd.com/es/products/software/adaptive-socs-and-fpgas/vivado.html [91] “Zynq-7000 SoC Data Sheet: Overview (DS190),” 2018. [92] Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278–2324, Nov. 1998, doi: 10.1109/5.726791. 124 References [93] Y. LeCun, C. Cortes, and C. Burges, “MNIST handwritten digit database.” Accessed: Oct. 26, 2024. [Online]. Available: https://yann.lecun.com/exdb/mnist/ [94] “MATLAB.” Accessed: Oct. 26, 2024. [Online]. Available: https://www.mathworks.com/products/matlab.html [95] B. J. Leiner, V. Q. Lorena, T. M. Cesar, and M. V. Lorenzo, “Hardware Architecture for FPGA Implementation of a Neural Network and Its Application in Images Processing,” in 2008 Electronics, Robotics and Automotive Mechanics Conference (CERMA ’08), Sep. 2008, pp. 405–410. doi: 10.1109/CERMA.2008.32. [96] W. Mansour, R. Ayoubi, H. Ziade, R. Velazco, and W. EL Falou, “An Optimal Implementation on FPGA of a Hopfield Neural Network,” Adv. Artif. Neural Syst., vol. 2011, no. 1, p. 189368, 2011, doi: 10.1155/2011/189368. [97] M. A. de A. de Sousa, E. L. Horta, S. T. Kofuji, and E. Del-Moral-Hernandez, “Architecture Analysis of an FPGA-Based Hopfield Neural Network,” Adv. Artif. Neural Syst., vol. 2014, no. 1, p. 602325, 2014, doi: 10.1155/2014/602325. [98] “[NeurONN]: Robot Obstacle Avoidance with Oscillatory Neural Networks (ONN) Implemented on an FPGA - YouTube.” Accessed: Sep. 06, 2024. [Online]. Available: https://www.youtube.com/watch?v=mO2lOBKSeDc [99] “[NeurONN]: ONN on FPGA Demonstrator - YouTube.” Accessed: Sep. 06, 2024. [Online]. Available: https://www.youtube.com/watch?v=eVfDEEpNKsQ [100]F. Rosenblatt, “The perceptron: A probabilistic model for information storage and organization in the brain.,” Psychol. Rev., vol. 65, no. 6, pp. 386–408, 1958, doi: 10.1037/h0042519. [101]L. F. Abbott, “Learning in neural network memories,” Netw. Comput. Neural Syst., vol. 1, no. 1, pp. 105–122, Jan. 1990, doi: 10.1088/0954-898X_1_1_008. [102]A. D. Bruce, A. Canning, B. Forrest, E. Gardner, and D. J. Wallace, “Learning and memory properties in fully connected networks,” in AIP Conference Proceedings, AIP, 1986, pp. 65– 70. doi: 10.1063/1.36221. [103]A. J. Storkey and R. Valabregue, “The Basins of Attraction of a New Hop eld Learning Rule”. [104]S. Schmidgall, R. Ziaei, J. Achterberg, L. Kirsch, S. P. Hajiseyedrazi, and J. Eshraghian, “Brain-inspired learning in artificial neural networks: A review,” APL Mach. Learn., vol. 2, no. 2, p. 021501, May 2024, doi: 10.1063/5.0186054. [105]C. Delacour and A. Todri-Sanial, “Mapping Hebbian Learning Rules to Coupling Resistances for Oscillatory Neural Networks,” Front. Neurosci., vol. 15, 2021, Accessed: Apr. 20, 2023. [Online]. Available: https://www.frontiersin.org/articles/10.3389/fnins.2021.694549 [106]T. Aonishi, “Phase transitions of an oscillator neural network with a standard Hebb learning rule,” Phys. Rev. E, vol. 58, no. 4, pp. 4865–4871, Oct. 1998, doi: 10.1103/PhysRevE.58.4865. [107]P. Maffezzoni, B. Bahr, Z. Zhang, and L. Daniel, “Analysis and Design of Boolean Associative Memories Made of Resonant Oscillator Arrays,” IEEE Trans. Circuits Syst. Regul. Pap., vol. 63, no. 11, pp. 1964–1973, Nov. 2016, doi: 10.1109/TCSI.2016.2596300. [108]J. Shamsi, M. J. Avedillo, B. Linares-Barranco, and T. Serrano-Gotarredona, “Hardware Implementation of Differential Oscillatory Neural Networks Using VO2-Based Oscillators and Memristor-Bridge Circuits,” Front. Neurosci., vol. 15, p. 674567, Jul. 2021, doi: 10.3389/fnins.2021.674567. [109]V. Folli, M. Leonetti, and G. Ruocco, “On the Maximum Storage Capacity of the Hopfield Model,” Front. Comput. Neurosci., vol. 10, Jan. 2017, doi: 10.3389/fncom.2016.00144. [110]A. Storkey and R. Valabregue, “A Hopfield learning rule with high capacity storage of timecorrelated patterns,” Electron. Lett., vol. 33(21), pp. 1803–1804, 1999. [111]S. Diederich and M. Opper, “Learning of correlated patterns in spin-glass networks by local learning rules,” Phys. Rev. Lett., vol. 58, no. 9, pp. 949–952, Mar. 1987, doi: 10.1103/PhysRevLett.58.949. 125 [112]L. Personnaz, I. Guyon, and G. Dreyfus, “Collective computational properties of neural networks: New learning mechanisms,” Phys. Rev. A, vol. 34, no. 5, pp. 4217–4228, Nov. 1986, doi: 10.1103/PhysRevA.34.4217. [113]L. Personnaz, I. Guyon, and G. Dreyfus, “Information storage and retrieval in spin-glass like neural networks,” J. Phys. Lett., vol. 46, no. 8, pp. 359–365, Apr. 1985, doi: 10.1051/jphyslet:01985004608035900. [114]P. Tolmachev and J. H. Manton, “New Insights on Learning Rules for Hopfield Networks: Memory and Objective Function Minimisation,” in 2020 International Joint Conference on Neural Networks (IJCNN), Jul. 2020, pp. 1–8. doi: 10.1109/IJCNN48605.2020.9207405. [115]B. Widrow and M. E. Hoff, “Adaptive switching circuits,” Defense Technical Information Center, Fort Belvoir, VA, Jun. 1960. doi: 10.21236/AD0241531. [116] A. Johannet, L. Personnaz, G. Dreyfus, J.-D. Gascuel, and M. Weinfeld, “Specification and implementation of a digital Hopfield-type associative memory with on-chip training,” IEEE Trans. Neural Netw., vol. 3, no. 4, pp. 529–539, Jul. 1992, doi: 10.1109/72.143369. [117]W. Krauth and M. Mezard, “Learning algorithms with optimal stability in neural networks,” J. Phys. Math. Gen., vol. 20, no. 11, pp. L745–L752, Aug. 1987, doi: 10.1088/03054470/20/11/013. [118] E. Gardner, “The space of interactions in neural network models,” J. Phys. Math. Gen., vol. 21, no. 1, pp. 257–270, Jan. 1988, doi: 10.1088/0305-4470/21/1/030. [119]G. Tanaka et al., “Spatially Arranged Sparse Recurrent Neural Networks for Energy Efficient Associative Memory,” IEEE Trans. Neural Netw. Learn. Syst., vol. 31, no. 1, pp. 24–38, Jan. 2020, doi: 10.1109/TNNLS.2019.2899344. [120]M. Jimenez-Traves, M. J. Avedillo, J. Nunez, and B. Linares-Barranco, “Enhancing Storage Capabilities of Oscillatory Neural Networks as Associative Memory,” in 2022 37th Conference on Design of Circuits and Integrated Circuits (DCIS), Pamplona, Spain: IEEE, Nov. 2022, pp. 01–05. doi: 10.1109/DCIS55711.2022.9970122. [121]“Google Colab.” Accessed: Oct. 26, 2024. [Online]. Available: https://colab.research.google.com/ [122]“Python.org,” Python.org. Accessed: Oct. 26, 2024. [Online]. Available: https://www.python.org/ [123]“Project Jupyter.” Accessed: Oct. 26, 2024. [Online]. Available: https://jupyter.org [124]M. Jiménez, J. Núñez, J. Shamsi, B. Linares-Barranco, and M. J. Avedillo, “Experimental demonstration of coupled differential oscillator networks for versatile applications,” Front. Neurosci., vol. 17, Dec. 2023, doi: 10.3389/fnins.2023.1294954. [125]D. A. Hodges and H. G. Jackson, Analysis and Design of Digital Integrated Circuits. McGraw-Hill, 1983. [126]“Digital Discovery - Digilent Reference.” Accessed: Sep. 13, 2024. [Online]. Available: https://digilent.com/reference/test-and-measurement/digitaldiscovery/start?srsltid=AfmBOorTcrYpA6m57wcjISAHy1NPSlVTE5kDNs42ZhkvQHvJcMOpeyh [127]J. Wu, L. Jiao, R. Li, and W. Chen, “Clustering dynamics of nonlinear oscillator network: Application to graph coloring problem,” Phys. Nonlinear Phenom., vol. 240, no. 24, pp. 1972–1978, Dec. 2011, doi: 10.1016/j.physd.2011.09.010. [128]S. Dutta, A. Khanna, and S. Datta, “Understanding the Continuous-Time Dynamics of Phase-Transition Nano-Oscillator-Based Ising Hamiltonian Solver,” IEEE J. Explor. SolidState Comput. Devices Circuits, vol. 6, no. 2, pp. 155–163, Dec. 2020, doi: 10.1109/JXCDC.2020.3045074. [129]E. Corti, C. Delacour, A. Todri-Sanial, and S. Karg, “Frequency Injection LockingControlled Oscillations for Synchronized Operations in VO2 Crossbar Devices,” in 2021 Device Research Conference (DRC), Jun. 2021, pp. 1–2. doi: 10.1109/DRC52342.2021.9467129. [130]O. Maher, N. Harnack, G. Indiveri, M. Sousa, B. Gotsmann, and S. Karg, “Solving optimization tasks power-efficiently exploiting VO2’s phase-change properties with Oscillating Neural Networks,” in 2023 Device Research Conference (DRC), Jun. 2023, pp. 1–2. doi: 10.1109/DRC58590.2023.10186951. 126 References [131]A. Mallick, M. K. Bashar, D. S. Truesdell, B. H. Calhoun, S. Joshi, and N. Shukla, “Using synchronized oscillators to compute the maximum independent set,” Nat. Commun., vol. 11, no. 1, p. 4689, Dec. 2020, doi: 10.1038/s41467-020-18445-1. [132]F. Hadlock, “Finding a Maximum Cut of a Planar Graph in Polynomial Time,” SIAM J. Comput., vol. 4, no. 3, pp. 221–225, Sep. 1975, doi: 10.1137/0204019. [133]R. M. Karp, “Reducibility among Combinatorial Problems,” in Complexity of Computer Computations, Springer, Boston, MA, 1972, pp. 85–103. doi: 10.1007/978-1-4684-20012_9. [134]W. Zhang, “Phase Transitions and Backbones of 3-SAT and Maximum 3-SAT,” in Principles and Practice of Constraint Programming — CP 2001, T. Walsh, Ed., Berlin, Heidelberg: Springer, 2001, pp. 153–167. doi: 10.1007/3-540-45578-7_11. [135]V. Choi, “Adiabatic Quantum Algorithms for the NP-Complete Maximum-Weight Independent Set, Exact Cover and 3SAT Problems,” ArXiv, Apr. 2010, Accessed: Sep. 27, 2024. [Online]. Available: https://www.semanticscholar.org/paper/Adiabatic-QuantumAlgorithms-for-the-NP-Complete-Choi/1fcb4d5749074714ca2eb0b56ab986ca66fd0b95 [136]P. Maritz and S. Mouton, “Francis Guthrie: A Colourful Life,” Math. Intell., vol. 34, no. 3, pp. 67–75, Sep. 2012, doi: 10.1007/s00283-012-9307-y. [137]P. Cheeseman, B. Kanefsky, and W. M. Taylor, “Where the really hard problems are,” in Proceedings of the 12th international joint conference on Artificial intelligence - Volume 1, in IJCAI’91. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., Aug. 1991, pp. 331–337. [138]T. Takemoto, M. Hayashi, C. Yoshimura, and M. Yamaoka, “A 2x 30k-Spin Multi-Chip Scalable CMOS Annealing Processor Based on a Processing-in-Memory Approach for Solving Large-Scale Combinatorial Optimization Problems,” IEEE J. Solid-State Circuits, vol. 55, no. 1, pp. 145–156, Jan. 2020, doi: 10.1109/JSSC.2019.2949230. [139]T. Takemoto et al., “A 144Kb Annealing System Composed of 9×16Kb Annealing Processor Chips with Scalable Chip-to-Chip Connections for Large-Scale Combinatorial Optimization Problems,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC), Feb. 2021, pp. 64–66. doi: 10.1109/ISSCC42613.2021.9365748. [140]M. Jerry, A. Parihar, A. Raychowdhury, and S. Datta, “A random number generator based on insulator-to-metal electronic phase transitions,” in 2017 75th Annual Device Research Conference (DRC), Jun. 2017, pp. 1–2. doi: 10.1109/DRC.2017.7999423. [141]J. Shamsi, M. J. Avedillo, B. Linares-Barranco, and T. Serrano-Gotarredona, “Effect of Device Mismatches in Differential Oscillatory Neural Networks,” IEEE Trans. Circuits Syst. Regul. Pap., pp. 1–0, 2022, doi: 10.1109/TCSI.2022.3221540. [142]J. Núñez, M. J. Avedillo, and M. Jiménez, “Solving Combinatorial Optimization Problems with Coupled Phase Transition based Oscillators”. [143]M. K. Bashar, A. Mallick, and N. Shukla, “Experimental Investigation of the Dynamics of Coupled Oscillators as Ising Machines,” IEEE Access, vol. 9, pp. 148184–148190, 2021, doi: 10.1109/ACCESS.2021.3124808.