scieee AI-readable full text Open interactive document viewer

Integración de principios en una unidad de hipoacusia

Pérez Zaballos, M.T.

Abstract

Programa de doctorado: Investigación aplicada a las ciencias sanitarias.

Full text

! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! UNIVERSIDAD DE LAS PALMAS DE GRAN CANARIA ESCUELA DE DOCTORADO PROGRAMA DE DOCTORADO: INVESTIGACIÓN APLICADA A LAS CIENCIAS SANITARIAS INTEGRACIÓN DE PRINCIPIOS FÍSICOS EN UNA UNIDAD DE HIPOACUSIA INTEGRATION OF PHYSICAL PRINCIPLES IN A HEARING LOSS UNIT TESIS DOCTORAL PRESENTADA POR: Dª MARÍA TERESA PÉREZ ZABALLOS DIRIGIDA POR: PROF. DR. ÁNGEL M. RAMOS MACÍAS PROF. DR. SANTIAGO RODRÍGUEZ FEIJOÓ DIRECTOR CODIRECTOR Ángel M. Ramos Macías Santiago Rodríguez Feijoó María Teresa Pérez Zaballos DOCTORANDO ! ! iv a la tortilla. Gracias a ellos ha sido posible que todas mis empresas lleguen a buen puerto, incluida esta tesis. Yo no soy nada sin mi familia. Ellos son mi motor. Por eso, quiero darles las gracias a todos y quiero felicitarnos, porque somos la mejor familia del mundo, porque aguantamos hasta en los momentos mas duros. Clemente; Vicente, Martina, Vicen y Dani; Balbina, José, Jose, Consu y Juan; la Abuela y el Abuelo y Dora, gracias. En especial, estoy agradecida a mis padres y mis hermanos, porque son lo que más quiero en el mundo. Gracias por levantarme cuando caigo y por agarrarme cuando me tambaleo. Juntos podemos con todo. ¿Y que puedo decirle a Carlos, que ha sido el sufridor último de los estragos y el estrés de escribir esta tesis?. Es el hombro firme donde me apoyo y mi alegría de cada día. Él es la razón de que quiera ser mejor y que intente dar siempre lo mejor de mí. Además, por razones que desconozco, siente especial cariño por mis momentos EUREKA, y me anima siempre a lanzarme y probar mis ideas. Con alguien así, ¿quien no iba a ser feliz? Soy muy afortunada de tenerlo a mi lado. Gracias por quererme tal y como soy. Y para finalizar, quiero agradecer a varias personas que me han ayudado de diferentes maneras: a Bernardo y a Mabel , por organizar las sesiones con los controladores; a Esther y a Rita, por apoyarme y escucharme tanto este año; a Miguel, por estimularme intelectualmente desde que tenía 16 años (no tiene ni idea de lo mucho que ha influido en mi vida) y a Guacimara, por pasar de voluntaria a gran amiga y consejera. A todos, gracias. Índice ! ! ÍNDICE Índice ! ! Índice ! ! Í N D I C E I. INTRODUCTION 1.- Motivation and overview of this dissertation 2.- The physics of sound 2.1.- Anatomy of a sound wave 2.2.- Intensity of a sound wave 2.3.- Fourier Analysis of a sound wave 3.- The physics and anatomy of the ear 3.1.- The outer ear 3.2.- The middle ear 3.3.- The inner ear 4.- Psychophysics 4.1.- Pitch perception 4.2.- Loudness perception 5.- The cochlear implant. 5.1.- Signal processing 5.2.- Signal processing strategies 5.3.- Pitch perception with cochlear implants 5.4.- Loudness perception with cochlear implants 3 5 5 8 11 14 15 16 18 24 25 28 33 34 36 41 42 II. EXPERIMENTAL WORK 6.- Psychoacophysics: Speech perception in noise studies 6.1.- Effects of High-Frequency Suppression for Speech Recognition in Noise in Spanish Normal-Hearing Subjects 6.1.1Introduction 47 47 48 Índice ! ! 6.1.2Objectives and hypotheses 6.1.3Subjects and Methods 6.1.4Results 6.2.- Effects of speech-in-noise training and high frequency suppression in Air Traffic Controllers. A control group study 6.2.1Introduction 6.2.2Objectives and hypotheses 6.2.3Subjects and Methods 6.2.4Results 7.- Electrically evoked psychoacoustics in cochlear implant recipients 7.1.- PsychoLAB: A Psychophysics experimental setup to measure electrode discrimination in Nucleus cochlear implant recipients 7.1.1Introduction 7.1.2Platform design 7.2.- Effect of the electrode-inner wall distance on pitch perception in cochlear implant recipients using straight and perimodiolar arrays 7.2.1Introduction 7.2.2Objectives and hypotheses 7.2.3Subjects and Methods 7.2.4Results 8.- Electromagnetism 8.1.- The effect of reference electrode position on power consumption in Cochlear Implants 8.1.1Introduction 8.1.2Objectives and hypotheses 8.1.3Material and Methods 50 50 56 59 59 60 62 65 70 71 71 73 82 82 85 86 92 106 106 106 109 110 82 Índice ! ! 8.1.4Results 8.1.4Results 116 III. DISCUSSION 9.- Psychophysics 9.1.- How the healthy ear understands speech in noise and the effects of training on it 10.- Electrically evoked psychoacoustics in cochlear implant recipients 10.1.- How physical and psychophysical variables affect electrode discrimination in cochlear implant recipients 11.- Electromagnetism 11.1.-The role of the reference electrode in cochlear implants power consumption 125 125 131 131 137 137 0 0 0 IV. CONCLUSIONS 143 V. RESUMEN EN ESPAÑOL 147 VI. REFERENCES 155 VIII. TABLE INDEX 177 IX. FIGURE INDEX 179 X. APPENDIX 193 ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! ! Introduction ! ! ! 1! I. INTRODUCTION Introduction ! ! ! 2! Introduction ! ! ! 3! 1.- MOTIVATION AND OVERVIEW OF THIS DISSERTATION In today's world, where interdisciplinarity plays a key role in research, it is vital to change the traditional schemes that separated professionals according to their degrees. To get a true scientific revolution, different fields of knowledge need to be mixed and hybridized, so that old problems are looked at from new perspectives. This is how innovative solutions emerge. This dissertation attempts to illustrate this concept. The addition of a physicist to a Hearing Loss Unit has yielded different solutions to clinical problems. The presence of a physicist in a clinical session that listens to the real needs of the health personnel can dramatically speed up innovation regarding surgery, cochlear implant programming, rehabilitation and patient evaluation. In this thesis, work has been done trying to cover all of these aspects. The thread of the studies presented here is the improvement of sound perception of cochlear-implanted patients under different conditions. The work has been divided into the different physical fields that were applied to solve each initial question. Three sound conditions were evaluated. First, speech-in-noise perception mechanisms are studied. Then, how pitch discrimination is affected by electrode proximity was evaluated. Finally, the improvement of sound in patients with high impedances was looked at. The first part focuses on psychoacoustics, a hybrid field between physics and psychology. Two studies were conducted. The first is a study on the relevance of high frequencies for the perception of speech in noise in normal hearing individuals. This showed that high speech components, above 8 kHz, help to understand words subjected to ambient noise. The second study is a continuation of the first, and explores the effect of long-term speech in noise training in air traffic controllers. This job sector develops speech perception abilities during the course of their work. The special acoustic conditions to which they are subjected train their brains to extract the maximum amount of auditory cues to be able to establish radio communications. These have the Introduction ! ! ! 10! Another variable that can be used to describe sound intensity is the sound pressure level (SPL). Intensity and pressure can be related by the equation: !=!! !" [2.4] Where P is sound pressure and is measure in Pascals (Pa), ρ is the density of air (1.2 kg/m3 at 20oC) and v is the speed of sound (331 m/s in air). Thus, SPL can be described as: (5) (6) !"#(!")=20 log!" ! ! ! [2.5] While the intensity of a sound is a very objective quantity that can be measured with sensitive instrumentation, the loudness of a sound is a subjective response that will vary with a number of factors. The sound intensity must be factored by the ear's sensitivity to the particular frequencies contained in the sound and will vary slightly depending on the individual. But this will be discussed in more detail in further chapters. For now, it suffices to say that, since humans do not perceive low and high frequency sounds as well as they do perceive middle frequencies around 2 kHz, three weighing systems for SPL have been established: A, B and C. The graph in figure 2.4 describes the sensitivity of each weighing system as a function of frequency. Using the dB(A) measuring system, the sound level meter will be less sensitive to very high and very low frequencies. It approximates the ear when hearing a 1kHz sound at 40 dB. Measurements made with this scale are expressed as dB(A), which is very useful for eliminating inaudible low frequencies. The intermediate B-contour approximates the ear for medium loud sounds and it is rarely used. Finally, the dB (C) measuring system does not filter out the lows and highs to such an extent as the other contours. It approximates the ear at very high sound levels and has been used for traffic noise surveys in noisy areas. (6) Introduction ! ! ! 11! Figure 2.4. Relative response of each weighing SPL system as a function of frequency. It can be seen how dBA is most sensitive to frequencies up to 1000 Hz, while dBB and dBC tend to flatten out sooner. (Own production) 2.3.- Fourier Analysis of a sound wave All sounds can be decomposed into superpositions of sinusoids of different frequencies, amplitudes and phases or, seen from the opposite perspective, simple sinusoids can combine to form any kind of complex sound wave. Figure 2.5 illustrates this principle. The decomposition is known as Fourier analysis and gives back the composition of the sound wave f(t) in term of its pure tones or frequency composition and the intensity with which they appear. This translates mathematically into a sum of cosines and sines of different frequencies: !!=!!+!!sin(2!∙!(2!"#)) ! !!!+!!cos(2!∙!(2!"#)) ! !!! [2.6] Where the expressions in the sums are called the Fourier coefficients, A0 is a constant that describes the average value of the function (i.e. the point around which it oscillates up and down) and An and Bn are the amplitudes of each frequency or tone component. (7) (8) Introduction ! ! ! 12! Figure 2.5. When two single frequency waves are combined, the result is a waveform that is the sum of them. (Own production) The transformation of a complex wave pattern in time into its frequency component distribution is known as Fourier transformation and it constitutes a basic tool in the field of acoustics, because it allows the scientist to characterise sounds. A Fourier transform shows that any wave can be decomposed into a sum of sinusoids. It is not the aim of this section to describe the mathematical background behind a Fourier transform, which is complex and rarely used analytically. Most real life waveforms are so complex that Fourier transforms are carried out by computers. These normally use an algorithm called the Fast Fourier Transform or FFT. Given a waveform function in the time domain (like that in figure 2.5, blue line) an FFT would give its profile in the frequency domain (figure 2.5 red lines). (8) These illustrations show the essential nature of the FFT. For a sine wave with a single frequency, the FFT consists of a single peak. Combining two sound waves produces a complex pattern in the time domain, but the FFT clearly shows it as consisting almost entirely of two frequencies. Thus, Fourier Introduction ! ! ! 13! analysis can greatly simplify acoustic problems and reveal aspects that are not clearly visible in the time domain. (8) Introduction ! ! ! 14! 3.- THE PHYSICS AND ANATOMY OF THE EAR The ear is the most complex transducer ever created. It is able to convert energy from sound pressure into mechanical movement and this, in turn, into electrical nerve impulses that travel all the way to the brain. The ear’s ability to carry out this process allows us to perceive the details of a complex sound. The ear detects the frequencies that make them up; each one of their individual intensities and, through the integration of both aspects, it can recognise features such as who is talking and what a person is saying even when noise is present. The ear has three main parts: outer, middle and inner ear (figure 3.1). Each part has a specific function in the process of detection and interpretation of sound. The first serves to pick up sound and prepares it to transfer the energy to the middle ear in the most efficient way, by matching air and bone impedance. The second part is responsible for transforming the energy of the sound wave (pressure variations) into mechanical displacements of the ossicles, which creates a compressive wave within the inner ear. The inner ear is filled with fluid, which bends hair cells cilia as it moves and in turn transforms this mechanical displacement into electrical impulses that are identifiable by the brain. Introduction ! ! ! 15! Figure 3.1. Anatomy of the ear. The pinna and external auditory canal form the outer ear, which is separated from the middle ear by the tympanic membrane. The middle ear houses three ossicles, the malleus, incus and stapes. Together they form the sound conducting mechanism. The inner ear consists of the cochlea, which transduces vibration to a nervous impulse and the vestibular labyrinth that houses the organ of balance. (Own production using a free license image from BruceBlaus) 3.1.- The outer ear The outer ear consists of the ear and ear canal. Its overall function is to act as a pre-amplifier to improve the sensitivity of the system and its action mechanisms are a direct result of the physical characteristics of sound. (9) The sound energy spreads spherically from the source. For a point source of sound, as the radii of the spheres of sound propagation get further away, they also become larger. Given that the surface of a sphere is proportional to the square of its radius (!=4!!!), the energy in the sound wave arriving at a small area (e.g. an ear drum) declines with the square of the distance from the source (!∝1/!!)). This relationship between sound intensity or energy and distance from the source is known as the inverse square law. (9) Introduction ! ! ! 16! Therefore, for a given sound intensity, a larger pinna will have a broader area and thus will capture sound more efficiently. The ear canal, on the other hand, acts as a broadband resonator tube, closed at one end. It has an average diameter of 0.7 cm and is 2.7 cm long (L). By acting as a tube closed at one end as in figure 3.2, it gives rise to standing waves at a fundamental resonant frequency that is four times its length and at odd multiples of it, called harmonics. The closed end is restricted to be a wave node and the open end an antinode. As seen in figure 3.2, this condition yields a wavelength of the fundamental mode that is four times the length of the air column inside the canal (10.8 cm). The restriction of the closed end prevents it from producing even harmonics. Since (!=!/! and !=4!), its resonant frequencies are given by the equation: ! !=!" !!;!=1,3,5… [3.1] So the first resonant frequency will be around 3 kHz. (9) (10) (11) ! Figure 3.2. The ear canal acts as a tube open at one end, which forces the closed end to be a node and the open end and antinode. This causes the minimum resonant frequency (1st harmonic) to be four times the length of the ear canal as. The next harmonic that follows these boundary conditions is an odd multiple of this resonant frequency. (Own production) 3.2.- The middle ear The main parts of the middle ear are the tympanic membrane and the three ossicles (malleus, incus and stapes). Together they amplify the sound wave Introduction ! ! ! 17! and perform the important function of impedance matching. This improves the efficiency of the transfer of the mechanical vibrations of the tympanic membrane (~ low mechanical impedance) to the oval window (~ 20 times higher impedance). (11) The tympanic membrane is a circular membrane of about 8-9 mm in diameter, 0.1 mm in thickness, with an area of 65 to 80 mm2 and a weight of 14mg. It has a conical shape and is attached to the inner part of the stapes. It establishes the boundary between the outer and middle ear, acting just like the membrane of a microphone. It converts the air pressure waves into mechanical vibrations that are transmitted to the ossicles. The tympanic membrane has the peculiar property of being a damped resonator (the system goes back very fast to its resting position) due to the high tension force to which it is subjected. This broadens the number of frequencies that can cause the system to resonate and it therefore has a large range that goes from 20 to 20,000 Hz. Because it is 20 times the size of the oval window (the membrane that connects the middle ear with the inner ear fluid cavity), it amplifies the incoming sound by a factor of 20, compared with the direct reception of the sound by the oval window alone. (12) The auditory receptors of the inner ear operate in a fluid environment, and the inner ear is really an "underwater" sound receiver. When sound in air strikes a fluid boundary (a boundary between media with different acoustic impedances) there is a theoretical loss of 99.9% of the energy in a sound wave in air due to reflection. This 99.9% loss is equivalent to 30 dB. In order to overcome this mismatch in the impedance between air and fluid, the middle ear is interposed between the tympanic membrane and the oval window. (13) Ossicles work like a composite lever that achieves a multiplication of force between the tympanic membrane and the oval window. Although they are among the smallest bones in the body they are capable of spreading force with great pressure because of their physical structure and layout within the inner ear. When these three bones work together in combination with two Introduction ! ! ! 18! small muscles, they form the ossicle chain. This chain transfers the force coming from the tympanic membrane (effective area of of 65 mm2) to the oval window (effective area of 3.2 mm2), thus increasing the pressure with which this force is applied. Since the pressure is the force per unit area (!=!/!), the relationship between the pressure in the tympanic membrane and the oval window will be: !=! !∙!! (Tympanic pressure) ! !=!!∙!! !! [3.2] !=! !∙!! (Oval window pressure) The factor At/Ao is approximately 20. It also happens that the lever arm formed by the malleus in rotating about its pivot is somewhat longer than that of the incus, giving another factor of about 1.3 in pressure increase. These two factors multiplied, yield about a 26x increase in pressure, which is about 29 dB, thus just about overcoming the theoretical 30 dB loss due to the air/liquid interface. Therefore, the middle ear matches the acoustic impedance between air and fluid, maximizing the flow of energy from the air to the fluid of the inner ear. (9) 3.3.- The inner ear The inner ear consists of a bonny labyrinth, divided into the vestibular system and the cochlea. It is a sealed cavity filled with a fluid called lymph. The cochlea is the organ in charge of sound transduction into electrical signals. The cochlea is spirally wounded up, forming two and a half turns, with a length of about 30 mm and a variable sectional area that ranges from 4 mm2 at its basal end, to 1 mm2 at its apical end. The cochlear cavity is divided by two membranes: the Basilar and Reissner membranes. Both divide it longitudinally into three channels. Figure 3.3 shows its main features. They connect at the apex of the cochlea in a region known as Helicotrema. The Reissner membrane separates the Scala Vestibuli from a smaller cavity Introduction ! ! ! 19! known as the cochlear duct. This, in turn, is separated by the basilar membrane (BM) from a third cavity called Scala Timpani. The BM is composed of a great number of taut, radially parallel fibres sealed between a gelatinous material of very weak shear strength. These fibres are resonant at progressively lower frequencies as one progresses from the basal to the apical ends of the cochlea. Four rows of hair cells lie on top of the BM, together with supporting cells. A single inner row is medial, closest to the central core of the cochlea. It has an abundant nerve supply carrying messages to the brain. The three outer rows receive mainly an afferent nerve supply. The assembly formed by the BM, the tectorial membrane, the hair cells and the nerve terminals is called the organ of Corti. Any natural displacement of the cochlear partition results in a rocking motion of these structures and consequently a lateral displacement of the inner hair cells. This movement causes depolarization and hiperpolarization of the hair cells and transmits the nerve impulse to the neural ends that go along the centre of the cochlea or modiolus forming a structure known as the spiral ganglion. This complex is the heart of the ear, as it converts mechanical vibrations into electrical currents than can be understood by the brain. (10) Introduction ! ! ! 26! mechanics of nature expands, so does the presence of physics in every field of science. With the exception of pure tones, pitch is not a simple function of the spectral content of a sound. Rather, pitch is related more closely to the repetition rate, or envelop repetition rate, of the sound with a range from 30 Hz to 5 kHz. Sounds with the same repetition rate and very different spectra often have the same pitch (e.g. a pure tone with a frequency of 500 Hz and a complex tone with high harmonics with a fundamental or first resonant frequency, F0 of 500 Hz). On the contrary, sounds with similar spectra can have very different pitches (e.g white noise with amplitude modulated at 100 or 200 Hz). Thus, the auditory system combines information across the cochlear location (frequencies and intensity), to derive the pitch of the sound. Moreover, pitch is also represented in terms of the precise timing of neural impulses in the auditory nerve and at higher centres in the auditory systems. Many sounds have acoustic waveforms that repeat themselves over time. These sounds are often perceived as having a pitch that corresponds to the repetition rate of the sound. This attribute of sound is one of the most relevant perceptual dimensions for speech communication, carrying prosodic information in western languages, but also semantic information in tonal languages such as Mandarin. Moreover, it enhances our ability to perceptually segregate sound sources, based on differences in fundamental frequency (F0). Conversely, it can be used to group together the individual sound components (harmonics) that arise from the same vibrating source. Pitch, therefore, defines and differentiates our acoustic environment. Theories of pitch perception There are currently two theories of pitch perception (see figure 4.1). The spatial theory is based on two main tenets. The first is that the stimulus is subjected to spectral analysis in the cochlea; so each frequency excites different locations along the basilar membrane and, consequently, neurons Introduction ! ! ! 27! with different central frequencies. This process is called tonotopic organization. The second tenet states that the perception of a stimulus is related to the pattern of excitation that it produces. For a pure tone, it is assumed that this corresponds to the place of maximum excitation. The first of these principles has been confirmed in a series of independent experiments, including direct observation of the movement of the MB. The second is still in dispute. (19) (15) Figure 4.1. Temporal and place coding of sound. In temporal coding (left), neurons fire action potentials in phase with the sound waves. Place coding (right) refers to the perception of pitch depending on the site of stimulation. This is possible because the auditory cortex that responds to sound is arranged tonotopically. (Own production) An alternative theory is the temporal theory, based on the assumption that the tone of a stimulus is related to its evoked neuronal firing time pattern. The peaks of neural firing tend to occur at a particular phase of the waveform in the MB (the cells phase-lock to stimuli). This causes the time intervals between successive peaks to approach integral multiples of the waveform period. Phase-locking becomes very weak for sinusoids with frequencies above approximately 5 kHz, as stated in the previous chapter, although the Introduction ! ! ! 28! precise upper limit in humans is not known. However, the tones produced by musical instruments, the human voice, and the most common sources of sound have fundamental frequencies below 5 kHz. (Moore 2013) Physiological measurements and simulations using computational models have demonstrated that the repetition rates of stimuli are very well represented by the pattern of phase locking in the auditory nerve. For many researchers, the time versus place argument has been replaced by the question of how and where in the auditory pathway the phase-locked activity is analysed. The maximum frequency to which a fibre will phase lock declines from the auditory nerve to the auditory cortex and it is thought that somewhere in the brainstem, synchrony representation is converted into a rate-place representation in which different neurons code for different pitches in terms of overall firing rate. (26) (27) However, there remain some fundamental questions to be answered conclusively, such as how phase-locked neural activity is transformed into a rate-place representation of pitch; where this transformation takes place and how this information is used in object and pattern identification. 4.2.- Loudness perception Loudness is the subjective intensity of a sound. It is subjective because it depends entirely on the individual. The term intensity is used here because the response indicates how strong the sound seems to the listener. This definition is vague, but in experiments on loudness it usually suffices to elicit consistent responses. These responses can then be studied to establish a relation between loudness and the experimentally manipulated physical variables. Sound intensity is the most important of them in determining loudness, but spectral variables such as signal frequency and bandwidth also have an effect on it. The duration and intermittency of a sound are among the significant temporal variables. Moreover, background sounds can also affect the perception of loudness. Therefore, sound intensity must be factored by the Introduction ! ! ! 29! ear's sensitivity to the particular frequencies, and their durations, contained in the sound. As a complete review of all the physical parameters that affect loudness perception would require a level of detail that is beyond the scope of this dissertation, just the main variables will be described in detail here. For a more detailed description, the author refers the reader to. (19) (28) Loudness level Loudness level indicates how loud a 1 kHz tone must be in order to sound as loud as a test tone. To determine the level of a given sound, the subject is asked to adjust the level of a 1 kHz tone until it has the same loudness as the test sound. They are presented alternatively rather than simultaneously. The resultant level of the 1 kHz tone is the loudness level of that tone and is given in phons. If instead the 1 kHz tone is fixed in level, and test sounds at different frequencies are adjusted to give a loudness match, an equalloudness contour is generated (figure 4.2). These standard equal loudness contours (ISO 226, 2003), are based on extensive measurements of several laboratories, as described by Suzuki and Takeshima (29) . This curve shows how strong each frequency sound must sound to equal the corresponding phon line. For example, if the line is at 30 dB SPL, a 125 Hz sound must have an SPL of 50 dB to sound as loud as a 1 kHz sound at 30 dB SPL. These contours tend to flatten at high loudness levels. This means that the rate of growth differs for tones of different frequency. The rate of growth is greater for low and very high frequencies than for middle ones. (30) (29) Introduction ! ! ! 30! Figure 4.2. Equal-loudness contours for loudness levels from 10 to 120 phones for sounds presented binaurally from the frontal direction. The absolute threshold curve (MAF) is also shown. (19) This is the basis behind the weighting systems seen in previous chapters. The A weighting is based roughly on the 30-phon equal-loudness contour. At high levels, the weighting becomes more linear, so the C weighting is used. The B weighting is used for intermediate levels, and it is based on the 70-phon equal-loudness contour. (19) Loudness models S.S Stevens was the pioneer in the development of scales of loudness and other sensory dimensions. He suggested that perceived loudness, L, was a power function of physical intensity I: !=!!!.![4.1] Where k is a constant that depends on the subject and the units used. According to this equation, given two sounds where one is 10 times more Introduction ! ! ! 31! intense than the other in the physical domain, the perceived loudness is only two times higher (10!.!≈2) which implies the existence of compression. (31) The power law relationship between intensity and loudness has been confirmed in a large number of experiments using a variety of techniques. However, there have also been criticisms. The techniques are very susceptible to bias effects like: 1) the range of stimuli presented; 2) the first stimulus presented; 3) the instructions to the subject; 4) the range of permissible responses; 5) the symmetry of the response range and 6) other factors related to experience, motivation, training and attention. (32) Unfortunately, the underlying mechanisms of loudness perception are not yet fully understood. A common assumption is that the intensity is in some way related to the overall neuronal activity evoked by a sound. In this scenario, the intensity of a sinusoidal tone would be determined not only by the activity of neurons with characteristic frequencies (CF) near the tone frequency, but also by the degree of activity of adjacent CF neurons. In other words, the volume may depend on a sum of neuronal activity through different frequency channels (known as critical bands). Models incorporating this basic concept have been proposed by Fletcher and Munson, Zwicker and Scharf and, more recently, by Moore et al. The models proposed by Moore are illustrated in figure 4.3. The first stage is a fixed filter to simulate the transmission of sound through the outer and middle ear. The next step is to calculate a pattern of excitation for sound studio. Then, there is a transformation from the excitation level (in dB) to the specific loudness, which is a kind of "sound density" that represents the volume per critical band. This transformation involves a non-linear compression; for example, a 10-fold increase in the sound intensity, which corresponds to an increase of 10 dB, causes less than a 10-fold change in specific loudness. Although the models are based on psychoacoustic data, this transformation can be understood as representing the way the physical excitation is transformed into neural activity. The specific loudness is probably related to the amount of neural activity in Introduction ! ! ! 32! the corresponding CF. The compression in the transformation partly reflects what occurs in the basilar membrane. The overall volume of a given sound, in sones, is assumed to be proportional to the total area under the specific pattern of loudness. One could think of this area as approximately proportional to the total neuronal activity evoked by a sound. Sound models of this type have been quite successful at explaining the experimental data on the sonority of simple and complex sounds. (30) (33) (34) Figure 4.3. Basic structure of Moore et al. loudness model. (Own production) Stimulus! Fixed!=ilter! for!transfer! of!outer/ middle!ear! Transformati on!of! spectrum!to! excitation! pattern! Transformati on!of! excitation!to! speci=ic! loudness! Calculation!of! Area!under! speci=ic! loudness! pattern! Introduction ! ! ! 33! 5.- COCHLEAR IMPLANTS Cochlear implants (CI) are devices designed to restore hearing to profoundly deaf people. Hearing aids are a common solution for people with moderate hearing thresholds; however, for higher thresholds (~ 90 dB), these do not work as successfully. In these situations, a cochlear implant is used to bypass the middle and inner ear and directly stimulate the auditory nerve from within the cochlea. Figure 5.1. Main components of a cochlear implant: a) speech processor; b) transmitter; c) receiver/stimulator; d) electrode array. (Own production using free image from National Institute on Deafness and Other Communication Disorders at the National Institutes of Health) Figure 5.1 shows the main parts of a cochlear implant. Most modern systems include four components: a portable external sound processor, a transmitter, an implanted receiver/stimulator and a set of intracochlear electrodes. The sound processor consists of a microphone that collects the sound signal; an a! b! c! d! Introduction ! ! ! 34! electronic processor, generally in the form of a digital signal processor (DSP), that encodes it into electrical pulses. The signal then goes to the radio frequency transmitter, which sends it to the transcutaneous receiver / stimulator through a transmitter coil. This transmission is done through inductance (magnetic field variations), which induces alternating current pulses in the receiver coil. The receiver / stimulator is a surgically implanted electronic device. It receives and decodes signals from the sound processor and generates electrical signals to selectively activate the intracochlear electrodes. Small electric current pulses are sent to these electrodes to stimulate nerve cells in the peripheral auditory system, thereby causing auditory sensations. The electrodes are placed in the form of an array, which is made up of 16 to 24 contacts (depending on the device). They are 12.1 to 27 mm away from each other depending on the manufacturer. Each electrode stimulates a different site in the cochlea, which corresponds to a frequency band matching the cochlear tonotopy. Thus, mainly a limited population of nerve fibres perceives the stimulus. This allows different sections of the cochlea to be independently stimulated, mimicking and bypassing the basilar membrane frequency filtering function. (35) (36) Sound is filtered in the speech processor by a bank of bandpass filters that do not overlap, creating a set of channels that contain information from different regions of the sound frequency spectrum. Each of these channels is able to send information to one electrode of the array. The resolution of a healthy human auditory system is 1 Hz in the 100 Hz region. This would be the equivalent of having 100 channels from 100 to 200 Hz. Instead, cochlear implants have 16 to 24 channels for the whole audible frequency range. This number depends on the processing algorithm of each specific cochlear implant manufacturer. (37) 5.1Signal processing The most important CI unit is the signal processing strategy used for the transformation of the electrical stimulation signal. The different signal Introduction ! ! ! 35! processing techniques that have been developed can be classified according to the sound stimuli information that they use: 1) the waveform, 2) the envelope, and 3) the spectral characteristics (fine structure) of the signal. The latest speech processors are based on the channel vocoder principle, which is used in telephone communication with much less bandwidth than that required to transmit the raw speech signal. A channel vocoder consists of an analyser and a voice synthesizer. CI speech processors use the signal analysing block of the vocoder. As shown in figure 5.2, a vocoder analyser filters the incoming speech signal into a number of contiguous frequency channels using a bank of bandpass filters. The output of each filter is then passed through an envelope detector, which consists of a full-wave rectifier and low-pass filter. This way, an estimation of the energy for each band is obtained as well as its evolution in time. In addition, the analyser makes the decision to allocate sound to that channel or not, and it also estimates its tone (F0). The dynamic range (DR) adaptation block transforms the acoustic DR for each channel into the electrical DR (defined as the span of intensities going from threshold to maximum comfortable level) necessary for each electrode. This transformation is specific for each patient and is different for each electrode. Finally, according to the stimulation rate (i.e. the number of pulses delivered to electrodes per second), the processor generates stimulation pulses representing the current level to be sent to each electrode at each time instant. In the majority of strategies, the stimulation pulses are generated in a way such that, at each moment, there is only one channel active, to avoid overlapping sensations caused by adjacent electrodes working at the same time. The stimulation pattern computed by the processor is transmitted to the CI receiver and the current pulses are then sent to the electrodes. The performance of cochlear implants depends greatly on several factors such as the patient's age, duration and causes of deafness, the number of surviving spiral ganglion cells, the number of functional electrodes, the speech processing strategy and so on. Introduction ! ! ! 42! 5.4Loudness perception with Cochlear Implants In designing a speech processor for an implant listener, one of the most important factors lies in the proper transformation of acoustic amplitude into electric amplitude. Normal acoustic hearing can process sounds over a range of 120 dB. Speech sounds in a normal conversation can range from 40 to 60 dB. However, implant listeners typically have dynamic ranges of only 6 to 30 dB. Thus, the normal acoustic hearing range must be considerably compressed for electric stimulation, in order to preserve the acoustic amplitude variations that transmit important speech information. (53) Loudness in electrically stimulated human listeners has been measured quantitatively using rating methods and magnitude estimation techniques. However, the existing data are not consistent across studies. Pfingst’s group showed that loudness function depends on stimulus frequency for simple sinusoidal and pulsatile stimuli in cochlear implants. They found that loudness was an exponential function of stimulus amplitude for high frequencies (> 300 Hz) and a power function for low frequencies (< 300 Hz). Some investigators have found that log loudness grows linearly as a function of log current, although some data show two stages, with loudness that grows gradually at low stimulation levels and then grows steeply at higher stimulation levels (figure 5.5). Other groups found that loudness grows linearly as a function of charge. There has not yet been a widely accepted quantitative description in implant research of the relation between loudness and electric stimulus level. It is not clear whether the discrepancies among the above-mentioned loudness functions were due to procedural differences or due to individual implant listener differences. (54) (55) (56) (57) (58) Introduction ! ! ! 43! Figure 5.5. Growth function of discharge rate versus electrical pulse intensity. [19] !!! 44! !!! 45! II. EXPERIMENTAL WORK !!! 46! Experimental Work ! !!! 47! 6.- PSYCHOACOPHYSICS: SPEECH PERCEPTION IN NOISE STUDIES In everyday listening situations, accurate speech perception relies on the capacity of the auditory system to process complex sounds in the presence of background noise. There is ample evidence of auditory training resulting in perceptual enhancements. However, there have been surprisingly few investigations of how training impacts speech-in-noise perception. Even less has been done in the field of how high frequencies affect this ability. In this chapter, two studies are described. The first focuses on the physical composition of sound, as a way to enhance speech-recognition in noise. More specifically, it investigates how speech’s high frequency content affects its understanding in noisy environments. The second study deals with how training can enhance this ability even when sound has been band-frequency limited. 6.1.- Effects of High-Frequency Suppression for Speech Recognition in Noise in Spanish Normal-Hearing Subjects (Published as: de Miguel, ÁR, Zaballos, MTP, Macías, ÁR, Barreiro, SAB, González, JCF, Plasencia, DP (2015). Effects of High-Frequency Suppression for Speech Recognition in Noise in Spanish Normal-Hearing Subjects. Otol Neurotol, 36(4), 720-726.) It is common for patients to ask why they hear but fail to understand or why they have trouble understanding voices. Part of the answer has to do with one’s hearing and part with one’s cognitive processing of the word. In this study, attention is focused on the first one. Hearing loss can affect the ability to understand the voices, especially when competing noise is present. With hearing loss, it isn’t about not hearing sounds at all. It is about hearing just parts of the sounds, maybe even just the vowels, which have lower Experimental Work ! !!! 48! frequencies. To illustrate this, let’s take the sentence: high frequencies are important to understand speech in noise. For someone with a moderate hearing loss in the higher frequencies, this might sound approximately like this: _i _ __e_ue__ie_ a_e i_o_a__ _o u__e__a__ __ee_ i_ oi_e. One would probably hear the words but not understand the meaning. Patients can sometimes complain about these problems but have an apparently normal audiometry. Thus, the aim of this study is to investigate whether frequencies above 8 kHz,-which are beyond the scope of a clinical audiometry and outside the bandwidth of operation of hearing prosthesis-, can have a significant effect on speech perception in noise. 6.1.1Introduction A healthy human ear can perceive frequencies from 20 to 20000 Hz. This range varies between individuals and degrades gradually with age, especially for high frequencies. Since the peak energy of fricatives spoken by female and child talkers tends to occur between 6.3-9 kHz, subjects with hearing loss (HL) in this range will have a limited audibility of such phonemes. Moreover, if there is important information contained in the high frequency end of the human hearing spectrum, then they will suffer from an important loss of information. In mild, moderate and high frequency hearing losses the only symptom may be subtle difficulty with word understanding, especially in situations where there is competing noise. (59) (60) (61) (62) (63) A series of studies by Stelmachowicz et al. reported that when the high frequency components of a sound stimulus are removed, the perception of the phonemes /s/ and /z/ are seriously compromised in children with Normal Hearing (NH) and with hearing loss (HL). They investigated the effects of lowpass filtering on the perception of /s/ in four groups of listeners: children and adults with NH and children and adults with HL. Test stimuli were produced by an adult female, amale, and a child and were low-pass filtered at various frequencies from 2 to 9 kHz. In general, children with NH performed more poorly than the adults, andsubjects with HL performed more poorly than Experimental Work ! !!! 49! their counterparts with NH. Importantly, in both cases mean performance for the female and child talkers improved up to a bandwidth of 9 kHz. These results suggest that the perception of fricatives may be difficult for listeners with HL, due to the limited bandwidth of current hearing prostheses. Nonetheless, the authors suggest that the effects of bandwidth may be task dependent. Since adults with post-lingual hearing loss are generally able to use semantic and syntactic cues in the perception of speech, the acoustic effects of removing high frequencies in speech may be masked in this age group. However, young children who are still learning speech and language may be more affected. Thus, it is important to establish how complex auditory conditions such as background noise can affect this ability of adults to compensate for a shorter audible frequency bandwidth. (63) (64) (65) A number of authors have also studied the influences of high frequency on sound quality perception in various auditory situations. Gabrielsson et al. showed that the perceived sound quality of a female voice depended on the audible frequency bandwidth and on the smoothness in frequency response of the reproduction or amplification system. Moore and Tan also reported significant degradation in speech sound quality when the cut-off frequency decreased below 10.8 kHz. (66) (67) These findings are also of vital importance in the area of Cochlear Implant (CI) development. Because such devices have a limited bandwidth of around 8kHz, their users cannot perceive auditory information from frequencies above this value. In spite of this, the majority of CI users have very good results in speech recognition tests in quiet. However, the same situation does not apply in more complex auditory conditions, such as in the presence of background noise. To get around this issue, CIs use a variety of strategies and noise filters, but even though some benefits have been reported, their performance is still not satisfactory. (68) (69) (70) (71) (72) (73) (74) Experimental Work ! !!! 50! 6.1.2Objectives, hypothesis and rationale for the study In order to understand why subjects with HL have difficulty in understanding speech in noise, it is first necessary to understand the processes allowing subjects with NH to direct attention to a desired source and comprehend it. However, the neural mechanisms that support speech discrimination in noise are not well understood. Therefore, this study aims to evaluate how high frequency perception contributes to speech perception in noise. To this end, an experiment was designed to study speech understanding degradation as a function of maximum stimulus frequency in subjects with NH. The hypothesis of this study is that high frequency filtering of the sound environment (words and noise) above 8 kHz causes a significant decrease in word recognition in noise. The motivation for this study has come mainly from the realization that hearing prostheses only reach 8 kHz. Which means that any sound above it is not amplified (in the case of hearing aids) or heard at all (in the case of CI). Despite the fine spectral and even temporal resolution obtained for low frequencies in the field of CI, which correspond to the fundamental components of speech, patients still struggle with understanding in noise situations. Therefore, the cause might not be restricted to the low frequency, but also to the high frequency range. If NH subjects do make use of this information, then it would be interesting to alter CI design to try and find a way to broaden their frequency range. 6.1.3Subjects and Methods This observational cross-sectional study was undertaken at the Hearing Loss Unit in the Otolaryngology Head Neck Department, University Hospital of Experimental Work ! !!! 51! Gran Canaria (Insular-Materno Infantil) (Universidad de Las Palmas de Gran Canaria). The corresponding Ethical Committee approved this study. Subjects A total of 29 subjects were selected, aged 19-55, with an average age of 30. All were native Spanish speakers. The subjects had normal hearing with pure tone thresholds better than 20 dB HL in all frequencies studied and with no other known pathology. None of them had prior experience with acoustic simulations or had participated in any hearing research experiment. All subjects came voluntarily and gave their informed consent. Speech material and background noise To study the effect of high frequency removal in speech recognition in noise, different amendments to the Spanish disyllabic speech lists were made. Frequency modifications and noise files were generated using MATLAB and Statistics Toolbox Release 2012b (The MathWorks, Inc., Natick, Massachusetts, United States). Six validated lists of 25 words each and spoken by a female speaker with a Spanish accent were used (Appendix II). They belong to the standard Spanish Disyllabic test, commonly used in hearing-loss assessments. Two groups were made of three lists each. In the first group, frequency components were unaltered and in the second group the words were band-pass filtered. For future reference, we will refer to them as groups A and B respectively. The audio files were saved in Waveform Audio File Format (.Wav) format with a sampling frequency (Fs) of 44 kHz. Fs determines how many times per second the analogue signal is digitally sampled by the recording system. According to the Sampling Theorem, for a band-limited signal to be reconstructed fully, this must be done at twice the rate of the highest frequency that it is of interest to record. Therefore, the maximum sound frequency that could be reproduced with this sampling rate was 22 kHz. (75) Experimental Work ! !!! 58! individual words. To examine these issues, correlations between the amount of alteration in a word and its mean score were analysed for group B lists (the filtered ones). The following hypothesis was evaluated: words whose high frequency spectrum had suffered the highest changes gave lower scores. Table 6.1 summarizes the results of the Pearson correlation for each list. The only statistically significant correlation was found for SNR = 5 ( p = 0.037). In the remaining cases, no correlation was found between these two variables (p = 0.487 for 0 dB SNR and p = 0.171 for -5 dB SNR). PEARSON CORRELATIONS 5 SNR Pearson Coefficient Sig 0.363 0.037 0 SNR Pearson Coefficient Sig -0.007 0.487 -5 SNR Pearson Coefficient Sig -0.198 0.171 Table 6.1. Pearson correlation results for the comparison of autocorrelation of words pre and post filtering and their total score. The only significant correlation was found for the case 5 dB SNR. (Own production) Experimental Work ! !!! 59! 6.2.- Effects of speech-in-noise training and high frequency suppression in Air Traffic Controllers. A control group study (Published as: Zaballos, MTP, de Miguel, ÁR, Plasencia, DP, González, MLZ, Macías, ÁR (2015). Effects of long-term speech-in-noise training in air traffic controllers and high frequency suppression. A control group study. J Int Adv Otol, 11(3), 212-7.) Learning electrically stimulated speech patterns is a new and difficult experience for many CI users. It does not suffice to have been given the tools to hear. These patients require extensive rehabilitation and auditory training to be able to make use of the information provided by the device, identify differences in sound parameters and to make sense of them. Specifically, speech-in-noise-understanding still presents a great challenge to most CI users. Given that CI and hearing aids alone do not restore these abilities, efforts are currently focused on the role of the brain in such situations and whether training can be an efficient therapy. In this study, air traffic controllers have been studied because they seem to have developed a greater ability to understand speech in noise than the average individual. This ability is a result of their continuous and long-term auditory training in noise under restricted frequency bands. This is an encouraging discovery for the field of cognitive sciences and auditory rehabilitation. Moreover, the effect of high frequency suppression for speech-in-noise understanding observed in the previous study has been further assessed for this job sector to evaluate possible learning effects. 6.2.1Introduction The ability to understand speech in noise requires both sensory and cognitive skills (78). The sensory part consists of the auditory system locking on to the target speech signal while excluding ambient noise and competing voices. This is achieved by organising auditory inputs into different groups by identifying shared characteristics, such as location and acoustical similarity. The identification of these groups is given by the relative stability of voice Experimental Work ! !!! 60! pitch and helps in grouping it separately from other voices. In addition, other signal-based cues like harmonics, location and timing, aid in group formation of speech. On the cognitive side, attention and working memory skills are key to a good speech-in-noise (SIN) ability. (79) (80) (81) There is evidence that such skills can be improved during an adult’s life, causing perceptual enhancements, plasticity in single neurons and in neural populations (82) (83). Studies of auditory perceptual learning reveal long-term neural changes in the adult auditory cortex (AC) of both animals and humans after intensive auditory training (84) (85). Current models propose that perceptual learning in adults depends strongly on top-down influences such as attention, reward, and task relevance. This demonstrates that the adult AC is a dynamic and adaptive processing centre. (86) (87) (88) (89) (90) (91) (92) (93) (94) (95) Natural auditory training occurs in job sectors with frequent daily radio communications. This is the case of Air Traffic Controllers (ATC). People that work in this field are constantly exposed to SIN, which is extremely hard to understand to an outsider. During their working life, ATC learn to identify, extract and comprehend conversations embedded in white noise. These conversations are normally restricted to the aeronautics field, which makes them an ideal population to study the effects and transferability of long-term auditory training. Unfortunately, the majority of studies on how training affects SIN perception used small stimulus sets and revealed that, while SIN perception can improve in artificial listening conditions, the benefit can be extended to untrained material is not clear. (96) (97) (98) (99) 6.2.2Objectives, hypotheses and rationale for the study As opposed to previous work where the patients are trained in situ and then tested, this study looks at how someone's job based listening experience can Experimental Work ! !!! 61! lead to enhanced performance on a laboratory-based test. Moreover, the role of high frequencies has also been evaluated because radio communications are limited, just like hearing prosthesis, to 8 kHz. Thus, this study has two main objectives. The first is to evaluate whether Air traffic Controllers enhanced ability to understand speech in noise could be tested in a speech-in-noise laboratory test. To this end, a control group of normal hearing listeners and a target group of ATC were assessed. The second is to compare their speech in noise performance in high frequency (>8kHz) filtered and non-filtered conditions to that of normal hearing individuals from the previous study (control group). The hypotheses are: 1) that ATCs understand speech in noise significantly better than an ordinary normal hearing individual; 2) the differences between filtered and non-filtered word recognition are lower than those for the control group. The first hypothesis is proposed on the basis that these individuals have undergone speech-in-noise training during adulthood. The second is constructed on the fact that their training material involved high frequency filtered radio communications (>8 kHz). The reason behind this study is that our aging society is bound to suffer speech-in-noise-understanding impairments. As individuals grow older, this ability degrades even before auditory losses become clinically relevant. Therefore, it is capital to come up with strategies to try and slow this process down as much as possible. Given the recent evidence that learning during adulthood is far greater than previously thought, the demonstration that auditory training under degraded sound quality can lead to better speech in noise understanding can have important implications in the way we treat hearing losses today. An effective rehabilitation program can improve people’s quality of life by easing their integration back into the sound complexity of society. Experimental Work ! !!! 62! 6.2.3Subjects and Methods This study was done at the Air Base of Gando Airport of Gran Canaria and the Psychoacoustics Laboratory of the Hearing Loss Unit, Otolaryngology Head and Neck Department, Complejo Hospitalario Universitario Insular Materno Infantil de Gran Canaria (CHUIMI). The Ethical Committee approved this study. The methods of this study are analogous to those for the previous study. So only the adds-on will be described. Subjects A total of 29 normal-hearing subjects (aged 19-55) from the previous study were used as control group and 48 ATC (aged 34-56), as the target group. All subjects were native Spanish speakers. The subjects had NH with pure tone thresholds better than 20 dB HL as measured by a validated clinical audiometry (from 0.5 to 8 kHz in steps of 0.5 kHz) and with no other known pathology. None of them had previously taken part in any auditory experiments before and they all signed an informed consent before commencement of the study. Age histograms for the ATC and control groups were illustrated in figures 6.6 and 6.7, and years working in the sector for ATC in figure 6.8. Experimental Work ! !!! 63! Figure 6.6. ATC age histogram. Figure 6.7. Control group age histogram. Experimental Work ! !!! 64! Figure 6.8. Years working as ATC histogram. Speech material and background noise The speech material was the same as for the previous experiment. Procedures The experiment procedure is analogous to that of the previous chapter. The data was recorded in an .xls file, Microsoft Excel 2011 (Microsoft. Microsoft Excel (Redmond, Washington, United States) and IBM SPSS Statistics for Macintosh, V. 21.0. (IBM Corporation, New York, United States) was used for the statistical treatment. The statistical analysis starts with a normality check for each specific condition using the Kolmogorov-Smirnov test for one sample. According to the result, a non parametric or a two tailed tstudent test was performed. Finally, correlations between age, years of Experimental Work ! !!! 65! experience in the Aeronautic field and test scores were analysed using a Pearson correlation test. The level of significance chosen was 5%. 6.2.4Results First, overall performance difference between the two groups was analysed. Figure 6.9 shows how the ATC group outperforms the control group in both conditions, -filtered and non-filtered-, and for all SNR conditions. Moreover, the difference between them becomes larger as noise increases. Normality check gave Gaussian distributions. The evaluation of significance returned pvalues < 0.01 for all cases, thus revealing significant mean differences between the two groups. For SNR 5 dB, this difference was 7.27% for list A and 13.25% for list B. For SNR 0 dB, these values rise to 16.95% and 11.45%. The most interesting difference is observed at SNR -5 dB. In this case, ATC outperform the control group by 17.03% and 22.52%. Figure 6.9. Success of ATC and control groups as a function of the test. The ATC group outperforms the control group in all SNR conditions and both in non-filtered (A=22 kHz) and filtered (B=8 kHz) conditions. Experimental Work ! !!! 66! Figures 6.10 and 6.11 illustrate the relation between years working as ATC and test performance and between age and test performance. The correlation study between age, years working in this profession and overall test results for each condition is shown in Table 6.2. No significant correlation was found between working experience and overall test results. A detailed analysis of each case revealed significant (p<0.05) inverse correlations between 8 kHz 5 SNR test results and both variables. 22 kHz 5 SNR also gave a significant, inverse link between age and test results (Table 6.3). However, these results do not have a physical justification. Figure 6.10. Success rate as a function of years working as ATC. Experimental Work ! !!! 67! Figure 6.11. Success rate as a function of age for the ATC group. Age Years working as an ATC Test results Age Pearson Correlation Sigma (bilateral) N - 0.759 0.000 288 -0.068 0.249 288 Years working as an ATC Pearson Correlation Sigma (bilateral) N - 0.002 0.973 288 Test results Pearson Correlation Sigma (bilateral) N - Table 6.2. Pearson correlation of age, years working as ATC and overall test results. No correlation was found between tests results and either age or experience. Obviously, in the vast majority of cases, the older the ATC, the longer they have worked, so there is a correlation between these two variables. Experimental Work ! !!! 74! Figure 7.1. Case of Use Diagram. The patient interacts with the researcher by providing answers for each of the test steps: T and C level calculation (MCL and THL, respectively), loudness balance and electrode discrimination test. Communication between the user interface in Visual Studio and the NIC Python library Because the NIC library was designed for Python, but the user interface was built in Visual Studio, it is necessary to develop a communication and depuration method between these two platforms. To this end, a file-based communication system was deployed. The data flow from Visual Studio to Python indicates the stimulus parameters: pulse gap, pulse width, rate, phase gap, intensity and duration of stimulus, and the test in progress. The flow from Visual Studio to Python gives information about the connection state of the device to the computer, when it is stimulating, etc. The (NIC) library for Python, developed by Cochlear LTD. was used to create and send electrical stimuli to the intracochlear electrodes. This is a specific Experimental Work ! !!! 75! library built to allow researchers to design customized stimuli by changing its parameters: pulse width, pulse rate, phase gap and inter-pulse gap. To eliminate the need for programming, a user interface was built using Visual Studio, which communicates with Python. In turn, Python scripts control the implant receiver/transmitter by sending instructions to the supplied processor. Thus, the platform development has two main parts: 1) Python functions and 2) user interface development and its communication with Python. Data Base The data base has been arranged in such a way that a patient can be assigned to a range of psychoacoustic tests that are controlled by the investigator and test setup and outcomes allocated to an individual patient. That is, the tests are nested within each patient. In figure 7.2, this structure is depicted. Testing generally starts with obtaining Threshold and Comfort levels and loudness balance testing and ends with alternative forced choice (AFC) experiments. It can be seen how for each patient, the user can define as many stimuli as desired, each of which can be associated with different C and T levels. In turn, each C and T level configuration can be matched with many loudness balance tests. Finally, for each of the loudness balance configurations, an unlimited number of alternative forced choice (AFC) experiments can be set up. Experimental Work ! !!! 76! Figure 7.2. Data Base schematics. Tests are nested within their previous required step, the variable patient being the main sorting variable. User Interface The user interface was developed using Visual Studio and is thus compatible with Windows 7 and Vista. The computer requirements are: 800 MHz 32/64 bits processor, 1GB of RAM memory, Microsoft. NET Framework 4.5 and Microsoft SQL Server 2008 R2. The application allows the design of psychoacoustic tests where various types of stimulation modes and electrodes can be evaluated without the need to script order commands in any programming language. The files can be exported to an excel file for later analysis. How it works Insert Patient The option Insert Patient on the Start page opens the window in figure 7.3, where demographic and clinical data can be inserted. The Comfort (C) and Threshold (T) level values from the patient’s standard map can also be Patient!! Stimulus!A! C!and!T! level!map!1! Loudness! Balance!1! AFC!1! AFC!n! Loudness! Balance!n! C!and!T! level!map!n! Stimulus!B! Experimental Work ! !!! 77! inserted. Finally, the location of the electrodes with respect to the modiolus can also be catalogued as perimodiolar, mid-modiolar and lateral wall. On this window, and all subsequent ones, there is an indicator of the communication state between the application and the processor (connected/not connected) for safety reasons. It will show whether the computer correctly detects the processor or not. Figure 7.3. Insert Patient. Here demographic, clinical data and electrode position with respect to the modiolus can be registered. Stimulus Definition Once the patient has been saved, the Existent Patient window replaces the New Patient, where the same information is displayed, but now the user can select the different tests to perform the study. A stimulus must be defined before the experiment can proceed. The Define Stimulus button opens a window, figure 7.4, that allows customization of the following parameters: stimulation mode, stimulus duration, inter-stimulus interval, pulse phase and gap durations and pulse frequency. The stimulus Experimental Work ! !!! 78! consists of a train of electrical biphasic pulses at a frequency specified by the pulse rate option. Stimulation mode refers to the type of ground electrode used: the reference electrode located in the mastoid region (1), the internal processor case (2) or a combination of the two (1+2). The inter-stimulus interval refers to the resting time between two stimuli. These settings will remain fixed for all tests of the study. Figure 7.4. Stimulus definition window. PPS: pulses per second; Stimulation mode: ground electrode location (1 for ground electrode in the mastoid, 2 for internal processor case and 1+2 for a combination of both); # PPS: number of pulses per second (defines stimulus duration); silence: defines time lapse between pulse trains. Dynamic Range Map The Dynamic Range (DR) of an electrode defines the current intensity span that produces auditory sensation and is determined by the T and C levels of each electrode. The window can be seen in Figure 7.5. The T level is calculated using the classical ascending and descending method of limits (116): a stimulus is presented at progressively increasing intensity until the patient reports hearing sensation and then it is lowered down again using smaller intensity steps until it is no longer heard. To find the C level, the stimulus is presented in progressively smaller steps as the intensity rises until Experimental Work ! !!! 79! the patient reports that the sound is no longer comfortable because it is too loud, but never painful. In this part of the test, a new map can be created or a previous session can be loaded by selecting one from the Load Session scroll. When two different stimuli are used in an experiment, two independent maps must be created using one stimulus at a time. Figure 7.5. Dynamic Range Map. C and T levels for each electrode are determined using the up-down method. Loudness Balance The second step is to loudness balance all electrodes so that the stimuli delivered to all electrodes are perceived as equally loud. In the Loudness Balance window shown in figure 7.6, the user can compare an electrode with 21 other electrodes. In addition, one can select the stimulus to be used (A or B). The user must select the percentage of DR intensity that the initial reference electrode will have. This will be the electrode that will be objectively fixed at the desired percentage, and all others will be perceptually balanced Experimental Work ! !!! 80! with respect to it. The number of times one electrode will be compared to another (repetitions) can also be selected. The method used is the confluence method. The patient is presented with two stimuli, the first corresponding to a reference electrode, and the second to a test electrode. The latter begins randomly above or below 10% of the DR percentage selected prior to test commencement. The patient selects which one of the two stimuli sounds louder using the window in figure 7.7 and the intensity of the test electrode will be lowered or raised until confluence is achieved. Then the patient has to state that they sound equal using the appropriate button. The stimuli will be repeated with the same settings to confirm that the patient perceives them as equal and then the system will automatically start again but with the test electrode being 10% different in the opposite sense. The test electrode final value is the average of the two approaches. Figure 7.6. Loudness Balance window. Adjacent electrodes are loudness balanced using the confluence method. Experimental Work ! !!! 81! Figure 7.7. Patient view for electrode loudness balance. The patient has to state which stimulus sounds louder. Alternative Forced Choice Experiment for Electrode Discrimination The chosen method for electrode discrimination is a three-interval, forcedchoice procedure (3FC) (figure 7.8). Two of the stimuli come from a selected reference electrode and a third one from a signal electrode. They are presented in random order and each test electrode is presented the number of times stated in the Number of Repetitions scroll. The patient has a user interface with three buttons (stimuli 1 to 3) and has to select the one that sounds different. The researcher will select a reference electrode and a number of electrodes to compare it with. Results are given as percentage correct scores. Figure 7.8. Alternative forced choice experiment for electrode discrimination. Experimental Work ! !!! 82! 7.2.- Pitch discrimination comparison in CI users using perimodiolar vs straight electrode arrays There is great interest in placing the electrode arrays as close as possible to the modiolus, in an attempt to achieve lower thresholds and hence reduce excess dissemination of current that can result in interaction between channels. Physiological studies and mathematical models suggest that such placement reduces the stimulation current required to activate a neuronal population, increases the dynamic range and results in more localized regions of neuronal excitation. (117) (104) (118) (119) For this reason, the aim of this study was to investigate the difference in electrical pitch perception between these two types of electrode and test it on subjects carrying both types of arrays. Using the software application described in the previous chapter, patients fitted with the two types of electrode arrays were tested for their ability to discriminate electrical pulses from adjacent and near-adjacent electrodes at low intensity level (matching 25% of the DR of electrode 11). 7.2.1Introduction Pitch perception is known to vary regularly and essentially monotonically with the longitudinal position of the electrode stimulated in most patients fitted with straight arrays. However, the predictive model developed by Frijns suggests that in some circumstances, there is a "cross-arousal", where axons from more apical areas are also excited. Thus, besides the local target population, they also stimulated its upper and lower spiral regions. This would result in abnormally low pitch perceptions, especially at high intensity levels. The study suggests that this effect would be more likely in a set of electrodes medially positioned and beyond the basal turn. It is therefore important to investigate how electrical pitch perceptions vary with transverse and longitudinal stimulation sites. (120) (121) (107) Experimental Work ! !!! 83! The results of several experiments in which electric field patterns for different electrode configurations have been investigated within the cochlea, reveal that the spatial gradient increases as the distance between the stimulating electrodes and neuronal tissue decreases. Therefore, due to the steeper slope of the electric field, an electrode that is closer to the neuronal endings requires less current injection to reach the threshold current discharge. The result would be lower for sets of electrodes placed proximally thresholds. In addition, electric fields with steeper slopes spread less to adjacent neuronal populations, which decreases the probability of interaction between channels. This can result in a better discrimination between the electrodes. (122) (123) (118) (124) (119) (108) (125) A number of studies have examined the effects of the proximity of the electrode pacing thresholds. Recent studies comparing the position of perimodiolar and lateral wall electrode arrays have shown lower thresholds in the former. Cohen, Saunders and Clark also measured the ability to discriminate in three subjects with perimodiolar arrays. In two of the three subjects, a portion of the electrode was located near the modiolus. They showed better pitch discrimination for electrodes in this region. These results suggest a better spatial selectivity and lower thresholds for electrodes located closer to the neuronal tissue. (126) (127) (128) (129) Shepherd et al. recorded lower electrical auditory brainstem responses (EABR) and a smaller slope of the growth function’s amplitude for electrodes placed close to the modiolus and to the spiral ganglion, as opposed to electrodes placed near the lateral wall or the middle portion of the scala tympani. Liang et al. showed that lower thresholds were associated with better spatial selectivity of the electrically stimulated auditory nerve fibers in cats. They also found great variability in the response of auditory neurons to electrical excitation and suggested that the nerve fibers closer to the stimulating electrode are the most selective. Finally, Cohen et al. described narrower width excitation profiles for implanted subjects with perimodiolar arrays than for the subjects implanted with straight arrays, as measured by Experimental Work ! !!! 90! dynamic range at random. The patient listens to the two sounds sequentially and must answer the question: which sound is louder? ", The answer being either 1st or 2nd. The response is recorded using the buttons on the interface to lower or raise the signal electrode by 1 CL until the patient perceives them as equally strong. At this point, the same sound is replayed to verify that this is the case and then the test is reset but starting with the opposite situation. That is, if the target electrode started 10% above 25%, this time it will begin below 10%. The equal loudness point is set as the average of both approaches. Subsequently, the target electrode will serve as the reference electrode for the next adjacent electrode. Hence, the same procedure is repeated until the apical end is reached and then the process starts again from electrode 11 towards the basal end of the array. The balance is set between the closest electrode pair for two reasons: 1) Discrimination is most difficult at this spatial distance than any other, so their stimuli need to be as similarly intense as possible and 2) because the patients reported great difficulties and even inability to establish equal loudness between two very different electrical pitches. Step 3: Test electrode discrimination Electrode discrimination was measured using the three-alternative-forcedchoice method. After selecting the central (reference) electrode and its two nearest neighbors on either side (signals), the number of times that each electrode will be tested with the central electrode was set to three. Thus, three electrical stimuli were sent to the implant, two of which came from the same electrode (reference) and one from the signal electrode. They were presented in random order. The patient had to indicate the stimulus that sounded different from the other two. After all electrodes were tested with the reference, the test finalizes and the researcher had to select the apical, most adjacent electrode as reference and its two nearest neighbours on either side as signals. The process was Experimental Work ! !!! 91! repeated until all electrodes had acted as references. This way all electrodes were tested against 4 nearest neighbours, except for the apical and basal most electrodes, which could only be tested to one side. A similar situation holds true for the ones before these, because they have two electrodes on one side but only one on the other. The results were given as percentage of correct answers. Imaging Plain X-ray Stenvers projections have been used in the past as a standard for electrode array location within the cochlea, but nowadays it has been abandoned (137). Many recent studies consider the benefits of high-resolution Computer Tomography (CT) scanning as a more exact method to visualize the structures of the temporal bone. The advantages of CT scanning include the elimination of respiratory and other motion artefacts and an obvious improvement in the diversity of tissue that can be observed (138). The main disadvantage of CT in the postoperative assessment of a cochlear implant is image degradation, due to partial volumes and the metallic artefacts that may interfere with the visibility of individual electrodes (139). In this study, a temporal bone CT scan with a z resolution of 0.062 mm was used. To overcome this limitation, a computer-based system was developed to obtain the distance between each electrode in the array and the inner wall of the cochlea. A MATLAB script was made to get an estimate using two images of a CT scan. High-density (i.e metallic) objects were highlighted in the first image, so that the contours of the implant became visible. In the second one, an average density with a high contrast was used to see the inner wall. Once the images were imported, the first image was used to mark a series of points (as many as necessary are created) that follow either: 1) the inner contour of the perimodiolar array, because the contour advance implant has half-banded electrodes that are only located in the inner side of the array; or 2) the middle portion of the straight array, because in this array the electrodes Experimental Work ! !!! 92! are a full ring. Once the points were set, the system generated a parametric, interpolating function to calculate the curved shape of the implant. Then the same procedure was repeated for the second image, but marking the inner wall instead. Once generated, the inter-electrode distance specification of the manufacturer is used to determine individual electrode location. Then the distance to the inner wall was calculated as the perpendicular distance to the tangent of the curve at the electrode location. The final measures are displayed on the second image (Appendix IV). Data analysis Analysis of the variables under study was done using IBM SPSS Statistics for Macintosh, V. 21.0. (IBM Corporation, New York, United States) and eViews . First, descriptive statistics for the endogenous and all exogenous variables were done. Then a linear regression model, a logit and a binomial logit model were calculated to study the correlation between the variable and its predictors. 7.2.4.- Results Four perimodiolar electrode array carriers (patients 2-4 and 9) and 4 straight array carriers (patients 5-8) completed the test. Patient 1 did not attend the CT scan appointment and patient 10 had an inappropriate MRI technique for an unrelated health issue and his array was displaced so he had to undergo a repairing surgery and so could not complete the test. Consequently, a total of 1968 trials were used for the analysis. Of all these, 1169 were correct answers and 799 were not. Trials were grouped into units of observations. A unit is the result of first averaging the success of a given reference electrode with each of its nearest and second nearest neighbours (4 for electrodes 3 to 20; 3 for electrodes 2 and 21 and 2 for electrodes 1 and Experimental Work ! !!! 93! 22) and then further averaging these into a single value. The success rate is thus the quotient between the number of successes by the number of trials, for a given reference electrode (6 trials for electrodes 1 and 22, 9 trials for electrodes 2 and 21 and 12 trials for the rest of electrodes). Its descriptive statistics (Table 7.2), the Jarque-Bera normality test, together with the histogram of frequencies are shown in figure 7.9. From the latter it is concluded that normality cannot be rejected at 4.60% significance (p=0.046) and hence the variable behaves has a near Gaussian distribution, with a success rate of 59.70%. Figure 7.9. Success rate histogram representation Descriptive Statistics Mean 0.597 Median 0.583 Maximum 1 Minimum 0 Standard Deviation 0.247 Asymmetry -0.373 Kurtosis 2.470 Jarque-Bera Test 6.141 Jarque-Bera associated probability 0.046 Table 7.2. Descriptive statistics. Experimental Work ! !!! 94! To identify which factors have the capacity to influence and modify the success rate, we started by contrasting whether the electrode number could be one of them. A Levene’s test for equality of variances of the mean success rate indicates that the variance is the same for all electrodes (p=0.35, the null hypothesis which states that the variances are different cannot be rejected). Consequently, an ANOVA contrast was also done, concluding that the equality of the means of the successes for all electrodes cannot be rejected (The F statistic value had an associated p= 0.82). A similar analysis of the variance was conducted for the patient factor. The result suggests the existence of statistically different results depending on the patient, both for mean and variance values (probability associated with the Levene statistic of equality of variances p=0.0001, and probability associated with the F statistic of Welch of equality of variances p=0.000). Accordingly, the empirical evidence supports the hypothesis that there is a patient effect that significantly affects success rate. Figure 7.10 shows the standard deviation plot of the mean success rate for each patient. ! Experimental Work ! !!! 95! Figure 7.10. Mean success rate values and standard deviations per patient (pat). As can be seen from the chart, there are four patients that stand out from the rest. Patients 6 and, most prominently, 8 have a large dispersion. In addition, patient 6 also has a low mean success rate, although not as low as patient 7, which is the third patient that stands out from the sample. On the opposite end is patient 4, with the highest success rate and the smallest dispersion between electrodes. Of these patients that have a differential behaviour, three of them, patients 6, 7 and 8, have straight electrode arrays (although so does patient 5). In fact, the data shows a statistically significant difference in success rates between patients implanted with straight and perimodiolar arrays. The probability associated with the Welch’s F statistic (because variances between groups were different) for the equality of the probability of success in the presence of different variances also gave a value of 0.000, being the probability of success 53% for straight electrode array carriers and 70% for perimodiolar electrode array carriers. This result provides empirical evidence in favour of the hypothesis that perimodiolar electrode arrays provide better electrode discrimination. The mean effect of distance on success rate in average terms is evident when these two variables are plot against each other (figure 7.11). Patients whose electrodes are further away from the inner wall are those with a lower success rate, on average. Experimental Work ! !!! 96! Figure 7.11. Relationship between the mean electrode-inner wall distance per patient and mean success rate. In any case, the above results must be taken as indicative,-not conclusive-, for several reasons. On the one hand, one cannot differentiate the patient effect from the distance effect on success rate. On the other hand, the results are based on this variable having normal behaviour. However, success is a binomial variable. The raw output data consisted on a matrix of ones (success) and ceros (failure). In fact, even though the Jarque-Bera contrast does not reject the null hypothesis of normality of the success rate, other contrasts reject such behaviour for significance levels of less than 1% (KolmogorovSmirnov, Shapiro-Wilk). A correlation and regression analysis was done to overcome these limitations and to search for greater precision in the measurement of the relationship between the distance of the electrode to the inner wall of the cochlea and the Experimental Work ! !!! 97! success rate. To this end, the following additional variables were also incorporated: 1. NRT: quality of the nerve measured in CL. A higher value indicates a worse state of the nerve, as more current needs to be delivered to the nerve end to elicit a measurable response. Thus, it is expected that the increase in the value of this variable results in a decrease in the probability of correctly discriminating between electrodes. 2. T level: minimum current needed to elicit hearing sensation, measured in CL. According to the literature, this variable is expected to have a positive relationship with success rate. 3. Impedance: measured in kΩ. Because larger impedances result in higher currents required to reach the neural ends, this variable is expected to have a negative effect on the success rate. To do this analysis, as indicated earlier, the sampling units are all the different electrode-patient combinations, i.e. electrode 1 of patient 1, electrode 1 of patient 2, etc. The descriptive statistics of the variables are shown in table 7.3. On average patients have been successful in 60% of the tests, have an electrode distance of 0.87 mm, a mean NRT of 183.61 CL, a T level of 126.29 CL and an impedance of 9.66 kΩ. The relative dispersion is moderate in all variables, being the variable distance the one with greater relative dispersion. Because the variable deterioration of the nerve (NRT) was not measured in 12 electrodes, the sample size had to be reduced to 158 elements. Table 7.4 contains the correlation matrix of the variables and their level of significance. It can be concluded that the variable of interest, the success rate, is significantly correlated with all explicative variables, and with the expected signs as well. There are also significant correlations between the different explanatory variables, which make it difficult to quantify their individual effect on the mean success rate. However, their values are not high enough to produce serious problems of multicollinearity. Experimental Work ! !!! 98! Success Rate Electrode-inner wall distance NRT T level Impedance Mean 0.60 0.87 183.61 126,29 9.66 Maximum 1.00 2.09 237.00 176.00 20.33 Minimum 0.00 0.11 109.00 95,00 4.23 Standard Deviation 0.25 0.50 22.06 18.56 3.24 Relative dispersion 0.41 0.58 0.12 0.15 0.34 No. of Observations 176 176 158 176 176 Table 7.3. Descriptive statistics of the variables: electrode-inner wall distance, NRT, T level and impedance. Success Rate Electrode-inner wall distance NRT T level Variance Inflation Factor Electrode-inner wall distance -0.207*** 1.23 NRT -0.294*** 0.421*** 1.87 T level 0.169** 0.260*** 0.521*** 1.64 Impedance -0.318*** -0.015 0.217*** -0.222*** 1.27 * Significant at 10%, ** significant at 5%, *** significant at 1% for bilateral contrasts Table 7.4. Pearson correlation of the variables: success rate, electrode-inner wall distance, NRT, T level and impedance. The determining factor when modelling the endogenous variable is the very nature of the variable itself. Success rate is bounded between 0 and 1 and its theoretical behaviour corresponds to a Bernouilli distribution. Let us remember that the success rate has been obtained as the quotient between the number of right answers divided by the number of trials per electrode pairs for a given reference electrode (6, 9 or 12). When calculated in this way the variable has a binomial behaviour. Therefore, the assumption of linear relationship with the explanatory variables, which implies a constant linear effect, is not strictly valid because linear models give values below zero and above one. Moreover, the variance of the success rate is not constant and so there is heterocedasticity. These properties limit the use of the standard method of multiple linear Experimental Work ! !!! 99! regression and its estimation by ordinary least squares and thus we must resort to alternative methods before giving final conclusions. Three predictive methods were used, all of them using the explanatory variables described above. First, success rate was modelled using the method of minimum chi-squares. To adapt the model to the nature of the variable of interest, two more models are proposed, both taking the logit transformation of the success rate as the endogenous variable. Linear probability model estimated by the method of least chisquares (MCSM) The first model describes the probability of success as a linear function of the predictive variables and the weighted least square estimation of the following form (140): !!=!! !! =!!+!!!!!+!!!!!+⋯+!!!!" +!! [7.1] This is the simplest method and the easiest to interpret, but is not without serious limitations. pi is the success rate (it is neither the number of successes, nor the number of tests); {x1,…,xk} are the set of predicting variables; ui is the random perturbation; are the parameters of the model that quantify the effect of each one of the predictive variables on the success rate and i relates to each of the sampling units (patient-electrode). Given the nature of the endogenous variable, the model has a problem of heteroscedasticity. Therefore, the data was weighted using a new variable, Pond1, defined as: !"#$1=!! !!(!!!!) [7.2] β i Experimental Work ! !!! 106! 8.- ELECTROMAGNETISM 8.1.- The effect of reference electrode position on power consumption in Cochlear Implants (Published as: Ramos-Miguel, A, Ramos-Macías, A, Artiles, JV, Zaballos, MTP (2015). The Effect of Reference Electrode Position in Cochlear Implants. J Int Adv Otol, 11(3), 222-8.) Despite the use of different electrode designs and coding strategies across Cochlear Implant (CI) companies, the variability in auditory abilities of CI recipients seems to be more similar across devices than for the same device across individuals. This suggests that significant recipient-dependent factors limit overall auditory ability at an individual level. Aspects such as neural survival are known to have a significant effect. Unfortunately, little can be done to improve patient performance in such cases. However, there might be other reasons related to the physical characteristics of the device, rather than the patient’s etiology. (144) In this study, the effect of Reference Electrode (RE) position on power consumption and stimulation intensity was investigated using two different approaches: 1) 2D numerical simulations and 2) real temporal bone measurements using a test implant developed by the Laboratory of Psychoacoustics of the Otoneurology Division, ENT Department of the Complejo Hospitalario Universitario Insular Materno Infantil (CHUIMI). 8.1.1Introduction The RE position in CI has been surprisingly understudied. There are no articles investigating a correlation between battery life and position of the RE. Most centers do not even use this electrode, with just the body of the implant Experimental Work ! !!! 107! being used. From a physical point of view, RE position has a bigger influence regarding power consumption than intracochlear electrode position. Despite using different electrode designs and coding strategies across CI companies, the variability in auditory abilities of recipients appears to be more similar across devices than for the same device across individuals. This suggests that at an individual level, significant recipient-dependent factors limit the overall auditory ability. Aspects, such as ossifications, are known to have a significant effect. Unfortunately, to improve patient performance in such cases split electrodes have to be used, and they require exquisite optimization of parameters to obtain good results!(144). Goehring et al. discovered that 12.4% of implantations intraoperatively demonstrate high impedances. In 8.2% of cases, impedances postoperatively remain high. (145) There is extensive literature wherein many variables concerning intracochlear electrodes are evaluated. In contrast, very few studies have explored the effects of position, number, or shape of RE, without which CI would not properly function (146). A field that has been explored in more depth is the comparison between monoand bipolar stimulation on the still controversial hypothesis that more restricted current fields can provide better speech perception. However, there are problems with bipolar stimulation: 1) Bipolar electrodes require higher currents for supra-threshold stimulation because of a current shunt from electrode to electrode and 2) high potentials and high current densities form around each of the two electrodes and supra-threshold excitation is possible near both the source and sink electrodes (147) (148) (149) (150) (151) (152). Therefore, power consumption is still a major drawback in this stimulation mode, thereby making monopolar the preferred choice of most audiologists. Therefore, it is important to evaluate whether RE position in monopolar stimulation can decrease power consumption and improve electrical field sharpness. Experimental Work ! !!! 108! In monopolar stimulation, the current is injected through a contact in the scala tympani and returned to a far-field electrode contact. Briaire and Frijns studied the current distribution in multiple electrode contacts along the scala tympani using this mode (106). The result demonstrated that electrical potential field patterns caused by monopolar stimulation are broad in nature and that potential relatively and slowly drop off as one moves away from the stimulating contact. High impedances in the tissues can cause losses in the electrical impulse quality, thereby requiring high power for the patient to efficiently detect this signal. It can even lead to the deactivation of some electrodes because their power consumption is so high that they hinder the global stimulation strategy. When impedance is too high, particularly at high current levels, then it is possible that the implant voltage will not be sufficient to generate the programmed electric current level. In this case, an electrode is said to be “out of compliance.” This leads to the situation where perceptual loudness does not increase with increases in the current level. To increase loudness, the width of stimulation pulses can be increased. In some cases, impedance is too high that this approach is not feasible because of too large pulse duration required. Hence, high impedance results in a decrease in battery life but also in a wider spread of excitation because of the large current intensity required. However, decreasing impedance is feasible and can solve the issue of battery consumption. How this might be done depends on the source of impedance. It could be achieved by reducing the space between the RE and signal electrode. The RE can be the extracochlear electrode of some types of implants, which can be placed in different parts of the mastoid, but also the body of the implant, used as ground electrode as well. By reducing the amount of tissue between the ground and intracochlear electrodes, impedance could be lowered, and the pulse intensity required could be obtained with less power. (153) (154) In CI surgery, not much attention is given to the specific position of the RE. The surgical technique historically suggests positioning the RE as far as Experimental Work ! !!! 109! possible from the electronic part of the implant to avoid interferences during stimulation. This part is currently isolated from this type of interference by Faraday cages. Therefore, RE can be currently positioned closer to the stimulation regions. (155) (156) This study aims to examine the variability in extracochlear electrode position as a contributing factor to power consumption in CI. This simple change can be of great advantage regarding the battery life for patients with cochlear ossifications. Because they require high currents to perceive electrical stimulation, a closer distance between the intraand extracochlear electrodes may enable the generation of larger current intensities without the need to upturn the pulse duration to increase auditory sensation. 8.1.2Objective, Hypotheses and rationale for the study The aim of this study is to examine variability in extracochlear electrode position as a contributing factor to variability in auditory ability in CI patients, in the context of two questions, 1) whether RE placement has a measurable effect on consumption, and 2) if RE position has a functional effect on stimulation intensity. The hypothesis is that the closer the RE is placed from the stimulation region, the lesser power consumption will be required to reach a point at a set intensity. The reason behind this study has been to try and increase the battery life of the cochlear implant and to improve implant functioning in patients with ossifications in the cochlea. Because they require high currents to perceive electrical stimulation, a closer distance between intra and extracochlear electrodes may allow the generation of more intense currents without the need to increase the pulse bandwidth to augment the auditory sensation. Experimental Work ! !!! 110! 8.1.3Material and methods Ethics committee approval was not required for this study because it did not involve human subjects. Classical Circuit Analysis To characterize the theoretical physical system a much-simplified electrical model was used. Three electrical formulas were used. The first is Ohm’s Law, which was postulated by Geor Simon Ohm in 1826. It states that a current I through a conductor between two points is directly proportional to the potential difference V across the two points, with resistance R being the constant of proportionality between them (157): !=!∙!![8.1]! In this case, V is the potential difference produced by CI between the signal and the RE, R is the impedance caused by the tissue between the electrodes, and I is the electrical current passing through it. Solving for I, it becomes clear that to raise this value, V needs to be increased, or R has to be lowered. Thus, the only way to do this without supplying extra energy is by reducing the resistance between the RE and the intracochlear electrode. The second equation is the electrical power formula: !=!∙!=!∙!! [8.2] It represents the rate of energy that is converted from the electrical energy of the moving charges to some other form, e.g., heat, mechanical energy, or energy stored in electric or magnetic fields. Power is measured in Watts (W) and is used as a measure of the efficiency of a device. It is directly Experimental Work ! !!! 111! proportional to the square of the electrical current and resistance of a circuit. (157) Thus, for a fixed I, the higher the R, the more the power required to be supplied by the circuit to keep I constant. Electrical hearing threshold occurs when the current I delivered to the nerve fibers reaches a minimum value. This value can be taken as the fixed value of I. Then, it becomes clear that the further away the RE is from the active electrode, i.e., the higher the resistance between them, the more power will be required to reach that threshold. The last formula used is the mathematical definition of the electrical resistivity of a material, which is the resistance opposed by a material to an electrical current flow across it. This formula gives the characteristic R of each biological tissue between the RE and an intra-cochlear electrode. !=!! ! [8.3] Where ! is the length of the conductor, S is its cross-section area, and ρ is the electrical resistivity coefficient of the material, which is measured in Ohms per meter (Ω·m). In this study, the values of ρ for endolymph (ρ=0.75 Ω·m), bone (ρ=20 Ω·m), and muscle (ρ=3.5 Ω·m) were obtained from the database that was created by IT’IS Foundation. (IT’IS Database Low-Frequency Conductivity) 2-D Numerical Simulations The first technique used was numerical simulation. The advantage of using this approach is that the physical phenomenon can be recreated in a completely controlled environment where all the variables and geometries can be defined, measured, and modified to match the real-life system. The simulation software used was Mathematica (Wolfram Research, Inc., Mathematica, Version 10.0, Champaign, IL (2014)). A 2D finite element Experimental Work ! !!! 112! simulation of electromagnetic phenomena for low frequency currents was build. A useful approach to the calculation of electric potentials (Voltage) is to relate that potential to the charge density. Since electric charge is the source of electric field, the electric field at any point in space can be mathematically related to the charges present. For multiple point charges, a vector sum of each of their electrical fields is required. For a continuous distribution of charge, vector sums are better handled through calculus. This is done through the divergence relationship. This states the divergence ∇∙ (the scalar derivative) of the electric field vector E is proportional to the charge density ! (158): ∇∙!=−!! !! [8.4] Where !!is the permittivity and !! is the charge density. Because the electric field is also the derivative of the potential, we obtain Poisson’s equation: !=∇∙∇V=∇!!=−! !! [8.5] In a charge-free region of space or when the frequency is low, the governing equation for the electromagnetic field is the static equation, which is Laplace’s equation. This is a differential equation with a smooth solution, without singularities. It is used to define the potential energy field caused by a given charge distribution: ∇!!=0 [8.6] The model’s task will then be to determine the solution to this equation for a biological model, which involves various materials and geometries. This Experimental Work ! !!! 113! equation gives the second derivative of V; however, because it is V itself that is needed, integration is required. However, while integrating, constants appear that cannot be numerically determined unless the boundary or initial conditions of the system are known. These are the values at the borders of the system’s space or the initial value of V before it starts moving through space and time. (158) Two types of conditions were used: Dirichlet and null Neumann. The first specifies the conditions that the electrical potential requires to take along the boundary. The second sets the value of its derivative along this boundary. Dirichlet conditions of 0 and 5 V were imposed on the RE and intracochlear electrode, respectively. The null Neumann condition was applied to the edges that limit the geometry of the system. (159) Another important step in designing the computer model is the geometry, which is illustrated in Figure 8.1. The model represents three electrode positions. They all share the same elements; however, the distance between the electrodes decreases by a factor of 2 between one case and the next. Figure 8.1 Geometries used in the numerical model. The inter-electrode bone layer thickness was modified to study its effect on impedance, and thus, electrical consumption. Each element was assigned the electrical resistivity values described in the circuit analysis section. Air was included in this model, with !=1.3×1012 Ω·m. (160) Experimental Work ! !!! 114! Temporal Bone Measurements A device was designed to stimulate and monitor electrical impulses that were delivered to a temporal bone while measuring impedance, current, and voltage. The purpose-built implant comprises one intracochlear and six extracochlear electrodes. Because only the effects of relative position between the RE and intracochlear electrode are of interest in this study, CI was simplified and comprised a single intracochlear electrode. In contrast, to assess extracochlear electrode currents along different regions at the same time, six of them where located along the temporal bone. They all have spherical ends with a mean diameter of 0.9 mm. All extracochlear electrodes can be set to be RE, one at a time, thus providing six different measurement positions. The remaining electrodes act as probes throughout a given stimulation trial. The system comprises a microcontroller (Arduino UNO, Somerville, Massachusetts, USA) that is connected to a computer via USB. It was programmed to transmit electrical signals, sample each of the probes, and store data. It also calculates impedances, currents, and power that are generated during stimulation. The board is connected to a breadboard with the electrodes and probes to be inserted into the temporal bone. Experimental Work ! !!! 115! Figure 8.2 Schematics of the implant circuit. The reference/probe electrodes were attached to the analog ports of the microcontroller board. Switchers S1–S6 define which probe to select as RE. S12 disconnects or reconnects the intracochlear electrode. Resistors R1–R7 all have a value of 1.5 kΩ. The circuit is based on a number of voltage dividers. The schematics are depicted in Figure 8.2. The two resistors in series form a voltage divider circuit. One end of the resistor pair is fed 5 V, while the other end is connected to the ground. Five volts that the Arduino provides are divided between the two resistors depending on their resistance. The resistor, which holds the greater resistance, gets more of the voltage, according to Ohm’s law formula. The voltage that falls across a component is directly proportional to the amount of resistance it contains (157): !!"#$ =!!"#"!"$%"!!!" ! !"# −!!"#"!"$%" [8.7] Using this principle, a model can be set up to determine the resistance on the basis of the voltage division. With this data, all electrical characteristics of the circuit can be defined, i.e., impedances, currents, and voltages. Discussion ! ! ! 122! Discussion ! ! ! 123! III. DISCUSSION Discussion ! ! ! 124! Discussion ! ! ! 125! 9.- PSYCHOPHYSICS 9.1.- How the healthy ear understands speech in noise and the effects of training on it The ability to focus and understand a speaker in a noisy environment is a critical social-cognitive capacity whose underlying mechanisms are still unclear. Many explanations have been proposed. However, none had looked at high frequencies as the possible cues for speech identification in noise. Thus, in the first psychophysical study of this thesis, the relation between speech recognition in noise and frequency bandwidth of the auditory stimulus was investigated. Many studies have already demonstrated the importance of frequency bandwidth in the perception of speech in quiet. For example, Stelmachowicz et al. have studied these effects on the perception of the phoneme /s/ in Normal Hearing (NH) children and adults. The results showed a correlation between these two factors for all groups. A clear example where a limited hearing bandwidth can be found is in presbycusis, which involves the rise of auditory thresholds for high frequencies (161). In these patients, the most common complain is the inability to understand Speech In Noise (SIN), while they can still follow a conversation in a quiet environment with little difficulty. These two findings (that high frequencies are important for speech perception in quiet, and that the rise in high frequency hearing thresholds hinder SIN ability) seem to point in the direction of high frequencies as important cues for speech perception, specifically in the case of background noise. A study by Stuart et al. analysed word recognition performance in 12 NH adults in continuous and interrupted broadband noise as a function of SNR with and without filtering at 2 kHz. However, their filtering threshold was lower than the actual 8 kHz that average implants have. Moreover, they focused primarily on the difference between continuous and interrupted noise, rather than the role of high frequencies. In a period where Cochlear Implants (CI) only reached around 4 kHz, Hornsby also demonstrated that better Discussion ! ! ! 126! performance was accomplished when the frequency bandwidth was increased to 7 kHz. (162) (163) The statistical outcome in this thesis is in accordance with the findings from Stuart and Hornsby, for NH subjects, when noise is either higher or lower than the speech signal. The results suggest that high frequency components above 8kHz may play a significant role in word recognition in noisy environments. However, for the situation where noise and speech have equal levels of intensity this was not the case. The characteristics of the experiment do not allow for a clear explanation as to why this happens. According to the results obtained, there seems to be a difference in how humans recognize speech in different noise situations. This is in accordance with a neurophysiological study by Wong et al. where they measured cortical activity of NH subjects while they identified words in two noise conditions: below and above the speech level. They found cortical noise-dependent activation of the bilateral middle and left posterior portions of the superior temporal gyrus. These results likely reflect demands in acoustic analysis, auditory-motor integration and phonological memory, as well as auditory attention. Unfortunately, they did not study what happened when noise and speech were at the same level, so the activation sites in this case are still a unknown. Thus future research points in the direction of combined psychoacoustics and neurophysiological experiments, where neural activity is recorded while speech material is presented using the three SNR of the present experiment. (69) Correlation studies between how much a word was altered by filtering and its score did not yield conclusive results. Significant correlation was only found for the case where noise was lower than speech. This was somewhat expected, since word recognition is a complex phenomenon. Vowels and consonants are perceived through different mechanisms. Nie et al. sustain that consonant recognition happens through temporal cues while vowel recognition relies on spectral cues. The results of the present experiment, in combination with the statistical outcome observed, seem to indicate that even Discussion ! ! ! 127! small additions of high frequency content to sounds can lead to an overall improvement in speech recognition in noise. In the light of these arguments, it is hypothesized that vowel recognition assumes an important role in speech recognition in noise. Moreover, results from Friesen et al. indicate that proper vowel recognition requires appropriate frequency allocation. The implications of their findings and those obtained in this thesis lead to the thought that for appropriate speech recognition in noise, it will be necessary to fulfil two requirements: 1) a correct frequency allocation and 2) the preservation of high frequencies. This is because if the audible frequency is increased, but the allocation of such frequencies is not accurate, then the result will not be reconstructed properly by the brain. (164) (165) The first study (section 6.1) thus showed how degraded high frequency thresholds can lead to decreased ability to understand SIN. Hereafter, the natural question to ask is: is there a way to prevent, or even reverse this problem? If the answer to this question were positive, this would mean that patients with hearing loss could significantly improve their SIN perception through an appropriate training protocol. The search for the answer leaded to the second study of this thesis (section 6.2). The ability of air traffic controllers (ATC) to understand SIN under limited frequency bandwidth conditions is extraordinary. However, this does not seem to be an innate ability. These individuals all report having progressively acquired this faculty throughout the years, but specially during their two-year training period. Thus, the second experiment of the series of psychophysical experiments in this thesis tried to objectify the effect of training on SIN understanding. Many studies have investigated the effect of SIN exercises. Song et al. studied training-related malleability using a program that incorporated cognitively based listening exercises to try to improve SIN perception in patients with hearing loss. Trained subjects exhibited significant improvements in SIN perception that were retained 6 months later. Discussion ! ! ! 128! Subcortical responses in noise demonstrated training-related enhancements in the encoding of pitch-related cues. This was the first time that short-term training was demonstrated to have the potential to improve the neural mechanisms for SIN perception. These results involve and define biological mechanisms that contribute to learning success, and they provide a conceptual advance to the understanding of the kind of training that can influence sensory processing in adulthood. (99) More recently, Sweethow and Sabes prospectively assessed the generalization of SIN training in a cohort of individuals with hearing loss (166).. Whitton et al. prospectively assessed the effect of signal-in-noise, audio-game training on SIN understanding with untrained materials (167). On the other hand, Fu and Galvin and Moore and Shannon have shown that targeted auditory training can further enhance the benefits of new implant devices and/or speech processing strategies (168) (169). These findings and those from this thesis suggest that SIN training could be beneficial both for CI users and for those with moderate hearing loss, since it equips patients with the cognitive tools to confront complex auditory situations. However, the only groups that have been studied and that have some kind of long-term training are musicians, who are auditory experts. For example Parbery-Clark et al. investigated the effect of musical training on SIN performance. They found that musical experience improves the ability to understand speech in challenging listening environments. The results also suggest that this enhancement is derived in part from musicians’ remarkable working memory and frequency discrimination.. (170) The results described in section 6.2 indicate that ATC, just like musicians, are more capable of identifying SIN. This ability is most apparent in the most adverse case, when the signal is lower than the noise. In all cases, they obtained better scores than the Control group. The proposed explanation is that this is due to the effect of training that occurs naturally during their daily activities. Even though the set of words used were completely unknown to Discussion ! ! ! 129! them, it seems that the skills learnt in one particular auditory situation are transferable to a new and different one, in this case, the laboratory test. This is an interesting result, since the first thing that subjects commented after being informed of the purpose of the study was that they ‘would probably not perform better because their “trick” is that they know which words to expect during radio communications’. However, it is clear from the results that this is not the case and they could identify more words than the standard NH population. Regretfully, the prevalence of this opinion was not quantified in this study. It would have been very interesting to see how much their beliefs correlate with the test results. The studies reviewed here favour the hypothesis that their subjects, and likely ATC as well, are better at focusing on the less degraded acoustic speech cues and then filling the gaps using cognitive skills in the presence of background noise. This way they require less time of adaptation and therefore miss less information. Given the promising results, it was decided to investigate possible correlations between years working and age. However, this analysis did not return clear results. The two cases where significant correlations were found (inverse correlations between 8 kHz 5 SNR test results and 22 kHz 5 SNR) do not seem meaningful in the context of the other tests. If a correlation of any kind is to be established between these two variables, then consistency across tests is required. Thus, the results observed are likely to be caused by the limited number of subjects, which is an unfortunate but common aspect in clinical studies, as opposed to pure physics studies, due to the difficulty of recruiting volunteers. In such cases, statistics needs to be treated with care. It is possible and very likely, given the time taken by subjects in the studies mentioned before (approximately 6 months1 year) (97), that the learning curve is steeper during their first year of training and approaches a plateau thereafter. Thus, another study emerges as a very interesting continuation of the present one: to test ATC in training at different stages of their learning period, or to follow them longitudinally. Maybe looking more closely into this period of their working life will reveal a significant correlation between SIN identification and training duration. Discussion ! ! ! 130! The study also looked at how high frequencies affected SIN understanding in this population group. Interestingly, both control and target groups had statistically significantly better results when the unfiltered lists were used. The results indicate that performance decreases when high frequency components of speech are removed. Still, the ATC group performed better in both conditions, thus indicating an auditory learning effect, but the present study could not acceptably correlate it with either age of working years. Nevertheless, the role of high frequencies is remarkable, since it is present in both trained and untrained subjects. This further supports that frequencies above 8 kHz can help humans to better understand SIN. To draw further conclusions, a more specific experiment needs to be designed, where groups of words with similar high frequency contents are selected to try and establish a relationship between word identification and its frequency content, because the speech material used has balanced phoneme content. The implications of these findings are profound because they can result in better rehabilitation therapies for people with hearing loss. Our aging society is bound to suffer SIN understanding impairments. As individuals grow older, this ability degrades even before auditory losses become clinically relevant. Therefore, it is capital to come up with strategies to try and slow this process down as much as possible. The results from section 6.2 have demonstrated that daily exposure to speech in noise can result in adaptation and enhanced performance in NH individuals. This is the first time that ATC are studied for this singular ability. In addition to the recent proofs that learning during adulthood is far greater than previously thought, we provide evidence that auditory training under degraded sound quality can lead to better SIN understanding and can have big implications in the way hearing loss is treated today. Hence, research in the line of this study will certainly aid in the development of efficient and effective training protocols and materials. Discussion ! ! ! 131! 10.- ELECTRICALLY EVOKED PSYCHOACOUSTICS IN COCHLEAR IMPLANT RECIPIENTS 10.1.- How physical and psychophysical variables affect electrode discrimination in cochlear implant recipients The aim of this study has been to measure the effect that the distance between the intrachoclear electrodes and the modiolus has on the ability of a patient to discriminate the origin of a certain stimulus, separating this effect from those of the NRT, T level and impedance, since these factors may also influence the ability to discriminate electrical pitch. The motivation for this study has been that a better ability to discriminate electrical pitch could help to design better training for CI patients, since it has been seen in section 9 how frequency resolution plays a major role. The results provide evidence that the ability to discriminate between electrodes depends on multiple factors. In this study, NRT, T level, impedance and distance were identified. They were all statistically significant in explaining electrode discrimination. Especifically, the statistical analysis provides empirical evidence in favour of the hypothesis that the electrode distance to the inner wall of the cochlea is a significant predicting variable on the electrode discrimination ability of the patient, measured as success rate. The results indicate that the greater the distance, the lower the success rate will be. Thus, distance to the inner wall will result in more difficulties to perceive electrode differences. The explanation can be found in the extensive research that has been conducted to describe current field patterns for various electrode configurations within the cochlea (171) (172) (173) (105). These experiments show that the spatial gradient of the electric field increases as the distance between the electrodes and the neural ends decreases. Thus, an electrode that is closer to a surviving neural element would require less current to reach