Full text
2014 3 Eva María Cavero Racaj Design and evaluation of echocardiograms codification and transmission for Teleradiology systems Departamento Director/es Instituto de Investigación en Ingeniería [I3A] Alesanco Iglesias, Álvaro García Moros, José Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Departamento Director/es Eva María Cavero Racaj DESIGN AND EVALUATION OF ECHOCARDIOGRAMS CODIFICATION AND TRANSMISSION FOR TELERADIOLOGY SYSTEMS Director/es Instituto de Investigación en Ingeniería [I3A] Alesanco Iglesias, Álvaro García Moros, José Tesis Doctoral Autor 2013 Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Departamento Director/es Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
DESIGN AND EVALUATION OF ECHOCARDIOGRAMS CODIFICATION AND TRANSMISSION FOR TELECARDIOLOGY SYSTEMS Eva MaCavero Racaj Supervisors: ´ Alvaro Alesanco Iglesias and Jos´e Garc´ıa Moros Arag´on Institute of Engineering Research (I3A) Communications Technology Group (GTC) PhD Dissertation Biomedical Engineering Doctoral Program Zaragoza, October 2013
A mis padres, Amalia y Santiago
Scientific Contributions Next peer-reviewed scientific publications have been derived from the research of this thesis. Publications in International Journals (JCR indexed) •E. Cavero, A. Alesanco, and J. Garcia, “Enhanced Protocol for Real-Time Transmission of Echocardiograms Over Wireless Channels,” IEEE Transactions on Biomedical Engineering, vol. 59, no. 11, pp. 3212–3220, Nov. 2012. •E. Cavero, A. Alesanco, L. Castro, J. Montoya, I. Lacambra, and J. Garcia, “SPIHT-Based Echocardiogram Compression: Clinical Evaluation and Recommendations of Use,” IEEE Journal of Biomedical and Health Informatics, vol. 17, no. 1, pp. 103–112, Jan. 2013. Submitted Publications in International Journals (JCR indexed) •E. Cavero, A. Alesanco, and J. Garcia, “Real-time Echocardiogram Transmission Protocol based on Regions and Visualization Modes ,” IEEE Journal of Biomedical and Health Informatics. •E. Cavero, A. Alesanco, and J. Garcia, “Low-complexity Ultrasound Image Format Based on Regions,” IEEE Journal of Biomedical and Health Informatics. Preparation Publications in International Journals (JCR indexed) •E. Cavero, A. Alesanco, and J. Garcia, “Real-Time Transmission of Echocardiogram: Display Recommendations and Enhanced Methods for the Transmission based on Visualization Modes,”. Publications in International Conferences and Proceedings •E. Cavero, A. Alesanco, and J. Garcia, “A new approach for echocardiogram compression based on display modes,” in 2010 10th IEEE International Conference on Information Technology and Applications in Biomedicine (ITAB), Nov. 2010. •E. Cavero, A. Alesanco, and J. Garcia, “Real-time transmission of 2d echocardiograms over wimax networks,” in Computing in Cardiology (CinC), 2012, Sept. 2012, pp. 997–1000. 13
•J. Garcia, A. Alesanco, and E. Cavero, “Cardiac signals coding and transmission in realtime mobile telecardiology applications,” in Computing in Cardiology (CinC), 2012, Sept., pp. 317–320. •E. Cavero, A. Alesanco, L. Trajkovic, C. Pattichis, and J. Garcia, “Proposal of real-time echocardiogram transmission based on visualization modes with wimax access,” in Computing in Cardiology (CinC), 2013, Sept. Publications in National Conferences and Proceedings •E. Cavero, A. Alesanco, and J. Garcia, “Nueva propuesta de compresi´on para ecocardiogramas basada en modos de visualizaci´on (A new proposal for echocardiogram compression based on visualization modes),” in CASEIB 2010 - XXVIII Congreso Anual de la Sociedad Espa˜nola de Ingenier´ıa Biom´edica, Nov. 2010. •E. Cavero, A. Alesanco, and J. Garcia, “Nueva propuesta de compresi´on para el almacenamiento de pruebas ecocardiogr´aficas (A new compression proposal for the storage of echocardiography exams),” in CASEIB 2011 - XXIX Congreso Anual de la Sociedad Espa˜nola de Ingenier´ıa Biom´edica, Nov. 2011. •E. Cavero, A. Alesanco, and J. Garcia, “Transmisi´on en tiempo real de Ecocardiogramas 2D sobre redes WiMAX (Real-time transmission of 2-D echocardiograms over WiMAX),” in CASEIB 2012 - XXX Congreso Anual de la Sociedad Espa˜nola de Ingenier´ıa Biom´edica, Nov. 2012.
Contents 1 Introduction 21 1.1 Motivation ......................................... 21 1.2 Tele-Echocardiography ................................... 23 1.2.1 Echocardiography ................................. 23 1.2.2 Challenges ..................................... 25 1.2.3 Working Scenarios ................................. 27 1.3 Thesis Approach and Objectives ............................. 28 1.4 Research Context ...................................... 29 1.5 Thesis Outline ....................................... 30 2 Tele-Echocardiography Realities: Background and What Can Be Improved? 31 2.1 Compression for Storage .................................. 31 2.1.1 DICOM ....................................... 33 2.1.2 JPEG 2000 ..................................... 35 2.2 Compression for Real-time Transmission ......................... 37 2.2.1 SPIHT ........................................ 38 2.3 Transmission Protocols ................................... 43 2.4 Error Control Methods ................................... 45 2.5 Wireless Transmission Technologies ............................ 46 2.5.1 Worldwide Interoperability For Microwave Access (WiMAX) ......... 46 2.6 Clinical Quality ....................................... 48 2.6.1 Mathematical Distortion Indices ......................... 49 2.6.2 Clinical Distortion Indices ............................. 49 2.7 Tele Ultrasound Systems Overview ............................ 50 2.8 Conclusions: Thesis approach ............................... 52 3 Echocardiogram Compression 55 3.1 Echocardiogram Databases and Characteristics ..................... 55 3.1.1 Stored Echocardiograms Database and Characteristics ............. 56 3.1.2 Real-time Echocardiograms Database and Characteristics ........... 57 3.2 Clinical Evaluation Methodology for Compression Recommendations ......... 62 3.2.1 First Phase: Semi-blind Test ........................... 62 3.2.2 Second Phase: Blind Test ............................. 63 15
3.3 Compression for Storage .................................. 67 3.4 Results and Discussion for Compression for Storage ................... 70 3.4.1 Stored Echocardiogram Tool: Format Converter, Display and Measure . . . . 70 3.4.2 Results for Compression of Stored Echocardiograms .............. 74 3.4.3 Discussion for Compression of Stored Echocardiograms ............. 75 3.5 Compression for Real-time Transmission ......................... 77 3.6 Results and Discussion for Compression for Real-Time Transmission ......... 78 3.6.1 Evaluation Setup for Compression Recommendations .............. 79 3.6.2 Results for Compression for Real-Time Transmission .............. 80 3.6.3 Discussion for Compression for Real-Time Transmission ............ 87 3.7 Conclusions ......................................... 89 4 Echocardiogram Transmission in Real-time 93 4.1 Clinical Evaluation Methodology for Display Recommendations ............ 94 4.2 Results and Discussion for Display Recommendations ................. 96 4.2.1 Evaluation Setup for Display Recommendations ................. 96 4.2.2 Results for Display Recommendations ...................... 97 4.2.3 Discussion for Display Recommendations .................... 98 4.3 Protocol for Echocardiogram Transmission in Real-time ................ 99 4.3.1 ETP Overview ...................................100 4.3.2 Control Packets Coding ..............................102 4.3.3 Data Packets Coding ................................107 4.3.4 ETP Working Procedure ..............................108 4.4 Error Control Method for Transmission over Wireless Channels ............111 4.4.1 SECM Working Procedure ............................111 4.4.2 SECM Configuration for ETP ...........................115 4.5 Results and Discussion for Echocardiogram Transmission ...............116 4.5.1 Evaluation Setup for Echocardiogram Transmission ...............116 4.5.2 Results for Echocardiogram Transmission ....................119 4.5.3 Discussion for Echocardiogram Transmission ..................127 4.6 Conclusions .........................................129 5 Conclusions and Future Work 131 5.1 Research Objectives Achieved ...............................131 5.2 Contributions and accomplished results .........................132 5.3 Future Work ........................................134 A Document Type Definition example for ETP 137 Bibliography 149
List of Figures 1.1 Echocardiogram modes. .................................. 23 1.2 Echocardiogram regions. .................................. 24 1.3 Typical scenarios of tele-echocardiography for real-time and store&forward systems. 27 2.1 Tele-echocardiography system structure followed in this Thesis. ............ 32 2.2 Calibration regions for the M mode of an echocardiogram acquired with an Agilent device. ............................................ 35 2.3 Structure of the JPEG 2000 encoder and decoder. ................... 36 2.4 Example of parent-offspring dependencies in the spatial-orientation tree. ....... 39 2.5 Parent-offspring dependency in 3-D SPIHT at the highest level. ............ 41 2.6 Bit stream of two different methods: (a) separate color coding and (b) embedded color coding. ........................................ 42 2.7 TCP/IP layer protocol stack. ............................... 43 3.1 Tele-echocardiography system structure followed in this Thesis: compression part. . 56 3.2 Echocardiogram regions for storage. ........................... 58 3.3 Echocardiogram regions for real-time transmission. ................... 60 3.4 Examples of frames for a sweep mode. .......................... 61 3.5 Semi-blind test: comparison of compressed echocardiogram with the original video. . 62 3.6 Blind test: part I. ...................................... 64 3.7 Blind test: part II. ..................................... 65 3.8 Blind test: part III. ..................................... 66 3.9 DTD file for the configuration of the stored echocardiograms. ............. 68 3.10 XML example for the Philips device. ........................... 69 3.11 Process of converting the stored image into the proposed storage file. ......... 71 3.12 Process of converting the proposed format into another format. ............ 71 3.13 Stored echocardiogram tool screen shots ......................... 73 3.14 Samples of echocardiogram images before and after compression with the proposed method for storage. ..................................... 76 3.15 Images of the B and M modes compressed at different rates with the proposed method. 88 4.1 Tele-echocardiography system structure followed in this Thesis: transmission and visualization parts. ..................................... 94 17
4.2 Semi-blind test: comparison of compressed echocardiogram with the recommended transmission rates. ..................................... 95 4.3 Samples of M mode images for the evaluation of the display recommendations. . . . 97 4.4 Samples of visualized bandwidth distribution for the evaluation of the display recommendations of a 2-D mode video of 1 minute duration. The bandwidth with 20 % of the time with 160 kbps is shown in dark blue. The bandwidth with 5 % of the time with 0 kbps is shown in sky blue. ......................... 98 4.5 Protocol flow diagram. ...................................101 4.6 Confirmation packets. ...................................107 4.7 Echocardiogram regions for real-time transmission. ...................107 4.8 Data packets. ........................................107 4.9 ETP transmitter and receiver flowchart. .........................110 4.10 SECM three states model. .................................112 4.11 SECM encoder for three states model. ..........................112 4.12 SECM decoder for three states model. ..........................113 4.13 ACK packets. ........................................113 4.14 SECM transmitter working procedure. ..........................114 4.15 SECM receiver working procedure. ............................114 4.16 Simulation scenarios for tele-echocardiography with WiMAX access. .........116 4.17 Effective bandwidth over the time without using SECM or ETP for echocardiogram number 3 is drawn in blue and the expected effective bandwidth to visualize the echocardiogram with sufficient clinical quality is drawn in gray. ............122 4.18 Effective bandwidth over the time without using SECM and using ETP for echocardiogram number 7. .....................................123 4.19 Effective bandwidth over the time without using SECM and using ETP for echocardiogram number 9. .....................................124 4.20 Effective and transmitted bandwidth over the time with ETP and SECM for echocardiogram number 7. The effective bandwidth is shown in blue and the transmitted bandwidth in green. ....................................126 4.21 Effective and transmitted bandwidth over the time with ETP and SECM for echocardiogram number 9. The effective bandwidth is shown in blue and the transmitted bandwidth in green. ....................................126
List of Tables 2.1 Some fields of the DICOM header from an Acuson device. ............... 34 2.2 Mobility, data rate and delay of 3G and beyond wireless access technologies. ..... 46 2.3 Ultrasound video transmission systems summary. .................... 51 3.1 Composition of the two types of echocardiogram. .................... 56 3.2 Database for the stored echocardiograms. ........................ 57 3.3 Acquisition devices and cardiac affection for the real-time echocardiograms database. 58 3.4 Database for the real-time echocardiograms. ....................... 59 3.5 Main regions present in the dataset echocardiograms for each device brand and mode. 59 3.6 Data type of the echocardiogram regions for real-time transmission. ......... 62 3.7 Parameters and compression results for the echocardiogram devices and modes of the database. ........................................ 75 3.8 Codec for every data type for compression for real-time transmission. ......... 79 3.9 Semi-blind test: CDI values for the B mode. ....................... 81 3.10 Semi-blind test: CDI values for the color Doppler mode. ................ 82 3.11 Semi-blind test: CDI values for the M mode. ...................... 83 3.12 Semi-blind test: CDI values for the pulsed/continuous Doppler mode. ........ 84 3.13 Selected transmission rates to be evaluated in the blind test. ............. 85 3.14 Blind test: CDI values for the four modes. ........................ 86 3.15 Recommended transmission rates per mode to obtain good clinical quality. ..... 89 3.16 Codecs and recommended transmission rates and bits per pixel for each type of region. 90 4.1 CDI values for the 2-D modes. .............................. 98 4.2 CDI values for the sweep modes. ............................. 99 4.3 Global parameters of the control data. ..........................103 4.4 Common configuration parameters for the regions of the control data. ........103 4.5 Region specific configuration parameters of the control data. ..............104 4.6 XML example for the configuration packets of the Philips Envisor device in the Figure 4.7. ..........................................105 4.7 XML equivalence example for an echocardiogram captured with a Sonosite device. . 106 4.8 Codification parameters expressions that correspond with the encoder in Figure 4.9. 108 4.9 SECM configuration parameters. .............................113 4.10 WiMAX configuration parameters. ............................117 4.11 WiMAX QoS parameters. .................................117 19
4.12 Is diagnosis possible for a 2-D ultrasound mode region with the fixed scenario and correction codes of 40% and 50 %? ............................118 4.13 Transmitted bandwidth in kbps for a 2-D ultrasound mode region with the fixed scenario. ...........................................119 4.14 ETP and SECM parameters for the different regions. S., P. Envisor and P. IE33 are the three available devices. ................................120 4.15 Transmitted bandwidth without using SECM. ......................121 4.16 Percentage of time with guaranteed clinical quality without using SECM. . . . . . . 121 4.17 Number of fragments without guaranteed clinical quality without using SECM. . . . 121 4.18 Transmitted bandwidth for various error control methods and ETP. .........125 4.19 Number of fragments without guaranteed clinical quality for various error control methods and ETP. .....................................127 4.20 Transmitted bandwidth and maximum number of fragments without guaranteed clinical quality for various error control methods and without using ETP. ......127
Acronyms ACK Acknowledge ACR American College of Radiology ADSL Asymmetric Digital Subscriber Line ANOVA Analysis of variance ARQ Automatic Repeat ReQuest ASCII American Standard Code for Information Interchange ASE American Society of Echocardiography AVC Advanced Video Coding BE Best-Effort BS Base Station BWA Broadband Wireless Access CBR Constant Bite-Rate CEZW Color Embedded Zero-tree Wavelet CLD Cross-Layer Design CSPIHT Color Set Partitioning in Hierarchical Trees CT Computed Tomography CVDs Cardiovascular diseases DICOM Digital Imaging and Communications in Medicine DS3 Digital Signal 3 DTD Document Type Definition EAE European Association of Echocardiography ECG Electrocardiogram ETP Echocardiogram Transmission Protocol EZW Embedded Zero-tree Wavelet FDD Frequency Division Duplex FEC Forward Error Correction FMO Flexible Macroblock Ordering 21
GGeneration HEVC High Efficiency Video Coding HL7 Health Level Seven HSUPA High-Speed Uplink Packet Access ICT Information and Communication technologies IEEE Institute of Electrical and Electronics Engineers IMT-Anvanced International Mobile Telecommunications-Advanced IP Internet Protocol ITU-T International Telegraph Union Telecommunication Standardization Sector JPEG Joint Photographic Experts Group JPIP JPEG 2000 Interactive Protocol JVT Joint Video Team KLT KahunenLo`eve Transform LIP List of isolated insignificant pixels LIS List of insignificant sets LSP List of isolated significant pixels NEMA National Electrical Manufacturers Association nrtPS Non-real-time Polling Service MAC Medium Access Control MPEG Moving Picture Experts Group MSE Mean Squared Error OCR Optical Character Recognition OFDM Orthogonal frequency-division multiplexing PACS Picture Archiving and Communications Systems PHY Physical PS Polling Service PSNR Peak Signal-to Noise Ration QoS Quality of Service RCVTP Reliable Clinical Video Transmission Protocol RLE Run Length Encoding ROHC RObust Header Compression ROI Region of Interest RS Reed−Solomon RTP Real-Time Transport Protocol
Chapter 1. Introduction 29 1.2.2 Challenges According to the EAE recommendations [12] the time allocated for a standard transthoracic study should be at least 30 min. A standard study will therefore require at least 60 GB of uncompressed data, the total size depending on the spatial and temporal resolution of each study. For example, a typical echocardiogram with a frame size of 720 x 576 pixels, 24 bits per pixel and 25 fps presents a raw data flow of about 250 Mbit per second.Furthermore, according to the EAE [12] it is recommended to store the studies for at least the duration of the patients’ life expectancy for subsequent analysis and review, although legal requirements affecting the duration of mandatory data storage have to be met. Consequently, compression must be applied for two purposes: reduction of storage requirements and reduction in the transmission rate. Unfortunately, lossless methods can only provide limited compression rates (up to 4:1) [15] which are insufficient. Therefore, lossy compression has to be used since it is able to reduce the data flow considerably, having compression rates up to 20:1 with better quality than the original analog echocardiograms [16] in the sense that noise is reduced. Moreover, in order to compress the echocardiogram efficiently, its visualization characteristics should be taken into account. Nevertheless, it is very important to note that lossy compression modifies the original video and may decrease its quality. The higher the compression rate, the higher the distortion [17,18]. In medical images, a minimum quality is required to be able to make an adequate diagnosis. There is a compromise between compression rate and clinical quality. Thus, the minimum recommended compression rates required for the compressed echocardiogram to achieve adequate clinical quality have to be provided. In order to quantify the clinical distortion introduced in the compressed signal, distortion indexes that quantify the real degradation in the diagnosis content of echocardiograms are necessary. A testbed that unifies and reflects the clinical evaluation procedure should be designed. The integration of the digital echocardiogram replacing traditional video tape recording in echocardiogram machines [19] presented several advantages that were addressed in [20,21]. The main advantage of the introduction of the digital echocardiogram was the possibility of including clinical compression for storage purpose. Clinical compression stores only the information that is important for the diagnosis. For echocardiograms, clinical compression was defined in [12,19,20], and the accuracy of the approximation was demonstrated in [22]. Consequently, instead of storing the whole video, the specialist cardiologist chooses several images of different heart visualizations or several images of the same heart visualization of at least one but preferably three cardiac cycles (from 14 to 64 frames). Clinical compression considerably reduces the storage space, although digital compression is also recommended. It is important to note that clinical compression is not applicable for echocardiogram transmissions in real-time, since it is the specialist cardiologist who makes the clinical compression while watching the transmitted echocardiogram. Thus, the whole echocardiogram has to be transmitted. Another important aspect to be taken into account is that the echocardiogram image may contain text regions with the patient’s personal data (see Figure 1.2). Thus the echocardiogram must be treated with special care since it identifies the patient. Consequently, the protection of these tests during transmission is required by several governmental regulations such as the HIPAA [23]
30 1.2. Tele-Echocardiography in the U.S., the PIPEDA [24] in Canada, the LOPD [25] in Spain and the Digital Signature Laws in several countries. Transmission presents more challenges for wireless than for wired channels because the former are band limited, time varying and error prone. This is a particular problem for medical video streaming applications [26]. However, mobile healthcare (m-health) has undergone an impressive development over recent decades [27,28] thanks to the improvement of wireless technologies. This emerging concept has seen the evolution of e-health systems from traditional desktop telemedicine platforms to wireless and mobile configurations, allowing access to telemedicine services anywhere. Two types of tele-echocardiogram systems can be distinguished: store&forward and real-time systems. The working procedures and main challenges addressed in this Thesis for each system are described below. •The storage&forward systems involve acquiring, encoding, storing and transmitting the echocardiogram at a convenient time for assessment offline, without time deadline requirements. First, while the echocardiogram is performed, the cardiologist selects the relevant images for the diagnosis (clinical compression). After the clinical compression, the acquisition device provides several images of different heart visualizations and a set of images from the same heart visualization (from 14 to 64 frames) for each echocardiogram instead of a video. Although a clinical compression is performed previous to storage, a lossy compression is also required to reduce the storage space and the transmission time, but without losing diagnosis quality. Since there are no time requirements for transmitting the echocardiograms, transmission is not a difficulty. Any reliable protocol can be used for the transmission of the whole echocardiogram without losing quality. Furthermore, store&forward telemedicine can be used with any bandwidth, although there is a trade-off between transmission time and bandwidth. In conclusion, this Thesis is focused on the encoding aspect of store&forward systems, which presents the most significant challenge. •The real-time transmission systems involve transmitting the echocardiogram video at the same time that it is acquired. Medical video streaming is the most demanding application in terms of bandwidth and delay, and hence requires compression for transmission while at the same time maintaining high image quality on reception in order to avoid losing diagnostic information. A bandwidth saving leads to better transmission performance, reducing the transmission time and error introduced by the channel. In order to provide an accurate diagnosis, it is not only necessary to have a compression method that guarantees clinical quality, but it is also essential to be able to guarantee the integrity of the video during the transmission process. It is well known that wireless channels are error prone. Consequently, digital transmission over wireless channels may be affected by erroneous bits that distort the reconstruction of the video on reception. Hence, error control methods are required in order to minimize the delay and errors in transmission. However, even if the echocardiogram is not visualized all the time with minimal acceptable quality, it may be possible to preserve the diagnostic information. An assessment by expert cardiologist is required in order to find out
Chapter 1. Introduction 31 the percentage of time that the echocardiogram can be visualized with lower than minimum quality but without losing diagnostic quality. 1.2.3 Working Scenarios With emerging wireless technologies, patients can access healthcare services not only from hospitals, but also from rural healthcare centers, ambulances, ships, trains, airplanes and homes. The tele-echocardiography scenarios dealt with in this Thesis are shown in Figure 1.3 and described below: Figure 1.3: Typical scenarios of tele-echocardiography for real-time and store&forward systems. •Telediagnosis of patients in remote areas with wireless access (real-time systems). The patient who lives in a remote area does not have to move to the hospital where the expert cardiologist works for the follow-up and early diagnosis of cardiovascular diseases. The echocardiogram is performed in a remote location by a sonographer and the echocardiogram is sent to the expert cardiologist who visualizes the echocardiogram in real-time and makes the diagnosis. •Emergency cases (real-time systems). The transmission of the echocardiogram video begins after the patient is transferred into the ambulance or other conveyance. The basic idea is to communicate ultrasound to an emergency physician, to further assist in the diagnosis and to prepare for the patient’s admission to the hospital. •Medical educational and collaborative evaluations (real-time and store&forward systems). This allows centers delivering cardiology education or collaborative lectures to show echocardiogram acquisition in real-time as well as stored echocardiograms to hospitals and health care centers. These applications provide fundamental principles, firsthand knowledge and evidenced-based methods for critical analysis of established clinical practice standards, and
32 1.3. Thesis Approach and Objectives comparisons with newer advanced alternatives. Various centers collaborate and share their perspectives based on their location, available staff, and available resources. •Storage of examinations (store&forward systems). The studies are stored for subsequent analysis and review or for onward transmission to another center in the event that the patient changes location. 1.3 Thesis Approach and Objectives The general approach of this Thesis is to research and make contributions to the field of ICT applied to the health area. Nowadays, research results in Communication Technologies and the Information Society are considered strategic. Moreover, the application of these technologies to the health area through new telemedicine services facilitates citizens’ access to the health system. Therefore, investigations in this area are highly relevant thanks to the potential benefits for patients, doctors and the health system as a whole. The main aim of this Thesis is to investigate telemedicine systems applied in cardiology environments, since cardiovascular diseases are the leading cause of death in the developing world mainly because they are not diagnosed sufficiently early. Thus, the overall objectives are: •The design, evaluation and recommendations for use of compression methods for storage and real-time transmission of echocardiograms. •The design and evaluation of protocols for transmission of echocardiograms in real-time, and recommendations for the echocardiogram visualization. On a deeper level, further detailed objectives can be mentioned. These are presented as follows, subdivided into the main topics of the Thesis, i.e. compression and transmission. Compression: •To conduct reviews on the state of the art in general aspects of compression methods for both image and video, and clinical quality. Specifically, to review the literature on compression techniques used in wireless medical video transmission systems. •To design a compression image format taking into account the characteristics of stored echocardiograms in order to save storage space while preserving clinical quality. •To design an efficient method to encode the echocardiogram in real-time taking into account its visualization characteristics in order to use less bandwidth while preserving its clinical quality. •To design an accurate echocardiogram evaluation methodology in order to provide a recommendation for echocardiogram compression. •To provide recommendations of use for the proposed compression method for real-time transmission and comparisons with the transmission rates described in the literature.
Chapter 1. Introduction 33 Transmission: •To conduct reviews on the state of the art in general aspects of wireless technologies and transmission protocols. Specifically, to review the literature on wireless medical video transmission systems. •To design a protocol to transmit the echocardiogram in real-time taking into account its characteristics in order to accomplish transmission using less bandwidth while preserving clinical quality. •To design an error control method for echocardiogram transmission that allows adequate diagnosis even over error prone channels. •To provide display recommendations for an accurate diagnosis of echocardiograms. •To study the transmission of echocardiograms in real-time over wireless networks with the designed techniques. To study quality parameters such as transmitted bandwidth and diagnostic validation. 1.4 Research Context This Thesis has been mostly developed within the following funded projects of the Telemedicine research line of the Communications Technologies Group (GTC), belonging to the Arag´on Institute of Engineering Research (I3A) of the University of Zaragoza: 1. DGA - PI029/09: “Analysis of echocardiogram coding and real-time transmission through communications networks”. 2. MCINN - TIN2008-00933/TSI: “New telemonitoring systems for e-Health services”. 3. MCINN - TIN-2011-23792/TSI: “Ontology-based interoperable architecture for patients telemonitoring and clinical decision support”. The research group collaborates with other entities, institutions, universities and hospitals, both national and international. Specially mention can be made of the following international cooperation partners relevant to this Thesis: −The e-Health Laboratory of the Department of Computer Science, of the University of Cyprus, through Professor Constantinos S. Pattichis, Dean of the School of Pure and Applied Sciences. −The Communications Network Laboratory in the School of Engineering Science at Simon Fraser University, Burnaby, British Columbia, Canada through Professor Ljiljana Trajkovic. Two research stages have been done in collaboration with these research groups. At a national level, specific mention can be made of the following clinical collaborations partners:
34 1.5. Thesis Outline −The Cardiology Service of the Lozano Blesa Clinic Hospital (Zaragoza) through Dr. Isaac Lacambra, Lena Castro and Jos´e Montoya. −The Cardiologist Service of the “Hospital viamed Montecanal” (Zaragoza) through Dr. Pedro Serrano. −The Cardiology Service of the Miguel Servet Hospital (Zaragoza) through Dr. Ernest Spitzer. 1.5 Thesis Outline The remainder of this Thesis is structured as follows: •Chapter 2 provides a detailed review of the state of the art of compression of image and video for storage and transmission in real-time purposes, transmission of ultrasound videos, wireless technologies and clinical quality. It also summarizes the most important characteristics of recent ultrasound wireless transmission systems. Finally, improvements are suggested that can be made to existing systems in order to achieve better results in compression, transmission and clinical evaluation. •Chapter 3 describes echocardiogram databases and characteristics for both storage and realtime purposes. It also contributes a clinical evaluation methodology for compression recommendations of echocardiograms, a compression technique for storage purpose and its results compared with other compression techniques, and finally a compression technique for realtime transmission purposes, including recommendations and compression results. The results are compared with the results of previous systems. •Chapter 4 describes an evaluation methodology for the display recommendations of medical images after transmission, specifically for echocardiograms that have been encoded with the method proposed in Chapter 3; display recommendations for the echocardiogram after compression for real-time transmission with the proposed method and recommended transmission rates listed in Chapter 3, a protocol for echocardiogram transmission in real-time; an error control method and its configuration for echocardiogram transmission using the proposed protocol, and finally results and discussion for real-time echocardiogram transmission over wireless channels. •Chapter 5 presents research objectives achieved, contributions and accomplished results of this Thesis, and future lines of research.
Chapter 2 Tele-Echocardiography Realities: Background and What Can Be Improved? The challenging issues in tele-echocardiography systems dealt with in this thesis are the key factors of encoding algorithms, transmission protocols, error control methods, wireless transmission technologies and clinical quality. The tele-echocardiography system scheme followed in this Thesis is depicted in Figure 2.1. The echocardiogram is acquired and can then be stored or transmitted in real-time. For both purposes, the echocardiogram, in video or image format, has to be compressed with the maximum compression rate while preserving the clinical quality. However, the data acquired for each purpose is different and therefore the compression methods must also be different. In the case of transmitting the echocardiogram, the protocols have to be carefully chosen in order to guarantee reception of the echocardiogram with minimal delay and without losing diagnostic information even if not all the information gets to the receiver in time. Furthermore, the wireless channel has to comply with the system requirements, acceptable delay and transmission rates. Sections 2.1,2.2,2.3,2.4,2.5 and 2.6 address the above-mentioned key factors in terms of the state of the art, improvements that can be made to achieve better results and the most relevant technologies that have been used in this Thesis. Finally, the most important characteristics of recent ultrasound wireless transmission systems are summarized in Section 2.7 and conclusions relating to the design premises used in this Thesis to improve tele-echocardiography systems are shown in Section 2.8. 2.1 Compression for Storage Hospitals and medical centers produce large volume of digital medical images that require considerable storage space. Standardization of storage format is critical to enable interoperability within and between centers and equipment from different vendors. A survey of the different medical image formats and compression techniques is presented in [29]. Digital Imaging and Communications in Medicine (DICOM) [30] is the most widely used and accepted standard for effective medical 35
36 2.1. Compression for Storage Acquisition & visualization without errors Compression for Storage Clinical quality Compression for real time Transmission Real-time transmission •Error control •Protocols Wireless channel raw video images Real-time reception •Error control Decoding Visualization with errors & storage Clinical quality video Figure 2.1: Tele-echocardiography system structure followed in this Thesis. imaging storage and transfer over large geographical areas, providing the basis for picture archiving and communications systems (PACS). Moreover, DICOM is the most extensively used standard for storage purpose, recommended by the EAE [12], the ASE [20] and the American College of Cardiology (ACR) [31]. It is also included in the majority of medical image devices [31–33]. For this reason, it is very important to design a compression format interoperable with DICOM and to take into account the DICOM characteristics, described below in Section 2.1.1. In order to enhance medical image compression while preserving diagnostic information, the concept of Region of Interest (ROI) has been adopted and widely used [34–38]. The challenge of this technique is to perform the image segmentation. In some medical image modalities an algorithm has been designed to obtain the regions automatically. However, obtaining the regions automatically is not always possible, and for almost all medical image modalities the regions of interest for each individual image must be defined by the cardiologist. In order to easily extract the echocardiogram regions, the facilities included in the acquisition devices to segment the images must be taken into account. The devices form the echocardiogram with the different regions: ultrasound image, ECG, auxiliary image and text. A proof that the ultrasound devices incorporate the division of the regions is the calibration DICOM header, which is described in the next Section 2.1.1. In the calibration header different regions are defined with different calibrations that correspond to the ultrasound image region and the ECG. The ROI regions are the ultrasound image and the ECG. The whole ultrasound image has to be selected as ROI. If a small ultrasound part is selected instead, the clinical quality is affected by the degradation quality of the non ROI part [39]. As regards the type of data, two types of region can be distinguished, image and text, since it is not efficient to compress the text as an image. In conclusion, better compression performances can be achieved if the facilities that the ultrasound devices incorporate to divide the image into regions are used to encode each region separately and take into account the data type of each region as well as its clinical importance. Image compression methods that use wavelet transforms have been successful in providing high compression ratios while maintaining good image quality: Joint Photographic Experts Group (JPEG) 2000 [40] and Set Partitioning in Hierarchical Trees (SPIHT) [41]. However, the DI-
Chapter 2. Tele-Echocardiography Realities: Background and What Can Be Improved? 37 COM standard includes the JPEG 2000 standard, but not SPIHT. Moreover, JPEG 2000 has been demonstrated to achieve very good results in the compression of medical images [42,43]. It is shown in [43,44] that with a compression rate of 1 bit per pixel (bpp) the diagnostic information is preserved for computerized radiography and ultrasound images, respectively. Consequently, JPEG 2000 can be used for medical image compression providing adequate clinical quality with 1 bpp. The main JPEG 2000 characteristics are described in Section 2.1.2. 2.1.1 DICOM 2.1.1.1 DICOM overview In response to the increased use of digital images in radiology ACR and the National Electrical Manufacturers Association (NEMA) formed a joint committee in 1983 to create a standard format for storing and transmitting medical images [30]. The committee published the original ACRNEMA standard in 1985. This has subsequently been revised and in 1993 the standard was renamed DICOM.DICOM is administered by the NEMA Diagnostic Imaging and Therapy Systems division and each year the standard is updated. Details of recent improvements can be found on [30]. The standard describes how to format and exchange medical images and associated information, both within the hospital and also outside the hospital. DICOM interfaces are available for connection of any combination of the following categories of digital imaging devices: (a) image acquisition equipment such as computed tomography, magnetic resonance imaging, computed radiography, ultrasonography, and nuclear medicine scanners; (b) image archives; (c) image processing devices and image display workstations; (d) hard-copy output devices such as photographic transparency film and paper printers. DICOM addresses five general application areas: 1. Network image management. 2. Network image interpretation management. 3. Network print management. 4. Imaging procedure management. 5. Off-line storage media management. DICOM is a message standard that facilitates interoperability of medical imaging equipment by specifying: 1. For network communications, a set of protocols to be followed by devices claiming conformance to the standard. 2. The syntax and semantics of Commands and associated information which can be exchanged using these protocols. 3. For media communication, a set of media storage services to be followed by devices claiming conformance to the standard, as well as a File Format and a medical directory structure to facilitate access to the images and related information stored on interchange media.
38 2.1. Compression for Storage 2.1.1.2 DICOM File Format A single DICOM file contains both a header (which stores information about the patient’s name, the type of scan, image dimensions, etc), as well as all of the image data. The header and the image data are stored in the same file. The image data follows the header information. Table 2.1: Some fields of the DICOM header from an Acuson device. Field Contents Filename [1x65 char] FileModDate “12-nov-2010” FileSize 2361370 Format “DICOM” FormatVersion 3 Width 1024 Height 768 BitDepth 8 ColorType “truecolor” FileMetaInformationGroupLength 204 MediaStorageSOPClassUID “1.2.840.10008.5.1.4.1.1.6.1” TransferSyntaxUID “1.2.840.10008.1.2.1” ImplementationClassUID “1.2.276.0.7230010.3.0.3.5.4” Modality “US” Manufacturer “SIEMENS” InstitutionName “HC LOZANO BLESA” ManufacturerModelName “ACUSON SC2000” PatientName [1x1 struct] PatientID “XXXXXXXXXXXX” PatientBirthDate “XX” PatientSex “X” HeartRate 88 SequenceOfUltrasoundRegions [1x1 struct] The size of the header varies depending on the acquisition device and image type. The DICOM elements required depend on the image type that are listed in Part 3 of the DICOM standard [45]. DICOM requires a 128-byte preamble (these 128 bytes are usually all set to zero), followed by the letters ’D’, ’I’, ’C’, ’M’. This is followed by the header information, which is organized in groups: general information, patient, study, series, frame of reference, equipment and image information. In Table 2.1 some fields of a header for an ultrasound device are shown. Of particular importance is the “Transfer Syntax Unique Identification” which reports the structure of the image data, revealing whether the data has been compressed or not. Another important field in the DICOM header included in the ultrasound is the regions calibration, see “SequenceOfUltrasoundRegions” in Table 2.1. It defines regions on the ultrasound image with different calibration and the calibration parameters in order to be able to perform measurements on the ultrasound regions. The calibration
Chapter 2. Tele-Echocardiography Realities: Background and What Can Be Improved? 45 Figure 2.5: Parent-offspring dependency in 3-D SPIHT at the highest level. single file for a video sequence can provide progressive video quality, i.e., the algorithm can be stopped at any compressed file size or let run until nearly lossless reconstruction is obtained. In 3-D SPIHT, sorting of pixels proceeds just as it would with 2-D SPIHT, the only difference being 3-D rather than 2-D tree sets. Once the sorting is done, the refinement stage of 3-D SPIHT will be exactly the same. On the 3-D subband structure, we define a new 3-D spatiotemporal orientation tree and its parent-offspring relationships. When the spatial and temporal filtering alternate, so that the decomposition is purely dyadic, a straightforward extension from the 2-D case is to form a node in 3-D SPIHT as a block with eight adjacent pixels, two extending to each of the three dimensions, hence forming a node of 2 x 2 x 2 pixels. The root nodes (at the highest level of the pyramid) have one pixel with no descendants and the other seven pointing to eight offspring in a 2 x 2 x 2 cube at corresponding locations at the same level. For nonroot and nonleaf nodes, a pixel has eight offspring in a 2 x 2 x 2 cube one level below in the pyramid. Figure 2.5 depicts these parent-offspring relationships in the case of a two-level dyadic 3-D decomposition with 15 subbands, produced by a once repeated spatial-horizontal, spatial-vertical, and temporal splitting, in that order. The number of temporal subband decomposition directly depends on the number of frames to be compressed at a time. The typical value for the temporal resolution is of sixteen frames, although a shorter filter with four or eight frames, such as the Haar or S+P filters, can be also used. Only two or three decompositions are possible with four or eight frames compressed at a time, respectively. However, with a spatial resolution of 352 x 288 pixels, for example, five spatial (dyadic) decompositions can be achieved with the high performance 9/7 biorthogonal filter [77]. The next step is compression of the coefficients into a bit stream. Essentially, it can be done by feeding the 3-D data structure to the 3-D SPIHT coding kernel. The 3-D SPIHT kernel will sort
46 2.2. Compression for Real-time Transmission the data according to the magnitude along the spatio-temporal orientation trees (sorting pass), and refine the bit plane by adding necessary bits (refinement pass). At the destination, the decoder will follow the same execution path conveyed by the received significance decision bits to recover the data. YYYYYYYYYYYYYYYY UUUU UUUU (a) YYYYYYYYYYYYYYYY UUUUUUUUUUUUUUUU VVVVVVVVVVVVVVVV (b) Figure 2.6: Bit stream of two different methods: (a) separate color coding and (b) embedded color coding. A simple application of the SPIHT to color video would be to code each color plane separately, as does a conventional color video coder. Then, the generated bit stream of each plane would be serially concatenated. However, this simple method would require allocation of bits among color components, losing precise rate control, and would fail to meet the requirement of the full embeddedness of the video codec since the decoder needs to wait until the full bit stream arrives to reconstruct and display. Instead, one can treat all color planes as one unit at the coding stage, and generate one mixed bit stream so that we can stop at any point of the bit stream and reconstruct the color video of the best quality at the given bit rate. In addition, we want the algorithm to automatically allocate bits optimally among the color planes. By doing so, we will still keep the claimed full embeddedness and precise rate control of 3-D SPIHT. The bit streams generated by both methods are depicted in the Figure 2.6, where the first one shows a conventional color bit stream, while the second shows how the color embedded bit stream is generated, from which it is clear that we can stop at any point of the bit stream, and can still reconstruct a color video at that bit rate as opposed to the first case. Let us consider a tri-stimulus color space with luminance Y plane such as YUV or YCrCb. Each such color plane will be separately wavelet transformed, having its own pyramid structure. Now, to code all color planes together, the 3-D SPIHT algorithm will initialize the LIP and LIS with the appropriate coordinates of the top level in all three planes. Since each color plane has its own spatial orientation trees, which are mutually exclusive and exhaustive among the color planes, it automatically assigns the bits among the planes according to the significance of the magnitudes of their own coordinates. The effect of the order in which the root pixels of each color plane are initialized will be negligible except when coding at extremely low bit rate. Note also that the wavelet transforms and sizes may be different among the three planes without affecting the method. In conclusion, 3-D SPIHT is an embedded subband-based video coder analogous to the 2-D spatial orientation trees in image coding. The video coder is fully embedded, so that different degrees of monochrome or color video quality can thus be obtained with a single compressed bit stream. The cost for this embeddedness is the coding delay (latency) to accept 16 frames into a buffer and a memory size of the order of the size of the coding unit to execute the 3-D SPIHT
Chapter 2. Tele-Echocardiography Realities: Background and What Can Be Improved? 47 algorithm. Precise rate control and self-adjusting rate allocations are automatically achieved. In addition, spatial and temporal scalability can be easily incorporated into the system to meet various types of display parameters requirements. 2.3 Transmission Protocols In order to transmit the coded information over Internet networks, Transmission Control Protocol (TCP)/Internet Protocol (IP) is the suite of communications protocols used to connect hosts on the Internet. TCP/IP provides end-to-end connectivity specifying how data should be formatted, addressed, transmitted, routed and received at the destination. It has four abstraction layers (see Figure 2.7) which are used to sort all related protocols according to the scope of networking involved: Application Transport Internet Link Figure 2.7: TCP/IP layer protocol stack. •The link layer contains communication technologies for a single network segment of a local area network. The link layer depends on the physical medium used for the transmission. •The Internet layer has the responsibility of sending packets across potentially multiple networks. Internet working requires sending data from the source network to the destination network. This process is called routing. In the Internet protocol suite, the Internet Protocol performs two basic functions: –Host addressing and identification: this is accomplished with a hierarchical IP addressing system. –Packet routing: this is the basic task of sending packets of data (datagrams) from source to destination by forwarding them to the next network router closer to the final destination. •The transport layer protocol establishes host-to-host connectivity, meaning it handles the details of data transmission that are independent of the structure of user data and the logistics of exchanging information for any particular specific purpose. Its responsibility includes end-to-end message transfer independent of the underlying network, along with error control, segmentation, flow control, congestion control, and application addressing (port numbers). End to end message transmission or connecting applications at the transport layer can be categorized as either connection-oriented, implemented in TCP, or connectionless, implemented in User Datagram Protocol (UDP).
48 2.3. Transmission Protocols •The application layer contains the higher-level protocols used by most applications for network communication. Data coded according to application layer protocols are then encapsulated into one or (occasionally) more transport layer protocols, which in turn use lower layer protocols to effect actual data transfer. Transport protocols are further distinguished in upper and lower layer, Real-Time Transport Protocol (RTP), and UDP/TCP respectively. TCP [78] provides reliable, ordered, error-checked to secure packet delivery to destination. These characteristics may lead to long delay and jitter, specially when high data flows are transmitted. Thus, TCP is highly adequate for applications where it is important to guarantee the reliability while the delay is less important. On the other hand, UDP [79] does not provide any error handling or congestion control mechanisms, allowing therefore packets to drop out. This properties made UDP highly adequate for applications where the packets arrive in time is more important than the integrity of the information. Therefore, UDP is the primarily established as the lower layer transport protocol for real-time transmission of video. However, as a minimal clinical quality is necessary to achieve an adequate diagnostic error control methods should be implemented in the upper layer. RTP [80] is widely used in clinical video steaming applications [49,50,57]. RTP provides endto-end delivery services for real-time video and audio transmission. RTP itself does not contain any mechanisms to ensure on time delivery. On the contrary, it relies on UDP or TCP for doing so. However, it does provide the appropriate functionality for carrying real-time content such as time-stamping and control mechanisms that enable synchronization of different streams with timing properties. Since the echocardiogram is composed of several regions with different type of data (text, image, video, audio) and different clinical importance and consequently every region can be compressed with different coding methods, different transmission methods can be applied for each region. RTP is not suitable for this proposal, since it does not provide delivery of text or images. Given the aforementioned, a protocol for end-to-end real-time transmission of echocardiograms over IP can be designed introducing different coding and transmission protocols for every region, UDP or TCP and further control error methods. The most suitable protocols will depend on the type of region and its clinical information. For example, in order to transmit a region with the ultrasound video, UDP is the adequate protocol introducing additionally an error control method in the application layer. However, in order to transmit the text present in the echocardiogram, TCP is more adequate, because the reliability of the transmission is required. In order to decrease header overheads, reduce packet loss and increase security over noisy wireless links, RObust Header Compression (ROHC) can be used for both UDP [81] and TCP [82] transport layers. This standard compresses IP and UDP headers to just 3 bytes, including the checksum field to discard erroneous packets, and the IP and TCP headers to just 10 bytes. Thanks to this standard, the redundancy introduced in the transmission is decreased, allowing small packets to be transmitted. This reduces packet loss without an excessive increase in the number of transmitted bits. It is important to take into account that some text regions may contain confidential information about the patient. For this reason, the text has to be protected so that it can only be accessed by authorized sanitary staff. An easy way to protect this information is by protecting all the
Chapter 2. Tele-Echocardiography Realities: Background and What Can Be Improved? 49 packets in which the information is contained. Typically, the text information is not protected in the clinical video streaming systems since the text is included in the video. Transport Layer Security (TLS) [83] is the most common choice for secure communications and is included in the medical standards DICOM and Health Level Seven (HL7). The primary advantage of TLS is that it provides a transparent connection-oriented channel. Thus, it is easy to secure an application protocol by inserting TLS between the application layer and the transport layer. 2.4 Error Control Methods In order to provide an accurate diagnosis, it is not only necessary to have a compression method that guarantees clinical quality, but it is also essential to be able to guarantee the integrity of the video during the transmission process. It is well known that wireless channels are error prone. Thus, digital transmission may be affected by erroneous bits that distort the reconstruction of the video on reception. Hence, the use of error control mechanisms for maintaining acceptable video levels in wireless communications channels are required [84]. However, for real-time applications dealing with multimedia data, reliable methods such as the incorporated in TCP are not recommended due to the resulting delay. Error control methods can be used in the application layer for the ROI regions in order to minimize distortion, but without introducing excessive delay. The two main error control techniques are Automatic Repeat ReQuest (ARQ) [85–87] and Forward Error Correction (FEC) [88–92]. Both techniques increase the original amount of bits and therefore there is a trade off between compression fidelity and protection. With retransmission techniques, only missing packets are retransmitted. If the channel delay is long or if several retransmissions of the same packet are required because of channel errors, the resulting delay would be intolerable for real-time applications. With FEC, redundant bits are added in the transmitted data and the decoder uses these added bits to correct the errors. The amount of redundancy embedded can be more than is necessary to correct the errors, using more bits than with the retransmission mechanism. However, if the amount of redundancy is less than is necessary, no error can be corrected. But this technique does not need a feedback channel and reduces the time needed to recover missing packets. The design and performance of a hybrid ARQ with concatenated FEC for real-time video streaming over wireless networks has been addressed in [93,94]. In [93], an adaptive technique in the Medium Access Control (MAC) layer of Worldwide Interoperability For Microwave Access (WiMAX) was used, but this technique is only valid for the WiMAX channel. In [94] a more general solution was proposed. Furthermore, the channel conditions can be taken into account to adapt the error control techniques to the channel conditions and use one or another techniques. For telemedicine applications, error resilient implementations for robust diagnosis performance have been addressed, for example, in [34,49–52,57,95]. A scalable video coding (SCV) employing spatiotemporal scalability is found in [34,57]. A flexible macroblock ordering (FMO) technique for variable quality slice encoding and redundant slices for resilience over error prone mediums, available in the baseline profile of H.264/AVC, were used in [50,52]. In [49] a ROI-based and prediction-based unequal channel error protection implementation to overcome transmission errors was presented. Another different approach is an optimized cross-layer design (CLD) based on a reinforcement learning algorithm for real-time medical video streaming [51,95]. In [51] an error
50 2.5. Wireless Transmission Technologies Table 2.2: Mobility, data rate and delay of 3G and beyond wireless access technologies. Wireless technology Max. speed Data rate1delay 3G UMTS [96] 300 km/h 220 - 384 Kbps <250 ms HSPA [97,98] 300 km/h 500 kbps - 2 Mbps <250 ms 3.5G HSPA+ [99,100] 300 km/h 1 - 4 Mbps <100 ms LTE [100,101] 500 km/h 1.5 - 5.8 Mbps <70 ms 4G LTE-Advanced [102] 500 km/h up to 100 Mbps <70 ms IEEE 802.16e [103] 120 Km/h up to 5.6 Mbps <70 ms WiMAX IEEE 802.16m [104] 350 Km/h up to 100 Mbps <70 ms 1Uplink data rates concealment technique for the ROI region was used. This last technique is not suitable for real time applications. However, a hybrid ARQ with a concatenated FEC method in which the channel conditions are taken into account has not been used in any of the reviewed telemedicine applications for clinical video streaming. 2.5 Wireless Transmission Technologies The development of more demanding telemedicine systems has been possible thanks to the evolution of mobile telecommunication systems. The evolution of mobile systems from 2 Generation (G) to 2.5Gand afterwards to 3Gfacilitates the provision of higher data rates and lower delays that enable the development of more responsive telemedicine systems. Furthermore, the later wireless technologies, 3.5Gand beyond, enable the transmission of high quality video resolution that requires more bandwidth. Before starting to design a telemedicine system, it is important to know the application requirements such as required bandwidth, delay, and speed in order to choose the appropriate access technology. Table 2.2 shows the main characteristics of the access technologies for 3Gand beyond, including maximum speed, typical up link data rates and typical delays. One of the most frequently used technologies in clinical video streaming applications is WiMAX [49,51,52,95] due to its improved uplink and downlink rates, increased coverage and throughput, mobility support enhancement, latencies reduction, quality of service (QoS) services and security enhancement. 2.5.1 Worldwide Interoperability For Microwave Access (WiMAX) The Institute of Electrical and Electronics Engineers (IEEE) 802.16 standard offers broadband wireless access (BWA) over long distance. WiMAX was firstly standardized for fixed wireless access by the IEEE 802.16-2004 [105] and then for mobile access with by the IEEE 802.16e [103] and 802.16m [104] standards. WiMAX is a promising technology to provide wireless services requiring high-rate transmission and strict QoS requirements in both indoor and outdoor environments. A
Chapter 2. Tele-Echocardiography Realities: Background and What Can Be Improved? 51 thorough overview of WiMAX standardization process and evolving concepts and technologies up to IEEE 802.16e standards appears in [106], while recent advances are described in detail in [107–109]. To provide flexibility for different applications, the standard supports two major deployment scenarios: •Last-mile BWA: In this scenario, broadband wireless connectivity is provided to home and business users in a wireless metropolitan area networks (WMAN) environment. The operation is based on a point-to-multipoint single hop transmission between a single base station (BS) and multiple subscriber stations (SSs). •Backhaul networks: This is a multihop (or mesh) scenario where a WiMAX network works as a backhaul for cellular networks to transport data/voice traffic from the cellular edge to the core network (Internet) through meshing among IEEE 802.16/WiMAX SSs. The WiMAX Forum, which is a nonprofit organization, encourages and supports IEEE 802.16based BWA. The main role of the WiMAX Forum is to standardize and maintain the process of testing and the certification program for compatibility and interoperability of IEEE 802.16 equipment [110]. The physical (PHY) and MAC layer protocols are well defined in the IEEE 802.16 standard, efficient radio resource management is still an open issue. Features of PHY and MAC layers are discussed next. 2.5.1.1 Physical Layer Features The physical layer of the IEEE 802.16 air interface originally operates at either the 1066 GHz (IEEE 802.16), 211 GHz band (IEEE 802.16a) [111]. Todays licensed deployment is typically in the range of 2.3, 2.5-2.7, 3.5, and 5.8 GHz, while 4Gfrequency bands will facilitate deployment between 450-3600 MHz [104]. Channel bandwidth allows great flexibility in the sense that it allows WiMAX operators to consider channel bandwidths between 1.25, 2.5, 5, 10, and 20 MHz (802.16e). In 802.16m scalable bandwidth between 5-40 MHz for a single radio frecuency carrier is considered, extended to 100 MHz with carrier aggregation to meet International Mobile TelecommunicationsAdvanced (IMT-Advanced) requirements. WiMAX employs a set of high and low level technologies to provide robust performance in both line-of-sight and non-line-of-site conditions. The primary features of the physical layer include adaptive modulation and coding (QPSK, 16-QAM, 64-QAM), hybrid ARQ, and fast channel feedback. WiMAX uses scalable orthogonal frequency division multiple access (SOFDMA) that divides the transmission bandwidth into multiple subcarriers. The number of subcarriers ranges from 128 for 1.25 MHz channel bandwidth and extends up to 2048 for 20 MHz channels. In this manner, dynamic QoS can be tailored to an individual application requirements. In addition, orthogonality among subcarriers allows overlapping leading to flat fading. In other words, multipath interference is addressed by employing orthogonal frequency-division multiplexing (OFDM) while available bandwidth can be split and assigned to several requested parallel applications for improved systems efficiency. The latter is true for both downlink and uplink. A Multiple input multiple output antenna system improves communication performance, including significant increases in data throughput and link range, without additional bandwidth or increased transmit power.
52 2.6. Clinical Quality 2.5.1.2 Medium Access Control Layer Features IEEE 802.16/WiMAX uses a connection-oriented MAC protocol, which provides a mechanism for the SSs to request bandwidth from the BS. Although each SS has a standard 48-bit MAC address, the main purpose of this address is for hardware identification. Therefore, a 16-bit connection identifier is used primarily to identify each connection to the BS.IEEE 802.16/WiMAX supports both frequency division duplex (FDD) and time-division duplex (TDD) transmission modes. The QoS scheduling is the most important feature in WiMAX systems, which makes it an ideal choice for QoS sensitive applications such as video content streaming. There are three major types of services supported with different QoS requirements: •Unsolicited grant service (UGS): This service supports CBR traffic. In this case the BS allocates a fixed amount of bandwidth to each of the connections in a static manner. UGS service is suitable for traffic with very strict QoS constraints for which delay and loss need to be minimized. A typical application is VoIP. •Polling service (PS): This service supports traffic for which some level of QoS guarantee is required. The amount of bandwidth required for this type of service is determined dynamically based on the required QoS performance and the dynamic traffic arrivals for the corresponding connections. It can be divided into two subtypes: –Real-time polling service (rtPS). This service is delay sensitive. Typical applications are audio and video streaming. –Non-real time polling service (nrtPS). This service can guarantee a certain throughput guarantee. A typical application is file transfer. •Best-Effort (BE) Service: This is for traffic with no QoS guarantee. The amount of bandwidth allocated to BE service depends on the bandwidth allocation policies for the other two types of service. In particular, the bandwidth left after serving UGS and PS traffic is allocated to BE service. Typical applications are web and email traffic. Mobility management is also address in 802.16e and current 802.16m standards. Established connections can move with speeds between 50-100 km/h for 802.11e and up to 350 km/h for 802.11m with adequate performance. 2.6 Clinical Quality The image quality in telemedicine systems is decreased by lossy compression and due to errors introduced by the network. In medical images, a minimum acceptable quality is required to be able to make an adequate diagnosis. In order to quantify the clinical distortion introduced in the transmitted signal, distortion indices should be used. There are two types of indices: objective or mathematical distortion indices and subjective or clinical distortion indices.
Chapter 2. Tele-Echocardiography Realities: Background and What Can Be Improved? 53 2.6.1 Mathematical Distortion Indices Mathematical distortion indices [112], thanks to their ease of use, can be useful for preliminary and fast measurement of quality. Some examples of these indices are: Mean Squared Error (MSE), Peak Signal-to Noise Ratio (PSNR) and SIMilarity Index [113]. Currently, the most commonly used objective image and video distortion metric even for medical images is PSNR [49,50,57]. PSNR is widely used because it is simple to calculate, has clear physical meanings, and is mathematically easy to deal with for optimization purposes. PSNR measured in decibels (dB) is given by PSNR = 10 ·log10 2552 MSE (2.1) where 255 is the maximum possible value that can be attained by the image signal of 8 bits. MSE is defined as MSE =1 M·N· M X m=1 N X n=1 |I(x, y)−J(x, y)|2(2.2) where M·Nis the frame dimension in pixels. I(x, y) is the source frame and J(x, y) is the compressed frame. However, PSNR has been widely criticized for not correlating well with perceived quality measurement [114,115] and for not being able to quantify the real clinical distortion [56]. 2.6.2 Clinical Distortion Indices New indices that include degradation in the diagnosis content of echocardiograms are necessary. There are two main types of test that can be included to calculate a clinical distortion index: •Semi-blind test: comparison of the compressed images with the original, without knowing the transmission rate used for the compressed images. Can the same diagnosis be made with both images? •Blind-test: evaluation of the compressed images without knowing the transmission rate used. In general, for medical images this evaluation consists of performing a complete diagnosis and comparing it to the results of the original images. There are several works in which medical opinion has been utilized to assess real clinical quality after codification [48,50,56] and transmission [17,18,48,50,57,116]. However, these assessments were not as accurate as the evaluation in [56]. In [56], a testbed that unifies and reflects clinical evaluation for echocardiogram procedures was developed using both a blind and a semi-blind test. This testbed gives a precise evaluation of the real degradation in compressed echocardiograms and provides a minimum recommended transmission rate required for each echocardiographic mode to achieve adequate clinical quality. However, it has a serious disadvantage in that it is burdensome and very costly in terms of time dedicated to evaluating the echocardiograms. Furthermore, if a guarantee of clinical quality for the visualized echocardiogram is desired, it will be necessary to evaluate a wide range of transmission rates and network scenarios with different channel conditions. In order to save time in the evaluation process, two different evaluations can be performed. In the first evaluation, a wide range of transmission rates can be evaluated with a method as accurate as
54 2.7. Tele Ultrasound Systems Overview that presented in [56], but saving time in the evaluation method. In order to achieve this, a two phase testbed can be designed. In the first phase, a fast and simple evaluation can be made using a semi-blind test for different transmission rates in order to select two transmission rates for each mode that are at the limits of acceptable clinical quality. These two rates can be evaluated in more detail in the second phase using a blind test in order to provide the recommended transmission rates. If the transmission occurs without errors and the whole echocardiogram is displayed at the recommended transmission rate, the same diagnosis as that of the original echocardiogram is guaranteed. However, if the recommended transmission rate is not received all the time, the same diagnoses may or may not be possible. Instead of assessing different channel conditions, as has been described in the literature, an evaluation in which the percentage of time that the echocardiogram is displayed with a lower transmission rate than the recommended rate can be evaluated. In this manner, it is possible to know for any channel if the diagnosis on reception will be possible with a single evaluation. As the starting point of the evaluation is the recommended transmission rate in the previous evaluation, for the second evaluation it may be sufficient to conduct a semi-blind test for different percentages of time that the echocardiogram is displayed with a transmission rate inferior to the minimum. 2.7 Tele Ultrasound Systems Overview Overviews of m-health systems and more recent wireless medical video transmission systems were addressed in [26,28], respectively. This section discusses the most relevant studies dealing with ultrasound video transmission over wireless channels, the main characteristics of which have been addressed throughout this chapter. Table 2.3 shows the most important parameters for the reviewed systems in order to compare the actual systems to each other and with respect to the system proposed in this Thesis. The most important parameters in a clinical video transmission over a wireless channel are the spatial and temporal (frame rate) resolution, the encoding and error resilient method, the transmission rate, clinical evaluation and the access channel. The video resolution affects the diagnostic capacity of the video and the transmission rate. The higher the resolution, the higher the quality and the transmission rate [52]. The resolution can be classified as low video resolution (lower than 480x480 and 15 fps) and original/high video resolution (equal to or higher than 480x480 and 15 fps). In Table 2.3, the systems are divided into low (light gray) and high (dark gray) resolution. Higher video resolution requires higher transmission rates, and consequently wireless technologies with high data transfer rates, 3.5Gand beyond. Other factors that affect the transmission rate are the encoding and error resilient methods used, thus it is very important to choose these methods correctly. The transmission rate has to be the minimum possible but without compromising the clinical quality of the video. For this reason, it is very important to perform an evaluation that defines the minimum transmission rate for the selected parameters in order to guarantee clinical quality, as in [56], where the lowest transmission rate is required for high resolution videos, see Table 2.3.
Chapter 3. Echocardiogram Compression 61 different cardiologist and from different patients. Table 3.2 shows the echocardiogram devices, the number of images available for each device, the typical image size of each device and the mean DICOM header size for each device. These devices are representative of the typical image distribution in echocardiogram images. The echocardiograms acquired with the Agilent and Siemens Acuson devices can be considered as one because both have a similar size and similar characteristics. Table 3.2: Database for the stored echocardiograms. Echocardiogram devices, number of echocardiogram images available, typical image size and mean DICOM header size in bytes of all the exams for each device. Device Number Typical DICOM Header Size of images Image Size (bytes) Agilent Sonos 22 600x430 10365 Siemens Acuson 23 576x456 19132 Philips Envisor 28 800x564 19132 Siemens CS2000 32 1024x768 19132 3.1.1.2 Stored Echocardiogram Characteristics There are three types of regions, as we can see in Figure 3.2: •Ultrasound: it is the most important because it contains the relevant information for the diagnosis. The ultrasound is always present and only appears once. In the case of having a video loop, it is the only region that changes over the time. •Auxiliary images: they surround and complement the ultrasound region. These auxiliary images are, for example, the ECG, the color label, other ultrasound images to complement the information of the main ultrasound image and some symbols regarding the configuration. •Text: it is always present in all the images and contains information such as patient information, date, time, configuration or measurements of the study. In the case of having a video loop, the only region that change over time is the ultrasound region, while the rest of the regions remain invariant. For each echocardiogram mode and device, the distribution and the size of the regions, the number of auxiliary regions and the text are different. Furthermore, some image regions contain color information that is relevant for the diagnosis and others not. For example, the Doppler modes have relevant information in color in the ultrasound image and in the color scale. 3.1.2 Real-time Echocardiograms Database and Characteristics 3.1.2.1 Database for Real-Time Echocardiograms Three cardiologists experienced in echocardiography recorded and stored original echocardiograms from patients with different diagnoses and with three different ultrasound devices (see Table
62 3.1. Echocardiogram Databases and Characteristics Figure 3.2: Echocardiogram regions of a Doppler mode (on the left) and a continuous Doppler mode (on the right) captured with an Agilent and Acuson devices respectively. The white solid line contains the ultrasound, the green dotted line the auxiliary images and the yellow dashed line the text. Table 3.3: Acquisition devices and cardiac affection for the real-time echocardiograms database. Device Patient Patient’s Diagnosis 1 Lateral myocardial infarction 2 Medial septal myocardial infarction Sonosite SonoHeart Elite 3 Ventricular septal defect 4 Atrial septal defect 5 Right ventricle hypertrophy and pulmonary insufficiency Philips ENVISOR C HD 6 Mitral and tricuspid regurgitation 7 Normal 8 Tricuspid regurgitation Philips IE33 9 Pulmonary insufficiency 3.3). Three sessions per device, each session corresponding with one patient, were recorded having a total of 200 videos. Each device had its specific characteristics and different qualities. First, a portable device was chosen (see SonoSite, Table 3.3) because the image quality is inferior and the manner of representation is quite different to that of non-portable devices. The other two chosen devices were of the same brand, but their quality and the form of representation were different. The image quality of both was superior to that of portable device. The selected echocardiograms were representative of typical and abnormal findings in the cardiovascular field. The physical conditions of the patients were diverse (large, medium, and slim build), as well as their ages (baby, young, middle-aged, and old). The acquired videos had a frame rate of 25 fps and a resolution of 720x576 pixels throughout the screen. This resolution is considered high, allowing make the same diagnosis as the original image. Each frame was encoded in YUV12 format [12 bpp, 8 bpp for the luminance (Y) component and
Chapter 3. Echocardiogram Compression 63 Table 3.4: Database for the real-time echocardiograms. Echocardiograms time and number of videos for each mode and total number of videos. B is the B mode, D the color Doppler mode, M the M mode and DP the pulsed/continuous Doppler mode. Patient Total time (minutes) % time # videos 2-D Sweep Stop B D M DP Total 1 28 68% 22% 10% 5 5 4 4 18 2 25 55% 27% 18% 6 6 4 4 20 3 36 64% 20% 16% 5 6 4 4 19 4 25 76% 9% 15% 8 8 5 4 25 5 30 83% 16% 1% 10 4 3 5 22 6 28 76% 10% 14% 7 9 4 5 25 7 27 52% 33% 15% 8 8 4 5 25 8 27 82% 10% 8% 7 6 24 4 41 9 30 91% 8% 1% 5 7 4 4 20 Table 3.5: Main regions present in the dataset echocardiograms for each device brand and mode. Mode 2-D modes Sweep modes Device US Text ECG US Text Auxiliary ultrasound SonoSite 4 4 4 4 Philips 4 4 4 4 4 4 4 bpp for the color components, chrominance (U and V)]. The echocardiographic sessions ranged from 25 to 36 minutes duration. The echocardiograms only contain the four basic operation modes according to the EAE described in Section 1.2. The operation modes are visualized one each time. Stops are introduced when measurements are performed or the mode is changing. The echocardiogram is composed of several fragments that correspond to a change of operation mode (videos). Table 3.4 shows for the available echocardiograms, the echocardiogram time distribution, total time and percentage of time for each mode, and the number of videos for the four modes. The modes are indicated with 2D, M, PCD and CD labels representing 2D, M, pulsed/continuous Doppler and color Doppler modes, respectively. The main regions for each device brand and visualization mode are shown in Table 3.5. 3.1.2.2 Real-Time Echocardiogram Characteristics The echocardiogram characteristics to be taken into account are the following: •The echocardiograms present different regions, as shown in Figure 3.3, each region having different visualization characteristics, diagnostic importance and data type.
64 3.1. Echocardiogram Databases and Characteristics Figure 3.3: Echocardiogram regions of a B mode (on the left) and a M mode (on the right) captured with a Sonosite device. The white solid line contains the ultrasound, the blue dotted and dashed line the ECG, the green dotted line the auxiliary images and the yellow dashed line the text. Each region is identified with a number (ID) on the figure. •According to their characteristics, five regions can be distinguished (see Figure 3.3): –Ultrasound: this is a video that contains the relevant diagnostic information. It is always present in the echocardiogram. This region changes with the mode. Furthermore, the ultrasound region can be divided into two according to its visualization characteristics: ∗2-D ultrasound: this represents a 2-D image of the heart in movement. Examples are the B and the color Doppler mode. ∗Sweep ultrasound: this represents a temporal evolution of one cut of the heart. Examples are M and pulsed/continuous Doppler mode. The ultrasound is displayed gradually. A new slice is visualized in each frame (see Figure 3.4). When all the screen is swept, it starts again from the beginning (see Figure 3.4b). Eventually, the sweep is stopped by the cardiologist in order to take measurements, so the same image is shown in the following frames. –ECG: this contains the ECG signal and can be seen as a sweep video (sweep ultrasounds) or as a signal. –Auxiliary video: this is a video located around the ultrasound that shows an auxiliary ultrasound to help to interpret the echocardiogram. –Auxiliary image: this is an image located around the ultrasound that shows an auxiliary image to help to interpret the echocardiogram, such as another ultrasound image or symbols regarding the configuration. It may be merely decorative. These images are always present in all the echocardiograms, and even more than onces. –Sound: the echocardiogram may also incorporate sound in order to listen to the heartbeat. –Text: this is always present in all the echocardiograms, normally more than onces. It contains information such as patient information and date, time, configuration or measurements of the study.
Chapter 3. Echocardiogram Compression 65 (a) Frame 3 of a sweep mode. (b) Frame 9 of a sweep mode. Figure 3.4: Examples of frames for a sweep mode. •The distribution and size of the regions change with the acquisition device and may change with the echocardiogram mode of operation. The represented or activated regions change according to the activated mode. For example, in Figure 3.3 the regions with identification numbers 1, 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14 are activated for the B mode. For the M mode the regions 2, 3, 4, 5, 6, 9, 10, 11, 12, 13, 14 are activated. The ultrasound region is always different for each mode, thus the activated mode depends on the ultrasound region. The other regions may be common for several modes, for example regions 3, 5, 8, 10, 11, 12, 13, 14 in Figure 3.3 are in both the B and M modes. •Depending on the data type of each region the regions that correspond with the activated mode of operation may change with time or may hardly ever change. The auxiliary image and text remain invariant and only change occasionally. However, the rest of the regions change each few seconds and thus synchronism between them is necessary. •Some regions contain color information relevant for the diagnosis while others do not. For example, the color Doppler modes have relevant information in color in the ultrasound region and in the color scale region. •The echocardiogram can be seen like a stored video. It can be stopped and run forward and back. This only affects the regions that change with time. In Table 3.6 is summarized the regions that may be presented in an echocardiogram, its data type and if the regions need synchronism between them.
66 3.2. Clinical Evaluation Methodology for Compression Recommendations Table 3.6: Data type of the echocardiogram regions for real-time transmission. Regions Data type Synchronism Ultrasound Video or sweep video Yes ECG Signal or sweep video Yes Auxiliary video Video Yes Auxiliary image Image No Sound Audio Yes Text Text No 3.2 Clinical Evaluation Methodology for Compression Recommendations This section presents an accurate clinical evaluation methodology for compressed echocardiograms and the four basic modes of operation according to the EAE recommendations. A two phase evaluation is proposed in order to simplify the evaluation process and save time compared to other accurate methodologies presented in the literature. The first phase is a semi-blind test to discriminate two transmission rates. The second phase is a more extensive blind test for the two transmission rates that are selected in the previous test. The evaluation leads to a recommendation for the compression of echocardiograms. Both tests are described in the following sections. 3.2.1 First Phase: Semi-blind Test The objective of this test is to determine in a fast and subjective way whether the cardiologist would be able or not to make the same diagnosis with the original and the compressed video. This test consists of three different parts, see Figure 3.5. The first part permits us to measure the cardiologist’s opinion about the similarity between the compressed echocardiogram and the original video. The second part is a question about whether the cardiologist’s diagnosis would be the same for both videos, the original and the compressed. The third part invites comments, if any. These questions are answered for each echocardiogram mode. 1. Measure of similarity between the original X mode video and the compressed one. 1: very different 2: different 3: acceptable 4: similar 5: identical 2. Would you give the same diagnosis with both videos for the X mode? YES NO 3. Comments. Figure 3.5: Semi-blind test: comparison of compressed echocardiogram with the original video.
Chapter 3. Echocardiogram Compression 67 In order to have an estimation that directly reflects the distortion in the diagnosis content of every mode in the echocardiogram, a clinical distortion index for the semi-blind test (CDISB) is calculated with all the information collected for each mode, cardiologist and transmission rate. This is defined as CDISB =1 2·(5 −C) 5+1 2·(1 −D) (3.1) where Cis the measurement of similarity between the original and the compressed video (1-5) and Dis the answer to the boolean question about the diagnosis (1-YES, 0-NO) for the mode being evaluated (see Figure 3.5). The CDISB is directly related to the clinical distortion of the echocardiogram. The lower the value of CDISB, the lower the distortion in the diagnosis content of the compressed echocardiogram. The CDISB could be grabbed into quality ranges thus making it possible to classify the echocardiograms. Cardiologists involved in the study considered that it would be more practical to split the CDISB values (0-0.9) into three ranges. The three defined ranges are: •The same diagnosis is possible (D= 1) with acceptable quality (C≥3): CDISB ≤0.2 •The same diagnosis is possible (D= 1) but with low quality (C < 3): 0.2 < CDISB ≤0.4 •The same diagnosis is not possible (D= 0) with low quality (C < 3): CDISB >0.7 It is important to note that there are no CDISB values between 0.4 and 0.7. It is because the cardiologists decided that it is not possible to have that the same diagnosis is not possible (D= 0) with acceptable quality (C≥3). This implies that if the same diagnosis is not possible the quality has to be unacceptable. The CDISB values are used to discard the transmission rates that are in the not acceptable range (CDISB >0.4) and to obtain the two transmission rates that are between acceptable or low quality but with the same diagnosis (CDISB ≤0.2 and 0.2≤CDISB ≤0.4) for each mode. In order to assess the test at least three cardiologist should participate and for each mode several videos of different patients and ultrasound devices should be evaluated. Once the CDISB has been calculated for all the patients, the two transmission rates are selected as follows: the upper rate is the first transmission rate with all the CDISB values of the different patients equal or lower than 0.2 and the lower rate is the immediately inferior. These two rates are the ones selected for a deep evaluation using a more extensive test, the blind test. 3.2.2 Second Phase: Blind Test This test corresponds to the same blind test proposed in [56] and consists of three different parts, which are much more extensive than in the semi-blind test. We have maintained this test because it is very complete, and now with the advantage that we just have to evaluate the transmission rates selected in the previous test, saving in this way a lot of time in the evaluation process. The objective of this test is evaluate whether the same diagnostic is possible or not with both videos, the original and the compressed video. In Figures 3.6,3.7 and 3.8 the three parts of the test for the four basic modes of operation are shown. The first part, Figure 3.6, shows the overall score given by cardiologists to the general video quality. The second part, Figures 3.7 and 3.8, consists of some questions that cardiologists have to complete based on the interpretations of the structures
68 3.2. Clinical Evaluation Methodology for Compression Recommendations and flows measured in a standard echocardiogram examination (the modes for each interpretation are shown in brackets). If these interpretations are the same as for the original video the same diagnostic is possible. Finally, there is a part for comments from the specialists, if any. 1a. General quality score for the B mode. 1: very bad 2: bad 3: tolerable 4: good 5: excellent 1b. General quality score for the M mode. 1: very bad 2: bad 3: tolerable 4: good 5: excellent 1c. General quality score for the color Doppler mode. 1: very bad 2: bad 3: tolerable 4: good 5: excellent 1d. General quality score for the pulse and continuous Doppler mode. 1: very bad 2: bad 3: tolerable 4: good 5: excellent Figure 3.6: Blind test: part I. The clinical distortion index for the blind test (CDIB) is defined as CDIB=max Qo−Qr 2·Qo ,0+P|sgn (I0i−Iri)| 2·K(3.2) where Qoand Qrare the general quality scores (see Figure 3.6) of the original and compressed videos for the mode being evaluated, respectively. For the scores rating special attention has been paid to the sharpness and definition of the edges of structures, their clear separation from other structures and the clarity in blood floods. Ioi and Iri are the interpretation of the ith parameter of the original and reconstructed videos, respectively. These interpretations are translated into numeric factors in order to be used in the equation. A zero value is assigned if the cardiologist can not answer the question due to bad video quality. The rest of the questions are assigned a value larger than 0. Kis the number of questions that contribute to each echocardiographic mode. Kis nine for the B mode, six for the M mode, five for the color Doppler mode and six for the pulsed/continuous Doppler mode. As described before at least three cardiologists should participate and for each mode several videos of different patients and ultrasound devices should be evaluated. As was established in [56], the maximum CDI value obtained by a compressed echocardiogram in order to be considered acceptable should be at most 0.25, being the CDI the mean between CDIBand CDISB. In other words, the minimum transmission rate to guarantee the clinical quality will be the lower with all the CDI values lower than 0.25. For example, let consider we evaluate a color Doppler mode
Chapter 3. Echocardiogram Compression 69 2. Provided with an interpretation. B and (B, M) Normal M study Aortic root Dilated (transversal diameter) Dissected (B, M) Normal Left atrium Dilated (transversal diameter) Intra-arterial mass-clot LVDD Normal Left (M) Dilated ventricle LVSD Normal (M) Dilated Hyperkinetic (B) Normal Global Light contractility Depressed Moderate Severe Asynergy Thin Normal (M) Light IVS CaModerate Left ventricular Severe hypertrophy Light Left AbModerate ventricle Severe wall Thin thickness Normal (M) Light PW CaModerate Left ventricular Severe hypertrophy Light AbModerate Severe Mitral Normal (B) Abnormal Valve Aortic Normal Morphology (B) Abnormal Tricuspid Normal (B) Abnormal Pericardial effusion Yes (B) No Ca: Concentric, Ab: Asymetric Figure 3.7: Blind test: part II.
70 3.2. Clinical Evaluation Methodology for Compression Recommendations 3. Provided with an interpretation. Doppler Normal Study Sten. Light (PCD) Moderate Pulmonary Max. velocity Severe flow Ins. Light Moderate Severe Yes Light Syst. Regurgitation Moderate (PCD, CD) Severa Mitral No Flow Normal E. wave, A. Pseudonormal Diast. pattern Relaxation (PCD) Abnormality Restrictive Normal Max. Sten. Light Syst. Velocity Moderate Aortic (PCD) Severa flow Yes Light Regurgitation Moderate Diast. (PCD, CD) Severe No Yes Light Tricuspid Regurgitation Moderate flow (PCD, CD) Severe No ASD Yes Septal (2D,CD) No defects VSD Yes (2D,CD) No Figure 3.8: Blind test: part III.
Chapter 3. Echocardiogram Compression 77 (a) Main window of the tool. (b) Edition window the tool. (c) Measurements window of the tool. Figure 3.13: Screen shots of the stored echocardiogram tool: compression, edition and measurements.
78 3.4. Results and Discussion for Compression for Storage Screen shots of the application are shown in Figure 3.13. The additional implemented functionalities are the following: •Visualization of the studies (Figure 3.13a). The tool is compatible with DICOM, images in other formats [29] and the proposed storage format. •A tool to modify, add and remove regions (Figure 3.13b). The tool is also able to change the text or add more regions with text. This functionality is very helpful since it allows the specialist doctor to add or delete information and consequently store more complete studies. •Measurements tool (Figure 3.13c). This is a very useful functionality to measure after the acquisition. The tool is able to charge the calibration parameters, if they are available in the DICOM header, or to perform the calibration and store these parameters in the DICOM header. It is also able to store the study including the measurements. 3.4.2 Results for Compression of Stored Echocardiograms The compression gain can be calculated previously to the compression process. The compression rate (CR) for the proposed compression format with respect to the raw image with and without including the DICOM header follows the expression: CR =widthT∗heightT∗N∗3∗8 ( + headerDICOM ) header +widthU∗heightU∗N∗bpp ( + headerDICOM )(3.3) widthTand heightTare the total size of the image, Nis the number of frames, 3 corresponds to the three color components, and 8 to the bits per pixels. The header is the header size of the proposed format including the XML file and auxiliary images, widthUand heightUare the size of the ultrasound region, and bpp are the bits per pixels or bit rate for the compressed ultrasound region. The headerDICOM is the DICOM header that may or may not be considered. The DICOM header is taken into account if the compression rate of the whole file needs to be calculated. However, the DICOM header is not included in the expression if the compression rate of the image format is to be evaluated. The following expression 3.4 is the percentage of gain in storage space of the proposed compression with respect to the compression study using JPEG 2000, which does not distinguish regions. The DICOM header may or may not be included. The ultrasound regions are compressed with the same quality for both methods. Gain =widthT∗heightT∗N∗bpp −header −widthU∗heightU∗N∗bpp header +widthU∗heightU∗N∗bpp ( + headerDICOM )∗100 (3.4) where the parameters are the same as those defined for CR. In order to evaluate the improvement of the proposed compression format compared to JPEG 2000, the echocardiograms database has been compressed with both methods. Table 3.7 shows for each echocardiogram device and mode available, the ultrasound size, average header size (XML file and auxiliary images) for one frame, the CR for the proposed format and the percentage of gain (G)
Chapter 3. Echocardiogram Compression 79 Table 3.7: Parameters and compression results for the echocardiogram devices and modes of the database. Device Mode Header ROI CR/CRDICOM G/GDICOM(%) (Bytes) Size 1 frame 16 frame 1 frame 16 frame M4401 576x295 31/19 - 28/16 - DP 5036 569x295 30/20 - 24/15 - D2746 404x417 29/21 32/31 20/14 32/32 Agilent/ Acuson B825 551x480 27/20 28/27 14/9 16/16 M6161 692x387 35/25 - 48/32 - DP 6838 641x369 39/25 - 61/37 - D3510 538x521 36/21 39/37 49/27 62/60 Philips B2729 442x521 39/27 42/41 62/41 75/73 M11489 1024x504 34/28 - 41/32 - DP 11637 1024x499 33/26 - 35/28 - D3227 845x739 28/22 29/29 19/14 23/23 Siemens B3481 824x739 29/23 31/30 22/17 27/27 of the proposed format with respect to JPEG 2000. The last two values have been calculated for the average of all the frames stored separately and considering 16 frames of the same visualization stored in the same file and sharing the configuration file. Furthermore, the results not taking into account the DICOM header and those taking into account the DICOM header are shown. The ultrasound regions have been compressed with 1 bpp (bit per pixel) for both formats and the auxiliary regions with 0.5 bpp for the proposed format. 1 bpp and 0.5 bpp have been selected because these bit rates have shown good clinical quality for ultrasound images compressed with JPEG 2000 [44]. In order to provide a visual inspection of the proposed method, Figure 3.14 shows the original images and the images after segmentation, compression and decompression. The ultrasound region has been compressed with 1 bpp and the auxiliary regions with 0.5 bpp for both images. The top images were acquired with an Agilent device corresponding to the Color Doppler mode. The PSNR of the ultrasound part is about 40 dB. The size for the proposed format is 24 kbytes and the original size is 774 kbytes, having a CR of 32. The bottom images were acquired with Acuson device corresponding to the M mode. The PSNR of the ultrasound region is about 35 dB. The size of the proposed format is 26 kbytes instead of 788 kbytes, having a CR of 30. 3.4.3 Discussion for Compression of Stored Echocardiograms According to the compression rate and gain expressions given in the previous section, the compression efficiency mainly depends on the header length (XML file and auxiliary images) and ultrasound size (ROI), and consequently depends on the acquisition devices. For example, in Table 3.7, the Agilent/Acuson device and B mode shows the worst compression results (CR = 27 for 1 frame). This is due to the ROI size (551x480 of 600x430). The ROI covers almost the whole image. On the other hand, the best compression results are for the Philips device and B mode (CR =
80 3.4. Results and Discussion for Compression for Storage Figure 3.14: Samples of echocardiogram images before and after compression with the proposed method. Images on the left are the original images and on the right the reconstructed images with the proposed method. The upper images were acquired with an Agilent device and the lower images with an Acuson device.
Chapter 3. Echocardiogram Compression 81 39 for 1 frame) due to its small ROI size (442x521 of 800x564) and header length (2729 bytes). Consequently, the smaller the ROI and header, the better the compression results. Another important factor is the number of frames (N). If more frames are compressed together, sharing the same headers, better compression efficiency is achieved. This is easily visualized if we compare the results taking into account and without taking into account the DICOM header (see Table 3.7). The compression results considering the DICOM header are worse than those not considering it. However, the compression results depend more on the DICOM header when only one frame is stored in a DICOM file. In conclusion, we can observe that for a typical quality the proposed format overcomes the JPEG 2000 standard for all the echocardiograms of the database. Furthermore, the obtained CR enhances the results of the ROI coding included in the JPEG 2000 standard, Maxshift method, that obtains a CR of 16 [35] for medical images with smallest ROI region than the ultrasound. Furthermore, although the compression results depend on the acquisition device and mode, a saving in the storage space is achieved for all the available devices and modes, which are representative of typical echocardiogram image distribution. It is expected that similar results will be obtained for other image codecs. The proposed format is specially designed for ultrasound images, but can also be applied to other medical image modalities. However, the compression performance depends on the image visualization. Thus, the compression results will have to be evaluated for each image modality. 3.5 Compression for Real-time Transmission An encoded algorithm is presented for each data type according to the previously presented echocardiogram characteristics. The encoding recommendations in Chapter 2are taken into account. As the ultrasound region is the largest and the most important for the diagnosis, it is necessary to take special care in the design of the compression method for this region. In a previous study [124] we proposed a compression method for the ultrasound region based on visualization modes (sweep and 2-D modes) and the use of SPIHT algorithms. The proposed approach was compared with the typically used video approaches, Xvid and H.264/AVC, showing better results for the compression by visualization modes in terms of PSNR for all the echocardiogram modes. For this reason, the visualization characteristics have been taken into account and SPIHT algorithms are proposed in this thesis for compression for real-time transmission. Some changes have been added to the proposal in [124] for the 2-D modes since the database was increased and the design was optimized for a wide set of echocardiograms. In [124], Run Length Encoding (RLE) was proposed for the compression of the color components and the 2-D modes. This algorithm is efficient for echocardiograms acquired with low color resolution devices such as the Sonosite device used in [124]. Nevertheless, it is not efficient for devices with high color resolution. The final encoded recommendations for every data type is summarize in Table 3.8 and described bellow. •Video. For the video compression, 3-D SPIHT [62] is proposed due to its good performance. The resolution time, time between frames encoded together, is the corresponding to 16 frames.
82 3.6. Results and Discussion for Compression for Real-Time Transmission As the available database has 25 frames per second, the resolution time for the database is of 0.64 seconds. •Sweep mode. The visualization characteristics of the sweep modes have been previously described. The image is visualized gradually. A new slice appears in each frame, remaining the rest invariant. The proposed approach compresses the new slice each frame. The slices are compressed with 2-D SPIHT due to the reasons named in Chapter 2. The images can contain color or not. If the image contains color 2-D CSPIHT [69] is proposed. On the contrary, if the image does not contain color, 2-D SPIHT [41] is proposed. The compression efficiency can be affected by the slice width. The slice height is large enough. A minimum slice width of 32 pixels has been considered suitable. If the slice width is less than 32 pixels, several slices are joined to reach this value and the rest of pixels form the next slice. This introduces a visualization delay in real-time applications, but it is always less than 32 pixels (the time delay depends on the device and the sweep speed). Time between slices encoding is always lower than 0.3 seconds for the ultrasound region. As each slice is coded separately, if an error occurs in the transmission it only affects to the quality of one slice. That is an important advantage compared with the MPEG-4 codecs that spread the error along the image and through several frames. •Image. The codecs for the images are the same as slice codecs for the sweep video, 2-D SPIHT for image without color and 2-D CSPIHT for color image. •ECG Signal. In [125] a real-time automatic ECG coding method that guarantees signal interpretation quality was designed. This coding method is based on 1-D SPIHT. It is the proposed solution in case of compress the ECG as signal. •Audio. The sound has to be encoded with an audio codec for real time, such as Opus [126]. Opus is a highly versatile audio codec. Opus can handle a wide range of audio applications, but it is particularly suitable for interactive real-time applications over the Internet due to its low delay (22.5 ms by default). It can scale from low bit-rate narrowband speech to very high quality. Opus supports constant and variable bitrate encoding from 6 kbit/s to 510 kbit/s, frame sizes from 2.5 ms to 60 ms, and certain sampling rates from 8 kHz (with 4 kHz bandwidth) to 48 kHz (with 20 kHz bandwidth). •Text. The text has to be encoded with a text codec, such as the ASCII 8-bit Unicode Transformation Format (UTF-8) [127]. The text is lossless coded. In order to guarantee clinical quality of the encoded regions, a clinical evaluation of the ROI regions is necessary. 3.6 Results and Discussion for Compression for Real-Time Transmission Previous results of clinical quality are not available for the proposed compression approach for real-time transmission of echocardiograms. Thus, a clinical evaluation following the methodology proposed in Section 3.2 was carried out for the compression method for real-time echocardiogram
Chapter 3. Echocardiogram Compression 83 Table 3.8: Codec for every data type for compression for real-time transmission. Data Type Codec Black&White Color Video 3-D SPIHT [62] Sweep video 2-D SPIHT [41] 2-D CSPIHT [69] Image 2-D SPIHT [41] 2-D CSPIHT [69] ECG Signal ECG SPIHT [125] - Audio Opus [126] - Text UTF-8 [127] - transmission and for different transmission rates. A precise evaluation of the real degradation in the compressed echocardiogram and a recommendation for the echocardiogram compression are provided. The recommended rates are compared with the transmission rates used in previous systems. 3.6.1 Evaluation Setup for Compression Recommendations The database described in 3.1.2.1 has been used to carry out the evaluation described in the previous Section 3.2. Since the ultrasound region is the region that contains the most relevant clinical information, the evaluation has been carried out only for this region. The resolution of the ultrasound region for the different devices were: 668x496 for the Sonosite, 644x488 for the Envisor, and 634x462 for the IE33. Since not all the session revealed clinical information and much of the time was taken up with the process of measuring cardiac features (removed from the original video to avoid providing hints for later evaluations), the sessions were edited so as to remove these parts. The resultant duration of each video mode was approximately 1 min in the B and color Doppler modes and 30 s in the M and pulsed/continuous Doppler modes. Each session contained at least four videos of each mode and up to ten videos, having a total of 61 videos for the B mode, 63 for the color Doppler mode, 37 for the M mode and 39 for continuous/pulsed Doppler mode. The number of videos per device and mode are shown in Table 3.4. Audio was excluded from this study. Three cardiologists experts in echocardiography interpretation, participated in the clinical evaluation. In order to evaluate the intra-observer variability, three sessions were evaluated twice for both tests and by the three cardiologists. The repeated sessions corresponded to different devices and rates. For the semi-blind tests, seven transmission rates were evaluated: 100, 150, 200, 250, 300, 400 and 500 kbps for the B and color Doppler modes, and 10, 15, 20, 25, 30, 40 and 50 kbps for the M and pulsed/continuous Doppler modes. These rates were chosen because their PSNR are considered adequate from the point of view of suitable clinical quality. In order to simplify the evaluation process, each pair of transmission rates and sessions was evaluated by two cardiologists. Thus, each cardiologist evaluated six sessions for each rate instead of nine. The total number of semi-blind tests carried out by each cardiologist was 180 (seven rates and six sessions, plus the three repeated sessions and 4 modes per session making a total of 180 tests). In each visualization the original and the compressed videos were visualized at the same time. Each cardiologist saw only two or three sessions (the four modes) per day so that this first evaluation lasted for approximately
84 3.6. Results and Discussion for Compression for Real-Time Transmission one month. Regarding the blind test, the three cardiologists evaluated two rates and the nine original sessions. Hence, the total number of blind tests carried out by each cardiologist was 120 (two rates plus the original session and nine sessions plus the three repeated sessions and 4 modes for session, making a total of 120 tests). Each cardiologist assessed a maximum of two sessions per day and six per week so that the evaluation took approximately two months. The evaluation in [56] took two months too, but only four transmission rates were evaluated versus seven in the proposed evaluation. Furthermore, the second test, which is very burdensome, is only carried out for two transmission rates in the proposed evaluation versus four in [56]. 3.6.2 Results for Compression for Real-Time Transmission The results for the semi-blind test are shown in Tables 3.9-3.12, one table for each mode (B, color Doppler, M, and pulsed/continuous Doppler, respectively). The CDISB values (mean ± standard deviation of scores obtained from two cardiologists) for the nine patients and the three devices at different transmission rates are listed. The CDISB rated as inadequate are shown in dark gray, the CDISB with unacceptable quality but for which the same diagnosis is possible in light gray, and the CDISB with acceptable quality and which the same diagnosis is possible in white. As it has been described in the evaluation methodology, Section 3.2, the two rates selected are between the two latter ranges (light gray and white). Therefore, the highest selected transmission rate is the lowest rate with all the CDISB in white, and the lowest selected transmission rate is the rate immediately inferior. The two transmission rates selected for each mode appear in Tables 3.13. There is a clear effect of the transmission rate on the CDISB values for all the modes. The higher the transmission rate, the lower the CDI. Analysis of variance (ANOVA) confirmed that the CDISB values obtained for the different rates show significant differences (p<0.05) for all the modes.
Chapter 3. Echocardiogram Compression 85 Table 3.9: Semi-blind test: CDI values for the B mode. Patient Transmission rate (Kbps) 100 150 200 250 300 400 500 1 0.85±0.07 0.25±0.07 0.20±0.00 0.20±0.00 0.20±0.00 0.20±0.00 0.05±0.07 2 0.90±0.00 0.80±0.00 0.25±0.07 0.15±0.07 0.15±0.07 0.15±0.07 0.10±0.00 3 0.80±0.14 0.75±0.07 0.20±0.00 0.20±0.00 0.20±0.00 0.15±0.07 0.10±0.00 4 0.25±0.07 0.20±0.14 0.20±0.14 0.15±0.07 0.10±0.00 0.15±0.07 0.05±0.07 5 0.35±0.07 0.20±0.00 0.15±0.07 0.10±0.00 0.15±0.07 0.15±0.07 0.05±0.07 6 0.85±0.07 0.25±0.07 0.15±0.07 0.20±0.00 0.15±0.07 0.05±0.07 0.05±0.07 7 0.25±0.07 0.15±0.07 0.15±0.07 0.15±0.07 0.10±0.00 0.15±0.07 0.05±0.07 8 0.30±0.00 0.25±0.07 0.15±0.07 0.15±0.07 0.20±0.00 0.15±0.07 0.05±0.07 9 0.25±0.07 0.20±0.00 0.10±0.00 0.15±0.07 0.15±0.07 0.05±0.07 0.05±0.07
86 3.6. Results and Discussion for Compression for Real-Time Transmission Table 3.10: Semi-blind test: CDI values for the color Doppler mode. Patient Transmission rate (Kbps) 100 150 200 250 300 400 500 1 0.25±0.07 0.20±0.14 0.15±0.07 0.15±0.07 0.10±0.00 0.15±0.07 0.05±0.07 2 0.25±0.07 0.15±0.07 0.15±0.07 0.10±0.00 0.15±0.07 0.15±0.07 0.05±0.07 3 0.20±0.00 0.20±0.07 0.10±0.00 0.15±0.07 0.15±0.07 0.10±0.00 0.10±0.00 4 0.25±0.07 0.25±0.07 0.20±0.00 0.15±0.07 0.10±0.00 0.15±0.07 0.05±0.07 5 0.20±0.00 0.20±0.00 0.15±0.07 0.10±0.00 0.15±0.07 0.10±0.14 0.05±0.07 6 0.75±0.07 0.25±0.07 0.05±0.07 0.10±0.00 0.05±0.07 0.00±0.00 0.00±0.00 7 0.20±0.00 0.15±0.07 0.15±0.07 0.15±0.07 0.05±0.07 0.10±0.00 0.05±0.07 8 0.25±0.07 0.25±0.07 0.15±0.07 0.15±0.07 0.15±0.07 0.10±0.14 0.00±0.00 9 0.25±0.07 0.25±0.07 0.15±0.07 0.15±0.07 0.15±0.07 0.05±0.07 0.10±0.00
Chapter 3. Echocardiogram Compression 93 assessed. The rest of images are compressed with 0.5 bpp, which presents a good quality in medical image compression [37]. The ECG signal recommendation is in [125]. The auxiliary video, which has a small size, and sound are compressed with typical compression rates. Table 3.15: Recommended transmission rates per mode to obtain good clinical quality. B D M DP 200 kbps 200 kbps 40 kbps 40 kbps 3.6.3.2 Codecs Comparison Section 2.7 lists the resolution, codecs and transmission rates for the most relevant ultrasound video systems. It can be seen that the recommended transmission rates using the compression proposed in this Thesis are lower than those used in the other systems, even for the transmission rates with low resolution video. The best compression result for high video resolution was presented in [56]. The recommended transmission rates were 768 Kbps for the B and M mode, and 256 Kbps for the color Doppler and pulsed/continuous Doppler. Hence, the proposed method requires less than 26% of the amount of data for the B mode, 78% for the Color Doppler mode, 5.2% for the M mode, and 16% for pulsed/continuous Doppler. We can conclude that the results obtained with the proposed technique clearly show an improved performance in terms of the transmission rate as compared with the other codecs presented in the literature for ultrasound videos (Xvid, H.264/AVC, Windows Media and HEVC). The low transmission rates are obtained mainly because the echocardiogram characteristics have been taken into account in the compression design and a clinical evaluation has been carried out in order to obtain the minimal recommended rates. This saving in the transmission rate will lead to better transmission performance. 3.7 Conclusions The main objective of this Chapter is to achieve the efficient compression of echocardiograms. This overall objective has been divided into three: to develop a clinical evaluation methodology for transmission rate recommendations, to design a compression method to store echocardiograms and to design a compression method to transmit echocardiograms in real-time. The main conclusions relating to these three specific objectives are listed below: •An evaluation methodology which is accurate but not very time consuming has been designed for compressed echocardiogram. The proposed method can be used to assess any video codec and to recommend a minimal transmission rate which guarantees clinical quality. For all the modes except the B mode it is only necessary to perform the semi-blind test because it gives the same result as the blind test. This evaluation methodology may be adapted to other ultrasound techniques or modes, or even to other medical image modalities.
94 3.7. Conclusions Table 3.16: Codecs and recommended transmission rates and bits per pixel for each type of region. Regions Data type Codec Compression Video 3-D SPIHT [62] 200 Kbps Ultrasound Sweep video 2-D SPIHT [41], 2-D CSPIHT [69] 40 Kbps Signal ECG SPIHT [125] 500 bps ECG Sweep video 2-D SPIHT [41], 2-D CSPIHT [69] 0.5 bpp Auxiliary video Video 3-D SPIHT [62] 100 Kbps Auxiliary image Image 2-D SPIHT [41], 2-D CSPIHT [69] 0.5 bpp Sound Audio Opus [126] 10 Kbps Text Text UTF-8 [127] 1-4 bytes per character •An image compression format for storage purposes has been proposed that takes advantage of the segmentation facilities of the device to enhance the compression performance. This compression format is easily integrable in the acquisition device without adding complexity to devices that already incorporate the DICOM standard. The proposed method shows better results than compression without using regions for all the available dataset, having a compression gain ranging from 14 % up to 75 % for a typical bit rate. Although the compression results depend on the acquisition devices, how the image is displayed, and the compression quality, the compression ratios obtained for the DICOM file ranged from 19 to 41 without losing diagnostic information. Furthermore, a tool that makes conversions between different image formats has been developed. This tool allows interoperability between medical centers and devices, and also enables echocardiogramsto be stored with the proposed format, thus saving storage space. Options to edit the image (edit/add/remove regions) taking advantage of the proposed format based on regions have been also included. •An echocardiogram compression method for real-time transmission based on regions and visualization modes has been designed. Codecs and transmission rate recommendations have been given for each region. Since the ultrasound region contains the most important information from a medical point of view, a comprehensive evaluation has been preformed for this region. Minimum transmission rates have been recommended for each mode in order to guarantee suitable clinical quality for the transmission and storage of echocardiogram videos using the proposed technique. The recommended transmission rates for the ultrasound regions are the following: 200 Kbps for the 2-D and the color Doppler modes, and 40 Kbps for the M and the pulsed/continuous Doppler modes. These are very good results in terms of bandwidth use, especially for the M and pulsed/continuous Doppler modes that have been obtained thanks to the fact that the compression method takes into account the stationary characteristics of the sweep modes and only a thin slice is compressed for each frame. The compression recommendation for the other regions has also been provided. Moreover, these results make possible the transmission of echocardiogram videos over 3Gwireless networks and beyond. The recommended transmission rates for previous ultrasound transmission systems are considerably higher than those given for the proposed method which allows better
Chapter 3. Echocardiogram Compression 95 transmission results.
Chapter 4 Echocardiogram Transmission in Real-time This Chapter deals with the second part of the tele-echocardiography systems (see Figure 4.1): transmission and display. Once the compression recommendations have been established, it is necessary to guarantee that the echocardiogram is received without diagnostic information being lost. If the transmission is completed without errors, and consequently the whole echocardiogram is visualized with the recommended transmission rate, the same diagnosis as that of the original echocardiogram is guaranteed. However, if errors occur in the transmission process, some echocardiogram parts will be visualized with a transmission rate lower than the recommended rate. In that case the diagnosis may be possible or not. For this reason, an evaluation for the echocardiogram display recommendations has been designed and carried out for the echocardiograms compressed with the method proposed in this Thesis for real-time transmission. The display recommendation allows us to know if the echocardiogram is visualized without losing diagnostic information for any transmission conditions instead of having to carry out different assessments for each channel condition. In addition, a protocol is required for the real-time end-to-end transmission of echocardiograms over IP that defines how to transmit each encoded region and the synchronism between them. In the case of WiMAX channels, which introduce packet losses, an error control method is also required for the regions with relevant clinical information. The echocardiogram database and characteristics for real-time transmission purposes used in this Chapter were described in the previous Chapter, Section 3.1. This Chapter is organized as follows. Section 4.1 describes the evaluation methodology for display recommendations of medical images after transmission, and specifically for echocardiograms that have been encoded with the method proposed in this Thesis for real-time transmission. Section 4.2 gives the display recommendations for the echocardiogram after compression with the proposed method and recommended transmission rates for real-time transmission as listed in Chapter 3. Section 4.3 describes the echocardiogram transmission protocol. Section 4.4 discusses the error control method and the configuration for each type of region. Section 4.5 provides the results and discussion relating to simulations of real-time echocardiogram transmission over WiMAX channels carried out both with the compression and transmission methods proposed in this Thesis and without them in order to evaluate the improvements of each proposed method. Finally, the conclusions of this Chapter are given in Section 4.6. 97
98 4.1. Clinical Evaluation Methodology for Display Recommendations Acquisition & visualization without errors Compression for Storage Compression for real time Transmission Real-time transmission •Error control •Protocols Wireless channel raw video images Real-time reception •Error control Decoding Visualization with errors & storage Clinical quality Transmission Display Figure 4.1: Tele-echocardiography system structure followed in this Thesis: transmission and visualization parts. 4.1 Clinical Evaluation Methodology for Display Recommendations This section presents an evaluation methodology to give recommendations for the display of echocardiograms after transmission. The starting point of the evaluation is the recommendations already given for the compression. If errors occur in the transmission, some parts of the echocardiogram are visualized with a transmission rate lower than the recommended rate, but this does not mean that the diagnosis is not possible. The maximum acceptable time during which the echocardiogram is visualized with a transmission rate lower than the recommended rate has to be determined. In order to give recommendations that are independent of the transmission channel, the worst case scenario has to be evaluated. The evaluation methodology has been designed taking into account the compression approaches followed in this Thesis for real-time transmission of the ultrasound regions. Different evaluations are proposed for each visualization mode: •2-D modes. The evaluation establishes the maximum percentage of time that the echocardiogram can be visualized with a transmission rate lower than the recommended rate. Different percentages of time and transmission rates have to be assessed for a 2-D video. •Sweep modes. For the sweep modes, the echocardiogram is compressed by slices. The important part for the diagnosis is when the sweep is stopped and the screen is filled with slices. For this reason, instead of evaluating the percentage of time with a lower than recommended transmission rate, the number of slices that can be visualized with a lower transmission rate is evaluated. Different numbers of slices and transmission rates have to be assessed for an image. This image is that used by the cardiologist to perform the diagnosis. The evaluation for the two types of visualization modes consists of a semi-blind test. In this case a blind test is not necessary because there is already adequate clinical quality resulting from the previously recommended transmission rates. The objective of this test is to determine whether or
Chapter 4. Echocardiogram Transmission in Real-time 99 not the cardiologist would be able to make the same diagnosis with the video or images compressed at the transmission rate recommended in Chapter 3and with the video or image with some parts compressed with a lower transmission rate due to transmission errors. This test consists of giving an opinion about the similarity between the two compared videos or images (see Figure 4.2). It is also established whether the diagnosis would be the same with both videos and images. The same diagnosis is not possible if the mark is lower than 3; otherwise, the same diagnosis is possible with either videos or images. 1. Measure of similarity between the two videos or images. 1: very different 2: different 3: acceptable 4: similar 5: identical Figure 4.2: Semi-blind test: comparison of compressed echocardiogram with the recommended transmission rates. A clinical distortion index (CDI) is calculated in order to have an estimation that directly reflects if the evaluated video or image having some parts with a lower than recommended transmission rate is of sufficient clinical quality for an adequate diagnosis. A CDI is calculated for each echocardigram operation mode, video or image, with different percentages of time of the total time or numbers of slices with lower than recommended transmission rates. This is defined as CDI =min {Di}(4.1) where Diis the measurement of the test in Figure 4.2 by different cardiologists and images or videos compressed with the conditions to be evaluated. The minimum Divalue is selected, because in the event that one of the evaluated visualizations is not correct, the conditions are not suitable for a valid diagnosis. The CDI value can be divided into two quality ranges: •CDI < 3: the same diagnosis is not possible and the quality is not sufficient for the evaluated videos or images. •CDI ≥3: the same diagnosis is possible and the quality is sufficient for all the evaluated videos or images. With the calculated CDI value, the maximum percentages of time or maximum number of slices with the different transmission rates lower than the recommended rate will be chosen for each mode. These values are the highest with CDI values higher or equal to 3. This is because the cardiologists have decided that the minimum acceptable value is 3. This evaluation methodology is valid for compression based on visualization modes. However, if no visualization modes are distinguished in the compression process, the methodology for 2-D modes can be applied for all other modes. These recommendations determine whether each fragment of the echocardiogram that corresponds with a change of operation mode has sufficient clinical quality. Each fragment has to be
100 4.2. Results and Discussion for Display Recommendations checked separately. If all the fragments are valid, the whole echocardiogram is valid. In the event that a mode is not suitably visualized, the whole mode has to be transmitted again, delaying the diagnosis process. 4.2 Results and Discussion for Display Recommendations The clinical evaluation proposed in Section 4.1 was carried out for the echocardiograms that are compressed with the method proposed in this Thesis and the recommended transmission rates listed in Chapter 3. Display recommendations are provided for each echocardiogram operation mode independently of the transmission channel. Thus, the display recommendations allow us to determine whether the echocardiogram fragments that correspond to a mode have adequate clinical quality without needing to perform an evaluation for each transmitted echocardiogram. 4.2.1 Evaluation Setup for Display Recommendations The database described in Section 3.1.2.1 has been used to carry out the evaluation. Since the ultrasound region is the region that contains the most relevant clinical information, the evaluation has been carried out for this region only. The echocardiograms were edited so as to remove the part without clinical information. The number of videos per device and mode are shown in Table 3.4. Each session contains videos of about 1 minute of duration for the 2-D modes and one image per video for the sweep modes. Two cardiologists expert in echocardiography interpretation participated in the clinical evaluation. In order to evaluate the intra-observer variability, three sessions were evaluated twice for both tests and by the two cardiologists. The repeated sessions corresponded to different devices, percentages of time and numbers of slices with transmission rates lower than recommended, and transmission rates lower than recommended. The echocardiograms were compressed with the proposed techniques and recommended transmission rates for each mode for the ultrasound region (200 kbps for the 2-D modes and 40 kbps for the sweep modes). We have to take into account that an error affects to all the bits encoded together, 16 frames for 2-D modes and 1 slice for the sweep modes. As the compression algorithms are embeddedness, for the bits that are encoded together, only the first error affects. When an error occurs the rest of bits are ignored and the region is decompressed with the received previous bits. For the sweep modes, the numbers of slices evaluated with transmission rates inferior to the recommend rate were 2, 4, 6, 7, 8, 9 and 10. The transmission rates were 20 kbps (50 % of the recommended transmission rate) and 0 kbps. The 0 kbps transmission rate simulates the case in which no packets are received. The slices with the lowest transmission rate were located in the middle of the screen to simulate the worst case, with the worst clinical quality in order to give recommendations that do not depend on the transmission channel. In Figure 4.3 two samples of images that have been evaluated for a M mode are illustrated. Figure 4.3a shows an example of an image with two slices with a transmission rate of 0 kbps and the rest of the slices with 40 kbps (recommended transmission rate). Figure 4.3b shows an example of an image with eight slices with a transmission rate of 20 kbps (50 % of the recommended rate) and the rest of the slices with a transmission rate of 40 kbps. The total number of slices for the whole screen depends on the device but is about 20 for the available database. For the 2-D modes, the evaluated percentages of time with transmission rates lower than the recommended rate were 5%,
Chapter 4. Echocardiogram Transmission in Real-time 101 20% and 50%. The transmission rates were 160 kbps (80 % of the recommended transmission rate), 100 kbps (50 % of the recommended transmission rate) and 0 kbps. The bandwidth with inferior quality is continuous and is located in the middle of each evaluated video to simulate the worst possible case. This is so that the recommendations do not depend on the transmission channel. A continuous distribution of the bandwidth with a lower transmission rate than the recommended rate is worse than a burst distribution. Figure 4.4 shows the bandwidth distributions over time for a video with 20 % of the time with 160 kbps (drawn in dark blue) and for another video with 5 % of the time with 0 kbps (drawn in sky blue). Each cardiologist saw only one visualization quality of the same mode per day so that the evaluation took approximately 17 days. (a) Image with two slices (in red) with transmission rate of 0 kbps. (b) Image with eight slices (in red) with transmission rate of 20 kbps. Figure 4.3: Samples of M mode images for the evaluation of the display recommendations. 4.2.2 Results for Display Recommendations The results of the test are shown in Tables 4.1 and 4.2, one table for each type of visualization mode (2-D and sweep modes, respectively). The CDI values rated as adequate are shown in dark gray and the CDI values with inadequate quality are in white. The maximum acceptable number of slices and percentages of time that the echocardiogram is visualized with a transmission rate lower than the recommended rate is the highest with a CDI value higher or equal to 3. There is a clear effect of the transmission rate, percentage of time and number of slices on the CDI values. The higher the transmission rate, the higher the CDI value. The higher the percentage of time and number of slices, the lower the CDI value. As regards inter-observer variability, both cardiologists were of the same opinion as to whether or not the video or image had adequate clinical quality (CDI value higher or equal to 3 or lower than 3). However, they had slightly differing opinions about the quality (in the same CDI range, but different values). The intra-observed evaluations show that there are no diagnostic differences between the repeated measurements.
102 4.2. Results and Discussion for Display Recommendations Figure 4.4: Samples of visualized bandwidth distribution for the evaluation of the display recommendations of a 2-D mode video of 1 minute duration. The bandwidth with 20 % of the time with 160 kbps is shown in dark blue. The bandwidth with 5 % of the time with 0 kbps is shown in sky blue. Table 4.1: CDI values for the 2-D modes. % time Transmission rate (kbps) B mode Doppler mode 160 100 0 160 100 0 53 3 3 4 3 3 20 3 2 2 3 3 2 50 2 2 1 2 2 1 4.2.3 Discussion for Display Recommendations Given the results in Tables 4.1 and 4.2 and the methodology in Section 4.1, the following display recommendations are given for each mode: •B mode. –Up to 5% of the time the B mode can be visualized with any transmission rate. –Up to 20% of the time the B mode can be visualized with a transmission rate of 160 kbps or higher. •Color Doppler mode. –Up to 5% of the time the color Doppler mode can be visualized with any transmission rate.
Chapter 4. Echocardiogram Transmission in Real-time 109 Table 4.6: XML example for the configuration packets of the Philips Envisor device in the Figure 4.7. <?xml version="1.0"?> <!DOCTYPE configuration SYSTEM "conf.dtd"> <configuration> <tsize w=’720’ h=’576’/> <device> Philips Envisor </device> <region id=’1’ mode=’Doppler’, reliability=’SECM,3,1,3,10,2,80,112,120’> <ultr color=’yes’fps=’25’ bitrate=’200’ resolution=’16’/> <position x=’77’ y=’0’/> <size w=’512’ h=’385’/> <mps>200</mps> </region> <region id=’2’ mode=’B,Doppler’> <ecg sf=’90.8’ bpp=’0.5’ ppslice=’128’/> <position x=’494’ y=’52’/> <size w=’454’ h=’32’/> </region> <region id=’3’ mode=’DP’, reliability=’SECM,1’> <ultr color=’no’type=’sweep’fps=’25’ pps=’121.43’ bitrate=’40’ resolution=’32’/ > <position x=’200’ y=’0’/> <size w=’544’ h=’321’/> </region> <region id=’4’ mode=’DP’> <img color=’no’bpp=’0.5’/> <position x=’77’ y=’211’/> <size w=’150’ h=’130’/> </region> <region id=’10’ mode=’all’> <text> PATIENT’S DATA </text> <position x=’0’ y=’0’/> </region> </configuration>
110 4.3. Protocol for Echocardiogram Transmission in Real-time Table 4.7: XML equivalence for real-time transmission configuration file and storage format file of the echocardiogram in Figure 3.3 captured with a Sonosite device. XML file for real-time transmission XML files for storage <?xml version="1.0"?> <!DOCTYPE configuration SYSTEM "conf.dtd"> <configuration> <tsize w=’720’ h=’576’/> <device> Sonosite </device> <region id=’1’ mode=’B’> <ultr color=’no’fps=’25’ bitrate =’200’ resolution=’16’/> <position x=’76’ y=’52’/> <size w=’573’ h=’420’/> </region> <?xml version="1.0"?> <!DOCTYPE configuration SYSTEM "roiformat.dtd"> <format> <tsize w=’720’ h=’576’/> <region> <pos x0=’76’ y0=’52’ x1=’659’ y1 =’472’><\pos> <roi><size>30083<\size><\roi> <\region> <\format> <region id=’2’ mode=’M’> <ultr color=’no’type=’sweep’fps =’25’ pps=’200’ bitrate=’40’ resolution=’32’/> <position x=’55’ y=’104’/> <size w=’512’ h=’456’/> </region> <?xml version="1.0"?> <!DOCTYPE configuration SYSTEM "roiformat.dtd"> <format> <tsize w=’720’ h=’576’/> <region> <pos x0=’55’ y0=’104’ x1=’567’ y1 =’560’><\pos> <roi><size>29184<\size><\roi> <\region> <\format>
Chapter 4. Echocardiogram Transmission in Real-time 111 Region ID Open/closed · · · Region ID Open/closed (8 bits) (1 bit) (8 bits) (1 bit) Figure 4.6: Confirmation packets. Figure 4.7: Echocardiogram regions of a color Doppler mode (on the left) and a pulsed Doppler mode (on the right) captured with a Sonosite device. The white solid line contains the ultrasound, the blue dotted and dashed line the ECG, the green dotted line the auxiliary images and the yellow dashed line the text. Each region is identified with a number (ID) on the figure. 4.3.3 Data Packets Coding AUDP connexion is used for the transmission of each encoded region. However, a high data flow has to be transmitted. Furthermore, for error prone channels it is recommended to transmit small packets to enable the easy recovery of lost packets [128]. In order to support segmentation, it is necessary to establish the maximum packet size (mps) for the data packets and to include a sequence number in each data packet. The data packets follow the structure shown in Figure 4.8. The length of the sequence number depends on the type of region: 8 bits for sound, ECG, auxiliary video and sweep ultrasound, and 16 bits for 2-D ultrasound. The sequence number is used to distinguish each datagram of each connection that corresponds with each region. In this way, the datagrams can arrive out of order. The size of the encoded region field depends on the type of region, but it always has a size equal or inferior to the mps whose value is selected according to the channel and type of region. It is not necessary to specify the total size of any type of codified packets because this is known thanks to the configuration, as we will see below. Seq no. Encoded region (8 or 16 bits) (≤mps bits) Figure 4.8: Data packets. The coding process for each region is shown in Figure 4.9a (labelled DATA). RTis the resolution time, referring to the time between data that is codified together. Bbits are the bits after the coding. Nare the packets after the fragmentation, having packets of Xbits. These packets
112 4.3. Protocol for Echocardiogram Transmission in Real-time Table 4.8: Codification parameters expressions that correspond with the encoder in Figure 4.9. Parameters Videos Sweep videos ECG Sound RTseconds fps ∗resolution resolution pps ppslice sf sampler B bits bitrate ∗RT+ 4 bitrate ∗RTbpp ∗h∗ppslice sampler ∗bitrate if B≤mps, 1 packet (no fragmentation) N packets if B > mps,jB mps kpackets (fragmentation) if B≤mps, (B+header) bits if B > mps, (N−1) packets with (mps +header) bits and X bits 1 packet with (B−(N−1) ∗mps +header) bits correspond to the data packet shown in Figure 4.8. Thus, for each region Npackets of Xbits are generated every RTseconds. It is important to highlight that when the packets are segmented, the last packet may have smaller size than the rest. These parameters, shown in Table 4.8, have different values depending on the type of region and the selected parameters in the control information (see Tables 4.3-4.5). For the 2-D modes and the auxiliary video a header of 4 bits is added to indicate how many frames are codified every RTseconds. This header is necessary because the last time that data of a region is codified the data can contain fewer number of frames than the rest, which have resolution frames. It is important to note that if some error control method is applied, the number of transmitted bits may increase and the data packets may change. 4.3.4 ETP Working Procedure The transmitter and the receiver protocol flowcharts are shown in Figure 4.9. If an error control method is added to some region, the flowcharts will change for the respective data connections. The protocol performance for the transmitter, see Figure 4.9a, is now described. •Control: the transmitter opens the control connection to send the initial configuration, where all the initial settings are defined. Then, a configuration packet is sent every time that the configuration changes, some region is added or deleted, the text or auxiliary image regions change, or the visualization ends. When a region with data connection is deleted, the data connection corresponding to that region is closed. After sending a configuration packet creating a new region with data connection, the data connection is opened after receiving the confirmation of the receiver that the data connection has been opened. After sending a configuration packet ending the transmission, all the connections are closed when the confirmation from the receiver arrives. The transmission has then finished. •Data: for every region with data connection that belongs to the activated mode, the data packets are sent through its data connection each RTseconds, except when the visualization is stopped. The protocol performance for the receiver, see Figure 4.9b, is described below.
Chapter 4. Echocardiogram Transmission in Real-time 113 •Control: the receiver opens the control connection and listens until a confirmation packet is received. The data packets cannot be received until the first confirmation packet arrives with the initial configuration and the data connections are opened. The configuration packets can indicate the end of the transmission in which case a control packet is sent with all the closing confirmations (see Figure 4.6) and all the connections are closed. When new regions are added, the data connections are opened and a confirmation packet is sent with the opening confirmations (see Figure 4.6). If regions are deleted, the connection for those regions are closed. The text or image regions can be changed. All the parameters and new text and image regions are given to the screen in order to visualize the echocardiogram. •Data: two types of data connections can be distinguished, those for ultrasound regions and those for other regions. The only difference between them is that when an ultrasound data packet is received, its mode is activated, and consequently the regions that belong to the mode are activated too. The received data packets are stored in the region buffer when a new data packet arrives. The buffer time starts the countdown when the first data packet of a mode is received. The buffer time value is sent in the configuration packets. When the buffer time expires each RTseconds, Npackets (see Table 4.8) are picked up for the activated regions to be decoded and the monitoring process begins for these regions. Each frame time the regions are picked from the monitoring buffer and are displayed on the screen.
114 4.3. Protocol for Echocardiogram Transmission in Real-time WAIT coding packet segmentation DATA send DATA · · · open control connection XML coding send CONTROL del? close data connections CONTROL open data connection/s close all connection/s CONTROL del no received close received open del yes change configuration RT region Z 1 packet B bits N packet X bits (a) ETP tranmitter flowchart. XML decoding end? Send close & Close all connections Open data connection/s Send ACK connection/s Delete connection/s Text image decoding CONTROL Set parameters end yes add region/s del region/s text, image parameters no LISTEN Open control connection received control packet LISTEN received control packet Active mode Buffer region 1 Region 1 ultrasound DATA received region 1 Decoding Monitoring buffer region 1 RT region 1 · · · Buffer region Z Region Z no ultrasound DATA received region Z Decoding Monitoring buffer region Z RT region Z screen text, image parameters region 1 (1 frame) region Z (1 frame) (b) ETP receiver flowchart. Figure 4.9: ETP transmitter and receiver flowchart. The dotted line represents the information flow between different parts, and the solid lines the execution procedure. The control connection is in white and the data connections in blue.
Chapter 4. Echocardiogram Transmission in Real-time 115 4.4 Error Control Method for Transmission over Wireless Channels Wireless channels are error prone, band limited and time varying. Consequently, it is necessary to introduce error control methods in order to guarantee clinical quality on reception. In [128] we proposed a Reliable Clinical Video Transmission Protocol (RCVTP). RCVTP was designed for clinical video encoded with a 3-D SPIHT algorithm, and the transmission of 2-D echocardiogram modes was tested over 3Gand WiMAX channels. In this Chapter, the State Error Control Method (SECM) is proposed based on the technique described in [128] but with some added improvements. Different configurations are proposed for the data connections of ETP that correspond to different regions. The characteristics of the encoded regions are taken into account. In [128] an error control protocol with two states that adapts to the channel conditions was proposed for the 2-D modes. The first state is used when there are few packet losses in the channel. In this state, the retransmissions mechanism is used because the use of error correction codes in these channel conditions would use a greater amount of bits than is necessary and the extra transmitted bits required would be unjustified. When more than a certain percentage of errors occur, the model turns into the second state in which both retransmissions and FEC techniques are used. The error correction code is now worth the extra amount of transmitted bits, because errors occur with the retransmissions mechanism. In this way, retransmissions adapt the bits used to the amount needed and the FEC code provides extra protection. The FEC uses Reed−Solomon (RS) code, which is a systematic block code [130]. 4.4.1 SECM Working Procedure After having explored error control methods developed over recent years, SECM is now proposed. A similar error control method to that proposed in [128] has been designed, but introducing a three states model (see Figure 4.10). Instead of using FEC and retransmissions techniques at the same time in the second state, only FEC is used. However, this has states with different FEC codes in order to adapt the bits used to the amount needed so as not to produce more delay as a result of the retransmissions. The percentage of blocks not successfully received is the number of blocks among the last 100 that have not been received successfully. •State 1. This is the initial state. A FEC code is used to provide protection in case of transmission errors. However, this FEC code may not be enough. If the percentage of blocks not successfully received is more than h21 the model turns into state 2. On the other hand, if errors do not occur in the h10 last blocks, the bandwidth is not efficiently used and the model turns into state 0. This state is the initial state because it is conservative, in case errors occur. •State 0. This is the state used when no errors have occurred in the last h10 blocks. In the event that errors start occurring and the percentage of blocks not successfully received exceeds a certain amount, the model turns into state 1. In this state retransmissions are used to recover the occasional errors and to use the bandwidth efficiently. •State 2. This is the state when the FEC code used in state 1 is not sufficient to recover the transmission errors and another FEC code is used with greater protection. In the event that
116 4.4. Error Control Method for Transmission over Wireless Channels errors reduce in number and the percentage of blocks that are not successfully received are less than h21, the model returns to the initial state, state 1. More states with FEC codes having greater protection can be added to the model. The last two states can even be removed, leaving only the first state with retransmissions. The three states model is the basic proposal. The transaction parameters (see h01,h10,h12 and h21 in Figure 4.10 for the three states model) are selected according to the display recommendations. Two more transaction parameters are added for each added state, and two transaction parameters are removed for each state removed from the three states model. Figures 4.11 and 4.12 show the coding and decoding respectively for SECM and the three states model. N,Kand resolution parameters are defined in Table 4.8 for ETP. Different methods are applied depending on the state. The RS code is applied to Npackets that correspond to a block, so that they do not cause more coding delay and to avoid the effects of burst errors. The RS code generates Kor Wpackets, depending on the state. Nis the number of packets of the coded video and (N−K) or (N−W) are the number of packets of the correction. N,Kor Wpackets are transmitted through the network, but Ypackets arrive at the receiver. If Nor more packets arrive correctly, the RS code is able to recover the lost packets and the video is visualized with the recommended transmission rate. Otherwise, the RS code is not able to recover the lost packets and the video is visualized with a lower transmission rate. The SECM parameters that have to be selected by the user are illustrated in Table 4.9. State 1 FEC 1 State 2 FEC 2 State 0 > h01 % blocks not successfully received h10 last blocks received without errors > h12 % blocks not successfully received < h21 % blocks not successfully received Figure 4.10: SECM three states model. SPIHT codec fragments of Xbytes state? RS FEC 1 RS FEC 2 RTtime resolution block Bbytes Npackets Kpackets Wpackets State 1 State 2 State 0 Npackets Figure 4.11: SECM encoder for three states model. In order to support retransmissions, a new type of data packet has to be introduced. The acknowledge (ACK) packets (see Figure 4.13) are used to acknowledge the reception of data packets and to indicate the actual state from receiver to transmitter. The length of the sequence number for the ACK depends on the type of data packets. For ETP the length depends on the type of region: 8 bits for the sound, ECG, auxiliary video and sweep ultrasound, and 16 bits for 2-D ultrasound.
Chapter 4. Echocardiogram Transmission in Real-time 117 RS Y < N SPIHT decodec SPIHT decodec Ypackets No Yes less than Npackets Npackets resolution frames Figure 4.12: SECM decoder for three states model. Table 4.9: SECM configuration parameters. Parameter Content States Number of states in the model. By default this value is 3. Transactions Transaction parameters between the states. This values correspond to h01,h10,h12 and h21 in Figure 4.10 for three states model. There are two parameters for each state except for the first state. No. packets Number of packets per resolution time that are send for each state. There are indicated as many number of packets as states in the model. This values correspond to N,Kand Win Figure 4.11. The SECM working procedure for transmitter and receiver are depicted in Figures 4.14 and 4.15. If SECM is used for the regions of ETP, a SECM transmitter and receiver have to be added for every region. The encoder and decoder are shown in Figures 4.14 and 4.15 for the three states model. The retransmission working procedure for the transmitter and the receiver, which works for state 1, are explained in the following paragraphs. ACK no. 1 ACK no. 2 ACK no. 3 ACK no. 4 State (8 or 16 bits) (8 or 16 bits) (8 or 16 bits) (8 or 16 bits) (8 bits) Figure 4.13: ACK packets. •Transmitter. The transmitter has three possible events: each RTsecond, when an ACK arrives and when retransmission time-out (RTO) expires. –For each frame resolution time (RT), the ACK count is updated to the maximum number of retransmissions possible, and the data packets are generated. In order to limit the amount of transmitted bits used in channels with a high incidence of error, in which many retransmissions may be required, a maximum number of retransmissions that can be made in the frame resolution time according to the bandwidth of the network has been established. Each data packet has a sequence number that identifies it. The data packets are sent and stored in a retransmission buffer awaiting acknowledgment. A RTO timer is set each time a data packet is sent.
118 4.4. Error Control Method for Transmission over Wireless Channels WAIT update state, RTT and RTO RTO deactivated update ack count RTO starts ENCODER Retransmission buffer send DATA T < Tup +TB Retransmission finished acks send <ack count? Update acks, retransmission RTregion Z send data packets store data packets ACK delete data packets RTO expires Yes No Yes No send data packets delete data packets Figure 4.14: SECM transmitter working procedure. LISTEN Screen Activate mode and ack transmission Buffer region Z DECODER Update state Monitoring buffer region Z RT region Z received region Z store data packets data packets region Z (1 frame) Figure 4.15: SECM receiver working procedure. –Each time an ACK packet is received, the state information is picked up from the ACK packet, the RTO timers associated to the acknowledged data packet are deactivated, the data packets are removed from the retransmission buffer, and the round-trip time (RTT), transmission time coming and going, is updated. The RTO time is calculated as follows RT Oi= 1.1·RT T ifRT T ≤RTOi−1; 0.9·RT Oi−1ifRT T > RT Oi−1. (4.2) Low RTO gives low delay, but RTO values lower than RTT will lead to the duplication of packets in reception and saturate the network. Thus, the RTO value is configured to decrease each time that an ACK packet is received, but only if the RTT is higher than the RTO.
Chapter 4. Echocardiogram Transmission in Real-time 125 4.17, no fragments are visualized with guaranteed clinical quality for the configuration without distinguishing modes. The number of fragments without guaranteed clinical quality are the total number of fragments for each echocardiogram (see number of videos in Table 3.5). The number of fragments without guaranteed clinical quality using ETP has dropped by at least 10 fragments, 53 % of the total fragments. Although the echocardiograms are visualized with guaranteed clinical quality for higher percentages of time when ETP is used than without using ETP, between 16 % and 25 % of the time the available echocardiograms are visualized without guaranteed clinical quality using ETP. Furthermore, between 2 and 9 echocardiogram fragments are not visualized with guaranteed clinical quality using ETP. Table 4.15: Transmitted bandwidth without using SECM. Configuration Bandwidth per echocardiogram (kbps) 123456789 without ETP 202 202 202 202 202 202 202 202 202 with ETP 147 123 139 159 174 159 121 171 187 Table 4.16: Percentage of time with guaranteed clinical quality without using SECM. Configuration Percentage of time per echocardiogram 123456789 without ETP 76.58 75.76 83.61 77.13 74.92 75.91 77.64 74.65 74.54 with ETP 90.47 82.94 93.22 87.73 92.2 88.52 94.21 93.77 80.7 Table 4.17: Number of fragments without guaranteed clinical quality without using SECM. Configuration No. of fragments per echocardiogram 12345678 9 without ETP 18 20 19 25 22 25 25 41 20 with ETP 56935852 9 In Figures 4.17,4.18 and 4.19 the effective bandwidth over the time is shown without using ETP for echocardiograms numbers 3, 7 and 9, respectively. The effective bandwidth when the fragments are not visualized with sufficient clinical quality is shown in red and when the fragments are visualized with sufficient clinical quality is shown in blue. The expected effective bandwidth for each fragment in order to visualize the echocardiogram with sufficient clinical quality is shown in gray. We can observe in these figures that the case without using ETP (see Figure 4.17) is the most affected by errors since modes are not distinguished and the highest bandwidth is transmitted independently of the operation mode. In the case where the operation modes are distinguished, (see Figures 4.18 and 4.19) the 2-D modes, the modes with the highest transmission rate, are more
126 4.5. Results and Discussion for Echocardiogram Transmission Figure 4.17: Effective bandwidth over the time without using SECM or ETP for echocardiogram number 3 is drawn in blue and the expected effective bandwidth to visualize the echocardiogram with sufficient clinical quality is drawn in gray. affected by the errors of the channel than the other modes with lower transmission rates. Note that there are fragments that are visualized with sufficient clinical quality although the effective bandwidth is lower than the expected bandwidth due to display recommendations. 4.5.2.2 Results for Echocardiogram Transmission with SECM Different error control methods for the 2-D ultrasound region are compared in this section in order to identify which is the most suitable. Different configurations and modifications to the basic three states SECM have been tested, using one-state, two-state and three-state models. Furthermore, FEC and retransmission techniques have been tested in the different states and even working both techniques together. When ETP is used, the error control method for the sweep ultrasound regions is SECM with one state model with retransmissions. The nine available echocardiograms have been transmitted using the mobile scenario, with and without using ETP and using the following error control methods: •One state: –FEC 1 (40% of protection). –FEC 2 (50% of protection). •Two states: –State 1: ACK, state 2:FEC 1 (40% of protection).
Chapter 4. Echocardiogram Transmission in Real-time 127 Figure 4.18: Effective bandwidth over the time without using SECM and using ETP for echocardiogram number 7. The effective bandwidth when the fragments are not visualized with sufficient clinical quality is shown in red and when the fragments are visualized with sufficient clinical quality is shown in blue. The expected effective bandwidth for each fragment in order to visualize the echocardiogram with sufficient clinical quality is shown in gray.
128 4.5. Results and Discussion for Echocardiogram Transmission Figure 4.19: Effective bandwidth over the time without using SECM and using ETP for echocardiogram number 9. The effective bandwidth when the fragments are not visualized with sufficient clinical quality is shown in red and when the fragments are visualized with sufficient clinical quality is shown in blue. The expected effective bandwidth for each fragment in order to visualize the echocardiogram with sufficient clinical quality is shown in gray.
Chapter 4. Echocardiogram Transmission in Real-time 129 –State 1: ACK, state 2:FEC 2 (50% of protection). –State 1: ACK, state 2: ACK + FEC 1 (40% of protection). –State 1: ACK, state 2: ACK + FEC 2 (50% of protection). •Three states: –State 1: ACK, state 2:FEC 1 (40% of protection), state 3:FEC 2 (50% of protection). Tables 4.18 and 4.19 show the transmitted bandwidth and the number of fragments without guaranteed clinical quality for the different error control methods and the nine available echocardiograms using ETP. When no fragments are visualized without guaranteed clinical quality, then 100 % of the time the echocardiogram is visualized with guaranteed clinical quality. As we can see, the proposed error control method with three states uses the lowest bandwidth displaying all the fragments with adequate clinical quality for the nine echocardiograms. All the fragments are correctly visualized with only one state and both FEC codes. The drawback of using only one state is that more bandwidth than necessary is used when the channel introduces few errors. In the case of the two states model, the only configuration which achieves all the fragments with adequate clinical quality is that using the most protective FEC code. When a less protective FEC code is used, less bandwidth is transmitted, but not all the fragments are visualized correctly for the nine echocardiograms. When retransmissions and the FEC techniques are used at the same time in the second state, as in [128], more bandwidth is used and more fragments without clinical quality are displayed than when using the FEC technique alone, because a longer buffer time is necessary to recover the errors with the retransmissions. The best results are obtained using SECM and the proposed three states model because the state model adapts to the channel conditions. All the fragments are visualized with adequate clinical quality and the transmitted bandwidth ranges from 154 kbps to 244 kbps for the available echocardiograms. Table 4.18: Transmitted bandwidth for various error control methods and ETP. Methods Bandwidth per echocardiogram (kbps) 123456789 FEC 1 203 168 191 222 241 221 163 238 260 1 state FEC 2 216 179 204 237 258 236 173 255 278 ACK, FEC 1 187 156 173 215 222 204 153 221 240 ACK, FEC 2 197 162 181 228 232 213 162 233 251 ACK, ACK+FEC 1 201 165 181 232 235 215 162 234 255 2 states ACK, ACK+FEC 2 207 170 187 245 244 225 170 244 265 3 states ACK, FEC 1, FEC 2 189 157 174 225 226 206 154 222 244 Figures 4.20 and 4.21 show the transmitted and effective bandwidth over the time for the proposed transmission method using ETP and SECM with three states for two different echocardiograms. If we compare these figures with Figures 4.18 and 4.19 we can see that more bandwidth is used in the modes with more errors, where the highest bandwidth is used due to the FEC code. Therefore, the three states model adapts to the channel conditions.
130 4.5. Results and Discussion for Echocardiogram Transmission Figure 4.20: Effective and transmitted bandwidth over the time with ETP and SECM for echocardiogram number 7. The effective bandwidth is shown in blue and the transmitted bandwidth in green. Figure 4.21: Effective and transmitted bandwidth over the time with ETP and SECM for echocardiogram number 9. The effective bandwidth is shown in blue and the transmitted bandwidth in green.
Chapter 4. Echocardiogram Transmission in Real-time 131 Table 4.19: Number of fragments without guaranteed clinical quality for various error control methods and ETP. Methods Number of fragments per video 12345678 9 FEC 1 00000000 0 1 state FEC 2 00000000 0 ACK, FEC 1 00000001 0 ACK, FEC 2 00000000 0 ACK, ACK+FEC 1 00100001 0 2 states ACK, ACK+FEC 2 00000201 0 3 states ACK, FEC 1, FEC 2 00000000 0 Table 4.20 shows the transmitted bandwidth and the maximum number of fragments without guaranteed clinical quality of the nine echocardiograms for various error control methods when ETP is not used (all the echocardiograms are considered as 2-D mode). The transmitted bandwidth is the same for all the echocardiograms since it is independent of the operation mode distribution without using ETP. Without using ETP more bandwidth is transmitted than using ETP, and consequently the echocardiogram visualization is more affected by the errors. The only error control methods that guarantee adequate clinical quality for the whole echocardiogram is when FEC 2 is used and with the 3 states model. For these two methods the bandwidth used is similar because for the three states method the model is in state 3 almost all of the time when the FEC 2 code is applied. The transmitted bandwidth for the three states model is 308 kbps for all the echocardiograms. Table 4.20: Transmitted bandwidth and maximum number of fragments without guaranteed clinical quality for various error control methods and without using ETP. Methods Bandwidth # fragments FEC 1 286 3 1 state FEC 2 310 0 ACK, FEC 1 287 3 ACK, FEC 2 307 2 ACK, ACK+FEC 1 313 2 2 states ACK, ACK+FEC 2 335 2 3 states ACK, FEC 1, FEC 2 308 0 4.5.3 Discussion for Echocardiogram Transmission The use of ETP allows a saving in the transmitted bandwidth, as is shown in Table 4.15, since the visualization characteristics are taken into account in the compression and transmission of the echocardiograms. The saving of bandwidth depends on the mode distribution of the echocardio-
132 4.5. Results and Discussion for Echocardiogram Transmission grams. The 2-D modes present the highest transmission rate while the sweep modes present a very low transmission rate. Hence the fact that a higher percentage of 2-D modes leads to higher transmitted bandwidth. The modes with lower transmission rates are less affected by the errors introduced by the channel, as is shown in Figures 4.17,4.18 and 4.19. Therefore, the lower the transmitted bandwidth, the higher is the percentage of time with guaranteed clinical quality, as shown in Tables 4.15 and 4.16. If the regions and modes are not taken into account, the highest bandwidth has to be used for all the modes, the echocardiograms being more affected by the transmission errors. The percentage of time with adequate clinical quality using ETP increases between 6 % and 19 %, while the number of fragments without adequate clinical quality using ETP decreases between 10 and 39 compared to not using ETP, as is shown in Tables 4.16 and 4.17, respectively. The percentage of time with guaranteed clinical quality depends on the error distribution and the time distribution of the operation modes, especially on the time distribution of the 2-D modes. The number of fragments that are visualized without guaranteed clinical quality not only depends on the error distribution but also on the number of fragments that correspond with 2-D modes. Although the echocardiograms are less affected by errors when ETP is used, there are still some parts that are visualized without guaranteed clinical quality when the channel introduces errors. Therefore, an error control method is necessary in order to guarantee clinical quality when transmission errors occur. SECM has been designed to adapt the error control method to the channel conditions with the proposed configuration. Error control methods increase the transmitted bandwidth in order to protect against errors. If FEC is used in a channel without errors, the bandwidth is inefficiently used. Retransmissions can cause long delay when many errors occur. Thus, in the proposed state model the retransmission method is used when few errors occur. When retransmissions are not sufficient, the FEC method is used with different correction codes. The selected configuration for SECM has been demonstrated to adapt to the channel conditions, as is shown in Figures 4.20 and 4.21. Furthermore, if the proposed SECM configuration is used, the echocardiograms are visualized with adequate clinical quality for a mobile network with WiMAX access and representative settings, as can be seen in Tables 4.19 and 4.20. However, without SECM, the available echocardiograms are visualized without guaranteed clinical quality for between 16.39 % and 25.46 % of the time, as shown in Table 4.16. If both techniques ETP and SECM are used together, a saving in the transmitted bandwidth is achieved compared to not using ETP. Furthermore, all the available echocardiograms are received with guaranteed clinical quality with the proposed SECM configuration. The bandwidth saving depends on the modes distribution of the echocardiogram and on the error distribution. For a mobile network with WiMAX access and representative settings, the transmitted bandwidth is between 154 kbps and 244 kbps for the available echocardiograms , as shown in Table 4.18. However, if ETP is not used and the echocardiogram is transmitted without distinguishing modes, the transmitted bandwidth is 308 kbps, as can be seen in Table 4.20. Thus there is a bandwidth saving of up to 154 kbps, 100 % of bandwidth saving. Without ETP the operation modes are not distinguished, all the modes are considered as 2-D modes and consequently more bandwidth is transmitted. Furthermore, blocks are not successfully received due to the fact that 2-D modes are more affected by errors and that the model for all the echocardiograms is in the last state (with more protection) almost all the time, which transmits more packets. A comparison of the results with previous works (see Table 2.3) shows that the bandwidths using
Chapter 4. Echocardiogram Transmission in Real-time 133 aWiMAX channel in the previous works are higher than that used with the proposed system, even for clinical videos with lower resolution. Comparing the results for clinical videos with a resolution similar to the echocardiograms used in this Thesis and having acceptable clinical quality on reception reveals that the saving in the bandwidth used in the present work is greater than 1 Mbps. The saving is due to several factors: •The compression method based on regions takes into account the visualization characteristics of the echocardiograms, hence the echocardiogram is compressed efficiently. •The compression recommendations give the minimal transmission rate to guarantee clinical quality for the ultrasound regions of each operation mode. •ETP allows the transmission by regions, and consequently compresses each region with its minimal recommended bandwidth. •SECM adapts the error control method to the channel conditions using the bandwidth more efficiently. •The display recommendations establish whether the echocardiogram is visualized with clinical quality even if errors occur, and as a result not all the errors have to be recovered to be able to visualize the echocardiogram with guaranteed clinical quality. 4.6 Conclusions The main objective of the work described in this Chapter is to achieve the transmission and the display of the echocardiograms without losing diagnostic information. In order to fulfill this objective, three tasks have been carried out: the design of a clinical evaluation for the display recommendations, a proposal of a protocol for real-time transmission of echocardiograms compressed by regions, and the design of an error control method. The main conclusions resulting from these three tasks are the following: •An evaluation methodology has been designed to give recommendations for the display of echocardiograms after transmission with errors. The methodology has been specially designed for the echocardiogram compression methodology proposed in this Thesis based on visualization modes. However, any medical video can be evaluated using the evaluation designed for the 2-D modes. The proposed evaluation is easy to perform because it starts from the previous recommendations given for the transmission rate for compression. The evaluation gives the display recommendations, the time and the range of the transmission rates for the visualization of the video with a lower transmission rate than the recommended rate due to transmission errors. It has a major advantage over making an evaluation for each visualized echocardiogram after transmission. With only one evaluation we are able to know if the echocardiogram is visualized with adequate quality without the need to carry out an assessment for each transmission. Display recommendations are given for the echocardiogram compressed with the proposed method. For the 2-D modes, the ultrasound part can be visualized at any transmission rate up to 5 % of the time. For the sweep modes, up to two image slices (about 10 % of the screen) can be visualized at any transmission rate.
134 4.6. Conclusions •A protocol for the end-to-end transmission of echocardiograms in real-time has been designed, known as the Echocardiogram Transmission Protocol (ETP). ETP allows transmitting the echocardiogram compressed by regions. The regions are transmitted separately, and consequently different transmission rates and error control methods can be used depending on the clinical importance of the region and on the network. ETP can be used for transmission in any network. The simulated transmissions have demonstrated that by transmitting the echocardiogram by regions, less bandwidth is used and the echocardiogram is visualized with adequate clinical quality for a greater percentage of time than when transmitting without considering the regions and modes. •An error control method based on states is proposed, known as the States Error Control Method (SECM). SECM adapts to the channel conditions using different states depending on the errors occurring. Furthermore, different numbers of states can be set depending on the data and the network characteristics. SECM has been demonstrated to adapt the bandwidth used to the channel conditions and to guarantee quality on reception. This method can be used for any data to be transmitted over error prone channels.
Appendix A Document Type Definition example for ETP <?xml version="1.0" encoding="UTF-8"?> <!ELEMENT configuration (tsize?,device?,delay?,end?,region+)> <!ELEMENT tsize EMPTY> <!ATTLIST tsize wCDATA #REQUIRED> <!ATTLIST tsize hCDATA #REQUIRED> <!ELEMENT device (#PCDATA)> <!ELEMENT delay (#PCDATA)> <!ELEMENT end EMPTY> <!ELEMENT region ((ultr|vid|img|text|sound|ecg), position?, size?, codification?, mps?)> <!ATTLIST region reliability CDATA #IMPLIED> <!ATTLIST region id ID #REQUIRED> <!ATTLIST region mode CDATA #REQUIRED> <!ATTLIST region delete (yes|no) "no"> <!ELEMENT ultr (calibration?)> <!ATTLIST ultr color (yes|no) "no"> <!ATTLIST ultr type (sweep|2D) "2D"> <!ATTLIST ultr fps CDATA #IMPLIED> <!ATTLIST ultr pps CDATA #IMPLIED> <!ATTLIST ultr bitrate CDATA #IMPLIED> <!ATTLIST ultr video (stop|forward|back) #IMPLIED> <!ATTLIST ultr resolution CDATA #IMPLIED> <!ELEMENT calibration (t?,d?,v?)> <!ELEMENT t (#PCDATA)> <!ATTLIST t measurements CDATA #IMPLIED> <!ELEMENT d (#PCDATA)> <!ATTLIST d measurements CDATA #IMPLIED> <!ELEMENT v (#PCDATA)> <!ATTLIST v measurements CDATA #IMPLIED> 141
142 <!ELEMENT vid EMPTY> <!ATTLIST vid color (yes|no) "no"> <!ATTLIST vid fps CDATA #IMPLIED> <!ATTLIST vid bitrate CDATA #IMPLIED> <!ATTLIST vid video (stop|forward|back) #IMPLIED> <!ATTLIST vid resolution CDATA #IMPLIED> <!ELEMENT img EMPTY> <!ATTLIST img color (yes|no) "no"> <!ATTLIST img bpp CDATA #IMPLIED> <!ATTLIST img codim CDATA #IMPLIED> <!ELEMENT text (#PCDATA)> <!ATTLIST text fsize (8|10|12) "8"> <!ELEMENT sound EMPTY> <!ATTLIST sound configuration CDATA #IMPLIED> <!ATTLIST sound channels CDATA #IMPLIED> <!ATTLIST sound bitrate CDATA #IMPLIED> <!ATTLIST sound sampler CDATA #IMPLIED> <!ELEMENT ecg EMPTY> <!ATTLIST ecg cod (1D|2D) "2D"> <!ATTLIST ecg der CDATA #IMPLIED> <!ATTLIST ecg block (128|256|512|1024) "128"> <!ATTLIST ecg fact (3|4|5|6) "4"> <!ATTLIST ecg sf CDATA #IMPLIED> <!ATTLIST ecg time CDATA #IMPLIED> <!ATTLIST ecg bpp CDATA #IMPLIED> <!ATTLIST ecg ppslice CDATA #IMPLIED> <!ATTLIST ecg width CDATA #IMPLIED> <!ELEMENT position EMPTY> <!ATTLIST position xCDATA #REQUIRED> <!ATTLIST position yCDATA #REQUIRED> <!ELEMENT size EMPTY> <!ATTLIST size wCDATA #REQUIRED> <!ATTLIST size hCDATA #REQUIRED> <!ELEMENT codification (#PCDATA)> <!ELEMENT mps (#PCDATA)>
Bibliography [1] (2013) World Health Organization. [Online]. Available: http://www.who.int [2] A. S. Go and et al., “Executive summary: heart disease and stroke statistics–2013 update: a report from the American Heart Association.” vol. 127, no. 1, pp. 143–152, Jan. 2013. [3] “Defunciones seg´un la Causa de Muerte,” 2013, instituto Nacional de Estad´ıstica (INE), http://www.ine.es/prensa/np767.pdf. [4] H. Oh, C. Rizo, M. Enkin, and A. Jadad, “What is eHealth (3): a systematic review of published definitions,” J Med Internet Res, vol. 7, no. e1, 2005. [5] R. Wootton and J. Craig, Introduction to Telemedicine: Royal Society of Medicine Press, 1999. [6] C. Pagliari, D. Sloan, P. Gregor, F. Sullivan, D. Detmer, J. P. Kahan, W. Oortwijn, and S. MacGillivray, “What is eHealth (4): a scoping exercise to map the field,” J Med Internet Res, vol. 7, no. e9, 2005. [7] A. G. Ekeland, A. Bowes, and S. Flottorp, “Effectiveness of telemedicine: a systematic review of reviews,” International Journal of Medical Informatics, vol. 79, no. 11, pp. 736–71, 2010. [8] C. Pattichis, E. Kyriacou, S. Voskarides, M. Pattichis, R. Istepanian, and C. Schizas, “Wireless telemedicine systems: an overview,” Antennas and Propagation Magazine, IEEE, vol. 44, no. 2, pp. 143 –153, apr 2002. [9] D. Giansanti and S. Morelli, “Digital tele-echocardiography: a look inside,” Ann Ist Super Sanita, vol. 45, no. 4, pp. 357–62, 2009. [10] G. Hooper, P. Yellowlees, T. Marwick, P. Currie, and B. Bidstrup, “Telehealth and the diagnosis and management of cardiac disease.” J Telemed Telecare, vol. 7, no. 5, pp. 249–56, 2001. [11] R. O. Bonow, D. L. Mann, D. P. Zipes, and P. Libby, Braunwald’s Heart Disease: A Textbook of Cardiovascular Medicine. Elsevier, 2011. [12] A. Evangelista and et al., “European Association of Echocardiography recommendations for standardization of performance, digital storage and reporting of echocardiographic studies,” European Journal of Echocardiography, vol. 9, pp. 438–448, 2008. [13] M. Picard and et al., “American Society of Echocardiography recommendations for quality echocardiography laboratory operations,” Journal of the American Society of Echocardiography, vol. 24, no. 1, pp. 1 –10, January 2011. 143
144 BIBLIOGRAPHY [14] S. M. Bierig, D. Ehler, M. L. Knoll, and A. D. Waggoner, “Minimum standards for the cardiac sonographer: A position paper,” Journal of the American Society of Echocardiography, vol. 19, no. 5, pp. 471–474, May 2006. [15] M. Ukrit, A.Umamageswari, and Dr.G.R.Suresh, “Article: A Survey on Lossless Compression for Medical Images,” International Journal of Computer Applications, vol. 31, no. 8, pp. 47– 50, October 2011, published by Foundation of Computer Science, New York, USA. [16] T. H. Karson, R. C. Zepp, S. Chandra, A. Morehead, and J. D. Thomas, “Digital storage of echocardiograms offers superior image quality to analog storage, even with 20:1 digital compression: results of the Digital Echo Record Access Study.” J Am Soc Echocardiogr, vol. 9, no. 6, pp. 769–78. [Online]. Available: http://www.biomedsearch.com/ nih/Digital-storage-echocardiograms-offers-superior/8943436.html [17] F. Y. Han, B. Soong, D. Watson, and J. Whitehall, “Realtime fetal ultrasound by telemedicine in Queensland. A successful venture?” J. Telemed. Telecare, vol. 7, no. 2, pp. 7–11, December 2001. [18] S. K. Yoo and et al., “Performance of a web-based, realtime, tele-ultrasound consultation system over high-speed commercial telecommunication lines,” J. Telemed. Telecare, vol. 10, no. 3, pp. 175–179, June 2004. [19] J. D. Thomas and M. L. Main, “Digital echocardiographic laboratory: where do we stand?” J Am Soc Echocardiogr, vol. 11, no. 10, pp. 978–83, 1998. [Online]. Available: http://www. biomedsearch.com/nih/Digital-echocardiographic-laboratory-where-do/9804104.html [20] J. D. Thomas, D. B. Adams, S. DeVries, D. Ehler, N. Greenberg, M. Garcia, L. Ginzton, J. Gorcsan, A. S. Katz, A. Keller, B. Khandheria, K. B. Powers, C. Roszel, D. S. Rubenson, and J. Soble, “Guidelines and Recommendations for Digital Echocardiography,” Journal of the American Society of Echocardiography, vol. 18, pp. 287–97, March 2005. [21] H. Feigenbaum, “Digital recording, display, and storage of echocardiograms,” J Am Soc Echocardiogr, vol. 1, no. 5, pp. 378–83, 1988. [22] S. Mobarek, Y. Gilliland, A. Bernal, J. Murgo, and J. Cheirif, “Is a Full Digital Echocardiography Laboratory Feasible for Routine Daily Use?” Echocardiography, vol. 13, no. 5, pp. 473–482, 1996. [23] “The Health Insurance Portability and Accountability Act (P.L.104-191),” 1996, Enacted by the U.S. Congress. [24] “The Personal Information Protection and Electronic Document Act,” 2000, Enacted in Canada for protection of health information against commercial use. [25] “Ley Org´anica de Protecci´on de Datos de car´acter personal (Organic Law for Protection of Personal Data),” 1999, Enacted in Spain. [26] A. Panayides, M. S. Pattichis, C. S. Pattichis, and A. Pitsillides, “A Tutorial for Emerging Wireless Medical Video Transmission Systems,” IEEE Antennas & Propagation Magazine, 2011.
BIBLIOGRAPHY 145 [27] R. Istepanian, S. Laxminarayan, and C. Pattichis, M-Health: Emerging Mobile Health Systems, ser. Topics in Biomedical Engineering. International Book Series. Springer, 2010. [Online]. Available: http://books.google.es/books?id=lZbBcQAACAAJ [28] N. Herscovici, C. Christodoulou, E. Kyriacou, M. Pattichis, C. Pattichis, A. Panayides, and A. Pitsillides, “m-Health e-Emergency Systems: Current Status and Future Directions - [Wireless corner],” IEEE Antennas and Propagation Magazine, vol. 49, no. 1, pp. 216–231, Feb. 2007. [Online]. Available: http://dx.doi.org/10.1109/map.2007.371030 [29] A. Bhagat and M. Atique, “Medical images: Formats, compression techniques and DICOM image retrieval a survey,” in Devices, Circuits and Systems (ICDCS), 2012 International Conference on, march 2012, pp. 172 –176. [30] (2013) DICOM (Digital Imaging and Communications in Medicine). [Online]. Available: http://medical.nema.org/ [31] H. Feigenbaum, “Digital echocardiography.” Am J Cardiol, vol. 86, no. 4A, pp. 2G–3G, 2000. [32] F. Miller, A. Vandome, and J. McBrewster, Digital Imaging and Communications in Medicine: Medical Imaging, File Format, Communications Protocol, Internet Protocol Suite, National Electrical Manufacturers Association, Picture Archiving and Communication System. Alphascript Publishing, 2009. [Online]. Available: http: //books.google.es/books?id=ytzYQgAACAAJ [33] J. D. Thomas, “The DICOM Image Formatting Standard: what it means for echocardiographers.” J Am Soc Echocardiogr, vol. 8, no. 3, pp. 319–27. [Online]. Available: http://www.biomedsearch.com/nih/DICOM-Image-Formatting-Standard-what/ 7640025.html [34] C. Doukas and I. Maglogiannis, “Region of Interest Coding Techniques for Medical Image Compression,” Engineering in Medicine and Biology Magazine, IEEE, vol. 26, no. 5, pp. 29 –35, sept.-oct. 2007. [35] P. Akhtar, M. Bhatti, T. Ali, and M. Muqeet, “Significance of ROI Coding using MAXSHIFT Scaling applied on MRI Images in Teleradiology-Telemedicine,” Journal of Biomedical Science and Engineering, no. 1, pp. 110–115, 2008. [36] C. Christopoulos, J. Askelof, and M. Larsson, “Efficient methods for encoding regions of interest in the upcoming JPEG2000 still image coding standard,” Signal Processing Letters, IEEE, vol. 7, no. 9, pp. 247–249, 2000. [37] M. Ansari and R. Anand, “Context based medical image compression with application to ultrasound images,” in India Conference, 2008. INDICON 2008. Annual IEEE, vol. 1, dec. 2008, pp. 28 –33. [38] K. V. Sridhar, “Implementation of prioritised ROI coding for medical image archiving using JPEG2000,” in Signals and Electronic Systems, 2008. ICSES ’08. International Conference on, 2008, pp. 239–242.
146 BIBLIOGRAPHY [39] K. Hamamoto and T. Nishimura, “Basic investigation on medical ultrasonic echo image compression by JPEG2000-availability of wavelet transform and ROI method,” in Engineering in Medicine and Biology Society, 2001. Proceedings of the 23rd Annual International Conference of the IEEE, vol. 3, 2001, pp. 2449–2452 vol.3. [40] D. Taubman and M. Marcellin, “JPEG2000: standard for interactive imaging,” Proceedings of the IEEE, vol. 90, no. 8, pp. 1336 – 1357, aug 2002. [41] A. Said and W. Pearlman, “A new, fast, and efficient image codec based on set partitioning in hierarchical trees,” Circuits and Systems for Video Technology, IEEE Transactions on, vol. 6, no. 3, pp. 243 –250, jun 1996. [42] D. H. Foos, E. Muka, R. M. Slone, B. J. Erickson, M. J. Flynn, D. A. Clunie, L. Hildebrand, K. S. Kohm, and S. S. Young, “JPEG 2000 compression of medical imagery,” pp. 85–96, 2000. [Online]. Available: +http://dx.doi.org/10.1117/12.386390 [43] A. Agarwal, A. Rowberg, and Y. Kim, “Fast JPEG 2000 decoder and its use in medical imaging,” Information Technology in Biomedicine, IEEE Transactions on, vol. 7, no. 3, pp. 184–190, 2003. [44] T. H. OH and R. BESAR, “Medical image compression using JPEG-2000 and JPEG: A comparison study,” Journal of Mechanics in Medicine and Biology, vol. 02, no. 03n04, pp. 313–328, 2002. [Online]. Available: http://www.worldscientific.com/doi/abs/10.1142/ S021951940200054X [45] (2011) Digital Imaging and Communications in Medicine (DICOM) Part 3: Information Object Definitions. [Online]. Available: http://medical.nema.org/Dicom/2011/11 03pu.pdf [46] M. Weinberger, G. Seroussi, and G. Sapiro, “The LOCO-I lossless image compression algorithm: principles and standardization into JPEG-LS,” Image Processing, IEEE Transactions on, vol. 9, no. 8, pp. 1309 –1324, aug 2000. [47] N. Tsapatsoulis, C. Loizou, and C. Pattichis, “Region of Interest Video Coding for Low bitrate Transmission of Carotid Ultrasound Videos over 3G Wireless Networks,” in Engineering in Medicine and Biology Society, 2007. EMBS 2007. 29th Annual International Conference of the IEEE, 2007, pp. 3717–3720. [48] S. P. Rao, N. S. Jayant, M. E. Stachura, E. Astapova, and A. Pearson-Shaver, “Delivering diagnostic quality video over mobile wireless networks for telemedicine,” Int. J. Telemedicine Appl., vol. 2009, pp. 1:1–1:9, Jan. 2009. [Online]. Available: http://dx.doi.org/10.1155/2009/406753 [49] M. G. Martini and C. T. E. R. Hewage, “Flexible Macroblock Ordering for Context-Aware Ultrasound Video Transmission over Mobile WiMAX,” International Journal of Telemedicine and Applications, vol. 2010, p. 14, May 2010. [50] A. Panayides, M. Pattichis, C. Pattichis, C. Loizou, M. Pantziaris, and A. Pitsillides, “Atherosclerotic Plaque Ultrasound Video Encoding, Wireless Transmission, and Quality Assessment Using H.264,” Information Technology in Biomedicine, IEEE Transactions on, vol. 15, no. 3, pp. 387 –397, may 2011.
BIBLIOGRAPHY 147 [51] C. Debono, B. Micallef, N. Philip, A. Alinejad, R. Istepanian, and N. Amso, “Cross Layer Design for Optimised Region of Interest of Ultrasound Video Data over Mobile Wimax,” Information Technology in Biomedicine, IEEE Transactions on, vol. PP, no. 99, p. 1, 2012. [52] A. Panayides, Z. Antoniou, Y. Mylonas, M. Pattichis, A. Pitsillides, and C. Pattichis, “HighResolution, Low-delay, and Error-resilient Medical Ultrasound Video Communication Using H.264/AVC Over Mobile WiMAX Networks,” Biomedical and Health Informatics, IEEE Journal of, vol. PP, no. 99, pp. 1–1, 2013. [53] (2013) The Moving Picture Experts Group website. [Online]. Available: http: //mpeg.chiariglione.org/ [54] (2012) Home of Xvid codec project. [Online]. Available: http://www.xvid.org/ [55] D. Marpe, T. Wiegand, and G. Sullivan, “The H.264/MPEG4 advanced video coding standard and its applications,” Communications Magazine, IEEE, vol. 44, no. 8, pp. 134 –143, aug. 2006. [56] A. Alesanco and et al., “A clinical distortion index for compressed echocardiogram evaluation: recommendations for Xvid codec,” Physiological Measurement, vol. 30, no. 5, pp. 429–440, June 2009. [57] P. Pedersen, B. Dickson, and J. Chakareski, “Telemedicine applications of mobile ultrasound,” in Multimedia Signal Processing, 2009. MMSP ’09. IEEE International Workshop on, oct. 2009, pp. 1 –6. [58] R. S. H. Istepanian, N. Y. Philip, and M. G. Martini, “Medical QoS provision based on reinforcement learning in ultrasound streaming over 3.5G wireless systems,” IEEE J.Sel. A. Commun., vol. 27, no. 4, pp. 566–574, May 2009. [Online]. Available: http://dx.doi.org/10.1109/JSAC.2009.090517 [59] G. Sullivan, J. Ohm, W.-J. Han, T. Wiegand, and T. Wiegand, “Overview of the High Efficiency Video Coding (HEVC) Standard,” Circuits and Systems for Video Technology, IEEE Transactions on, vol. 22, no. 12, pp. 1649–1668, 2012. [60] A. Panayides, Z. Antoniou, M. Pattichis, C. Pattichis, and A. Constantinides, “High efficiency video coding for ultrasound video communication in m-health systems,” in Engineering in Medicine and Biology Society (EMBC), 2012 Annual International Conference of the IEEE, 2012, pp. 2170–2173. [61] A. Panayides, Z. Antoniou, M. Pattichis, and C. Pattichis, “The use of H.264/AVC and the emerging high efficiency video coding (HEVC) standard for developing wireless ultrasound video telemedicine systems,” in Signals, Systems and Computers (ASILOMAR), 2012 Conference Record of the Forty Sixth Asilomar Conference on, 2012, pp. 337–341. [62] B.-J. Kim, Z. Xiong, and W. Pearlman, “Low bit-rate scalable video coding with 3-D set partitioning in hierarchical trees (3-D SPIHT),” Circuits and Systems for Video Technology, IEEE Transactions on, vol. 10, no. 8, pp. 1374 –1387, dec 2000.
148 BIBLIOGRAPHY [63] L. Kaur, R. Chauhan, and S. Saxena, “Performance improvement of the SPIHT coder based on statistics of medical ultrasound images in the wavelet domain.” J Med Eng Technol, vol. 29, no. 6, pp. 297–301. [64] H. Lalgudi, A. Bilgin, M. Marcellin, and M. Nadar, “Compression of fMRI and ultrasound images using 4D SPIHT,” in Image Processing, 2005. ICIP 2005. IEEE International Conference on, vol. 2, 2005, pp. II–746–9. [65] B. Ramakrishnan and N. Sriraam, “Compression of dicom images based on wavelets and spiht for telemedicine applications.” [66] S. Cho, D. Kim, and W. A. Pearlman, “Lossless Compression of Volumetric Medical Images with Improved Three-Dimensional SPIHT Algorithm.” J. Digital Imaging, vol. 17, no. 1, pp. 57–63, 2004. [Online]. Available: http://dblp.uni-trier.de/db/journals/jdi/jdi17.html# ChoKP04 [67] B. Ramakrishnan and N. Sriraam, “Compression of DICOM images based on wavelets and SPIHT for telemedicine applications.” [68] V. K. Bairagi and A. Sapkal, “Automated region-based hybrid compression for digital imaging and communications in medicine magnetic resonance imaging images for telemedicine applications,” Science, Measurement Technology, IET, vol. 6, no. 4, pp. 247–253, 2012. [69] A. Kassim and W. S. Lee, “Embedded color image coding using SPIHT with partially linked spatial orientation trees,” Circuits and Systems for Video Technology, IEEE Transactions on, vol. 13, no. 2, pp. 203–206, 2003. [70] J. Shapiro, ““Embedded image coding using zerotrees of wavelet coefficients”,” Signal Processing, IEEE Transactions on, vol. 41, no. 12, pp. 3445 –3462, dec 1993. [71] Y. Chen and W. A. Pearlman, “Three-Dimensional Subband Coding of Video Using the Zero-Tree Method,” in Proc. SPIE, 1996, pp. 1302–1309. [72] B.-J. Kim and W. Pearlman, “An embedded wavelet video coder using three-dimensional set partitioning in hierarchical trees (SPIHT),” in Data Compression Conference, 1997. DCC ’97. Proceedings, mar 1997, pp. 251 –260. [73] R. K. Kouassi, J.-C. Devaux, P. Gouton, and M. Paindavoine, “Application of the karhunenloeve transform for natural color images analysis,” in Signals, Systems amp; Computers, 1997. Conference Record of the Thirty-First Asilomar Conference on, vol. 2, 1997, pp. 1740–1744 vol.2. [74] K. Shen and E. Delp, “Color image compression using an embedded rate scalable approach,” in Image Processing, 1997. Proceedings., International Conference on, vol. 3, 1997, pp. 34–37 vol.3. [75] M. Saenz, P. Salama, K. Shen, E. Delp, M. S´aenz, P. Salama, K. Shen, and E. J. Delp, “An evaluation of color embedded wavelet image compression techniques,” in SPIE Conference on Visual Communications and Image Processing 99, 1999, pp. 282–293.
BIBLIOGRAPHY 149 [76] R. S. D. W. B. M. Santhi, “Inter Color Correlation Based Enhanced Color SPIHT Coder,” European Journal of Scientific Research, vol. 57, no. 4, pp. 592–600, 2011. [77] M. Antonini, M. Barlaud, P. Mathieu, and I. Daubechies, “Image coding using wavelet transform,” IEEE Transactions on Image Processing, vol. 1, no. 2, pp. 205–220, Apr. 1992. [Online]. Available: http://dx.doi.org/10.1109/83.136597 [78] J. Postel, “Transmission Control Protocol,” RFC 793 (Standard), Internet Engineering Task Force, September 1981, updated by RFCs 1122, 3168. [Online]. Available: http://www.ietf.org/rfc/rfc793.txt [79] ——, “User Datagram Protocol,” Internet Engineering Task Force, RFC 768, August 1980. [Online]. Available: http://www.rfc-editor.org/rfc/rfc768.txt [80] H. Schulzrinne, S. Casner, R. Frederick, and V. Jacobson, “RTP: A Transport Protocol for Real-Time Applications,” RFC 3550, July 2008. [81] G. Pelletier and K. Sandlund, “RObust Header Compression Version 2 (ROHCv2): Profiles for RTP, UDP, IP, ESP and UDP-Lite,” RFC 5225 (Proposed Standard), Internet Engineering Task Force, April 2008. [Online]. Available: http://www.ietf.org/rfc/rfc5225.txt [82] G. Pelletier, K. Sandlund, and L.-E. Jonsson, “RObust Header Compression (ROHC): A Profile for TCP/IP (ROHC-TCP,” RFC 6846 (Proposed Standard), Internet Engineering Task Force, January 2013. [Online]. Available: http://tools.ietf.org/html/rfc6846 [83] T. Dierks and E. Rescorla, “The Transport Layer Security (TLS) protocol version 1.2.” RFC 5246, 2008. [84] Y. Wang, S. Wenger, J. Wen, and A. Katsaggelos, “Error resilient video coding techniques,” Signal Processing Magazine, IEEE, vol. 17, no. 4, pp. 61–82, 2000. [85] M. Schier and M. Welzl, “Optimizing Selective ARQ for H.264 Live Streaming: A Novel Method for Predicting Loss-Impact in Real Time,” Multimedia, IEEE Transactions on, vol. 14, no. 2, pp. 415–430, 2012. [86] ——, “Content-aware selective reliability for DCCP video streaming,” in Multimedia Computing and Information Technology (MCIT), 2010 International Conference on, 2010, pp. 53–56. [87] A. Husz´ak and S. Imre, “Source controlled semi-reliable multimedia streaming using selective retransmission in DCCP/IP networks,” Comput. Commun., vol. 31, no. 11, pp. 2676–2684, Jul. 2008. [Online]. Available: http://dx.doi.org/10.1016/j.comcom.2008.02.033 [88] L. Rizzo, “Effective erasure codes for reliable computer communication protocols,” SIGCOMM Comput. Commun. Rev., vol. 27, no. 2, pp. 24–36, Apr. 1997. [Online]. Available: http://doi.acm.org/10.1145/263876.263881 [89] X. Yang, C. Zhu, Z. G. Li, X. Lin, and N. Ling, “An unequal packet loss resilience scheme for video over the Internet,” Multimedia, IEEE Transactions on, vol. 7, no. 4, pp. 753–765, 2005.
150 BIBLIOGRAPHY [90] N. Thomos, S. Argyropoulos, N. Boulgouris, and M. Strintzis, “Robust Transmission of H.264/AVC Video using Adaptive Slice Grouping and Unequal Error Protection,” in Multimedia and Expo, 2006 IEEE International Conference on, 2006, pp. 593–596. [91] S. Cho and W. A. Pearlman, “Multilayered protection of embedded video bitstreams over binary symmetric and packet erasure channels,” J. Vis. Comun. Image Represent., vol. 16, no. 3, pp. 359–378, Jun. 2005. [Online]. Available: http://dx.doi.org/10.1016/j.jvcir.2004.08. 001 [92] P. Frossard, “FEC performance in multimedia streaming,” Communications Letters, IEEE, vol. 5, no. 3, pp. 122–124, 2001. [93] M. Chatterjee, S. Sengupta, and S. Ganguly, “Feedback-based real-time streaming over WiMax,” Wireless Communications, IEEE, vol. 14, no. 1, pp. 64–71, 2007. [94] A. Majumda, D. Sachs, I. Kozintsev, K. Ramchandran, and M. Yeung, “Multicast and unicast real-time video streaming over wireless LANs,” Circuits and Systems for Video Technology, IEEE Transactions on, vol. 12, no. 6, pp. 524–534, 2002. [95] A. Alinejad, N. Philip, and R. Istepanian, “Cross-Layer Ultrasound Video Streaming Over Mobile WiMAX and HSUPA Networks,” Information Technology in Biomedicine, IEEE Transactions on, vol. 16, no. 1, pp. 31 –39, jan. 2012. [96] 3GPP, “Overview of the GPP Release 4,” V1.1.2, 2010. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [97] ——, “Overview of the GPP Release 5,” V0.1.1, 2010. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [98] ——, “Overview of the GPP Release 6,” V0.1.1, 2010. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [99] ——, “Overview of the GPP Release 7,” V0.9.16, 2012. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [100] ——, “Overview of the GPP Release 8,” V0.2.10, 2013. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [101] ——, “Overview of the GPP Release 9,” V0.2.9, 2013. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [102] ——, “Overview of the GPP Release 10,” V0.1.8, 2013. [Online]. Available: http: //www.3gpp.org/ftp/Information/WORK PLAN/Description Releases/ [103] “IEEE standard for local and metropolitan area networks-Part 16: air interface for fixed broadband wireless access systems,” IEEE Std. 802.16e-2005, May 2006. [104] “IEEE standard for local and metropolitan area networks-Part 16: air interface for fixed broadband wireless access systems. Amendment 3: advanced air interface,” IEEE Std. 802.16m-2011, May 2011.