Full text
Universidade do Minho Escola de Engenharia Daniel Filipe Coelho Rocha Automated Garment Recognition and Assessment System Using Computer Vision and Artificial Intelligence for Blind People outubro de 2024 Automated Garment Recognition and Assessment System Using Computer Vision and Artificial Intelligence for Blind People Daniel Rocha UMinho | 2024
Daniel Filipe Coelho Rocha Automated Garment Recognition and Assessment System Using Computer Vision and Artificial Intelligence for Blind People Tese de Doutoramento Programa Doutoral em Engenharia Eletrónica e de Computadores Trabalho efetuado sob a orientação do Professor Doutor Vítor Hugo Mendes da Costa Carvalho Professora Doutora Filomena Maria da Rocha Menezes de Oliveira Soares outubro de 2024
i DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição CC BY https://creativecommons.org/licenses/by/4.0/
ii DECLARAÇÃO Nome: Daniel Filipe Coelho da Rocha Endereço eletrónico: [email protected] Telefone: 917966638 (alternativo) Bilhete de Identidade/Cartão do Cidadão: 14147147 Título da tese: Automated Garment Recognition and Assessment System Using Computer Vision and Artificial Intelligence for Blind Orientadores: Professor Doutor Vítor Hugo Mendes da Costa Carvalho e Professora Doutora Filomena Maria da Rocha Menezes de Oliveira Soares Ano de conclusão: 2024 Designação do Ramo de Conhecimento do Doutoramento: Engenharia Eletrónica e de Computadores DE ACORDO COM A LEGISLAÇÃO EM VIGOR, NÃO É PERMITIDA A REPRODUÇÃO DE QUALQUER PARTE DESTA TESE/TRABALHO. Universidade do Minho, 28/10/2024 Assinatura: __________________________________________________
iii STATEMENT OF INTEGRITY I hereby declare having conducted my thesis with integrity. I confirm that I have not used plagiarism or any form of falsification of results in the process of the thesis elaboration. I further declare that I have fully acknowledge the Code of Ethical Conduct of the University of Minho. University of Minho, 28th October 2024 Full name: Daniel Filipe Coelho Rocha Signature: ____________________________________________________________________
iv ACKNOWLEDGEMENTS Throughout this journey, I faced numerous challenges that made me question and doubt myself. However, these obstacles only served to strengthen my resilience and increase my determination to succeed, surpassing myself each day. But none of this would have been possible without the right people by my side, and for that, I want to express my gratitude to all those who accompanied me on this journey. To my supervisor, Professor Vítor Carvalho, I am deeply grateful for his guidance, valuable teachings, and for the excellence, strength, and resilience he has always demonstrated. His enthusiasm in the pursuit of knowledge was a constant source of inspiration and direction throughout this journey. To my co-supervisor, Professor Filomena Soares, I extend my thanks for her advice, high standards, and professionalism, as well as for the excellent feedback that allowed me to improve every day. To Professor Celina P. Leão, I express my gratitude for her support and generous sharing of knowledge in the field of statistics, always available with enthusiasm and joy to help. To all the professionals, especially those at ACAPO Braga, who were always willing to collaborate on the project, my sincere thanks. I would like to highlight the contribution of Mr. Filipe Azevedo, secretary of ACAPO Braga, for his full availability, dedication, and the enthusiastic way he embraced the project, always proposing improvements to benefit the blind community. To my friend Toninho, I am deeply grateful for the support and help throughout this journey. For the valuable scientific knowledge he shared with me and for the precious hours he spent with me during the toughest moments, proving that the impossible was, indeed, possible. To you, my deepest thanks for everything! To my friend João, my thanks for his constant support, his willingness to share his vast experience and knowledge, and for always being ready to help. To my family, this achievement is also yours. Together, we faced days of challenges and overcoming, and it was your unconditional support that gave me the strength to keep pursuing this dream. None of this would have been possible without the tireless dedication and sacrifices of my wife, who gave up so much to allow me to fly freely, always offering her best. Lastly, I dedicate this achievement to my daughter. From the moment you were born, my journey took on a new purpose: to be an example for you and to show you that anything is possible when we believe. Watching you grow during this time gave me the responsibility to inspire you, and I hope that one day, when you read these words, they will motivate you to pursue your own dreams with the same determination.
v AGRADECIMENTOS Ao longo deste percurso, enfrentei inúmeros desafios que me fizeram questionar e duvidar de mim mesmo. No entanto, foram esses obstáculos que fortaleceram a minha resiliência e aumentaram a minha vontade de conquistar, superando-me a cada dia. Porém, nada disto seria possível sem as pessoas certas ao meu lado, e por isso, quero expressar a minha gratidão a todos os que me acompanharam nesta jornada. Ao meu orientador, Professor Vítor Carvalho, agradeço por todo o acompanhamento, pelos ensinamentos valiosos, pela excelência, força e resiliência que sempre demonstrou. O seu entusiasmo na procura de conhecimento foi uma fonte constante de inspiração e orientação ao longo de todo este percurso. À minha coorientadora, Professora Filomena Soares, agradeço os conselhos, a exigência e o profissionalismo, bem como o excelente feedback que me permitiram melhorar a cada dia. À Professora Celina P. Leão, expresso a minha gratidão pelo apoio e pela partilha generosa de conhecimento na área de estatística, estando sempre disponível, com entusiasmo e alegria, para ajudar. A todos os profissionais, especialmente à ACAPO de Braga, que se mostraram sempre disponíveis para colaborar com o projeto, o meu sincero agradecimento. Gostaria de destacar a colaboração do Sr. Filipe Azevedo, secretário da ACAPO de Braga, pela sua total disponibilidade e dedicação, e pelo entusiasmo com que abraçou o projeto, sempre a propor melhorias que beneficiem a comunidade cega. Ao meu amigo Toninho, agradeço profundamente pelo suporte e ajuda ao longo deste percurso. Pelo valioso conhecimento científico que partilhou comigo e pelas horas preciosas que me dedicou nos momentos mais difíceis, demonstrando que o impossível era, afinal, possível. A ti, o meu muito obrigado por tudo! Ao meu amigo João, o meu agradecimento pelo apoio constante, pela disponibilidade em partilhar a sua vasta experiência e conhecimento, e por estar sempre pronto a ajudar. À minha família, esta conquista é também vossa. Juntos enfrentámos dias de desafios e superação, e foi o vosso apoio incondicional que me deu forças para continuar a perseguir este sonho. Nada disso teria sido possível sem a dedicação incansável e os sacrifícios da minha esposa, que abdicou de tanto para me permitir voar livremente, sempre oferecendo o seu melhor. Por último, dedico esta conquista à minha filha. Desde o momento em que nasceste, a minha jornada ganhou um novo propósito: ser um exemplo para ti e mostrar-te que tudo é possível quando acreditamos. Ver-te crescer durante este percurso deu-me a responsabilidade de te inspirar, e espero que, ao leres estas palavras um dia, elas te motivem a perseguir os teus próprios sonhos com a mesma determinação.
vi ABSTRACT Clothing management is one of the most significant challenges faced by blind and visually impaired people. By leveraging advanced technologies, such as computer vision and artificial intelligence (AI), this PhD work aimed at developing and validating a mechatronic device, the iSight prototype, able to assist in the identification of clothing types, colours, and conditions, primarily designed for blind and visually impaired users. Valuable insights were obtained from a nationwide survey, conducted in collaboration with the Association of the Blind and Amblyopes of Portugal (ACAPO), which revealed a high demand for technological solutions in garment management, with 95.7 % of respondents expressing willingness to adopt new technologies. The iSight prototype integrates a smart wardrobe equipped with an image acquisition system and controlled lighting, ensuring optimal image quality capture for accurate analysis. The device is controlled via a user-friendly mobile application, designed to be highly accessible and intuitive. The system's functionality was rigorously tested with 15 participants from ACAPO, encompassing both blind and low-vision people. The prototype's accuracy in identifying clothing categories and colours was highly acknowledged by users, with 60 % finding it very precise in identifying categories and 80 % in identifying colours. Additionally, 86.7% of participants rated the system's ability to detect stains and identify Near Field Communication (NFC) tags as highly effective. Finally, the obtained results demonstrated high levels of user satisfaction. Extensive statistical analyses confirmed significant positive correlations between the iSight functionalities and in users' confidence, self-esteem, well-being, and independence. These findings highlight the prototype's potential to significantly enhance the daily lives of visually impaired people. Key areas for further improvement were also identified by the users, namely the reduction of menu complexity and addition of detailed fabric information. In conclusion, this PhD work demonstrates the feasibility and effectiveness of integrating computer vision and AI technologies into a mechatronic device (smart wardrobe) to support blind and visually impaired people in garment identification and management. The iSight prototype offers a robust and user-friendly solution that significantly improves users' quality of life, offering room for future refinement and paving the way towards a truly disruptive product for smart and effective clothing management. KEYWORDS: Artificial Intelligence (AI), Assistive Technology, Blind People, Clothing Identification and Modifications, Computer Vison, Mobile Application, Smart Wardrobe.
xiii Figure 4.19: Example of a defect undetected by the YOLOv5m6 model that was subsequently identified with the aid of data augmentation: a) original image; b) predicted image from model YOLOV5m6 without augmentation; c) Predicted image from model YOLOV5m6 with augmentation (Rocha, Pinto, et al ., 2023). ........................................................................................................................................................ 74 Figure 4.20: Example of a misinterpretation of a defect: a) original image; b) predicted image from model YOLOV5m6; c) predicted image from model YOLOV5l6 (Rocha, Pinto, et al ., 2023). .......................... 75 Figure 4.21: Examples of predicted images from the YOLOv5l6: a) single stain detection; b) multiple stain detection; c) multiple hole detection; d) hole detection near the seam (Rocha, Pinto, et al ., 2023). .... 76 Figure 5.1: iSight Smart Wardrobe Prototype. .................................................................................... 80 Figure 5.2 - Schematic layout of the smart wardrobe prototype. ......................................................... 81 Figure 5.3 Nema 17 stepper motor. .................................................................................................. 82 Figure 5.4: Picture of the driver motor A4988 used in the developed prototype. ................................. 83 Figure 5.5: Picture of a Raspberry Pi 4 Model B used in the developed prototype. .............................. 84 Figure 5.6: V3 Oficial Raspberry Pi, 12MP, 120°, camera module, used in the developed prototype. . 84 Figure 5.7. ITEAD PN532 NFC reader used in the developed prototype. ............................................ 86 Figure 5.8. LED strip 5M SMD 5050, 300LEDS 60LEDS/M 12V IP20 used in the developed prototype ( Fita LED 5m SMD 5050 300leds 60leds/m 12v IP20 - Branco Frio , 2024). ..................................... 87 Figure 5.9. Picture of the 12V 60W 5A power supply incorporated in the developed prototype. ........... 88 Figure 5.10. Chip label NFCTAG213 applied to the developed prototype. ........................................... 88 Figure 5.11. Purpose-built branded clothing hanger developed for the iSight prototype. ...................... 89 Figure 5.12. Schematics of the three main phases and communication included in the iSight prototype. ........................................................................................................................................................ 90 Figure 5.13: Flowchart of wardrobe server. ....................................................................................... 91 Figure 5.14: General view of the wardrobe interior: a) LED strips applied on the left panel; b) camera placed in the middle of LED strips. .................................................................................................... 92 Figure 5.15: Reflective illumination in hole defect. This image shows how light reflected from a flat surface is affected by non-flat features like holes, which can redirect light outside the lens's acceptance angle, creating dark areas that reveal surface defects. ................................................................................. 93 Figure 5.16: Interior of the wardrobe showing the distance from the camera to the hanger and internal components. .................................................................................................................................... 94 Figure 5.17: Motor supporting the hanger. ........................................................................................ 95 Figure 5.18: NFC tag attached to the clothing label. .......................................................................... 96
xiv Figure 5.19: NFC Reader. ................................................................................................................. 96 Figure 5.20: Schematic diagram illustrating the integration and connectivity of system components with the Raspberry PI controller ................................................................................................................ 97 Figure 5.21: iSight mobile interface block diagram. ........................................................................... 99 Figure 5.22: iSight main menu. ...................................................................................................... 100 Figure 5.23: Read NFC Tag ............................................................................................................ 101 Figure 5.24: Closet Menu. .............................................................................................................. 102 Figure 5.25: Clothing list for upper clothes category. ....................................................................... 103 Figure 5.26: Detailed information corresponding to a selected clothing item. ................................... 104 Figure 5.27. Overview of the “Add new clothing item” menu: a) Description, Category, Colour and Type fields; b) NFC Tag, Size, Pattern, Washing, Style and Printed fields. ................................................. 105 Figure 5.28: Submenu of the Categories and Colours Menu. ........................................................... 106 Figure 5.29. Example of a returned result from colour analysis. ....................................................... 107 Figure 5.30. Submenu included in the Modifications menu.............................................................. 108 Figure 5.31. Example of a returned result from the Modifications Analysis. ...................................... 109 Figure 5.32. Wardrobe menu details. .............................................................................................. 110 Figure 6.1: Workflow of the iSight prototype testing and validation process. ...................................... 113 Figure 6.2: ACAPO members interacting with the iSight prototype during the testing phase, showcasing hands-on engagement: a) Participant using the mechatronic device; b) Participant using the iSight mobile application. .................................................................................................................................... 116 Figure 6.3: Distribution of professional background of participants. ................................................. 119 Figure 6.4: Age distribution of the time that visual impairment occurred among participants. ........... 121 Figure 6.5: Distribution of conditions. .............................................................................................. 122 Figure 6.6: Visual representation of common causes of visual impairment (Burton et al ., 2021) ....... 123 Figure 6.7: Word cloud summarizing the visually impaired users’ feedback. ..................................... 149 Figure 6.8 - Word cloud summarizing the feedback of blind users. ................................................... 150 Figure 6.9 - Word cloud summarizing the feedback of low vision users............................................. 151
xv LIST OF TABLES Table 2.1: Literature overview on fashion image classification works (Rocha, Soares, et al ., 2023a). .. 18 Table 2.2: Summary of advanced image segmentation techniques for clothing identification .............. 20 Table 2.3: Literature overview on textile fabric defect detection including the author, year, method, dataset, defect classes and metrics. .................................................................................................. 22 Table 4.1: Summary of available datasets for fashion category classification (Rocha, Soares, et al ., 2023a). ............................................................................................................................................ 51 Table 4.2: Main characteristics of pre-trained models (Rocha, Soares, et al ., 2023a). ........................ 54 Table 4.3: Hyperparameters of model experiments (Rocha, Soares, et al ., 2023a). ............................ 54 Table 4.4 - Test performance results (Rocha, Soares, et al ., 2023a). ................................................. 55 Table 4.5: Inference time by network architecture (Rocha, Soares, et al ., 2023a)............................... 56 Table 4.6: Test performance results with augmented data (Rocha, Soares, et al ., 2023a). ................. 56 Table 4.7: Classification report of the GoogLeNet network (Rocha, Soares, et al ., 2023a). .................. 57 Table 4.8: Hyperparameters (Image Size, Optimizer, Learning Rate, and Batch Size) of model experiments. .................................................................................................................................... 61 Table 4.9 - Main performance results of model for clothing mask segmentation (Precision, Recall and AP at IoU = 0.50). ................................................................................................................................. 61 Table 4.10: Hyperparameters for fine-tuned Mask R-CNN (Rocha, Soares, et al ., 2023a). .................. 67 Table 4.11: A report on the evaluation of Common Objects in Context (COCO) (Rocha, Soares, et al ., 2023a). ............................................................................................................................................ 67 Table 4.12: A summary of the losses associated with the model (Rocha, Soares, et al ., 2023a). ........ 67 Table 4.13: A distribution of class defects, focusing on the two main class defects of interest in this study (Rocha, Pinto, et al ., 2023). .............................................................................................................. 69 Table 4.14: Configuration of HyperParameters in Model Testing: Image Size, Optimizer, Learning Rate, and Batch Size (Rocha, Pinto, et al ., 2023). ...................................................................................... 72 Table 4.15: Main results from the fine-tuning of the models without data augmentation for defect detection (Precision, Recall and AP at IoU = 0.50) (Rocha, Pinto, et al ., 2023). ................................................ 72 Table 4.16: Main performance results of the models after introducing data augmentation for defect detection (Precision, Recall and AP at IoU = 0.50) (Rocha, Pinto, et al ., 2023). ................................. 73 Table 4.17: Performance results of each model with data augmentation and defect classification (Precision, Recall and AP at IoU = 0.50) (Rocha, Pinto, et al ., 2023). ................................................ 74
xvi Table 4.18: Inference time on test dataset for the different YOLOv5 models tested (Rocha, Pinto, et al ., 2023). .............................................................................................................................................. 76 Table 5.1: Nema 17 motor specifications ( Motor de Passo Nema 17 p/ Impressora 3D - 42BYGH48-23D , 2024). .............................................................................................................................................. 82 Table 5.2. Description of the A4988 specifications ( Driver Para Motor - A4988 , 2024). ...................... 83 Table 5.3: Module camera raspberry pi 4 specifications ( Raspberry Pi Camera Module 3 , 2024). ...... 85 Table 5.4: Specification of lens ( Raspberry Pi Camera Module 3 , 2024) . .......................................... 85 Table 5.5: Specifications of ITEAD PN532 NFC module ( PN532 , 2024). ............................................ 86 Table 5.6. NFCTAG213 label specifications ( NTAG213 , 2024)........................................................... 89 Table 6.1: Fisher's Exact Test results for associations with overall satisfaction. ................................ 134 Table 6.2: Spearman's rank correlation results. ............................................................................... 135 Table 6.3: Mann-Whitney testing results for ease of navigation and overall experience across different levels of comfort with technology. .................................................................................................... 137 Table 6.4: Spearman correlation analysis results for technology familiarity with ease of navigation and overall experience. .......................................................................................................................... 138 Table 6.5: Spearman's correlation results between different functionalities of the iSight prototype. ... 140 Table 6.6: Spearman's rank correlation results for frequency of technology use and comfort level. ... 142 Table 6.7: Spearman's rank correlation results for iSight functionality and increased confidence, selfesteem, well-being, and independence. ........................................................................................... 143 Table 6.8: Ficher´s Exact test results for iSight functionality and increased confidence, self-esteem, wellbeing, and independence. ............................................................................................................... 144 Table 6.9: Thematic analysis of user feedback for the iSight prototype. ............................................ 146 Table 6.10: Distribution of feedback categories by visual impairment type. ...................................... 150 Table 6.11: Usability preferences by visual impairment type. ........................................................... 151 Table 6.12: Expectations for innovations and commercial adoption by visual impairment type. ......... 152 Table 6.13: Suggestions for improvement by visual impairment type. .............................................. 152 Table 6.14: Impact on daily life and future use by visual impairment type. ....................................... 153 Table 6.15: User feedback and corresponding affected components in chapter 4 and 5. .................. 156 Table 6.16: Summary of statistical tests, hypotheses, and results *[ corresponds to “Not significant” and denotates none correlation/significance, while corresponds to “Significant” and denotes substantial impact/correlation]. ......................................................................................... 157
xvii ABBREVIATIONS AADVDB Associação De Apoio Aos Deficientes Visuais Do Distrito de Braga ACAPO Associação de Ambliopes e Cegos de Portugal AI Artificial Intelligence ANN Artificial Neural Network AP Average Precision API Application Programming Interface CAVI Support Centre for Independent Living CEICSH Comissão de Ética para a Investigação em Ciências Sociais e Humanas CMYK Cyan Magenta Yellow Black CNN Convolutional Neural Network CPU Central Processing Unit DC Direct Current EU European Union GPU Graphics Processing Unit GUI Graphical User Interface Hex Hexadecimal HSI Hue Saturation Illumination HSL Hue Saturation Luminance HSV Hue Saturation Value IDE Integrated Development Environment IIC Inter-Integrated Circuit IoT Internet of Things IOU Intersection Over Union IP Ingress Protection LED Light Emitting Diode MSCOCO MicroSoft Common Objects in Context MSVI Moderate to Severe Vision Impairment NFC Near Field Communication FDENet Frequency Domain Enhancement Network QR Quick Response
xviii RFID Radio Field Identification RGB Red Green Blue RNID Regulamento Nacional de Interoperabilidade Digital ROI Region Of Interest RPN Region Proposal Network SDG Sustainable Development Goal SIFT Scale Invariant Feature Transform SGD Stochastic Gradient Descent SPI Serial Peripheral Interface SSD Single Shot MultiBox Detector STA Statistical Feature SVM Support Vector Machine UART Universal Asynchronous Receiver/Transmitter UID Unique Identifier URL Uniform Resource Locator VOC Visual Object Classes WHO World Health Organization
1 INTRODUCTION Chapter Overview The purpose of this chapter is to present the problem statement, the motivations, and the scope of the research. An overview of assistive technologies for clothing management for blind people is presented first. In the following section, the present work's objectives and ethical considerations are discussed. Finally, the results of the developed scientific activity are presented, as well as the structure of the dissertation. 1 Introduction 1.1 Problem Statement, Motivations, and Scope 1.2 Objectives 1.3 Contribution to Knowledge and Results of the developed scientific activity 1.4 Results of the developed scientific activity 1.5 Thesis Structure
Chapter 1 – Introduction 2 1.1 Problem Statement, Motivations, and Scope Blindness is a condition that affects thousands of people worldwide. An analysis of the Vision Atlas from the International Agency for the Prevention of Blindness reveals that worldwide, 43 million people are blind and 295 million have moderate-to-severe visual impairments ( Magnitude and Projections - The International Agency for the Prevention of Blindness , n.d.). In addition, the World Health Organization (WHO) reports that at least 2.2 billion individuals worldwide suffer from a near or distance vision impairment, with nearly half of these cases being severe ( Blindness and Vision Impairment , n.d.). In Portugal, according to the 2011 census (INE, 2012), there are approximately 900.000 people with vision difficulties, of whom approximately 28.000 are blind. Additionally, 23% of population state that have difficulties in seeing, even when using glasses or contact lenses. Such visual impairment, e.g. blindness, can significantly affect a person's psychological and cognitive functioning. Several studies have demonstrated that vision impairment is associated with a variety of negative health outcomes and a diminished quality of life (Chia et al ., 2006; Langelaan et al ., 2007). Blind people can benefit from the use of assistive technology by reducing the negative effects of blindness and improving their quality of life. Sustainable Development Goal (SDG) 10 (Reduce Inequalities) is a fundamental objective in the development of inclusive technologies, which aims to reduce inequalities both within and between countries, promoting social, economic, and political inclusion for all people ( Goal 10 | Department of Economic and Social Affairs , 2024). Hence, it is imperative to address the existing disparities and barriers that affect this community. To meet the unique needs of blind people and enhance their overall quality of life, inclusive technologies should be developed that cater to their specific needs and enhance their overall quality of life. The advancement of these technologies allows people with blindness to bridge the gap and participate fully in society on an equal basis, enhancing their mobility, navigation, and overall independence, thereby promoting inclusion. Despite the proliferation of smart devices and the advancement of cutting-edge technology for blind people, most of the research has focused on navigation, mobility, and object recognition, ignoring aesthetic considerations (Bhowmick & Hazarika, 2017; Elmannai & Elleithy, 2017; Messaoudi et al ., 2022). Figure 1.1 illustrates this trend by providing an overview of research efforts across various domains in terms of number of publications between 2007-2023. Data was obtained from the SCOPUS database, using the following keywords search engine: 'Navigation AND Blind People', 'Mobility AND Blind People', 'Object Detection AND Blind People', 'Computer Vision AND Blind People', and 'Clothes AND Blind People'.
Chapter 1 – Introduction 3 Figure 1.1: Bar chart illustrating the number of publications between 2007-2023 in various research areas concerning blind people. The search strategies employed include: 'Navigation AND Blind People', 'Mobility AND Blind People', 'Object Detection AND Blind People', 'Computer Vision AND Blind People', and 'Clothes AND Blind People'. Data was obtained from the SCOPUS database. The analysis of the data presented in Figure 1.1 indicates that navigation and mobility-related technologies have received significant attention in recent years, with an increase in object detection and computer vision research from 2018 onwards. It is important to note, however, that the proportion of resources allocated to aesthetic considerations, as reflected in the "Clothes" category, remains relatively low throughout the period under study. Clothing and style preferences for different occasions form an integral part of a person's identity (Johnson et al ., 2014). This has a significant impact on how they perceive themselves as well as how they are perceived by others (Adam & Galinsky, 2012; Johnson et al ., 2014). Blindness is often overlooked when it comes to aesthetics. Even though many believe that aesthetics is only concerned with visual appeal, it can hold significant relevance in daily life. Aesthetics can significantly affect a person's sense of identity, self-expression, well-being and self-esteem. Nevertheless, blind people may experience insecurity and stress when it comes to dressing-up due to the inability to recognize the condition of the garments. The inability to perceive visual cues or even colours can make dressing up a daily challenge. Additionally, blind people may be more likely to experience clothing staining and tearing due to their inherent difficulties in handling objects and performing daily tasks. For instance, detecting stains as soon as possible is essential to avoid them becoming permanent or difficult to remove. To address this gap, it is essential to develop technologies that cater to the aesthetic needs of blind people. By integrating innovative solutions
Chapter 1 – Introduction 4 for clothing, we can enhance the self-confidence, comfort, style, and self-esteem of blind people. This will not only enhance their overall well-being but will also contribute to their social inclusion and equal participation in society. In spite of the potential for technological solutions in the future, significant challenges remain to be overcome. Due to their visual impairment, these people face challenges in detecting minor irregularities or stains in clothing textures or fabrics, necessitating reliance on others for assistance. This is the premise behind the motivation of the present work - enable blind people to have an equal sense of confidence in what they wear, without permanently requiring human assistance. Sadly, this perception is still lacking in the lives of blind people, and clothing still presents a daily challenge for them. To address this absence of support, this work aimed at developing an innovative solution that combines state-of-the-art research and technology with conventional apparatus and furniture of common use. This resulted in the development of a wardrobe equipped with ideal lighting conditions and image capture capabilities that allows detecting stains and defects in clothing items, as well as identifying their colour and category using Artificial Intelligence (AI) algorithms. By combining this system with a mobile application, the user can manage and select clothing intelligently, which can propel his or her self-esteem, confidence, well-being, and self-confidence on a daily basis, filling a technological gap in the aesthetics and image of blind people. The validation of the developed work was conducted through a collaboration with the Associação de Cegos e Amblíopes de Portugal (ACAPO), the Portuguese Association for the Blind and Amblyopic, and the Associação de Apoio aos Deficientes Visuais do Distrito de Braga (AADVDB), the Association of Support for the Visually Impaired of Braga, which helped to identify key areas for improvement. 1.2 Objectives This work aimed to design and create an innovative solution that merges cutting-edge research and technology with everyday devices. The result is a wardrobe designed with optimal lighting and image capture features that can detect stains and defects in clothing, as well as identify colours and categories using Artificial Intelligence algorithms. This endeavour should promote the autonomy of blind people in their daily activities, thereby enhancing their overall well-being. To achieve this objective, specific goals were set, particularly: G1. Conduct a nationwide survey distributed by ACAPO to assess user interest, needs, and receptivity to the study; G2. Develop an algorithm for the acquisition, processing, and analysis of images using AI to identify
Chapter 2 – Literature Review 11 2.1 Assistive Technologies for Blind People The rapid advancement of technology has significantly enhanced support systems for visually impaired people. This chapter reviews and analyses various existing solutions and projects that align with the technological framework intended for this research. The objective is to provide a comprehensive characterization of these technologies, while also identifying potential gaps and areas for further development and innovation. By critically evaluating these existing solutions, this chapter seeks to establish a foundation upon which new, improved methods can be developed to better support the visually impaired community. 2.1.1 Solutions for the Visually Impaired in Clothing Selection Visually impaired individuals face significant challenges in selecting appropriate clothing, which is essential for personal dignity and confidence. Recent research has focused on developing technological solutions to assist visually impaired individuals with clothing-related tasks. Through technological advancements, a number of solutions have been developed to assist the visually impaired. Among these solutions are mobile applications, Internet of Things (IoT) systems, tactile labelling systems, and inclusive fashion design, all of which offer unique advantages and features. To simplify the task of online shopping in the garment sector, a system was proposed by Yang Huiqiaoand Peng (2015) that makes online shopping easier for the visually impaired. This idea took shape after several interviews with visually impaired students from the Special Education College of Peking University. These interactions highlighted challenges such as excessive advertising, unintuitive interfaces, and insufficient subdivisions of clothing categories. Based on this study, the authors proposed a web application to address the identified issues and help the visually impaired make informed decisions when purchasing clothing. Essential information about each item (name, type, photo, price, and sales volume) is displayed immediately below the respective photo, simplifying the search and reading process. Additionally, Stangl et al . (2018) developed the BrowseWithMe system, an AI-powered online shopping assistant designed to make online clothes shopping more accessible for visually impaired people. This system converts product web pages into a structured representation, allowing users to interactively ask for specific information about products. The study found that BrowseWithMe yields accurate image descriptions and makes the online shopping more accessible for visual impaired. Gatis Filho Sílvio José Vieira and de Assumpção Macedo (2018) introduced a system that uses NFC tags and Quick Response
Chapter 2 – Literature Review 12 (QR) codes on clothing to help visually impaired individuals match outfits. The system provides audio feedback on clothing combinations through an Android app, enabling users to select appropriate outfits independently. An audio assistive technology that helps blind individuals identify clothing patterns and colours is presented by N.Swathi and Jyothi (2020). Yang et al . (2014) developed a camera-based prototype system that recognizes clothing patterns and colours, providing verbal descriptions to visually impaired users. Participants provided positive feedback about the system, highlighting its ability to promote greater independence in their daily lives. Another example, proposed by (Allam Mahmoud and ElShaarawy, 2022) uses the Sequential Minimal Optimization method for clothing classification. Similarly, the system proposed by Yang et al . employs Statistical Feature (STA) and SIFT (Scale Invariant Feature Transform) algorithms for clothes pattern recognition. Stangl et al . (2018) introduced Vision4All, a deep learning fashion assistance solution that helps visually impaired users to identify clothing attributes such as colours, categories, textures, fabrics, styles, graphics, and text-based content on clothes. A prototype with an endoscopic camera attached to a fingertip, allowing for image capture and recognition of colours and patterns on clothing surfaces, was proposed by Medeiros et al. (2017). Rini J.and Thilagavathi (2015) proposed a system that recognizes four types of patterns (checkered, striped, irregular, and plain) and eleven colours using a camera for image capture and a Support Vector Machine (SVM) algorithm for classification. A method that converts visual information obtained from clothing into verbal exposition using Deep Neural Networks was presented by Tateno et al. (2020). This system captures photos of clothing combinations, which are then pre-processed to remove the person's face and extract key features. After processing the images, a mathematical model is applied to implement the system holistically, enabling it to automatically recognize and characterize the presented clothing items. Inclusive fashion design aims to enhance the clothing selection experience for visually impaired individuals by incorporating tactile and functional elements. Dassoler et al . (2023) explored the development of inclusive fashion using sensory elements in clothing to allow visually impaired individuals to create tactile representations of shapes. This approach promotes individual expression and autonomy, cognitive and creative development, and a more enriching clothing experience. Despite the aforementioned research, there are several commercial solutions available in the market, both as standalone devices and mobile applications. Regarding commercial devices, the Colorino is a colour identifier that detects approximately 150 different colours and announces them aloud. It features three volume levels and an option to use headphones. Additionally, the Colorino functions as a light detector, emitting distinct acoustic signals to differentiate between natural and artificial light sources ( Colorino – Color and Light Detector - Caretec , 2024). A similar device, the ColorTest 2000, not only
Chapter 2 – Literature Review 13 identifies colours like the Colorino but also it reads dates and times and detects whether a household light is on or off ( ColorTest 2000 - Assistive Technology at Easter Seals Crossroads , 2024). An improved version of the Colorino, the Color Star Pro, offers a more compact and lightweight design for easier use. It features advanced speech output and enhanced measurement accuracy. Capable of detecting over 1,700 colour nuances, patterns, and contrasts, it uniquely translates these features into musical tones. The device is also highly effective in identifying light and its colours, making it an invaluable tool for individuals with visual impairments. Additionally, the Color Star Pro provides colour information in Red Green Blue (RGB) and Hue Saturation Luminance (HSL) formats, recognizes colour patterns, and can determine light temperature, further increasing its functionality ( Color Star Pro – Color and Light Detector - Caretec , 2024). The ultimate version, the Color Star, is a sophisticated and portable colour identification device that can distinguish over 1,700 unique colour shades, conveyed through clear spoken voice output. It is also adept at measuring contrast, detecting Light Emitting Diode (LED) light colours, gauging ambient light intensity, and identifying patterns. This versatile device provides a comprehensive analysis of colour and light in various environments ( Color Star – Talking Color and Light Detector , 2024). In terms of mobile applications, the V7 Aipoly is a mobile application that enables real-time identification of objects, texts, and colours, and communicates with the user through audio( V7 Aipoly , n.d.). Similarly, Colour Identifier is an iPhone application available on the App Store that detects the colour of any pixel in a photo or image, displaying colours in RGB, CMYK (Cyan, Magenta, Yellow, Black) and HEX (Hexadecimal) formats along with their names ( Color Identifier: Color Picker on the App Store , 2024). Mobile applications can also play an essential role in human proximity, helping to combat social exclusion and directly addressing SDG 10. For example, Be My Eyes is a mobile application designed to assist visually impaired people through internet-based video calls. This application facilitates real-time assistance by connecting visually impaired users with volunteers who are available to help. If a visually impaired individual needs to verify the expiration date on a carton of milk, they can utilize the application to initiate a video call. A volunteer receives a notification and, if available, responds to provide the necessary assistance ( Be My Eyes - Trazendo Visão Aos Cegos e Pessoas Com Visão Reduzida , 2024). This procedure can also be applied to tasks such as selecting clothing items or identifying their colours. It is noteworthy that the functionality of this application is entirely dependent on the availability of volunteers. Muhsin et al. (2023) conducted a comprehensive review of substitutive assistive technologies for people with visual impairments, categorizing them based on the type of feedback (visual, haptic, or auditory). The study highlighted the limitations in user experience and the challenges of translating numerous
Chapter 2 – Literature Review 14 research prototypes into practical, widely adopted solutions. It emphasized the need for more effective assistive aids and the importance of distinguishing between the needs of individuals with partial vision, colour blindness, and those who are totally blind. Focusing on specific groups would allow for the development of more tailored and effective assistive technologies. 2.1.2 Smart Wardrobe Solutions Some research has focused on developing smart wardrobe systems to assist blind individuals in managing their clothing. These systems typically employ technologies such as Radio Field Identification (RFID) (Goh et al. , 2011), NFC (Alabduljabbar, 2022), and IoT (Kathiravan, 2021) for garment identification and tracking. Goh et al. (2011a) propose a smart wardrobe system using RFID technology to help users manage their clothing and make decisions based on preferences for colour, style, events, and emotions. Targeting busy entrepreneurs and colour-blind individuals, the system tracks garment movement within the wardrobe and stores clothing data from RFID tags. This technology aids users in making better wardrobe decisions and simplifies clothing management. Alabduljabbar (2022) introduced an IoT-based smart clothing system which employs NFC technology to help visually impaired people manage their closets independently, allowing to select and find appropriate clothing within their closets, through the use of smartphone to scan NFC tags. The system performed reasonably well and met the usability requirements for visually impaired people. A speech recognition software module for a smart closet, designed as assistive technology for visually impaired individuals was presented by Carolina Sauceda, Peña Ingrid and Alejandro Luna Gómez (2021). This system allows users to manage their clothing through voice commands, enabling tasks such as garment insertion, searching by description, and retrieval from the wardrobe. The module, developed for web-based environments, shows promise in enhancing the quality of life for visually impaired people. The increasing demand for solutions within the concept of smart wardrobes has prompted the emergence of numerous solutions in the market. Each solution exhibits unique capabilities and can be subcategorized into physical and virtual implementations. Although these solutions are not specifically focused on blind individuals, they highlight the growing interest and potential in developing assistive technologies that can be adapted to meet the needs of visually impaired users. In 2019, the Fashion Application Programming Interface (API) mobile application ( Our Technology - Fashion Taste API , 2024) was developed to manage wardrobes using QR code scanning. This solution allows users to create a virtual wardrobe for total
Chapter 2 – Literature Review 15 management, including adding or removing clothing items. The addition of items is based on scanning a QR code on the clothing label using the smartphone camera with the installed mobile application. Beyond wardrobe management, this smart wardrobe concept provides clothing usage advice, suggests new purchases based on user needs, and helps planning outfits. The technology is offered through a mobile application featuring an intuitive interface, dashboards, and statistical data showing the most worn items and suggested outfits. The Smart Closet solution ( Smart Closet: Your Personal Stylist , 2024), introduced in 2018, also uses a mobile application for wardrobe management based on clothing purchases from various stores. This application allows users to fully manage their wardrobe items and obtain detailed statistics, such as the most worn pieces. However, this solution lacks “intelligence”, as it does not use machine learning or computer vision algorithms to autonomously identify clothing items or provide fashion advice. The Stylebook Closet App ( Stylebook Closet App: A Closet and Wardrobe Fashion App for the IPhone and IPad , 2024) serves as a comprehensive tool for planning and organizing wardrobes. Users can upload images of their clothing items and create stylish outfits with minimal effort. The app includes a background removal feature to maintain a clean virtual wardrobe. Additional features such as a calendar, packing list, and inspiration board allow users to schedule and track their outfits, prepare packing lists for trips, and analyse their wardrobe usage. Analogous to Stylebook, Pureple combines simplicity with functionality. Its intuitive interface permits users to categorize their clothes, plan outfits, and pack for trips efficiently. A distinctive feature is its ability to allow users to open their closets to other users on the app for styling assistance ( Pureple - Outfit Organizer App , 2024). Similarly, the Chicisimo app assists users in cataloguing their clothes and offers outfit inspiration from fashion enthusiasts worldwide. For instance, if a user seeks new ways to style a classic red dress, the app provides images of various styling options used by women globally ( Chicisimo - Build with Fashion Shoppers’ Taste Data , 2024). Cladwell offers users daily outfit suggestions and AI-powered styling advice based on the clothes already present in their wardrobe. This app facilitates the creation of a minimalist closet filled with capsule pieces. Unlike other wardrobe apps, users do not need to catalogue their wardrobe items manually. Instead, they select from pre-built capsule wardrobes, and the Cladwell wardrobe is populated with generic items from the chosen capsule. Users can then upload pictures of similar items from their actual wardrobe to replace the generic ones ( Cladwell | Simplify Your Life With A Capsule Wardrobe , 2024). Acloset is a digital closet application that allows users to manage all their fashion items in one place, as well as buy and sell clothing items that are no longer needed. It enables the creation of outfit ideas, provides outfit recommendations and style analytics, and facilitates the buying and selling of pre-loved items at reasonable prices (Acloset - AI Fashion Assistant on the App Store, 2024).
Chapter 2 – Literature Review 16 2.2 Clothing Category Identification In the last few years, deep learning techniques have arisen as a great method to solve problems in computer vision, such as image classification, object detection, face recognition and language processing, where convolutional neural networks (CNNs) play an imperative role (Voulodimos et al ., 2018a). This is particularly important for clothing category identification, as CNNs can extract important features from clothing images, enabling accurate identification and classification. The CNNs have exhibited excellent results and advances in image recognition since 2012 (D. Bhatt et al ., 2021; Patel et al ., 2022), when AlexNet (Krizhevsky et al ., 2012) was introduced on the ImageNet Large Scale Visual Recognition (ILSVRC) (Deng et al ., 2009; ImageNet Large Scale Visual Recognition Competition (ILSVRC) , 2024). The ImageNet competition consists of evaluating several algorithms for large-scale object detection and image classification, allowing researchers to compare detection across a variety of object classes. In recent years, several CNNs have been presented, such as: VGG (Simonyan & Zisserman, 2015), GoogLeNet (Szegedy et al ., 2014), SqueezeNet (Iandola et al ., 2016), Inception (Szegedy et al ., 2015), ResNet (K. He et al ., 2015), ShuffleNet (Ma et al ., 2018), MobileNet (Howard et al ., 2019; Sandler et al ., 2018), EfficientNet (Tan & Le, 2019), RegNet (Radosavovic et al ., 2020), among others, using these networks for different image classification problems. In line with this premise, some researchers turned to the fashion world, making use of the more recent advances in computer vision to explore diverse areas such as fashion detection, fashion analysis and fashion recommendation, achieving promising results (W.-H. Cheng et al ., 2021). Given the core scope of this thesis, only fashion classification is covered throughout the work. A literature survey allowed identifying several studies that attempted to handle the classification of fashion images. Most of the authors evaluate their models based on top-k accuracy regarding clothing attributes recognition, normally with top-3 and top-5 scores, which means that the correct label is among the top k predicted labels. The research of Chen et al. (2015) presented a network for describing people based in fine-grained clothing attributes with an accuracy of 48.32%. Similarly, Liu et al. (2016) introduced the FashionNet which learns clothing features by jointly predicting clothing attributes and landmarks. Predicted landmarks are used to pool or gate the learned features maps. The authors reported a top-3 classification accuracy of 93.01% and a top-5 of 97.01%. Another method to detect fashion items in a given image using deep convolutional neural networks, was performed by Hara et al. (2014), presenting a mean Average Precision (mAP) of 31.1%. Likewise, Corbière et al. (2017) proposed other method based on weakly supervised learning for classifying e-commerce products, presenting a top-3 category accuracy of 86.30% and a top-
Chapter 2 – Literature Review 17 5 92.80% accuracy. Later, Wang et al. (2018) proposed a fashion network to address fashion landmark detection and category classification with the introduction of intermediate attention layers for a better enhancement in clothing features, category classification and attribute estimation. In their work, accuracies of 90.99% and 95.75% were reported for top-3 and top-5, respectively. In another study by Li et al. (2019), the authors presented a two-stream convolutional neural network with one branch dedicated for landmark detection and the other one for category and attribute classification, allowing the model to learn the correlations among multiple tasks and consequently improvement in the results. Accuracies of 93.01% and 97.01% for top-3 and top-5 were reported, respectively. The fashion classification model proposed by Cho et al. (2019) allowed improving the performance taking into account the hierarchical dependences between class labels reaching accuracies of 91.24% and 95.68% in top-3 and top-5, respectively. A multitask deep learning architecture was then proposed by Lu et al. (2016) that groups similar tasks and promotes the creation of separated branches for unrelated tasks, with accuracy results of 83.24% and 90.39% for top-3 and top-5, respectively. Seo and Shin (2019) proposed a Hierarchical Convolutional Neural Network (H-CNN) for fashion apparel classification. The authors demonstrated that hierarchical image classification could minimize the model losses and improve its accuracy, with a result of 93.3%. The research of Fengzi et al. (2020) applied transfer learning using pertained models for automatically label uploaded photos on ecommerce industry. The authors reported an accuracy of 88.65%. Additionally, a condition CNN was proposed by Kolisnik, Hogan, and Zulkernine (2021), based on branching convolutional neural networks. The proposed branching can predict hierarchical labels of an image and the last label predict in the hierarchy is reported with an accuracy of 91.0%. A new algorithm for clothing category classification to address the complexity of diverse clothing styles was proposed by X. Liu et al. (2023). CloNet improves top-1 classification accuracy by 0.8% and reduces model size to one-third compared to state-of-the-art models. Furthermore, Shi et al. (2023) introduce Frequency Domain Enhancement Network (FDENet), a network for clothing category classification that captures advanced features like texture and contour information. FDENet uses spectrum enhancement and depthwise separable convolutions to improve feature extraction and reduce redundancy. It achieves a 1.3% improvement in top-1 accuracy on Deepfashion compared to current models.
Chapter 2 – Literature Review 18 Table 2.1 summarizes the data of the aforementioned works, including the used datasets. Table 2.1: Literature overview on fashion image classification works (Rocha, Soares, et al ., 2023a). Author Dataset Year Accuracy Hara et al. (2014) Fashionista 2014 mAP: 31.1 % Chen et al. (2015) Street-data 2015 Top-1: 48.31 % Liu et al. (2016) DeepFashion 2016 Top-3: 82.58 % Lu et al. (2016) DeepFashion 2016 Top-3: 83.24 % Corbière et al. (2017) DeepFashion 2017 Top-3: 86.30 % Wang et al. ( 2018) DeepFashion-C 2018 Top-3: 90.99 % Li et al. (2019) DeepFashion-C 2019 Top-3: 93.01 % Cho et al. (2019) DeepFashion 2019 Top-3: 91.24 % Seo and Shin. (2019) Fashion-MINIST 2019 Top-1: 93.33 % Fengzi et al. (2020) Fashion Product Images 2020 Top-1: 88.65 %1 Kolisnik, Hogan, and Zulkernine(2021) Fashion Product Images 2021 Top-1: 91.0 %1 X. Liu et al. (2023) DeepFashion 2023 Top-1: 68.7 %2 Shi et al. (2023) DeepFashion 2023 Top-1: 71.22 %2 1 Results only reported for fashion classification accuracy. 2Results only reported for DeepFashion dataset. It becomes clear that there has been a great effort to build efficient methods for fashion category classification. However, as the works listed in Table 2.1 highlight, there is a lack of focus on developing systems to aid visual impaired people, and there is yet no solution capable of covering all the difficulties stated by a blind person, namely an automatic system for clothing type identification. 2.3 Clothing Segmentation Image segmentation is the process of identifying and annotating various objects within images to classify them into distinct categories (K. K. Singh & Singh, 2012). This method has a wide range of applications, including driver assistance systems and disease detection. Within the realm of segmentation, there are two primary types: semantic segmentation and instance segmentation. Given that one of the main goals of this work is to identify clothing items to assist visually impaired people, it is crucial to perform preprocessing actions when capturing images of clothing. This ensures that the image presented to the user is as “clean” as possible. Therefore, it is necessary to segment the image and remove the background so that the colour detection algorithm can achieve the best results by analysing only the pixels relevant
Chapter 2 – Literature Review 19 to the clothing item. Below, a brief review of existing implementations in clothing image segmentation is presented. Yingheng and Yueqi (2020) presented a method that aids fashion analysts and consumers in identifying types of clothing. Segmentation results are enhanced using a system that incorporates Multi-Attention Mask R-CNN, enabling the extraction of detailed information from clothing items. This implementation is particularly effective in delineating boundaries between clothing items, outperforming other Mask R-CNN implementations. Another solution, described by Khurana et al . (2018) focuses on the problem of identifying and detecting clothing items in images. This method differentiates between various types of clothing that may have similar patterns or textures. The implementation consists of two modules: the first utilizes a Fully Convolutional Network for spatial boundary segmentation, and the second detects specific features to aid in recognizing boundaries. The results demonstrate an improvement over existing methods, reducing confusion in images with visually similar clothing items made of different materials. DeepFashion2, a method designed to extract the maximum amount of data from photos of clothing items, was proposed by Ge et al. (2019). This method addresses the limitations of the DeepFashion system, aiming for accurate segmentation and identification of all clothing items in each photo. Moreover, it allowed generating a dataset comprising ca . 801,000 annotated images with identified types and styles. These images were segmented using masks and bounding boxes. The novel model, Match R-CNN, is based on Mask R-CNN (K. He, Gkioxari, Dollar, et al ., 2017). The results indicate a slight improvement in Average Precision compared to the Mask R-CNN and Match R-CNN models. Complex backgrounds can make image segmentation a demanding task. DeepLabV3+, proposed by Wang et al. (2021), is a model that segments clothing items within complex backgrounds where the boundaries might be ambiguous. A new Neural Network architecture was redesigned to improve segmentation performance, showing better adaptability to the boundaries of each clothing item. After training the model and testing it on the dataset, the results showed that accuracy (97.26%), mean intersection over union (IoU) (93.23%), and average precision (AP) (90.56%) values were slightly higher than those obtained with the previous DeepLabv3+ model. Tackling another issue, the method described by X. Zhang et al . (2020) recognizes sets of clothing items and identifies the year they belong to. Here, a CNN was used to segment the human body and classify the type of clothing within the segmented areas without background interference. A dataset of 9,339 images spanning eight years was used to validate the model, demonstrating good efficacy. Various segmentation models were tested, with SegNet achieving an IoU value of approximately 0.951, closely
Chapter 2 – Literature Review 20 matching the proposed method IoU value of 0.951. The proposed method also yielded the best classification results, with an accuracy of 0.805 compared to 0.785 using the ResNet101 model. Similarly, T. Yang et al. (2021) presented a method for segmenting images of clothing items captured in different positions allows for effectively detecting the silhouettes of shirts using instance segmentation with Mask R-CNN. Approximately 9,000 images were collected, segmented, and classified by various attributes. The experimental results showed that the model performed well in detecting short-sleeved shirts and collars, achieving AP values of 0.95 and 0.94, respectively. Overall, the conducted literature review on clothing segmentation underlines the significance of advanced segmentation techniques in enhancing the accuracy and effectiveness of algorithms for identifying and analysing clothing items. Table 2.2 summarizes the mentioned works on advanced image segmentation techniques for clothing identification. Table 2.2: Summary of advanced image segmentation techniques for clothing identification Method Authors Key Features Applications Performance Fully Convolutional Network Khurana et al. (2018) Spatial boundary segmentation, feature detection Identifying and detecting clothing items Reduces confusion in images with similar patterns/textures DeepFashion2 Ge et al. (2019) Extracts maximum data, masks and bounding boxes Accurate segmentation and identification of clothing items Slight improvement in AP compared to Mask R-CNN and Match R-CNN Multi-Attention Mask R-CNN Yingheng and Yueqi (2020) Enhances segmentation results, detailed information extraction Fashion analysis, consumer identification Outperforms other Mask RCNN implementations in delineating boundaries SegNet X. Zhang et al. (2020) Human body segmentation, clothing type classification Recognizing sets of clothing items, year classification IoU: 0.951, Classification Accuracy: 0.805 DeepLabV3+ Wang et al. (2021) Segments within complex backgrounds, redesigned NN architecture Adaptable segmentation performance Accuracy: 97.26%, Mean IoU: 93.23%, AP: 90.56% Instance Segmentation with Mask R-CNN T. Yang et al. (2021) Segments images of clothing in different positions Detecting silhouettes of shirts AP: 0.95 (short-sleeved shirts), AP: 0.94 (collars)
3 THEORICAL CONCEPTS Chapter Overview This chapter outlines the methodologies used in the present work. It begins with an overview of blindness and continues with a discussion on Artificial Intelligence, Machine Learning (ML), and Deep Learning (DL), emphasizing their roles in developing “intelligent” systems. The theoretical concepts of Neural Networks and Transfer Learning are presented, followed by the evaluation metrics used to assess model performance. 3 Theorical concepts 3.1 Blindness 3.2 Artificial Intelligence, Machine Learning and Deep Learning 3.3 Neural Networks 3.4 Transfer Learning 3.5 Evaluation Metrics
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 28 3.1 Blindness The human eye can be broadly compared to a photographic camera in terms of its functioning. Light reflected from objects passes through the cornea, pupil, and lens, eventually reaching the retina ( The Anatomy of the ‘Camera’ Eye | UCL Institute of Ophthalmology - UCL – University College London , 2024). Here, specialized cells encode the visual information, which is then transmitted to the brain via the optic nerve, as illustrated in Figure 3.1. Figure 3.1: Constituents of the human eye, illustrating the pathway of light from the cornea to the retina and the optic nerve, and comparing its functioning to a photographic camera( The Anatomy of the ‘Camera’ Eye | UCL Institute of Ophthalmology - UCL – University College London , 2024) . Among the five senses dominated by most humans, vision plays a particularly prominent role. A disability, by definition, refers to the partial or total loss of a part of the human body or its function, whether physical or psychological, which hinders an individual's ability to perform common daily activities (Enoch et al ., 2019; Khorrami-Nejad et al ., 2016). Visual impairment is one of the significant types of disabilities and can be categorized into two main forms. The first is low vision, which refers to visual impairments that cannot be fully corrected by standard means such as glasses or contact lenses. The second form is blindness, which denotes the total loss of vision (Medeiros et al ., 2017). Visual impairment, whether partial or total, significantly affects an individual's quality of life, limiting their ability to engage in daily activities independently. This limitation points out the importance of technological advancements aimed at supporting visually impaired people. Innovative solutions such as assistive devices and applications have been developed to enhance their autonomy and improve their quality of life. Technological aids range
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 29 from basic tools, like magnifying glasses and braille readers, to advanced digital applications that leverage artificial intelligence and machine learning. For instance, mobile applications that provide real-time audio descriptions of the user's surroundings or smart devices that assist in identifying objects and navigating spaces are becoming increasingly prevalent (Ghazal et al ., 2019; Najm et al ., 2022). These technologies not only offer practical support but also contribute to the social inclusion of visually impaired people by enabling them to perform tasks that would otherwise require assistance. In summary, the eye functions similarly to a camera by capturing and transmitting visual information to the brain. Vision is a critical sense that greatly influences an individual's daily functioning. Visual impairments, whether in the form of low vision or blindness, present significant challenges that can be mitigated through the development and utilization of various assistive technologies. 3.2 Artificial Intelligence, Machine Learning and Deep Learning Artificial intelligence is the development of computer systems that can perform tasks traditionally handled by human intelligence, e.g , speech recognition, understanding of natural language, and decision-making (Soori et al ., 2023). The fields of machine learning (ML) and deep learning (DL) are subfields of artificial intelligence, Figure 3.2, that use techniques of automatic learning to train computer models from data over time, resulting in improved performance. Figure 3.2: Interconnection between the fields/subfields of AI/ML/DL. ML involves training computer models to learn from data and make predictions (Sarker, 2021b), whereas DL, a subset of ML, employs deep neural networks with multiple layers or stages through which data is processed for building a data-driven model to discover data structures automatically (Sarker, 2021a).
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 30 3.3 Neural Networks Neural networks are a fundamental component of deep learning, modelled after the structure and function of the human brain. They typically consist of multiple layers, including an input layer, one or more hidden layers, and an output layer as depicted in Figure 3.3. The hidden layers are crucial for learning complex patterns and representations from the data (Kriegeskorte & Golan, 2019). Figure 3.3: Deep neural network with two hidden layers. During training, neural networks adjust their parameters (weights and biases) through a process called backpropagation. Backpropagation is essential for the network to gradually improve its ability to make accurate predictions or perform specific tasks (Lecun et al ., 2015). This learning process involves minimizing the difference between the network's predictions and the true targets, which is typically achieved through optimization algorithms. Optimization algorithms are vital in this process, with commonly used methods including stochastic gradient descent (SGD), Adam, and RMSprop (Kingma & Ba, 2014; Ruder, 2016). These algorithms iteratively adjust the learning rate and parameters to improve convergence and find the optimal set of weights and biases that minimize a predefined loss function. The loss function quantifies the discrepancy between predicted and actual values. The deep architecture of neural networks, characterized by multiple hidden layers, enables them to capture hierarchical features and intricate relationships within the data (Lecun et al ., 2015). For example, in image recognition tasks, lower layers may detect basic features such as edges and textures, while higher layers recognize more complex structures like shapes and objects. Neural networks have demonstrated remarkable versatility and effectiveness across a wide range of applications, from natural language processing and speech recognition to medical diagnosis and autonomous driving. This adaptability stems from their ability to
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 31 model non-linear relationships and represent high-dimensional data comprehensively. 3.3.1 Convolutional Neural Networks Convolutional Neural Networks are a class of deep learning models specifically designed for image recognition and computer vision tasks (Krizhevsky et al ., 2012). They have demonstrated remarkable success in various applications, including image classification, object detection, and image segmentation (Minaee et al ., 2022; Rawat & Wang, 2017; Voulodimos et al ., 2018b; Z.-Q. Zhao et al ., 2019). CNNs are composed of multiple layers, each tasked with extracting different types of information from the input data. The architecture of CNNs comprises two primary modules: convolution and classification, as illustrated in Figure 3.4. Figure 3.4: The architecture of a Convolutional Neural Network (CNN) comprises an input layer, several alternating convolutional and max-pooling layers, followed by a fully-connected layer and a classification layer (Alom et al ., 2019) During the convolution stage, the network employs filters to extract specific elements from the images, abstracting and dividing the image into distinct parts. The processed data is then passed through a pooling layer, which reduces the dimensionality and computational load by downsampling the feature maps. In the final stage, the fully connected layer classifies the information obtained from the previous layers. By integrating these three stages—convolution, pooling, and classification—CNNs acquire the ability to identify and recognize regions of interest within each image. One of the significant advantages of CNNs is their reduced need for pre-processing compared to other classification algorithms, as they can optimize filters through learning (X. Zhao et al ., 2024). Each neuron in a CNN calculates a value based on the inputs and a specific activation function. This value is adjusted using a weight vector, which is updated during the network's learning process. These weight vectors
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 32 represent the features that the network aims to prioritize. 3.3.2 Object Detection Object detection, which involves identifying and localizing multiple objects within an image, is a fundamental task in computer vision. Several CNNs architectures have been developed to address this complex task (S. Agarwal et al ., 2019). These architectures can be broadly categorized into one-stage and two-stage detectors based on their architecture and processing steps. One-stage detectors, such as YOLO and Single-Shot MultiBox Detector (SSD), perform object detection in a single pass through the network. They divide the image into a grid and predict bounding boxes and class probabilities directly from the full image. This approach is typically faster and more suitable for real-time applications, though it may sacrifice some accuracy compared to two-stage detectors(Soviany & Ionescu, 2018; Y. Zhang et al ., 2021). Since its initial release in 2015, YOLO has evolved significantly, reaching its eighth version, YOLOv9 (C.- Y. Wang et al ., 2024). Unlike region-based algorithms, YOLO directly predicts bounding box locations and class probabilities (Redmon et al ., 2016). The entire image is divided into a grid, with each grid cell predicting the bounding boxes and class probabilities for objects detected within it as depicted in Figure 3.5: In the YOLO architecture, the input image is segmented into an S × S grid, with each grid cell predicting B bounding boxes characterized by their position, size, and confidence score. The final label of a high-confidence bounding box is determined by the class probability map (Redmon et al ., 2016).
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 33 Figure 3.5: In the YOLO architecture, the input image is segmented into an S × S grid, with each grid cell predicting B bounding boxes characterized by their position, size, and confidence score. The final label of a high-confidence bounding box is determined by the class probability map (Redmon et al ., 2016). This approach allows YOLO to perform object detection in a single pass through the network, enabling highly efficient and real-time object detection. Each iteration of YOLO has brought improvements in accuracy, speed, and architecture, with YOLOv9 benefiting from the latest advancements in deep learning and computer vision research (Terven et al ., 2023; C.-Y. Wang et al ., 2024). SSD, introduced by Liu Wei and Anguelov (2016) presents a method for detecting objects in images using a single deep neural network. This approach discretizes the output space of bounding boxes into a set of priors with varying aspect ratios and scales for each feature map location. During prediction, the network generates confidence scores for each prior, indicating the presence of objects of interest, and adjusts the priors to better fit the object shapes. Additionally, the network combines predictions from multiple feature maps with different resolutions to naturally handle objects of various sizes. This one-stage detector is renowned for its balance between speed and accuracy. A comparison of SSD and YOLO architectures is depicted in Figure 3.6.
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 34 Figure 3.6: Comparison of two single-shot detection models: SSD and YOLO. The SSD model augments the base network with additional feature layers that predict adjustments to default boxes of varying scales and aspect ratios, along with corresponding confidence scores (Liu Wei and Anguelov, 2016). Two-stage detectors, such as R-CNN, Fast R-CNN, and Faster R-CNN, operate in two steps. The first stage generates region proposals that are likely to contain objects. The second stage processes these proposals to refine the bounding boxes and classify the objects. This approach generally achieves higher accuracy but can be slower and more computationally intensive. The R-CNN (Girshick et al ., 2014), Figure 3.7, was one of the first successful CNN-based methods for object detection. It introduced a multi-step pipeline consisting of Region Proposal, Feature Extraction, and Object Classification and Localization. While R-CNN made significant contributions to the field, its computational and memory inefficiencies limited its practical use in real-time and large-scale applications. Figure 3.7: The multi-step pipeline of R-CNN, including region proposal, feature extraction, and object classification and localization (Girshick et al ., 2014) .
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 35 To address the inefficiencies of R-CNN, Fast R-CNN was developed. This architecture optimizes the object detection pipeline, resulting in faster and more efficient processing (Girshick, 2015). Key improvements include feature sharing, a unified architecture, and end-to-end training. Faster R-CNN, Figure 3.8, further extends and improves upon the original R-CNN architecture, by introducing a novel Region Proposal Network (RPN) that shares convolutional features with the object detection network (Ren et al ., 2015). Figure 3.8: Faster R-CNN is a unified network for object detection, where the RPN module functions as the 'attention' mechanism within the network (Ren et al ., 2015). This innovation addresses the computational and memory inefficiencies of R-CNN. Faster R-CNN also serves as the foundation for Mask R-CNN, which adds instance segmentation capabilities. RetinaNet, proposed by Lin et al. (2018) introduces the Focal Loss, a loss function that addresses the class imbalance problem by focusing more on hard-to-classify samples. This architecture strikes a balance between the speed of one-stage detectors and the accuracy of two-stage detectors. Object detection continues to evolve with new methods and improvements, driven by ongoing advancements in deep learning and computer vision. The choice of the ideal architecture depends on the specific application requirements, such as accuracy, speed, and the capability to operate in real-time. 3.3.3 Image Segmentation Image segmentation, a closely related but distinct task from object detection, involves partitioning an image into multiple segments or regions to locate objects and boundaries accurately. Unlike object detection, which provides bounding boxes around detected objects, image segmentation identifies the
Chapter 3 – Theorical Concepts ___________________________________________________________________________ 36 exact pixels belonging to each object, providing a more detailed understanding of the scene. Instance segmentation, a subtype of image segmentation, not only detects objects but also segments each object's pixels to provide a pixel-level mask. This approach is particularly useful in applications where precise object boundaries are required (Sharma et al ., 2022) . Several popular deep learning models are used for image segmentation. U-Net (Ronneberger et al ., 2015) employs a U-shaped architecture for the efficient segmentation of medical images, demonstrating high efficacy with small datasets and delivering precise segmentation results. Fully Convolutional Networks (FCNs) are capable of processing images of any size and generating spatial maps by substituting fully connected layers in conventional CNNs with convolutional layers, enabling pixel-by-pixel segmentation of entire images. SegNet (Badrinarayanan et al ., 2016) features an encoder-decoder network structure utilized in tasks such as scene understanding and object recognition. The encoder captures contextual information from the image, while the decoder leverages this context for precise localization and segmentation of objects. DeepLab (L.-C. Chen et al ., 2017) is distinguished by its use of atrous convolutions, which capture multi-scale context through multiple parallel filters. Mask R-CNN, an extension of Faster R-CNN, is a prominent architecture for instance segmentation. Mask R-CNN enhances Faster R-CNN's object detection by predicting pixel masks for detected objects. It begins with a backbone that extracts feature maps from the input image, capturing hierarchical and contextual information. The RPN generates region proposals (bounding box candidates) for potential objects. These proposals are processed through the Region of Interest (ROI) Align layer to extract fixed-size feature maps, which are then passed through two branches: the region of interest classifier and the bounding box regressor (K. He, Gkioxari, Dollar, et al ., 2017). The classifier predicts the probabilities of object classes, and the regressor refines the bounding box coordinates. An additional branch, the mask prediction head, produces a binary mask for each object instance, representing pixel-level segmentation and distinguishing the object from the background as illustrated in Figure 3.9
Chapter 4 – System Development ___________________________________________________________________________ 43 characterization of the participant, such as age, gender, academic qualifications and current occupation. The second part of the questionnaire, Identification of the type of disability , enabled the analysis of conditions that caused blindness in the participants, as well as the age group in which these conditions occurred. The third part, Clothing , includes 6 questions to identify the importance and concerns that blind people have regarding the actions they need to do to get dressed, to choose clothes and how they do shopping. In this subject, some questions were presented as statements. The answers were based on a 5-point Likert scale of agreement, importance or frequency. The fourth part, Technology , includes 3 questions that aim to assess the knowledge, usability and satisfaction that blind community has in relation to the available supportive technologies. However, it is also possible to assess the proximity that this community has within this technology. The fifth and last part, Research and development , includes 3 main questions and other 2 complementary. In this section, it is intended to assess the availability and the interest in developing a technology that supports the blind community identifying and combining clothes. Moreover, this section aims to assess if blind people want this technology to be available. The suggestions and opinions of the participants were taken into consideration in the final version. 4.1.4 Implementation The data collected using google forms from July to September 2020 and the participation was voluntary and anonymous. The data was gathered in a table and analysed using the statistical software SPSS (Statistical Package for the Social Sciences) version 22. At this initial stage of analysis of the collected data, a descriptive statistic in which data was translated into percentage values was considered. 4.1.5 Participants Characterization The sample of was essentially composed by blind associates of ACAPO, who were asked to fill a purposebuilt survey, delivered by email. A total of 26 participants replied to the survey, of which 23 full responses fulfilled all the set conditions, and with only 3 being excluded due to not fulfilling the aforementioned conditions. In the present study, all the analyses were made considering the full responses. To obtain the characterization of the participants in this survey, 4 variables were defined: age, gender, academic qualifications and profession.
Chapter 4 – System Development ___________________________________________________________________________ 44 It was observed that there was no equal distribution between age groups, with individuals between 45 and 54 years old presenting 52.2% of the answers, the youngest participant being 18 years old, and the oldest being 65 years old. Due to the wide range of participant’s age, during the analysis of the survey, the variable age was grouped into two main groups, one for participants being 44 or less years old, and the second to participants being 45 years old or more. Regarding gender and education, most of the participants were female (73.9%), and had higher education (43.5%), with only 26.1% having, at least, nine years of basic education. In the analysis of the occupations, it was found thar most of them were unemployed (30.4%), which highlights the difficulty of this population to get integrated in the labour market. F&B (food & beverage) and administrative services were the occupation areas with greater representativeness, accounting for 13% each. Other categories include all retired people and students, representing 17%. 4.1.6 Identification of the Type of Disability The second part of the questionnaire contains two questions enabling each participant to identify the causes and types of conditions that caused the participants' loss of vision, as follows: QII.1: Since what age are you visually impaired? QII.2: What caused your disability? On the basis of the received responses, the majority (69.6%) became blind during the first year of life. This shows that these participants do not have basic visual notions about the “world” around them. An identical result was obtained regarding the reason for the loss of vision, that is, 65.2% had lost their vision due to congenital disease, and only 8.7% as a result of an accident. The remaining 26.1% refers to acquired diseases or conditions. Congenital disease occurs during pregnancy, which explains the high number of participants with vision loss before reaching one year old. Hence, the conditions that originated blindness or vision loss were grouped in two main categories, congenital disease and acquired disease, with glaucoma and retinopathy being the main causes of vision loss, respectively. 4.1.7 Aesthetics and Clothing The way we dress has a big impact on our daily lives and it is crucial to our well-being. This perception ranges from the importance that society gives to the way we dress to the importance of choosing clothing for the workplace, to help convey the right and appropriate image for the context in which we are inserted in. So, in the third part of the questionnaire, three statements were considered:
Chapter 4 – System Development ___________________________________________________________________________ 45 QIII.1: Vision is one of the senses that dominates the life of human beings. It allows them to know and perceive the world around them, while giving meaning to objects, concepts, ideas and tastes. QIII.2: How to dress and the style we prefer for different occasions is part of someone's identity. Blind people do not have this sense, and dressing can be a difficult and stressful task. With the advancement of technology, it is important to minimize all the limitations of a blind person whilst organizing their own clothing. QIII. 3: The lack of knowledge about the colours, the type of pattern or the condition of the garments makes this a daily challenge, in which the current resources are not the best. And the following three questions: QIII.4: Are you worried about matching garments? QIII.5: How often do you need help buying clothing? QIII.6: Where do you buy your clothes? Due to the relatively small size of the sample, for all statements, the scale was recoded so that only two groups were possible: "agree" and "disagree", showing a positive or negative response for each statement. Fig.1 shows the distribution of results for the different statements. Figure 4.1: Positive and negative response distribution for each statement (Rocha et al ., 2022). The majority of participants recognize that the sense of vision is crucial to the attribution of meanings and recognition of everything that surrounds us. The importance of aesthetics as a definition of our identity is also in agreement among the participants, as well as the daily difficulties experienced by participants due to the lack of resources. As expected, all participants have reinforced the importance of combine clothing for them, the majority of which (73.9%) saying that they frequently seek help to shop their clothes, mainly from family members, support centre professionals and CAVI (Support Centre for Independent Living).
Chapter 4 – System Development ___________________________________________________________________________ 46 When asked where they buy their clothes, all participants referred physical stores, with 39.1% also referring to do shopping online. 4.1.8 Technology Technology has evolved considerably in this field; however, it is necessary to understand whether there is enough technology to minimize the obstacles faced by people with blind condition or vision loss, in their daily lives. So, in the fourth part of the questionnaire, after the identification of the device and system used, two questions were considered: QIV.1: Are you aware of any technology for clothing identification? QIV.2: How satisfied are you with the technology available on this topic? Most of participants (65.2%) demonstrated to be aware of existing clothing identification technologies. Smartphone apps were pointed as main source of aid to detect colours. As an alternative, the same users opt for Colorino ( Colorino – Color and Light Detector - Caretec , 2024). However, only 26.1% referred to be satisfied with the existing technology, which reinforces the importance of the topic addressed on this work contributes to the motivation behind it. Overall, ca. 89% of participants referred they used smartphones, 72% used a laptop, 33% used mobile phones and 11% tablet devices. Most smartphone users referred they had the IOS operating system. 4.1.9 Research and Development Aesthetics has an impact on the image and appearance of our body, directly influencing our well-being, self-esteem and optimism in “day-to-day” life. The research and development conducted in this work sought to include these characteristics. Therefore, in the fifth part of the questionnaire, three main questions were considered: QV.1: Would it be important to develop something that would make it possible to identify and combine clothing? QV.2: How likely is it to use a technology that includes the act of identifying and combining clothing? QV.3: Did you need help in completing this survey? According to Figure 4.2, most participants (91.3%) agreed that a framework should be developed to help identifying and matching garments. In line with the previous answer, almost all participants (95.7%) were willing to use technology to help them identifying and combining clothes. When choosing the platform to
Chapter 4 – System Development ___________________________________________________________________________ 47 use the technology, the smartphone and the laptop were the preferred choices, accounting for 43.2% and 35.1%, respectively. Lastly, only 17.4% of participants stated that they sought some help to fill out the questionnaire. Figure 4.2: Results distribution for QV.1 and QV.2 (Rocha et al ., 2022). 4.1.10 Conclusions of the survey Based on the collected and analysed information, it becomes important to develop tools that allow the identification and combination of garments. Moreover, ca . 95.5% of the respondents are available to use technology to help them in this task. Although the sample dimension used in the analysis could be considered small (n=23) to be representative for the blind population at national level, the authors believe that the collected results reinforce the central importance of aesthetics on clothing for blind people, who are becoming increasingly inserted in the labour market, following global and European guidelines of inclusion ( Reduzir as Desigualdades • ODS - BCSD Portugal , n.d.). Through the survey undertaken by the participants, it is possible to conclude that not knowing colours, patterns, and the overall condition of clothes is a major challenge that blind people face every day. This comes to emphasize the need to identify and combine clothing. Despite of all the difficulties, the existing offer of different technologies and support is very limited. So far, mobile phone applications and Colorino remain the leading sources of support used by blind people. Analysing the comments, participants have demonstrated their enthusiasm and willingness to collaborate in this research. 4.2 System Requirements
Chapter 4 – System Development ___________________________________________________________________________ 48 After analysing the results of the conducted survey and summarizing the key aspects of a system that could help blind people choose and characterize their clothes, especially with regard to their well-being, the system architecture and requirements were designed. The design phase involved identifying critical functionalities and components that would effectively address the needs highlighted by the survey respondents. The proposed system(Rocha et al ., 2021) aims to develop a comprehensive solution that includes the following key requirements: • Clothing Classification: The system should accurately identify and categorize clothing items based on type (e.g., shirts, dresses). • Colour Detection: It should precisely determine the colours of clothing items to assist users in making appropriate wardrobe choices. • Defect Detection: The system must identify defects such as stains and holes to ensure that users are aware of the condition of their clothing. • User-Friendly Mobile Interface: The mobile application should be highly accessible and intuitive, allowing blind and low-vision users to navigate and use the system with ease. • Smart Wardrobe Integration: The system should include a smart wardrobe equipped with an image acquisition system and controlled lighting to capture high-quality images of clothing items. • NFC Tagging: The system should support NFC tagging to facilitate the organization and retrieval of clothing information. • Clothing Management: The application should enable users to manage their wardrobe, including adding new items, updating information, and verifying item details. Once the system requirements were defined, it was possible to outline the methodology for its conception, as explained in the following section. 4.3 System Overview The development of the iSight prototype was structured around three main components: (i) the collection of data and the application of deep learning algorithms for garment identification and modification detection, (ii) the design and implementation of the mechatronic device, specifically the automated wardrobe system, and (iii) the creation of a user-friendly mobile application. The workflow illustrating the main steps of this methodology is presented in Figure 4.3.
Chapter 4 – System Development ___________________________________________________________________________ 49 Figure 4.3: Schematic representation of the workflow comprised in the iSight prototype development. The following subsections provide a detailed account of each component. Initially, the data collection process and the deployment of advanced deep learning algorithms are discussed, focusing on their role in accurately identifying garments and detecting modifications. Subsequently, the design, construction, and integration of the automated wardrobe system is elaborated (Chapter 5), highlighting its importance as the core mechatronic element of the iSight prototype. Finally, the development of the mobile application is explored, with an emphasis on its user-friendly interface and seamless integration with the automated wardrobe, ensuring accessibility and ease of use for visually impaired users. 4.4 Deep Learning Approaches This subsection details three experiments using deep learning techniques. The first experiment addresses the classification of clothing categories using CNNs. The second experiment expands this approach by incorporating both segmentation and classification of clothing categories, in addition to extracting colour information. Lastly, the third experiment focuses on identifying modifications in clothing items.
Chapter 4 – System Development ___________________________________________________________________________ 50 4.4.1 First Experiment - Clothing Category Classification using CNN In the first experiment, state-of-the-art deep learning techniques were leveraged through transfer learning to perform category classification using CNNs (Rocha et al ., 2019). The pipeline for applying deep learning to clothing category classification involves several key steps. The process begins with the selection, preparation, and pre-processing of a dataset. In parallel, several data augmentation techniques are employed to examine their effects on the results with and without its use. Finally, transfer learning is implemented, using state-of-the-art deep learning models to perform category classification using CNNs. Figure 4.4 illustrates this workflow. Figure 4.4: Workflow methodology for the classification of clothing categories (Rocha, Soares, et al ., 2023a). 4.4.1.1 Dataset preparation To apply deep learning to clothing category classification, a dataset was needed before being fed into the neural networks. As a first step, it was necessary to identify a dataset that could meet this work’s requirements. Considering that all images taken by the blind people are collected in a controlled environment and with one item of clothing at a time (Rocha et al ., 2020), and based on the authors' previous research (Rocha et al ., 2021), the authors were able to conclude that a reasonable quantity of data was required for training and obtaining more accurate results. For this reason, a research survey was conducted to identify available datasets - Table 4.1.
Chapter 4 – System Development ___________________________________________________________________________ 51 Table 4.1: Summary of available datasets for fashion category classification (Rocha, Soares, et al ., 2023a). Dataset Year # of photos Description DeepFashion-C (Z. Liu, Luo, et al ., 2016) 2016 289,222 Annotated with clothing bounding box, pose variation type, landmark visibility, clothing type, category, and attributes. Fashion Landmark Dataset (Z. Liu, Yan, et al ., 2016) 2016 123,016 Annotated with clothing type, pose variation type, landmark visibility, clothing bounding box, and human body joint. FashionMinist (Xiao et al ., 2017) 2019 70,000 Grayscale image dataset associated with a label from 10 classes. DeepFashion2 (Ge et al ., 2019) 2019 491,000 A versatile benchmark of four tasks including clothes detection, pose estimation, segmentation, and retrieval. Fashion Product Images (Fashion Product Images Dataset | Kaggle, n.d.) 2019 44,400 Annotated with gender, master category, subcategory, article type, base colour, season, year, usage and product description. Based on the characteristics of each dataset, it was decided to use the Fashion Product Images Dataset, which is a smaller resolution version of the dataset. This dataset provides a variety of attributes that are relevant to the current and future needs of the project, such as category, colour, season, among others. However, despite the large number of annotated features included in the dataset, it still presents unbalanced data between article types, which were referred to as “categories” in this work. Figure 4.5 shows the clothing categories for which at least 500 images were considered in their distribution, i.e. each entry on the Xx axis represents a specific clothing category.
Chapter 4 – System Development ___________________________________________________________________________ 52 Tshirts Shirts Casual Shoes Watches Sports Shoes Kurtas Tops Handbags Heels Sunglasses Wallets Flip Flops Sandals Briefs Belts Backpacks Socks Formal Shoes Perfume and Body Mist Jeans Shorts Trousers Flats 0 1000 2000 3000 4000 5000 6000 7000 8000 Records Category Figure 4.5: Distribution of the number of records for each article type/category (Fashion Product Images dataset) (Rocha, Soares, et al ., 2023a). For a comprehensive and unbiased comparison between different categories, the dataset was carefully curated by selecting exactly 500 records from the initial complete range for each category as shown in Figure 4.5. This approach avoids imbalanced data, ensuring that each category in the dataset is adequately represented. A representative item from each category is illustrated in Figure 4.6, allowing for a better understanding of the dataset composition and diversity.
Chapter 4 – System Development ___________________________________________________________________________ 59 4.4.1.4 Conclusions of First Experiment This study presents an analysis of clothing type identification and stain detection for blind individuals using deep learning models. Through comprehensive benchmarking, it was demonstrated that a deep learning model, particularly GoogleLeNet, can achieve up to 91% F1-Score in identifying clothing type categories, representing a significant improvement over previous literature. The study highlighted the effectiveness of transfer learning, where fine-tuning all network layers and leveraging pre-trained weights yielded superior results compared to traditional methods that add new head layers. Augmented data techniques, such as random horizontal flipping, were shown to further enhance model performance, while random rotations negatively impacted results. This underscores the importance of carefully selecting augmentation strategies to optimize model accuracy. It was also observed that the depth of the architecture ( e.g., ResNet-18 vs. ResNet-50) did not significantly affect validation accuracy, though it did influence the computational efficiency. ResNet-18, for instance, provided the fastest inference times without compromising accuracy. Despite the promising results, there are limitations to this study. One notable constraint is the reliance on images of clothing worn by models, which may introduce redundancy and lead to misclassifications between similar clothing articles, such as casual shoes and sport shoes, or tshirts and tops. To build on these findings, further tests are necessary to determine the most effective types of data augmentation for the models used. Additionally, addressing the limitations related to image sources could enhance the model's robustness and generalizability. In summary, the fine-tuned GoogleLeNet model, enhanced with appropriate data augmentation techniques, outperforms existing methods and offers a viable solution for clothing type identification for blind individuals. This study’s results align with the accuracy levels reported Kolisnik, Hogan, and Zulkernine (2021), despite differences in class specifications, demonstrating the model’s potential for practical applications. 4.4.2 Second Experiment - Clothing Category Segmentation and Classification and Colour Extraction The second experiment extended the first by introducing clothing segmentation. While the first experiment focused on image classification, this one evaluates the problem as instance segmentation, refining predefined categories to accommodate different backgrounds and enabling the analysis of pixel colours as a result of segmentation. Hence, this process can be used to classify clothing categories and
Chapter 4 – System Development ___________________________________________________________________________ 60 characterize individual pixels based on their colours. The workflow is depicted in Figure 4.8. Figure 4.8 – Workflow for clothing segmentation and colour extraction. 4.4.2.1 Dataset collection The developed dataset consists of 2,000 images that are equally distributed across eight categories of garments: dresses, jackets, pants, polos, shirts, shoes, shorts, and t-shirts. The present dataset is an enhancement and expansion of the previous one, incorporating a refined selection of the most significant categories. In addition to refining the existing categories, new images were sourced from personal wardrobes and included diverse backgrounds in order to increase the dataset's robustness. As a result, each garment category is fully represented under a variety of conditions, improving its applicability in real-life scenarios. A sample of the images added to the dataset is depicted in Figure 4.9. Figure 4.9 – Samples of images added to the dataset from personal wardrobes: a) jacket, b) jeans, c) dress, and d) shorts. A split of 70/20/10 was used to divide the dataset into training, validation, and test sets, ensuring an equal distribution of data across all categories. As the objective of the study was not only to classify
Chapter 4 – System Development ___________________________________________________________________________ 61 garments, but also to segment them to extract pixel-level colour information, the study was approached as a problem of segmentation. For this purpose, the YOLOv8s model was employed, and the network was fine-tuned. While training, a variety of data augmentation techniques were applied, such as changes in hue, value, and saturation, horizontal flips, and scaling. As a result of these augmentation strategies, the model was enhanced in terms of robustness and generalizability. Based on testing and experimentation throughout the training process, Table 4.8 presents the hyperparameters that yielded the best results. Table 4.8: Hyperparameters (Image Size, Optimizer, Learning Rate, and Batch Size) of model experiments. Parameters Value Image Size 640x640 pixels Optimizer SGD Learning Rate 0.01 Batch Size 16 The results of this study, including performance metrics for each category, are presented in Table 4.9. Precision indicates a low percentage of false positives relatively to the total number of pixels correctly identified as part of an object mask. Recall refers to the percentage of pixels correctly identified as belonging to an object mask compared with the total number of pixels actually belonging to that mask - a low percentage indicates few false negatives. A mAP at IoU = 50 measures the average precision of the model's predictions when the intersection over union (IoU) is, at least, 50%. Table 4.9 summarizes the performance results. Table 4.9 - Main performance results of model for clothing mask segmentation (Precision, Recall and AP at IoU = 0.50). Category Precision Recall AP at IoU = 0.50 AP all 0.941 0.896 0.965 0.949 Dress 0.849 0.917 0.941 0.906 Jacket 1 0.911 0.99 0.955 Pants 0.989 0.875 0.97 0.96 Polo 0.953 0.837 0.96 0.952 Shirt 0.863 0.875 0.925 0.916 Shoes 0.996 1 0.995 0.995 Shorts 0.833 0.833 0.964 0.955 T-shirt 0.917 0.917 0.971 0.954 Table 4.9 presents the performance results for the YOLOv8s model in the task of segmenting clothing across various categories, as well as detailed metrics for each type of garment, focusing on segmentation masks. For the overall performance of the model across all categories ("all"), the model achieved a
Chapter 4 – System Development ___________________________________________________________________________ 62 precision of 0.941, a recall of 0.896, and an AP for an IoU of 0.5 of 0.965. These results indicate a high overall performance of the model when it comes to segmenting masks. In the dress category, the model had a precision of 0.849, a recall of 0.917, and a mAP of 0.941. Even though these results indicate good performance, the precision is slightly lower than in other categories. The jacket category demonstrated an excellent performance with a precision of 1.0 and a recall of 0.91, resulting in a mAP of 0.99, demonstrating an excellent ability of the model in identifying and segmenting jackets. In the pants category, a precision of 0.989, a recall of 0.875 and a mAP of 0.97 were obtained, indicating almost a perfect performance in the mask segmentation process. The polo category had a precision of 0.953, recall of 0.837, and mAP of 0.96. Despite the high precision, the lower recall indicates that the garment may not be fully detected. Shirts ("shirt") achieved a precision of 0.863, a recall of 0.875, and a mAP of 0.925, showing balanced performance, but with room for improvement. It is noteworthy that shoes had an almost perfect precision of 0.996, a recall of 1.0, and a mAP of 0.995, indicating excellent performance in this category. In the case of shorts, the precision was 0.992, the recall was 0.833, and the mAP was 0.964. The low recall indicates that some instances may not have been detected despite the high precision. Lastly, the t-shirts category demonstrated good overall performance with a reasonable balance between precision and recall, being 0.887 and 0.917 respectively, also achieving a mAP of 0.971. According to Figure 4.10, the confusion matrix for the YOLOv8s model provides detailed insight into its performance in the segmentation of clothing across various categories.
Chapter 4 – System Development ___________________________________________________________________________ 63 Figure 4.10 – Illustration of the relationship between predicted and true values in the confusion matrix of each clothing category. This matrix – Figure 4.10 - provides an evaluation of where the model performs well and where it fails. The lower value for shorts may be attributed to misclassification with dresses and shirts, likely due to the similar poses of the models wearing these garments. Furthermore, t-shirts and polos are sometimes confused probably due to their similar shapes. Figure 4.11 illustrates some results from segmentation.
Chapter 4 – System Development ___________________________________________________________________________ 64 Figure 4.11: Example of inference images with category and respective segmentation mask: (a) Original image of pants; (b) Segmented image of pants; (c) Original image of a shirt; (d) Segmented image of the shirt. After segmentation, the extracted mask was used to determine the colour of the clothing item. A pixellevel analysis of the resulting mask was conducted since the model generally performs well in classification. This involves extracting the segmentation mask of the clothing and converting it to the HSV colour space to determine the exact colour of the garment. Figure 4.12 illustrates the entire workflow.
Chapter 4 – System Development ___________________________________________________________________________ 65 Figure 4.12: Workflow for colour extraction. Using the black-and-white mask of the segmented region, a bitwise AND operation was performed on the original image. This process produces an image containing only the region of interest, namely segmented clothing, in the RGB colour space. This image is converted into a HSV colour space in order to extract the colour components. Due to the separation of the chrominance component (Hue) from the saturation and value components, the HSV colour space offers advantages when selecting a desired hue as compared to the RGB colour space. The hue of a particular colour is defined as an angular value between 0 and 360 degrees, with 0 degrees representing pure red, 120 degrees representing green, and 240 degrees representing blue. Since the context of the problem does not require distinguishing between light and dark shades or specific colours, the primary colours to be detected were defined. Therefore, specific colour intervals were used to define the Hue component analysis, including black, white, red, green, yellow, blue, orange, brown and pink. The colour extraction process, although not quantified by specific metrics such as accuracy or precision, contributes significantly to the overall functionality of the system by accurately identifying and distinguishing various clothing colours. By converting the segmented regions into the HSV colour space, it is possible to isolate the hue component, which simplifies the identification of the primary colours regardless of variations in lightness or saturation. Despite the absence of quantitative metrics, the qualitative assessment confirms that the defined hue intervals effectively cover the intended primary colours. This internal validation supports the system's capability to provide accessible and precise clothing identification. 4.4.2.2 Conclusions of Second Experiment Following fine-tuning with data augmentation techniques, the YOLOv8s model was effective at segmenting various categories of clothing. Almost all categories showed high performance metrics for segmentation
Chapter 4 – System Development ___________________________________________________________________________ 66 masks, with jackets and shoes achieving near-perfect results. Even so, there is still room for improvement in categories such as dresses and shorts, where recall could be improved. Following the segmentation process, the mask of the segmented region was used to analyse the colours present in the clothing item. Due to the fact that it is not necessary to differentiate between detailed colour nuances in this context, nine essential colours were defined. Therefore, the HSV colour space was used, in which the Hue component was defined to encompass the defined colours. 4.4.3 Third Experiment - Identification of Modifications on Clothing The third experiment aimed at identifying defects in clothing, i.e. , holes and stains, by using computer vision advances, namely deep learning algorithms. This experiment was divided into two stages, the first one aiming at the creation of a small dataset for stain detection using a neural network, and the second one at expanding the dataset and introducing other types of defects, as well as using a different neural network for defect detection and classification. 4.4.3.1 First Stage During the initial phase of the first stage, images were captured in order to establish a dataset. Note that, at the time of the study, there was no publicly available database on clothing with stains. Then, a neural network was trained using transfer learning techniques to detect the presence of stains, as shown in Figure 4.13. Figure 4.13: Workflow of the methodology used for detecting stains (Rocha et al ., 2023a). As a result, a dataset with approximately 104 images was compiled. The garments in this dataset may display multiple stains that are distributed across various sections, resulting in a total of approximately 300 stains. These images were taken from personal wardrobes where coffee and wine stains were intentionally applied to clothing, as it is shown in Figure 4.14.
Chapter 4 – System Development ___________________________________________________________________________ 67 Figure 4.14: Stains distribution on clothing from the dataset, a) coffee stain; b) wine stain; c) multiple stains; d) backside stain (Rocha, Soares, et al ., 2023a). Despite garments having two types of stains, the annotation process was carried out considering only one class, namely, "stain". Due to the annotations in the dataset and the limited amount of data, the transfer learning method was employed. By applying transfer learning, mask_rcnn_R_50_FPN_3x with ResNet50 was utilized as the backbone. A Mask R-CNN for stain detection was implemented using the Detectron2 library,(Wu et al ., 2019) the hyperparameters are presented in Table 4.10. Table 4.10: Hyperparameters for fine-tuned Mask R-CNN (Rocha, Soares, et al ., 2023a). Parameters Value Iterations 360 Leaning rate 0.001 Batch size 1 As an extension of Faster R-CNN, Mask R-CNN incorporates instance segmentation as an enhancement. By adding pixel-level segmentation, this network has an advantage over Faster R-CNN, which labels every pixel that belongs to a detected object. The main results of the network performance and model losses from Mask R-CNN are presented in Table 4.11 and Table 4.12, respectively. Table 4.11: A report on the evaluation of Common Objects in Context (COCO) (Rocha, Soares, et al ., 2023a). Method AP AP at IoU = 0.50 AP at IoU = 0.75 Bounding Box 0.549 0.857 0.672 Segmentation 0.540 0.858 0.674 Table 4.12: A summary of the losses associated with the model (Rocha, Soares, et al ., 2023a). Total Loss Loss Classification Loss Box Regression Loss Mask
Chapter 4 – System Development ___________________________________________________________________________ 68 0.480 0.053 0.154 0.240 According to the COCO evaluator, the results presented in Table 4.11 are promising, particularly for the AP at IoU = 0.50 of 0.857, despite the small quantity of data. It is noteworthy that the highest verified loss was 0.240 (Table 4.12), indicating that the segmentation of the stain presented a more challenging task. The evaluation of the model allowed to conclude that the misclassifications were mainly related to the detection of brand logos in the clothing and the low contrast between the stain and the clothing colour, as shown in Figure 4.15. Figure 4.15: Example of an item of clothing that has been misclassified (Rocha, Soares, et al ., 2023a). This first stage presents an innovative method for detecting stains from a clothing image. The presented method, using a dataset of clothing with stains caused by wine and coffee, demonstrates the ability of a deep learning algorithm to accurately identify and locate stains on clothing. Despite the obtained results look promising, it is expected that they could be further improved when a larger dataset is used. 4.4.3.2 Second Stage After evaluating the results of the initial stage, which included only a limited number of images depicting only one type of defect, i.e. stain, it became apparent that a broader approach could improve the results. As part of the study, it was decided to explore the inclusion of a new defect type, holes, along with a substantial expansion of the dataset. This expansion was intended to not only diversify the types of considered defects, but also enrich the dataset to provide a more comprehensive description of the problem. Moreover, it employed a different neural network architecture that addressed the problem through object detection rather than segmentation. In addition, techniques of data augmentation were
Chapter 4 – System Development ___________________________________________________________________________ 75 YOLOv5m6 model. However, considering the context of this application, prioritizing recall over precision may be more beneficial. In other words, it is preferable for the model to have fewer false negatives than false positives. Figure 4.20 illustrates an example of a false positive, where the buttonholes are misinterpreted as a defect. This highlights the importance of having representative images that include such scenarios. Figure 4.20: Example of a misinterpretation of a defect: a) original image; b) predicted image from model YOLOV5m6; c) predicted image from model YOLOV5l6 (Rocha, Pinto, et al ., 2023). Based on the performance of the models across all experiments, the model YOLOv5l6 exhibits the best generalization to unseen data, i.e. , test dataset, when compared to other models. Figure 4.21 displays predicted images that encompass various scenarios, including variations in illumination, backgrounds, multiple defects, as well as challenging areas.
Chapter 4 – System Development ___________________________________________________________________________ 76 Figure 4.21: Examples of predicted images from the YOLOv5l6: a) single stain detection; b) multiple stain detection; c) multiple hole detection; d) hole detection near the seam (Rocha, Pinto, et al ., 2023). The main constrain of using this model in a practical context is the computational cost. Such impact was evaluated through the calculation of the inference time for the test with the dataset. Table 4.18 exhibits the results of the inference time on the test dataset. Table 4.18: Inference time on test dataset for the different YOLOv5 models tested (Rocha, Pinto, et al ., 2023). Model Inference Time (s) YOLOV5s6 0.0092 YOLOV5m6 0.0112 YOLOV5l6 0.0157 These results suggest that, despite incurring computational costs, all models are believed acceptable due to the negligible required time. Thus, the findings indicate that implementing object detection technology with augmented data may be a successful strategy for identifying defects in clothing. This study stands out from previous research work (L. Cheng et al ., 2023), as it evaluates defect detection on clothing as a whole, instead of focusing on zoomed-out images of defects on stretched textiles and without background.
Chapter 4 – System Development ___________________________________________________________________________ 77 Upon comparing the new dataset created in this work with those from textile industry it becomes clear how challenging the task presented in this work is, mainly due to the presence of certain features such as buttonholes, which could potentially be interpreted as defects. Furthermore, this approach was proved to be effective in high demanding contexts, namely with wrinkled textiles, various backgrounds, different illumination, and diverse partners. 4.4.3.8 Conclusions of Third Experiment and future work The detection and classification of clothing defects was successfully carried out with a deep learning approach. An enhanced dataset was constructed with new types of stains and with holes. Through the fine-tuning of three models from the YOLOv5 object detector, a total of three experiments was carried out. Data augmentation was demonstrated to be essential for a better generalization of the model, allowing to achieve higher precision results. Still, recall values demonstrate that the model can be improved to minimize false negatives. Maximum precision, recall, and average precision (AP) values of 0.915, 0.543, and 0.747, respectively, were achieved with the YOLOv5l6 model for defect detection and classification. The detection of holes was found to be more challenging than the detection of stains, which emphasizes the importance of integrating the findings of this study in an automatic wardrobe that could take multiple cameras capture from different clothing perspectives. The dataset built in this work demonstrates that object detection technology can be used to accurately detect and classify defects on clothing autonomously. Moreover, it represents the first step for the creation of a mobile application that can effectively detect multiple defects on clothing, based on the integration of these findings in an automated closet system as a future step.
5 SMART WARDROBE PROTOTYPE Chapter Overview This chapter presents a detailed description of the smart wardrobe prototype and its key components, followed by an overview of the mobile application developed to control the system. The aim of this chapter is to outline the design and essential elements necessary for the development of a mechatronic prototype — a smart wardrobe system capable of identifying clothing characteristics such as type, colour, and alterations within a controlled environment. 5 Smart Wardrobe Prototype 5.1 Prototype System Design 5.2 Main Components of the Developed Prototype 5.3 Mobile Application Interface
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 79 5.1 Prototype System Design The primary objective is to simulate a prototype closely resembling real-world scenarios, from which the entire hardware infrastructure can be seamlessly transferred to individual wardrobes. Therefore, it is essential to carefully select adaptable components to ensure the prototype's effectiveness. Hence, the prototype must meet a number of specifications that are essential to the integration and functionality of the product, namely: • Identify the type, colour, and modifications of the garment; • Implement a NFC reader to extract each garment's unique identifier; • Maintain a controlled lighting environment to ensure accurate assessment; • Automate the rotation of garments to evaluate both sides; • Include a camera for acquiring images; • Establish TCP/IP communication to control all hardware components remotely. To prove the concept (Rocha et al ., 2020) and devise a readily deployable solution, a compact IKEA wardrobe measuring 80 cm by 50 cm by 30 cm (height x width x depth) was initially selected for the evaluation of appropriately sized garments (Silva et al ., 2023). The system, however, faced some limitations. The constrained dimensions of the wardrobe restricted its use to garments of diminutive proportions, primarily tailored for children. Further, the imaging process was hampered by partial capture, which necessitates the rotation of the camera and the subsequent application of a stitching algorithm to incorporate different segments of the image. Also, as the effectiveness of the NFC reader depends on the orientation of the garment, a variation in the accuracy of the read occurred. Lastly, the wardrobe's irremovable support infrastructure complicated the process of garment replacement, making a quick garment change a difficult task. As a result, these challenges led to the design of a new prototype architecture, as shown in Figure 5.1.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 80 Figure 5.1: iSight Smart Wardrobe Prototype. The dimensions of the smart wardrobe prototype depicted in Figure 5.1, 123 cm by 120 cm by 63 cm (height x width x depth), enabled its effectiveness to demonstrate the proposed solution and its transportation for testing purposes. In the next subsection a detailed description of its components is provided. 5.2 Main Components of the Developed Prototype Before detailing the individual components of the smart wardrobe prototype, an overview is provided in Figure 5.2, which illustrates the schematic layout of the prototype and highlights the interconnections between key hardware components, including the Raspberry PI, DC motor, camera, lighting system, and NFC reader.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 81 Figure 5.2 - Schematic layout of the smart wardrobe prototype. This schematic, Figure 5.2, provides a representation of the Smart Wardrobe Prototype, highlighting the key components and their roles within the system. These components are outlined and described in detail below: • Raspberry PI: Acts as the central unit, managing the operation and coordination of all other hardware elements. • Camera: Captures images of the garments to facilitate visual analysis and identification. • Lights: Provide controlled illumination, ensuring clear and consistent image capture. • DC Motor: Rotates the garments, allowing the camera to assess them from multiple angles. • NFC Reader: Uniquely identifies each garment using Near Field Communication (NFC) technology. The system is ultimately accessed and controlled through a TCP/IP communication interface, enabling remote operation and integration with external devices. In the following subsections, a detailed analysis of each component is provided. 5.2.1 DC Motor Nema 17 stepper motor, model 42BYGH48-23D ( Motor de Passo Nema 17 p/ Impressora 3D - 42BYGH48-23D , 2024) - Figure 5.3 - is a type of electric motor widely used in applications requiring precise positioning, such as rotating a shaft at a specific angle. Due to its exceptional cost-effectiveness,
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 82 this motor has a high torque and a limited step angle of 1.8 degrees, which guarantees accurate and reliable performance. Figure 5.3 Nema 17 stepper motor. Table 5.1 provides the specifications of the DC motor used in the prototype. Table 5.1: Nema 17 motor specifications ( Motor de Passo Nema 17 p/ Impressora 3D - 42BYGH48-23D , 2024). Specifications Description Torque 55 N.cm Step 1.8 degrees Number of Steps 200 Phase Current 1.7A Number of Phases 2 Number of Wires 4 Shaft Diameter 5mm Motor Length 48mm 5.2.2 Drive Motor The driver motor A4988( Driver Para Motor - A4988 , 2024) - Figure 5.4is a widely used stepper motor driver renowned for its efficiency in controlling bipolar stepper motors. It provides an effective means of controlling the current and polarity of the motor coils, enabling precise motion control.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 83 Figure 5.4: Picture of the driver motor A4988 used in the developed prototype. Equipped with current limitation and adjustable settings, along with overheating protection and five different microstep resolutions (up to 1/16-step), this motor is designed to meet the demands of various motion control applications. Table 5.2 describes the main specifications of the A4988 driver motor. Table 5.2. Description of the A4988 specifications ( Driver Para Motor - A4988 , 2024). Specifications Description Control Method Step and direction control Logic Voltage 3-5.5V Motor Output Voltage 8-35V Microstep Resolutions Full-step, Half-step, 1/4-step, 1/8-step, 1/16-step Current Control Adjustable, allowing setting of maximum output current using a potentiometer, enabling utilization of voltages above nominal motor voltage for higher step rates Voltage Regulator Built-in Protection Features Current overload and short-circuit protection The specifications presented above make the A4988 stepper motor driver used in a wide range of motion control systems, offering versatility, reliability, and precise control over stepper motor operations. 5.2.3 Raspberry PI The Raspberry Pi 4 Model B ( Raspberry Pi 4 Model B , 2024) - Figure 5.5 - is a single-board minicomputer. This central processing unit is equipped with a 1.5GHz quad-core Arm® Cortex®-A72 CPU and 8GB of RAM. Because of its high-performance CPU and ample memory, the Raspberry Pi 4 Model B is ideally suited for processing-intensive applications. The advanced graphics capability and high-resolution video
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 84 output capabilities allow for robust multimedia processing. Moreover, both USB 3.0 and USB 2.0 ports, true Gigabit Ethernet, and dual HDMI outputs provide versatile connectivity and fast data transfers. Figure 5.5: Picture of a Raspberry Pi 4 Model B used in the developed prototype. 5.2.4 Camera The Official Raspberry Pi Camera Module V3 ( Raspberry Pi Camera Module 3 , 2024) - Figure 5.6 - equipped with a 12MP sensor, is designed to offer high-resolution imaging capabilities suitable for a variety of applications. This module utilizes the Sony IMX708 sensor ( Raspberry Pi Camera Module 3 , 2024), which boasts advanced features and specifications tailored for enhanced performance. Figure 5.6: V3 Oficial Raspberry Pi, 12MP, 120°, camera module, used in the developed prototype. The camera module's high resolution of 12MP ensures detailed image capture, while the sensor's compact size (7.4mm diagonal) and pixel dimensions (1.4μm × 1.4μm) contribute to high sensitivity and image quality. Table 5.3 details the camera specifications.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 91 Figure 5.13: Flowchart of wardrobe server. Figure 5.14 illustrates the initial state of the wardrobe when it is powered on. The wardrobe’s interior is entirely covered with chroma key fabric to provide a uniform colour for image acquisition, enhancing image consistency. The LED strips located on the left panel are powered by a 12V power supply, as shown in Figure 5.14a. Positioned centrally is the official Raspberry Pi Camera Module V3, as depicted in Figure 5.14b.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 92 Figure 5.14: General view of the wardrobe interior: a) LED strips applied on the left panel; b) camera placed in the middle of LED strips. The light source is essential for ensuring the quality of images captured by the camera, providing sufficient illumination for the clear identification of clothing items regardless of environmental conditions. In this technique, both the sensor (camera) and the light source are positioned on the same side of the object being captured. This method is typically employed to obtain detailed information about the texture and other surface characteristics of the object. This type of illumination can be categorized into two subtypes: one that primarily leverages specular reflections and another that predominantly utilizes reflected rays constituting diffuse reflection. In scenarios where light reflected by a flat surface is captured by the optical system, non-flat characteristics such as holes can reflect the light outside the lens's maximum acceptance angle, thereby revealing dark features, as schematized in Figure 5.15.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 93 Figure 5.15: Reflective illumination in hole defect. This image shows how light reflected from a flat surface is affected by nonflat features like holes, which can redirect light outside the lens's acceptance angle, creating dark areas that reveal surface defects. At the same time, the server, i.e. the Raspberry Pi mounted on the external left panel ( Raspberry Pi 4 Model B , 2024), is initialized and begins awaiting HTTP requests from the graphical user interface (GUI), i.e. mobile interface. If the request is to capture an image, the camera is initialized to execute the capture. As explained before, this specific camera was selected based on its 120° lens wide diagonal field of view, which allows for a minimal distance to effectively capture the entirety of the garment. This wide-angle lens ensures comprehensive coverage facilitating detailed imaging of the clothing item, ensuring no part of the garment is omitted in the capture process.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 94 Figure 5.16: Interior of the wardrobe showing the distance from the camera to the hanger and internal components. Upon receiving the rotation request, the motor situated on the top right side (Figure 5.16 and Figure 5.17), which supports the hanger (Figure 5.16 and Figure 5.17), is activated to rotate 180 degrees. This feature enables the system to capture garments from both sides, enhancing operational efficiency and flexibility. The ability to rotate the hanger motor by 180 degrees is crucial for processes that require access to both sides of the garment, such as inspection of modifications.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 95 Figure 5.17: Motor supporting the hanger. Another potential request that can be handled by the system involves reading the NFC tag attached to each garment. The placement of these tags on garments has been designed to be minimally invasive. Thus, the tag can be found either on the original clothing label, as illustrated in Figure 5.18.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 96 Figure 5.18: NFC tag attached to the clothing label. To facilitate this, an NFC reader has been integrated into the wardrobe system (Figure 5.19). This reader is tasked with retrieving the unique identifier (UID) code from the NFC tags, as per the user's request. Specifically, an NFC reader is strategically positioned on the exterior left side of the wardrobe to ensure seamless reading of the NFC tags affixed to the garments, as demonstrated in Figure 5.19. Figure 5.19: NFC Reader. The integration and connectivity of all system components with the Raspberry Pi controller are comprehensively illustrated in the schematic diagram provided in Figure 5.20.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 97 Figure 5.20: Schematic diagram illustrating the integration and connectivity of system components with the Raspberry PI controller This diagram (Figure 4.41) offers a detailed overview of the setup, highlighting the interfacing of the NFC reader, camera module, and motor rotation with the Raspberry PI, thereby elucidating the system's operational framework. The data captured from image processing and NFC tag reading is transmitted from the Raspberry PI to the GUI. Additionally, the mobile interface processes the output from the wardrobe server and, when necessary, communicates with AI models for inference.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 98 5.3 Mobile Application Interface The purpose of this section is to explore the development of a mobile interface, emphasizing compliance with accessibility standards. 5.3.1 Accessibility In the European Union (EU), the standard that encompasses accessibility requirements for the web and mobile applications is the EN301549 standard ( Introdução à Norma Europeia EN 301 549 - Acessibilidade.Gov.Pt , 2024). This standard is essentially a copy of the WCAG 2.1 'AA' compliance. In Portugal, these requirements are included in the Regulamento Nacional de Interoperabilidade Digital (RNID) as decreed by D.L. No. 83/2018(<i>DL n.o 83/2018 - Acessibilidade Dos Sítios Web e Das Aplicações Móveis - Acessibilidade.Gov.Pt</i>, 2024). According to this decree, "Accessibility" refers to the principles and techniques to be observed in the design, construction, maintenance, and updating of websites and mobile applications to make their content more accessible to users, especially people with disabilities. In compliance with Article 5, this accessibility entails several requirements regarding information and navigation which this study must also adhere to in order to be accessible to the blind. Therefore, in the context of information, the following aspects are considered: 1. Presentation of information and interface components in a manner perceivable by users. 2. Ensuring the functionality of the user interface, making sure that the components and navigation are actionable. 3. Information and the operation of the user interface must be easily understood. 4. Interpretation by assistive technologies to be reliably interpreted by a wide range of user agents. Adhering to these principles ensures that the iSight mobile application aligns with both EU and Portuguese accessibility standards, thereby facilitating an inclusive user experience for visually impaired individuals. 5.3.2 Mobile Interface Features Given the intended user base of Portuguese-speaking people with visual impairments, the application has been meticulously designed in Portuguese and developed using a cross-platform language. This strategic choice facilitates comprehensive testing and deployment across both iOS and Android devices, ensuring broad accessibility and functionality. To address the initial project requirements, a detailed block diagram
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 99 representing the entire layout of the application was developed and it is depicted in Figure 5.21. Figure 5.21: iSight mobile interface block diagram. The application is exclusively accessible to people with disabilities, leveraging screen readers or Braille displays as graphical interfaces between the blind user and the computer. These interfaces interpret the information displayed on the screen, enabling seamless interaction and navigation. Although developed using a cross-platform language, the application was primarily designed and tested for iOS devices to ensure robust performance and accessibility. Upon examining Figure 5.21, it becomes apparent that following the login process, six primary elements emerge: Tag, Clothes, Categories, Colours, Modifications, and Wardrobe. To navigate the application, users must enable VoiceOver, a gesture-based screen reader that allows iPhone use without visual input. This feature provides auditory descriptions of on-screen content, ranging from battery levels to caller identification and the current application in use. When the user touches or swipes the screen, VoiceOver announces the name of the element under the finger, including icons and
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 100 text, and enables interaction with buttons and navigation to other screens using specific VoiceOver gestures. The iSight main menu is presented below in Figure 5.22. Figure 5.22: iSight main menu. The first menu (Figure 4.43), NFC Tag, is used to read tags associated with clothing items. Upon selection, the user is informed of the need to ensure the tag is correctly placed on the clothing item, as illustrated in Figure 5.23. Once confirmed, the user can press the button to read the tag.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 107 Figure 5.29. Example of a returned result from colour analysis. Despite involving two HTTP requests - one for image capture and another for inference - the pipeline is swift, and the user is notified of the successful operation through a haptic alert on their device. For category requests, the pipeline remains the same, differing only in the inference response. After receiving the result, the user can always add the clothing item, being redirected to the form navigation panel as previously described. The Modifications menu is similar to the Category and Colour menus, but the later allows the analysis of clothing items from both sides using a button to rotate the item - Figure 5.30.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 108 Figure 5.30. Submenu included in the Modifications menu. The pipeline for this process differs in the model utilized, employing a specific model designed for modifications inference. This model is tailored to detect and analyse defects such as stains or holes in the clothing item. Upon completion of the inference process, the results are presented to the user, providing detailed information about the number of detected stains or holes on the garment, as illustrated in Figure 5.31. This detailed feedback enables the user to make informed decisions regarding the condition and maintenance of their clothing items, ensuring they can address any issues promptly. This functionality is particularly useful for maintaining the quality and appearance of their wardrobe, further enhancing the overall user experience.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 109 Figure 5.31. Example of a returned result from the Modifications Analysis. The final menu pertains to the wardrobe settings - Figure 5.32. In this menu, users can configure IP addresses and ports for communication with the wardrobe server and the AI model server through the Settings submenu. The NFC Tag submenu allows for reading NFC tags, similar to the first menu, but utilizing the NFC reader present in the wardrobe.
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 110 Figure 5.32. Wardrobe menu details. Although this feature may seem redundant, it is useful when the mobile device's NFC reader is unavailable. In summary, designing an application that specifically addresses the needs of blind people and incorporating advanced technologies such as NFC, AI models, this study represents a significant advancement in accessibility and usability for blind users. The integration of these technologies ensures that users can independently manage their wardrobe, make informed choices, and efficiently organize their clothing items. This comprehensive approach not only enhances the user experience but also empowers blind people by providing them with the necessary tools for greater autonomy and self-reliance in their daily lives. Chapter Summary This chapter offers a comprehensive overview of the development and components of the smart wardrobe prototype, detailing its design, core elements, and the associated mobile application interface aimed at
Chapter 5 – Smart Wardrobe Prototype ___________________________________________________________________________ 111 improving wardrobe management for visually impaired individuals. The prototype system was engineered to identify clothing characteristics such as type, colour, and modifications through a mechatronic setup, which is remotely controlled via a mobile application. The initial prototype encountered limitations related to spatial constraints and image capture difficulties. Consequently, the design was revised to a larger model. The revised prototype features a chroma key fabric lining to facilitate uniform image capture and incorporates a DC motor (Nema 17) for rotating garments to enable viewing of both sides. Key components of the developed prototype include the Nema 17 stepper motor, an A4988 motor driver, a Raspberry Pi 4 Model B, a high-resolution 12MP camera, an NFC reader utilizing the PN532 chip, a controlled LED strip for lighting, and a 12V power supply. NFC NTAG213 tags were employed to label clothing items, allowing for the reading of unique identifiers. The system operates through a Raspberry Pi, which serves as a local server handling requests from a mobile interface. These requests encompass capturing images, reading NFC tags, and controlling the motor for garment rotation. A wide-angle camera aids in capturing comprehensive images, while the motor rotates the hanger 180 degrees to provide a full view of both sides of the clothing. The NFC reader is strategically positioned to ensure efficient tag reading. An accessible and user-friendly mobile application was developed to enable blind users to independently organize and identify clothing items. Ensuring accessibility was a primary objective; therefore, the application adheres to standards such as EN301549 in the EU and WCAG 2.1 AA requirements. In Portugal, these standards are reinforced by the Regulamento Nacional de Interoperabilidade Digital (RNID) under DL No. 83/2018, ensuring that the application is perceivable, operable, understandable, and compatible with assistive technologies such as screen readers and Braille displays. The mobile interface was designed with a focus on a Portuguese-speaking user base with visual impairments, employing a cross-platform development approach for compatibility with both iOS and Android devices. Features include VoiceOver for iOS, facilitating interaction through auditory cues and gestures. The application supports NFC tag reading, closet management, and clothing categorization via AI analysis based on colour and type. Users interact with the system through the mobile application, which communicates with the Raspberry Pi via HTTP requests. This includes image capture, NFC tag reading, and motor control. The application processes the data and, when necessary, interfaces with AI models to provide detailed feedback on clothing items, including identification and recommendations based on user preferences. In summary, the iSight prototype is designed to offer an accessible and practical solution for wardrobe management, with the potential to significantly enhance independence and improve the daily lives of visually impaired users.
6 EVALUATION OF PROTOTYPE Chapter Overview The following chapter describes the results of a survey conducted to validate the iSight prototype. Members of ACAPO ( Associação dos Cegos e Amblíopes de Portugal ) participated in a structured testing protocol followed by a detailed questionnaire. The goal was to gather feedback on the usability, accessibility, and overall importance of the iSight prototype for the visually impaired community. 6 Evaluation of Prototype 6.1 Ethical Considerations 6.2 Testing Protocol 6.3 Questionnaire Methodology 6.4 Results and Analysis 6.5 Statistical analysis and relevant findings
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 113 Chapter 6 presents the evaluation and validation of the iSight prototype, designed to assist blind individuals in managing their clothing. The chapter starts with ethical considerations, ensuring participants' safety and well-being through ethical approvals and informed consent. The testing protocol involved a structured process where participants were introduced to the prototype, engaged in hands-on testing, and guided through scenarios to evaluate specific functionalities like colour and category evaluation, detection of modifications, adding new clothing items, and database verification. After the testing, participants provided immediate feedback and completed a comprehensive questionnaire. The questionnaire methodology involved developing a detailed set of questions, initially tested with ACAPO board members, to gather in-depth feedback on the prototype. The results and analysis section presents the findings from the questionnaire, covering demographics, types of visual impairment, technology use, accessibility of the mobile application, usability of the prototype, and the perceived importance of the iSight system. Statistical analysis was conducted to explore relationships such as the correlation between ease of navigation and user satisfaction, the impact of technology familiarity on ease of navigation and overall experience, and the effectiveness of the prototype in identifying clothing features. The analysis also examined user suggestions for improvements and compared feedback between blind users and those with low vision. The chapter concludes with a summary of the statistical tests, hypotheses, results, and significance levels, highlighting key findings and areas for further development. This comprehensive evaluation demonstrates the iSight prototype's potential to enhance the independence and quality of life for visually impaired users. Figure 6.1 summarizes the main steps and iterations in Chapter 5: Figure 6.1: Workflow of the iSight prototype testing and validation process.
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 114 This workflow illustrates the structured process followed in evaluating and validating the iSight prototype, from ethical considerations and testing protocols to detailed statistical analyses and user feedback, providing a clear overview of the chapter's content and methodology. 6.1 Ethical Considerations To ensure the safety and well-being of the participants, ethical considerations were paramount. A consent form is included in Appendices 2 and 3, which documents that the testing protocol and the questionnaire were submitted to the Ethics Committee for Research in Social and Human Sciences (Comissão de Ética para a Investigação em Ciências Sociais e Humanas - CEICSH) of the University of Minho and received approval under code CEICSH 185/2023. The corresponding ethical approval document is provided in Appendix 4. Prior to participating in the study, each participant signed an informed consent form acknowledging their understanding of the study's purpose, procedures, potential risks, and their right to withdraw at any time without any consequence. As a result of this consent, participants were made aware of their involvement and the nature of the study, adhering to ethical standards and protecting their privacy and rights. 6.2 Testing Protocol To ensure a comprehensive evaluation of the iSight prototype, a structured testing protocol was followed. Participants were given detailed instructions on how to use the prototype, which included the following steps: 1. Introduction to the Prototype: each participant was provided with an overview of the iSight prototype, including its purpose and its features. 2. Hands-on Testing: participants were given a set period to interact with the prototype independently, allowing them to explore its functionalities and interface. 3. Guided Scenario: To evaluate the iSight prototype, participants were guided through a series of tasks designed to test its functionality in the following real-world scenarios: a. Categorization and Colour Evaluation - Participants were asked to request an assistant to provide a piece of clothing for colour evaluation. They were instructed to place the clothing item in the wardrobe, ensuring it was parallel to the sides. Participants then
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 115 navigated to the "Colours" menu and captured an image of the clothing item, verifying the colour evaluation results with the assistant. Optionally, participants could repeat the process in the "Category" menu to determine the category of the clothing item and verify these results with the assistant as well. b. Evaluation of Modifications on a Clothing Item: participants requested an assistant to provide a clothing item with modifications, such as stains or holes. They placed the item in the wardrobe, ensuring it was parallel to the sides. Participants accessed the "Modifications" menu and captured an image of the clothing item, subsequently confirming the modification results with the assistant. c. Adding a Clothing Item: participants requested an assistant to provide a piece of clothing. They placed the clothing item in the wardrobe, ensuring it was parallel to the sides. Participants opened the "Clothing" section and selected "Add New Clothing Item." After capturing the image, they verified the automatic filling of characteristics such as category and colour, completing any additional information as desired. Participants also had the option to add the NFC tag automatically using the provided button. d. Verification in the Database: participants requested an assistant to provide a previously added clothing item. From the main menu of the application, they accessed the "NFC Tag" function, brought the tag of the clothing item close to the NFC reader, and confirmed that the item was registered in the database. Alternatively, they could use the wardrobe's NFC reader, located in the upper left corner, accessed via the "Wardrobe" menu followed by "NFC Tag”. 4. Feedback Session: after completing the tasks, participants discussed their experiences with a facilitator, providing immediate feedback on any issues or suggestions for improvement. Upon completing the testing protocol, participants were asked to complete a comprehensive questionnaire to document their experiences and evaluations. Figure 6.2 shows participants engaging with the iSight prototype during the hands-on testing phase.
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 116 Figure 6.2: ACAPO members interacting with the iSight prototype during the testing phase, showcasing hands-on engagement: a) Participant using the mechatronic device; b) Participant using the iSight mobile application. 6.3 Questionnaire Methodology To evaluate the effectiveness and usability of the iSight prototype, a comprehensive questionnaire was developed. The detailed questionnaire is provided in Appendix A5. Before distribution to the participants, a pilot test was conducted with two members of the ACAPO board. This pilot testing allowed for improvements and corrections in the semantic formulation of the questionnaire, ensuring clarity and effectiveness in capturing the necessary data. The questionnaire was designed to gather detailed feedback on various aspects of the iSight prototype. A total of 15 participants voluntarily took part in the study. The responses were analysed using SPSS 22, and a Likert scale was employed to measure levels of agreement, importance, or frequency across various questions. In the Likert scale used, a lower score indicated a negative opinion, while a higher score indicated a positive opinion. The questionnaire comprised 36 questions divided into six main sections, described as follows: I. Personal Identification: this section gathered demographic information about the participants, including age, gender, academic qualifications, and current occupation. a) b)
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 123 Figure 6.6: Visual representation of common causes of visual impairment (Burton et al ., 2021) III. Technology The responses in this third section highlighted the current state of technology adoption among the blind community. It provides insights on their familiarity with supportive technologies, their satisfaction levels, and the frequency of use. QIII.1: Do you currently use any application dedicated to clothing? The responses indicate that the majority of participants, 67 %, do not currently use any applications dedicated to managing clothing, while 33 % do use such applications. This suggests that there is a significant opportunity for the iSight prototype to fill a gap in technology usage among visually impaired people in the context of clothing management. For those who indicated usage of clothing-related applications, further details were sought on the specific applications used. One participant reported using colour detectors on an Android mobile device, and four participants reported using the Colorino device. This additional information highlights that among those using clothing-related technology, the Colorino device is the most commonly used. Colorino is a well-known colour identifier for visually impaired people,
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 124 which helps users distinguish colours in their clothing. The use of colour detection applications and devices emphasizes the importance of colour identification in clothing management for the visually impaired community. Interestingly, among the users of these technologies, only one user had a low literacy level. This suggests that the adoption of such devices is not significantly hindered by literacy barriers, indicating their user-friendly nature and accessibility. Additionally, the majority of these people are still engaged in professional activities and have not yet retired. This detail highlights the continued need for practical and efficient solutions to support visually impaired people in their active professional lives, ensuring they can manage their clothing choices effectively and maintain their professional appearance. QIII.2: To what extent do you feel comfortable using technological devices such as smartphones and computers? The responses varied, with the majority indicating a moderate level of comfort. Specifically, 40 % of the participants reported feeling very comfortable, 53.3 % felt somewhat comfortable, and 6.7 % felt somewhat uncomfortable. These data suggest that most participants feel, at least, somewhat comfortable using technological devices, with a substantial portion feeling very comfortable. This overall comfort with technology is highly motivating for the adoption of the iSight prototype, as it indicates that most users are likely to be receptive to using a new technological solution for clothing management. Understanding these comfort levels helps in designing user interfaces and features that cater to the varying degrees of technological proficiency among the participants. QIII.3: How often do you use technological devices in your daily life? The frequency of technological device usage was overwhelmingly in favour of daily use. Fourteen participants indicated daily use ( ca . 93.3 %), while only one participant representing ca. 6.7 % reported occasional use. These data indicate that most participants use technological devices daily. This high frequency of daily use further supports the potential acceptance and integration of the iSight prototype into their daily routines. The prevalence of daily technology use among participants highlights their familiarity and comfort with digital tools, making them well-suited for adopting new technological solutions like the iSight prototype. QIII4.: How would you rate your ability to use software and applications specifically designed for blind people, such as screen readers?
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 125 The responses showed that while a significant number of participants (46.7 %) rate their skills as average, there is also a considerable proportion with very high (26.6 %) and high (13.3 %) proficiency levels. Only 13.3 % rated their skills as low. This range of abilities underscores the importance of designing the iSight prototype to be accessible and user-friendly for users with varying levels of technical proficiency. By ensuring that the prototype is intuitive and easy to navigate, it can cater to the needs of both highly skilled users and those with less experience using specialized software for the visually impaired. QIII.5: To what extent do you believe that technologies improve your quality of life and independence? Ten participants, representing ca. 66.7 %, felt that technologies improve their quality of life and independence very significantly, while five participants, 33.3 %, believed that technologies improve their lives significantly, with only one participant, 6.7 %, felling that the improvement is moderate. This high level of perceived benefit from technology underscores the critical role that technological solutions like the iSight prototype can play in enhancing the daily lives and autonomy of visually impaired people. Designing the iSight prototype to meet these expectations can further reinforce its positive impact and ensure its widespread adoption and use. QIII.6: How often do you seek to learn and explore new technologies adapted to your needs? The interest in learning and exploring new technologies adapted to visually impaired people’s needs was high. The gathered data reveal that a significant majority of participants, 73.3 %, are always looking to learn and explore new technologies, while 20 % do so regularly, and only one participant, representing 6.7 %, rarely seeks new technologies. This high level of engagement with adaptive technologies suggests that the iSight prototype has a receptive audience eager to integrate new solutions into the daily lives of the participants. This enthusiasm for exploring new technologies highlights the potential for positive reception and sustained use of the iSight prototype among the visually impaired community. QIII.7: How do you evaluate the support and accessibility of technologies for blind people? The support and accessibility of technologies for visually impaired people received mostly positive feedback. Specifically, 53.3 % rated the support as satisfactory, while 40 % rated it as very satisfactory, and only 6.7 % had no opinion. This feedback underscores the importance of continuing to develop and refine assistive technologies to meet the needs and expectations of visually impaired users. Ensuring that the iSight prototype offers robust support and high accessibility will be crucial for its success and adoption
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 126 among this group of users. QIII.8: To what extent do you feel that technologies facilitate greater social and professional inclusion in your life? The role of technologies in facilitating greater social and professional inclusion was overwhelmingly positive. The obtained data reveal that 93.3 % of the participants believe that technologies contribute very significantly to their social and professional inclusion, while the remaining 6.7 % consider the impact to be just significant. This highlights the critical role that technology plays in enabling visually impaired people to participate more actively in social and professional settings. The strong endorsement of technology's role in inclusion supports the continued development and implementation of innovative solutions like the iSight prototype to further enhance these opportunities for visually impaired users. QIII.9: Do you consider technologies to be a useful tool for your education or training? The usefulness of technologies as tools for education or training was highly endorsed. The collected data indicate that all participants find technologies to be beneficial for their education or training, with 46.7 % rating them as extremely useful and 53.3 % as very useful. This strong endorsement highlights the importance of integrating educational and training functionalities into the iSight prototype. By ensuring that the prototype supports educational and training needs, it can provide substantial benefits to visually impaired people, enhancing their learning experiences and professional development opportunities. IV. Accessibility of the iSight Mobile Application Feedback on the accessibility and usability of the iSight mobile application was gathered through detailed questions in the fourth section. Participants provided valuable suggestions and opinions, which were crucial for refining the application's user interface and ensuring it met the high standards of accessibility required by the visually impaired community. QIV.1: How do you evaluate the ease of navigation in the mobile application? The responses were generally positive, indicating a favourable user experience. A total of 66.7 % of participants rated the navigation as easy, while 33.3 % rated it as very easy. This feedback suggests that the mobile application's interface is user-friendly and accessible for visually impaired users. The high ratings for ease of navigation indicate that the design and layout of the application effectively met the needs of the target user group. Ensuring ease of navigation is crucial for the successful adoption and
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 127 continued use of the application, as it directly impacts the overall user experience. The positive evaluation of navigation within the mobile application highlights the importance of maintaining simplicity and intuitiveness in the user interface design. This will help users efficiently perform tasks and access features without unnecessary complexity or frustration. The feedback also provides valuable insights for further refinement and improvement, ensuring that the application remains accessible and user-friendly for all users. QIV.2: Does the mobile application provide image descriptions or textual alternatives for visual elements adequately? The responses indicated a strong overall satisfaction with the application's provision of image descriptions and textual alternatives. Specifically, 53.3 % of participants rated the descriptions as excellent, 26.7 % as good, and 20 % as adequate. No participants rated the descriptions as poor or very poor. The positive feedback on the application's image descriptions and textual alternatives emphasizes the effectiveness of these features in enhancing accessibility. Providing clear and comprehensive descriptions allows visually impaired users to fully engage with visual content, making the application more inclusive and usable. The high ratings for image descriptions and textual alternatives suggest that the application meets or exceeds the expectations of most users in this regard. This aspect of the application's accessibility is crucial, as it ensures that all users, regardless of visual ability, can access and understand the content presented within the app. The feedback also highlights the importance of continuing to prioritize and refine these features to maintain high standards of accessibility and user satisfaction. QIV.3: How do you evaluate the clarity and organization of the textual content in the mobile application? The data indicate that the majority of participants found the textual content to be either very clear or clear. Specifically, 53.3 % of participants rated the content as very clear, while 46.6 % rated it as clear. There were no responses indicating that the content was unclear. The high ratings for clarity and organization of textual content suggest that the application effectively communicates information in a way that it is easily understandable for visually impaired users. This clarity is essential for ensuring that users can navigate the application and access the information they need without confusion or difficulty. The positive feedback regarding the textual content highlights the importance of maintaining high standards for clarity and organization. Clear and well-organized content helps users efficiently finding and comprehending the information they need, enhancing their overall experience with the application. This feedback is valuable
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 128 for ongoing improvements, ensuring that the application continues to meet the needs of its users effectively. QIV.4: How do you evaluate the overall user experience of the mobile application? The responses show a high level of overall satisfaction with the mobile application. Specifically, 66. 7% of participants rated their experience as very satisfactory, while 33.3 % rated it as satisfactory. No participants rated their experience as unsatisfactory or very unsatisfactory. The high ratings for overall user experience suggest that the mobile application successfully meets the needs and expectations of visually impaired users. The positive evaluations indicate that the application provides a satisfactory and user-friendly experience, which is crucial for ensuring continued use and acceptance. The feedback on the overall user experience emphasizes the importance of maintaining a high level of user satisfaction through continuous improvements and updates. By addressing any potential issues and incorporating user feedback, the application can further enhance its usability and effectiveness, ensuring it remains a valuable tool for visually impaired people. V. Usability of the iSight Prototype The usability of the iSight prototype was assessed through questions focused on the ease of use and utility of the device. Participants' feedback indicated how intuitive they found the prototype and its effectiveness in meeting their needs. The fifth section also explores the preferred methods of technology delivery and availability for the blind community. QV.1: Evaluate in the ease of use of the iSight prototype, which includes both the mobile application and the smart wardrobe. The responses were overwhelmingly positive, indicating a high level of user-friendliness and accessibility. A total of 53.3 % of participants found the iSight prototype to be very easy to use, while 46.7 % found it just easy to use. No participants reported difficulties or rated the ease of use as less than easy. This feedback suggests that the iSight prototype successfully meets the usability requirements of its target user group, ensuring that the combination of the mobile application and smart wardrobe is intuitive and straightforward for visually impaired users. The high ratings for ease of use indicate that the design and functionality of both the mobile application and the smart wardrobe are well-integrated and accessible. This is crucial for encouraging the adoption and consistent use of the iSight prototype, as users are more likely to engage with technology that they find easy to navigate and operate.
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 129 QV.2: Evaluate the accuracy of the iSight prototype in identifying the category and colours of clothing items. For the accuracy in identifying the category of clothing items, 60 % of participants rated it as "very precise," while 40 % rated it just as "precise." For the accuracy in identifying the colours of clothing items, 80% of participants rated it as "very precise," while 20 % rated it as "precise." Most participants found the iSight prototype to be highly accurate in identifying both the category and colours of clothing items. However, a small number of participants (20 %) reported a divergence in their responses, rating the accuracy as "precise" for category identification and "very precise" for colour identification. The remaining participants provided consistent ratings for both category and colour accuracy. This suggests that while the iSight prototype generally performs well in both areas, there is a perceived higher accuracy in colour identification compared to category identification among some users. QV.3: Evaluated the iSight prototype's ability to detect stains and identify NFC tags. The responses for detecting stains were overwhelmingly positive, with the majority, 86.7 %, rating it as "highly effective." Similarly, the identification of NFC tags also received high ratings, with most participants, 86.7 %, finding it "highly effective." There was a strong correlation between the participants' ratings for these two functionalities, indicating that users generally found both features to be highly effective. Specifically, participants who rated the iSight prototype as "very precise" for both category and colour identification consistently rated its ability to detect stains and identify NFC tags as "highly effective." Participants who rated category identification as "precise" but colour identification as "very precise" also showed high ratings for detecting stains and NFC tags, though with some minor variations. Those who rated both category and colour identification as "precise" tended to rate stain and NFC tag detection as "effective," though some rated these capabilities higher. This consistency across responses suggests that the iSight prototype performs reliably across different functionalities, providing users with an effective tool for managing their clothing. The strong performance in colour identification and NFC tag detection is particularly noteworthy, as it highlights the prototype's ability to handle complex tasks with high precision. In summary, the usability and accuracy of the iSight prototype are highly rated by participants, indicating that the prototype meets the needs and expectations of visually impaired users. The positive feedback on ease of use and accuracy underscores the effectiveness of the design and functionality of the iSight prototype. Continuing to refine these aspects based on user feedback will be important for further enhancing the user experience and ensuring the successful adoption of the technology. The strong
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 130 correlation between responses across different functionalities reinforces the prototype's overall reliability and effectiveness in assisting visually impaired people in managing their wardrobe. The additional comments and suggestions provide valuable insights for future improvements, ensuring the system remains user-centric and responsive to the needs of its users. VI. Assessment of the Importance of iSight This assessment aims to evaluate the perceived importance of the iSight prototype among users, specifically examining its effectiveness in facilitating the clothing selection process, identifying specific characteristics and modifications in garments, and enhancing wardrobe organization. By analysing user feedback and satisfaction, the study seeks to elucidate the prototype's impact and pinpoint areas for potential enhancement. QVI.1: Evaluate the importance of the iSight prototype through two statements concerning its functionality. S1: "iSight facilitates the process of selecting clothes, while allowing identification of characteristics and modifications to the original state." S2: "iSight organizes my clothes effectively." The responses were compellingly positive. For the statement S1, 66.7 % of participants strongly agreed, and 33.3 % just agreed. Similarly, for the statement S2, 66.7 % of participants strongly agreed, while 33.3 % just agreed. These results indicate a high level of agreement among participants regarding the importance and effectiveness of the iSight prototype in both facilitating the selection of clothes and organizing them effectively. The majority of participants strongly agree that the prototype significantly aids in these aspects, suggesting that the system meets critical needs for visually impaired users in managing their wardrobe. The consistent positive responses across both statements reflect the perceived value of the iSight prototype in improving daily routines and enhancing the independence of visually impaired people. This feedback underscores the system's potential to make a meaningful impact on the lives of its users by providing practical and reliable solutions for clothing management. QVI.2: To assess whether the use of the iSight prototype increased their feelings of confidence, self-esteem, well-being, and independence through four statements. S1: “The use of iSight increased my sense of confidence”. S2: “The use of iSight increased my sense of self-esteem”. S3: “The use of iSight increased my sense of well-being”. S4: “The use of iSight increased my sense of independence”.
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 131 The responses indicate a very positive impact of the iSight prototype on these aspects. For S1, 66.7 % of participants strongly agreed, whereas 33.3 % just agreed. For S2, 66.7 % of participants strongly agreed, while 33.3% agreed. For S3, 66.7 % of participants strongly agreed, and 33.3 % agreed. Finally, for S4, 66.7 % of participants strongly agreed, and 33.3 % agreed. There is a high level of consistency in the responses across all four statements, with all participants either strongly agreeing or agreeing that the iSight prototype positively impacted their confidence, self-esteem, well-being, and independence. This suggests a robust and consistent positive perception of the prototype's benefits. All participants who strongly agreed on one statement also strongly agreed on the other statements. Similarly, participants who agreed on one statement also agreed on the other statements. This uniformity in responses indicates that participants consistently experienced an increase in confidence, self-esteem, well-being, and independence through the use of the iSight prototype. The lack of divergent responses further supports the conclusion that the iSight prototype has a comprehensive positive effect on users' lives. The uniform positive feedback across all four statements underscores the significant impact of the iSight prototype on enhancing the quality of life for visually impaired users. The increase in confidence, self-esteem, wellbeing, and independence highlights the prototype's effectiveness in addressing the everyday challenges faced by these users. QVI.3: To assess the importance of the iSight prototype in identifying clothes, colours, and stains/dirt. Participants uniformly recognized the iSight prototype as a crucial tool in their daily lives. For the importance of iSight in identifying clothes, 66.7 % of participants found it extremely important, and 33.3 % found it very important. Similarly, for identifying colours, 66.7 % of participants found it extremely important, and 33.3 % found it very important. Regarding the identification of stains/dirt, 66.7 % of participants found it extremely important, and 33.3 % found it very important. These responses indicate that participants uniformly found the iSight prototype to be an extremely important tool in identifying clothes, colours, and stains/dirt. The high percentage of participants rating it as "Extremely Important" reflects the prototype's significant role in aiding their daily routines and enhancing their independence. This consistency in high ratings across all three areas highlights the comprehensive utility of the iSight prototype in assisting visually impaired people in managing their clothing effectively. QVI.4: Whether they believe that the iSight prototype would improve their quality of life in relation to clothing management.
Chapter 6 – Evaluation of Prototype ___________________________________________________________________________ 132 The responses indicate that the majority of participants believe that iSight would significantly enhance their quality of life. For the statement, "iSight improves my quality of life in relation to clothing management," 80.0 % of participants indicated that it greatly improves their quality of life, while 20.0 % indicated that it significantly improves their quality of life. These responses reflect a strong consensus among participants regarding the beneficial impact of the iSight prototype on their daily lives. The majority also believes that the iSight prototype greatly improves their quality of life, while the remaining participants still view it as a significant improvement. This indicates that the iSight prototype is perceived as a valuable tool that can substantially aid visually impaired people in managing their clothing more effectively, ultimately enhancing their overall quality of life. QVI.5: About the specific aspects of their daily lives where the iSight system has a positive impact. The options included facilitating the choice of clothes for specific occasions, helping to identify the colours of clothes more effectively, enabling the identification of stains/dirt on clothes, contributing to greater autonomy in managing their wardrobe, and other unspecified benefits. The responses were highly positive, with every participant selecting all the options, except for one participant who did not select the first option. QVI.6 Do you have any specific suggestions or improvements you would like to see implemented in the system to make it even more useful in your daily life? To provide a comprehensive overview and avoid redundancy, feedback from section IV (Question: "Please share any additional comments or specific suggestions to improve the mobile application based on your experience.") and section V (Question: "Please share any additional comments or specific suggestions to improve the system based on your experience.") were combined. The aggregated responses highlight the innovative nature of the iSight prototype and its potential to significantly enhance the daily life for visually impaired people. Users appreciate the application’s ability to aid in clothing selection, identification of characteristics, and facilitating organization, noting that it could help them dress independently and efficiently. Many users suggested making the application faster by including fewer menus to streamline the process. The utility of the application is evident, with participants emphasizing its usefulness in daily routines and expressing eagerness to have such a system at home. Several users suggested enhancements to the application's features, such as including detailed information about clothing fabrics, reading label characteristics, and identifying stains more accurately. The application’s ability to distinguish colours and similar items, like shoes, was also highlighted as a critical area for improvement. Despite