scieee AI-readable full text Open interactive document viewer

Machine Learning algorithms to address the polarity and stigma of mental health disclosures on Instagram

Merayo Álvarez, Noemí,Ayuso Lanchares, Alba,González Sanguino, Teresa Clara

Abstract

Producción Científica

Full text

1 of 17 Expert Systems, 2025; 42:e13832 https://doi.org/10.1111/exsy.13832 Expert Systems ORIGINAL ARTICLE OPEN ACCESS Machine Learning Algorithms to Address the Polarity and Stigma of Mental Health Disclosures on Instagram NoemíMerayo1 | AlbaAyusoLanchares2 | ClaraGonzálezSanguino3 1Signal and Communications Theory and Telematics Engineering, School of Telecommunications Engineering, Universidad de Valladolid, Valladolid, Spain | 2Department of Pedagogy, Faculty of Medicine, Universidad de Valladolid, Valladolid, Spain | 3Department of Psychology, Faculty Education and Social, Universidad de Valladolid, Valladolid, Spain Correspondence: Noemí Merayo ([email protected]) Received: 9 July 2024 | Revised: 16 December 2024 | Accepted: 23 December 2024 Funding: This work was supported by Universidad de Valladolid. Keywords: Instagram| machine learning| mental health| natural language processing| sentiment analysis| social networks| stigma ABSTRACT This research explores the social response to disclosures and conversations about mental health on social media, which is a pioneering and innovative approach. Unlike previous studies, which focused predominantly on psychopathological aspects, this study explores how communities react to conversations about mental health on Instagram, one of the favourite social media platforms among young people, breaking new ground not only in the Spanish context, but also on a global scale, filling a gap in international research. The study created a novel corpus by collecting and labelling comments on Instagram posts related to celebrity mental health disclosures, categorising them by polarity (positive, negative, neutral) and stigma. Additionally, the research implements machine learning algorithms to detect stigma and polarity in mental health disclosures on Instagram. While traditional techniques like Support Vector Machine (SVM) and RF (Random Forest) displayed decent performance with lower computational loads, advanced deep learning and BERT (Bidirectional Encoder Representation from Transformers) algorithms achieved outstanding results. In fact, BERT models achieve around 96% accuracy in polarity and stigma detection, while deep learning models achieve 80% for polarity and 87% for stigma, very high accuracy metrics. This research contributes significantly to understanding the impact of mental health discussions on social media, offering insights that can reduce stigma and raise awareness. Artificial intelligence can be used for more responsible use of social media and effective management of mental health problems in digital environments. 1 | Introduction Social networks have become one of the most widespread communication channels today and allow for a constant flow of information that reflects the attitudes, trends and opinions of our society in real time. In fact, there are currently 4.76 billion social network users worldwide (Datareportal 2023), equivalent to approximately 60% of the world's population. Looking at social networks in terms of monthly active users, the latest data suggests that Facebook remains the world's number one social network with nearly 3000 million users. In this context, Instagram has also consolidated its position among the top social media platforms, ranking fourth with 2 million users behind Facebook, Youtube and Whatsapp, with an average time per user per month of 12 h. When examining social media preferences based on age and gender, individuals aged 16–24 and young women aged 25–34 prefer Instagram as their top social platform. Indeed, in January 2023, nearly twothirds of Instagram's total audience were 34 years old or younger (51% of the total audience were between 13 and 17 years old; 33.7% between 18 and 24 years old, and 31.3% between 25 and 34 years old) (Statist2023). Furthermore, some recent studies show that This is an open access article under the terms of the Creative Commons Attribution-NonCommercial License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited and is not used for commercial purposes. © 2025 The Author(s). Expert Systems published by John Wiley & Sons Ltd. 2 of 17 Expert Systems, 2025 Instagram is the SN most used by young people (Oden and Porter2023). Similarly, the significance of mental health in society has grown in recent years, as the World Health Organisation reports a global increase in mental health issues. In 2019, 970 millions of people worldwide were living with a mental disorder, with anxiety and depression being the most prevalent conditions (World Health Organization 2024). Moreover, an European study from October 2023 revealed that 46% of EU citizens had experienced an emotional or psychosocial issue in the last 12 months, such as feelings of depression or anxiety (European Commission2023). When analysing the mental health landscape of younger populations, the situation becomes particularly concerning. According to UNICEF, in 2019, approximately one in seven adolescents globally, representing 166 million individuals (89 million boys and 77 million girls), were estimated to be affected by mental illness (UNICEF2021). Half of all mental disorders in young people develop before age 14, and 75% by their midtwenties. In fact, 3% of 12to 17yearolds affirm experiencing depression and 32% report anxiety. This issue extends to young adults, with 33.7% of those aged 18 to 25 reporting some form of mental illness. Particularly, in 2023, approximately 1 in 5 children and young people aged 8 to 25 in the UK were estimated to have a probable mental disorder, with prevalences of 20.3% in the 8 to 16 age group, 23.3% among those aged 17 to 19, and 21.7% in the 20 to 25 age group (NHS Digital 2023). In the same year (2023), 20.17% of 12–17 year olds in the USA reported at least one depressive episode during that year (Mental Health America2023). A recent study conducted in Spain claims that 42.1% of individuals have experienced depression, and 14.5% have had suicidal ideation or attempted suicide, with an average age of diagnosis at 26 years (Fundación Mutua2023). However, despite the commonality and importance of mental health, it is possible to say that mental health stigma still exists in our society, and despite the progress still is an issue necessary to address (Gronholm and Thornicroft2022). This construct refers to thoughts (beliefs, myths, attributions), emotions (such as reactions of fear or pity) and negative behaviours (usually discrimination, or desire to distance oneself), shared by the society towards a specific group, in this case people with mental health problems (Corrigan and Watson2002). In Spain, stigma towards mental health is present in the general population (González Sanguino etal.2023), and certain stigmatising beliefs and attributions have been similarly found in adolescents (González Sanguino etal.2024). On the face of it, leveraging social media to talk about mental health should promote acceptance and reduce stigma, as some studies show how highimpact posts by celebrities can promote awareness and helpseeking by reaching large audiences (Gronholm and Thornicroft2022; Jain, Pandey, and Roy2017; Lee2019; Lee, Yuan, and Wohn2021). However, due to the sheer volume and immediacy of opinions expressed on these platforms, it has become a complex phenomenon. As a result, it is unclear whether these social media posts are actually fostering acceptance and good quality knowledge or inadvertently perpetuating further stigmatisation and/or trivialization of mental health (Pavlova and Berkers 2022; Robinson etal.2018). In this social environment of widespread adoption of social networks, together with the constant increase of information and opinions in real time, the application of automated techniques becomes highly relevant. Thus, Artificial Intelligence (AI) allows these actions to be carried out jointly, specifically the branch of Natural Language Processing (NLP) (Mäntylä, Graziotin, and Kuutila 2018) through what is known as Sentiment Analysis. Sentiment analysis combines natural language processing and computational linguistics to explore the meanings of words and their context, with the aim of understanding the underlying emotional tones. This technology can be applied to discern emotional reactions in comments posted on social networks, enabling realtime trend tracking, understanding current or future behaviours. However, the linguistic nature of comments posted on social media platforms exhibits significant disparities compared to conventional language use (MartínezCámara et al. 2014), often posing substantial challenges. These unique features, including brief messages, missing context, grammatical errors, and a casual writing style, complicate the application of effective sentiment analysis techniques. Furthermore, in sentiment analysis, two main approaches are primarily used: lexiconbased techniques and supervised learningbased techniques. Lexiconbased techniques require dictionaries where words are labelled with emotional responses, such as polarity or associated emotions. However, in Spanish, the number of dictionaries is limited, and it is necessary to develop specific dictionaries for each area of application (Redondo etal.2007). Additionally, this method should be supplemented with techniques for identifying negation or language ambiguity, which adds complexity to these strategies, especially on social networks (Taboada etal.2011). In contrast, supervised learning techniques require labelled corpora, which are examples of opinions or comments that have been previously annotated. This allows machine learning algorithms to learn from this data and make efficient predictions (Arco etal.2021; Shah etal.2023). This method provides the benefit of allowing various machine learning algorithms to be flexibly used on the same dataset. Furthermore, certain studies, like the one mentioned in (Srivastava, Bharti, and Verma2022), have shown that supervised methods can outperform lexicalbased techniques in specific scenarios, so our research will follow the machine learning approach. Despite the multiple advantages that AI could provide, research integrating AI to address mental health and stigma in social networks is limited. NLP brings important benefits to mental health research by autonomously uncovering important insights and patterns in the data that might go unnoticed or unavailable to mental health experts, and might even be overlooked in manual reviews. In addition, expressions related to mental health often exhibit greater emotional complexity and subtlety, as individuals in these contexts tend to use ambivalent or metaphorical language to describe their emotional state. This characteristic requires the development of more advanced NLP models to accurately interpret these expressions. In addition, discussions about mental health are often marked by social stigma. This influence can lead people to hide their true feelings or use language that does not accurately reflect their experience, adding an additional level of complexity to the analysis of these conversations. Even more, the existing 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 3 of 17 studies that assess stigma and polarity in social networks comments relate to specific diagnoses such as schizophrenia, bipolar disorder and Alzheimer's disease (Budenz etal.2019; Jilka etal.2022; Oscar etal.2017). Other studies have focused on reactions to antipsychotic medication (Mon et al. 2021), and other studies are indirectly related to mental health issues such as obesity (Bograd, Chen, and Kavuluru2022) or Covid19 (Xue etal.2020). In this way, other research (Budenz etal.2019) uses supervised learning to analyse the existence of stigma in social networks towards bipolar disorder, one of the most stigmatised mental illnesses. Specifically, the study allows us to characterise stigma from supportive messages about bipolar disorder and their repercussion and impact on Twitter. Regarding psychosis, the conducted research of (Jilka etal.2022) proposes to identify stigmatising tweets on Twitter related to schizophrenia by applying algorithms such as SVM (Support Vector Machine), RF (Random Forest) among others. Finally, the authors of (Oscar et al. 2017) use machine learning to analyse the existence of stigma on Twitter regarding Alzheimer. All these studies have been carried out on the Twitter social network and only one of them was developed in the Spanish context (Mon etal.2021). In that Spanish research, the authors aimed to investigate Twitter conversations concerning antipsychotic drugs in order to gain insights into public reactions and identify the most frequently discussed areas of clinical interest related to their usage. Consequently, this context depicts a scenario where the mental health and social networks are revealed as an exceptional environment, in terms of their relationship and impact, and allows us to understand public opinion of mental health, the emotional responses it elicits and the presence of stigma. Thus, the main contributions of our research are: (a) Design a novel corpus labelled with polarity (positive, negative, neutral) and stigma from Instagram post comments on celebrity mental health disclosures. This dataset can be accessed on GitHub (Merayo, Ayuso, and GonzálezSanguino2023), allowing researchers to use it. (b) Modelling machine learning algorithms to predict polarity and stigma in social networks in case of disclosure of mental health problems, specifically anxiety or depression. Thus, this research is innovative on several levels. It proposes machine learning analysis of mental health on Instagram for which there is barely any precedent, since Studies on AI have mainly concentrated on Twitter. Secondly, this is a novel approach, as previous research has focused primarily on identifying psychopathology rather than examining the community response to it (Ahmed etal.2022; Birnbaum etal.2017; Fodeh etal.2019; Guntuku etal.2017; Joshi and Kanoongo2022; Lejeune etal.2022). Thirdly, this research holds significant social relevance as it examines the emerging trend of public reactions to celebrities discussing their mental health issues, influencing millions of individuals, rather than focusing on hashtags or general comments on a topic. This can facilitate more effective campaigns and actions, leading to increased awareness and favourable effect on society as a whole, along with particular groups, including mental health professionals, and organisations, promoting responsible use of social media and making informed decisions to address mental health on these platforms. This document is structured as follows. Section2 describes the methodology for creating the dataset from Instagram social media posts. Section3 details the classification models implemented. Section4 reveals the results of these models. Finally, Section5 explains the main conclusions. 2 | Design of the Mental Health Corpus in Social Networks 2.1 | Selection of Posts on Instagram A search was conducted for primary posts (made directly by the author) containing disclosures or conversations about their mental health problems by Spanish influencers on Instagram (with over 100,000 followers). To carry out this process, we have searched for publications from different profiles of the people with the most followers in Spain, as well as reviewing press and television news that usually announce this type of publication. This search covered publications from September 2020 to December 2022. After an initial review, most of the posts were made by women, and as we were unable to have a gender balance in the posts, we decided to include the male gender in the Instagram posts as an exclusion criterion to avoid possible bias in the analysis. We found around 20 posts by highimpact female influencers on Instagram and selected the 10 with more responses or comments. All posts had a similar format, with one or more photos accompanied by text talking about mental health problems, such as depression or anxiety. A couple of posts announced that they were withdrawing from the social network due to their mental health problems, and in another couple of posts, the image showed the person crying. Unusually, one of the posts consisted of a promotional video in which an influencer talked about her mental health issues to promote a product. Once the Instagram posts were selected, all comments in response to them were collected using IGCommentExport (“One Click Comment Extractor for IG”) (Chrome Web Store2023), a tool to export Instagram comments to CSV (Comma Separated Values) format. A total of N = 21,151 comments were collected. Regarding ethical considerations and data privacy, the study has been approved by the ethics and deontology committee of the University of Valladolid (PI 233365) and we anonymised all Instagram comments (discarding @mentions, usernames and URLs). The final selection of Instagram posts, together with the name of the influencers, the number of followers and the responses associated with each post are in the Supporting Information. 2.2 | Description of the Labels in the Corpus: Polarity and Stigma Following a manual observation of the dataset, and in line with previous literature on manual labelling of comments (Budenz et al. 2019; Mon et al. 2021; Bograd, Chen, and Kavuluru 2022; Tomar, Mathur, and Suman 2022; Delanys etal.2020) we set up guidelines for the different labelling categories: polarity and stigma. Regarding the polarity associated with a comment, it consists of giving a positive, negative or neutral/undefined value to the comments in response to the disclosure or description of the symptomatology in the post. Positive polarity reflects understanding, encouragement or 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 4 of 17 Expert Systems, 2025 even admiration of the post. For example, “Cheer up, we love you”. Negative polarity is assigned when the person expresses negative opinions, usually by questioning the post with ironic, sarcastic or even derisive and derogatory comments. For example, “how you show that you don't know what depression or anxiety is, shame on you!”. Neutral or undefined polarity is assigned in cases where no clear opinion is detected or can be interpreted in both directions. For example, “take medication, it will help you” “and your partner?” About stigmatisation, stigmatising responses to comments are behaviours in which negative beliefs and emotions towards mental health problems are expressed. Stigma manifests in a variety of forms including rejection and anger against the person, which may extend to contempt or mockery, belittling their problem. For example, “What a desire to draw attention to yourself”; “you're so inconsistent and seeking the limelight”. Because socially we know that “stigma is wrong” many rejection comments are made in an ironic or sarcastic way. For example, and how do you write on insta?. Additionally, anger is shown by arguing that such posts “trivialise or commercialise” mental health. For example, “don't come and tell me your false stories of overcoming, without even knowing what it is to work…”. Other times the stigma manifests itself as pity or sorrow for the person. For example, “It breaks my heart”. 2.3 | Process of Categorising the Corpus The labelling process was divided into three phases: an initial phase with a pilot corpus (N = 787 comments), a second phase focused on the development of the corpus with all the comments of the selected posts (N = 21,151) and a third phase with the final corpus (N = 2287). The same methodology was followed in the first two phases: once the comments were collected, the corpus was cleaned, and then two independent experts were responsible for labelling each category. A third expert then reviewed the comments to resolve discrepancies. In the third and final phase, a final corpus is built from the large corpus to apply machine learning algorithms (N = 2287). In this way, the labelling process in the pilot corpus represented a first stage carried out on a random subset of comments (N = 787, including emoticons). The process was divided into the following stages: (a). Initial data cleaning: comments in other languages, with acronyms only (e.g., “TQ”, “I love you” in English) and those lacking coherence, e.g., “cuideseBR” or “gusta ver tu” (“like see you”) in English, were deleted to maintain the sample's relevance. Additionally, comments labelling other people who have replied to the same post have been removed, except when the author of the post is labelled and relevant information is give (e.g., “@dulceida I hope all is well”). This results in a final sample of N = 573 comments. (b). Handling Emoticons: emoticons have been excluded to focus solely on the linguistic effects. (c). Expert labelling: two specialists independently labelled the clean sample without knowing each other's assessments, while a third expert examined the inconsistencies that emerged in the clean sample. In this pilot corpus, we identified discrepancies in only 2.43% (N = 14) of the labelled comments, which showed a very satisfactory interrate reliability among experts who categorised (Hallgren2012). These discrepancies occurred predominantly in the neutral polarity categories, as some comments contained messages with both positive and negative polarities (“What a pity! I'm so sorry about what happened to you, lots of encouragement”). The third reviewer found that the message of sympathy prevailed in these cases, so it was categorised as positive polarity. Satisfactory results from this initial process (pilot corpus) provide a solid basis for replicating the results in a larger sample, ensuring consistency of labelling and maintaining data quality for future analysis. Regarding the second phase, which corresponds to the whole corpus, the corpus was labelled with all comments (N = 21,151). The same procedure as in the pilot study was followed: a. Initial data cleaning: this process reduced the whole corpus to a final sample of N = 15,213 comments. To guarantee the suitability of the sample, comments that consisted solely of acronyms (e.g., LOL), that were not written in English or that were incomprehensible were discarded in this process. In addition, comments that only served to name other users without further content were also removed. b. Handling Emoticons: emoticons have been removed. c. Expert labelling: the final sample was labelled by two experts separately, while a third expert evaluated any inconsistencies. If the two initial experts could not agree on the assigned label, a third reviewer was required. If this third reviewer also could not resolve the tie, the comment was excluded from the corpus to avoid possible errors. To carry out the labelling process, the experts were psychologists or trained persons who used a labelling guide, elaborated with examples and descriptions of the different categories of our corpus (polarity, stigma). This labelling guide is also freely accessible and available in a Github repository (Merayo, Ayuso, and GonzálezSanguino2023). The percentage of discrepancy in this second phase was around 2.3% (489 comments). The greatest disparity was observed between neutral and positive polarity in those comments where advice was given (“Ayyy, take as much time as you need, health comes first”). In these cases, the third reviewer categorise these comments with neutral polarity. Finally, the third phase was associated with the final corpus. Once the categorisation of the full corpus was completed, a thorough selection process was undertaken to develop a representative corpus that would be appropriate for our IA algorithms. A key consideration in data corpora is class balance; when addressing a classification problem, an imbalance where one class has significantly more data than another can lead classification models to favour predictions for the majority class. This imbalance adversely affects the algorithm's performance and predictive capability. As a result, the original set of 15,213 comments from the extensive corpus was systematically condensed into 2287 comments in the following steps: 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 5 of 17 a. Removal of redundant comments: messages with identical content were removed, keeping only a single example for each recurring topic. b. Distribution: we evaluated the distribution of comments among different categories to ensure a more equitable representation of polarity and stigma in the dataset, ensuring that models would make accurate predictions in all classes. Although there are more comments of positive polarity than negative or neutral, and more nonstigmatising than stigmatising comments, the distribution is significantly more balanced compared to the whole corpus. c. Comments randomly chosen: after determining the target distribution percentages, comments from all posts were randomly picked to be part of the final dataset. 2.4 | Corpus Statistics Table1 displays the frequencybased descriptive statistics for each category of the corpus, and Figure1 presents the descriptive statistics by percentage. The most predominant polarity is positive (66.7%), as well as nonstigmatising responses (80.1%). Besides, it is observed that there are more negative polarity comments than stigmatising ones. This is because many comments that include disclosures of mental health problems are not stigmatising but their emotional polarity is negative (e.g., “I cried with you when I saw your tears… even though I don't know you in person, tell you that my hand will always be with you”). Likewise, there are also neutral messages that include advice or reactions without a specific emotional tone (e.g., “Good bless you”; “real life”) that also do not meet the requirements to be categorised as stigmatising. 3 | Classification Models 3.1 | Support Vector Machine The Support Vector Machines (SVM) algorithm is used in both classification and regression problems. Its goal is to find an optimal separation hyperplane that maximises the distance between data classes, in our case polarity and stigma categories, in a feature space (Chollet2021; Noble2006). The most important configurable hyperparameters in SVM include the Kernel, which determines the type of transformation used to separate data in a highdimensional space, and the regularisation parameter C, which controls the tradeoff between achieving a wider margin and minimising errors in classification. Proper tuning of these hyperparameters is essential to achieve optimal performance in SVM models and ensure accurate classification of the data. 3.2 | Random Forest The Random Forest (RF) algorithm relies on building multiple decision trees during training and combining their results to make more accurate and robust decisions (Probst, Wright, and Boulesteix2019). Each tree in the forest is trained on a random sample of the data and uses a random selection of features to make decisions. Typical configurable hyperparameters in Random Forest include the number of trees and the maximum number of features to consider in each split (max_features). Proper tuning of these hyperparameters can significantly influence the performance and generalisability of the model. TABLE 1 | Descriptive statistics by frequencies. Variable Frequency Polarity Negative (P) 588 Positive (N) 1526 Neutral (NEU) 173 Total 2287 Estigma No 1833 Yes 454 Total 2287 FIGURE 1 | Descriptive statistics by percentage of the dataset for polarity and stigma categories (a) Polarity (b) Stigma. 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 6 of 17 Expert Systems, 2025 3.3 | Hybrid CNNLSTM Model The proposed hybrid deep learning architecture consists of the following layers (Figure2): 1 Embedding layer ⟶ 1 Dimension convolutional layer (Conv1D) ⟶ 1 MaxPooling Layer ⟶ 1 LSTM (Long ShortTerm Memory) layer ⟶ 1 Dense Layer. Our model benefits from the potential offered by the combination of recurrent (Recurrent Neural Network, RRN) and convolutional (CNN, Convolutional Neural Network) layers. As the convolutional layer assumes certain tasks, it reduces the processing load on the LSTM (recurrent layer), improving the effectiveness of that layer. This model will classify the mental health comments into three polarities (P, N, NEU) and in two stigma categories (Yes, No). In the following, each layer of the model will be described in detail: 1. Embedding layer: This layer transforms texts, Instagram comments in our case, into numerical vectors so that they can be interpreted by neural networks. This process is carried out using word embedding, in which each word is depicted as a vector. The goal is to assign similar values to words that share a certain semantic relationship. Thus, two techniques can be applied: learning word embeddings in conjunction with the problem to be solved (using the problem corpus) or loading embedding vectors from precomputed databases of word embeddings. The first option was chosen because precalculated dictionaries are designed in a general way and their effectiveness depends to a large extent on their similarity to the words in our specific corpus. In contrast, learning directly from our corpus will be optimal as it provides a relevant source of information. The embedding layer functions like a dictionary that maps integer indices, which correspond to specific words, to dense vectors. It accepts integers as input, searches them in its internal dictionary, and returns the associated vectors, much like a dictionary lookup. Initially, when the embedding layer is set up, its internal word vectors, or weights, are randomised. Throughout the training process, these vectors are adjusted using backpropagation, resulting in a distinct structure that is tailored to the specific task by the end of training. Therefore, the embedding layer takes a twodimensional tensor as input and returns a threedimensional tensor that can be further processed by a convolutional layer. 2. 1D (Dimension) Convolution Layer (Conv1D): This layer employs filters on the data to identify local features within the input. Its function is to identify important patterns using convolution operations, reducing the workload for the subsequent RNN layer. In essence, it streamlines the processing for the RNN by removing intermediate stages through text pattern detection. In addition, the ReLU (Rectified Linear Unit) function shall be used as an activation function, which shall be applied after convolution. Additional key parameters include the number of filters (with each filter designed to capture a specific pattern in the input data) and the filter size (kernel size), which determines the dimensions of the filters based on the length of the window. 3. MaxPooling layer: This layer is used after the convolution layer to reduce the dimensionality of the features extracted by the previous layers of the network, while preserving the most relevant information. The MaxPooling layer transforms a data matrix into a smaller matrix, retaining the key elements of the original. 4. LSTM layer: This RNN layer is employed to identify longrange patterns in the input data. In particular, LSTM layers improve the functionality of traditional recurrent networks by combining both longterm and shortterm memory capabilities. The parameters to be adjusted are the number of neurons, a dropout rate and a recurrent dropout rate. Both parameters are adjusted to prevent overfitting and enhance the model's generalisation during training. These parameters indicate the proportion of neuron units that are randomly “deactivated” during each training step, to prevent them from becoming too dependent on neighbouring neurons. 5. Dense Layer: This layer takes the features learned by the previous layers and produces the final output of the network, crucial for producing the final model predictions. Since we use the model for classification, the number of neurons in the output layer corresponds to the categories into which we want to classify the data, that is in multiclass classification problems there will be one neuron per class. Therefore, in our case we will put three neurons to categorise polarity (P, N, NEU) or two neurons to categorise stigma (Yes, No). Finally, the softmax activation function will be utilised to transform FIGURE 2 | Configuration of layers in the proposed hybrid CNNLSTM network model. 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 7 of 17 the outputs into a probability distribution, as it is frequently used in multiclass classification tasks. Finally, in the context of classification problems, there are other crucial hyperparameters influencing model convergence, including optimizers and loss functions. In our particular scenario, we chose the categorical crossentropy loss function, which is a commonly used option for classification tasks, especially when dealing with three or more labels as in our scenario. Additionally, the ADAM optimizer was selected due to its versatility, strong performance, and widespread utilisation in similar problem domains. As mentioned above, the word embedding process can be carried out in two ways: learning word embeddings jointly with the problem you want to solve (using the corpus of the problem) or by loading embedding vectors from precomputed databases. The first option was chosen because of the advantages that this technique provides. To accomplish this, a vocabulary is generated from the corpus by employing the Information Gain (IG) technique, which emphasises terms that occur most frequently. IG is favoured over absolute frequency since it evaluates how often a word appears in a particular class compared to its occurrence in other classes, while absolute frequency only quantifies total occurrences without considering class differences. To determine the IG of a word, we first compute its entropy (Larose and Larose2014; Witten, Frank, and Hall2011), shown in 1: where pi is the probability that the word wi appears in the class ci . Then, we calculate the IG following 2: In this equation, C represents the set of classes and X is the subset of texts in which the word wi is found. To calculate H(C) , the probabilities of each category in the corpus are determined, while to calculate H(C,X) the likelihood of a word occurring or not occurring in the corpus needs to be calculated, along with its probabilities of occurrence and nonoccurrence in each category. 3.4 | BERT Model BERT is a transformerbased language model that understands the context of words in both directions, enhancing tasks like natural language processing and comprehension. Employing BERT for sentiment analysis requires initially training the model on a substantial dataset before finetuning it on a dedicated dataset. This process enables the model to gain a broad understanding of language, which it can then refine to capture the nuances of sentiment analysis within a specific field, such as mental health in social media contexts. In our case, we have used a pretrained linguistic model for social networking text in Spanish, called RoBERTuito, trained following the RoBERTa guidelines on 500 million tweets (Pérez etal.2021). This model surpasses other pretrained language models for Spanish. Moreover, the Transformers library offers the Trainer class to finetune any of the pretrained models on a particular dataset. This approach allows us to identify the best training parameters, such as the learning rate (which accelerates model convergence), batch size (the number of samples before updating weights), and epochs (which indicate the number of iterations over the training dataset). 4 | Experiments and Results This section describes the main results of the experiments. To identify the ideal configuration, we will search for the most suitable hyperparameter values for each model in order to optimise various performance metrics. Although accuracy reflects the overall proportion of instances that the model has correctly classified, it is essential to consider additional metrics that assess the model's performance for each individual class. Metrics like precision, recall, and F1score are particularly valuable when dealing with imbalanced datasets. A key aspect of model evaluation is having two separate datasets: the training set, used for training the model, and the validation set, used for evaluating its performance. In our case this has been split in a typical ratio of 70%–30%, respectively (Oneiros2023). On the other hand, we used the crossvalidation technique to identify the optimal parameter sets for the model being trained. Specifically, We utilised kfold crossvalidation, a method that entails performing k iterations in which the model is trained and assessed k times. Additionally, we implemented the EarlyStopping technique to maintain the model's generalisation ability. This method halts training once the validation loss reaches its minimum, preventing further overfitting. We implemented our classification models in Python (version 3.1) using Keras (Oneiros2023) and TensorFlow (Google Brain Team2023), and executed all models on the Google Colab platform. 4.1 | Data PreProcessing and Encoding Text preprocessing consists of cleaning and preparing textual data to obtain a semantically richer representation that facilitates its computational representation. Thus, the following series of preprocessing techniques have been implemented: • To convert to lower case: to reduce duplicate words. • To eliminate mentions (), hashtags (#), and URLs: as they do not contribute any valuable information. • To delete punctuation marks. • To minimise the repetition of characters: for example, change “Siiiiii” to “Si” (“yes” in English). • To standardise slang/jargon: for example, change “tb” to “también” (“furthermore” in English). The following step involves narrowing down the vocabulary used by the classification models (feature reduction) by applying the techniques outlined below: • Remove stopwords: words that carry no significant meaning on their own, such as articles, adverbs, prepositions, (1) H = N−1 ∑ i=1 pilogi(pi) . (2) IG(C,X)=H(C)−H(C,X). 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 8 of 17 Expert Systems, 2025 conjunctions, and certain verbs. They are usually very frequent words in natural language and depend on the language. • Apply stemming: a text normalisation technique that reduces words to their root. This technique removes affixes from words, which can lead to invalid words. For instance, the Spanish words “pensando” (meaning “thinking” in English) and “pensamiento” (meaning “thought” in English) will be shortened to “pensa.” We will employ the SnowballStemmer from NLTK (Python Software Foundation2021) for this purpose in Spanish. In addition, the BERT model (RoBERTuito) (Pérez et al. 2021) includes a specific data preprocessing consisting of: character repetitions are capped at three, usernames are replaced with a designated token, hashtags are substituted with another token, and emojis are converted into their textual descriptions using a specialised library. However, RoBERTuito was evaluated using both data preprocessing methods, and the performance results were quite similar. The subsequent step involves tokenization, a fundamental step in text processing, that consists of breaking text into discrete units called “tokens”, in our case words. Tokens provide discrete units that computers can work with to understand and parse text more effectively. In this case we use the TweetTokenizer tokenizer (NLTK Project2023) for the SVM, RF and hybrid CNNLSTM models. On the contrary, BERT algorithms use their own tokenizer. The goal is to find the most meaningful but smallest representation. The next step is to transform the texts into number (feature extraction process), since machine learning models and their inputs have to be numbers. In our case we have to transform two parts: on the one hand the tokenised and normalised messages (Instagram posts), and on the other hand the labels that correspond to the categories (polarity and stigma in this case). To convert the messages into numerical format, a dictionary has been established where each word is assigned a specific index vector (as described in the previous section on word embedding). For the labels, One Hot Encoding was used, which encodes various classes as a matrix. In this matrix, a “1” is placed in the column corresponding to the class of the text (Instagram message), while “0” is used for all other classes. Therefore, for the polarities (P, N, and NEU), we will create a matrix with three columns. A similar approach is used for categorising stigma labels (Yes and No), resulting in a twocolumn matrix. For the SVM and RF models, instead of using individual columns for each variable with binary values (0 or 1), we use a single global variable to represent one output, as these models generate only a single result. Here, polarity is indicated as P, N, or NEU, unlike the earlier binary encoding of “0” or “1”. 4.2 | Polarity Results in the Mental Health Corpus 4.2.1 | Results SVM Model The optimal hyperparameters in SVM will be searched in the next order: kernel type and regularisation parameter (C). The optimisation process begins with a kernel scan, which reveals that the RBF kernel achieves the highest accuracy at 63%. In contrast, the Poly and Sigmoid kernel types yield lower results of 62%. We then proceeded to optimise the C regularisation parameter, where the best value is 1.4, achieving an accuracy of 65%. However, as we increase the value of C, starting at 10, the model exhibits the worst accuracies, reaching values of 59% and 57% for 100 and 200, respectively. In contrast, if we continue decreasing the value of C below 1, the model's accuracies remain more stable but lower, reaching levels of 63% from C = 0.01 to C = 0.001. In summary, the values of the hyperparameters that optimise the SVM model are RBF for the kernel type and 1.4 for C (regularisation parameter), reaching a final accuracy of 65%. In addition to accuracy, it is important to evaluate the model in each class separately through other metrics such as precision, recall and F1 score (Table2). The results show that the P class is the best predicted, reaching a precision of 68%. In contrast, the NEU and N classes show lower performance in all metrics. 4.2.2 | Results of the RF Model In the case of RF, the hyperparameter search will focus on determining the ideal number of trees to employ. Additionally, RF selects random feature subsets to optimise splits, making the hyperparameter max_features crucial in deciding how many features should be considered. Therefore, the hyperparameters will be searched in the following order: max_features and the number of trees. First, the accuracy of the model was evaluated with different values of max_features. Using Log2, the accuracy achieved was 66%. Using the square root (Sqrt), the accuracy improved to 69%, so Sqrt was selected. Then, we proceed to find the best number of decision trees to address the problem. Figure3 shows that the optimal value is 600 decision trees, obtaining an accuracy of 71%. In addition to global accuracy, Table3 shows the results of precision, recall and F1score for each class separately. It is observed that the P class is the class best predicted by the RF model, as all metrics show very good results. As for the N class, a precision close to 60% is achieved (higher than the SVM model). Finally, the NEU class has a low precision of 40%, which is reasonable since it is quite difficult to detect neutral comments when we express opinions on social networks. As expected, the RF classification model improves SVM performance for all polarities. 4.2.3 | Results of the Hybrid CNNLSTM Model To train and evaluate the performance of the model, different types of tests have been performed to see the variations in the accuracy metrics involved. These tests are: adjustment of the number of filters and neurons, the dropout rates and the learning rate for the Adam optimizer. Next, tests were made with the reduction of the total number of unique words chosen TABLE 2 | Summary of the SVM model results for precision, recall and F1score metrics considering three polarity classes (P, N, NEU). Label Precision Recall F1score P68% 90% 77% N42% 19% 26% NEU 50% 2% 4% 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 9 of 17 from the corpus and finally an adjustment of the batch size parameter. All this to adjust the hyperparameters and avoid overfitting the developed model. The testing phase was carried out using several examples of performance metrics: precision, recall and F1score. As a first stage of the training phase, the training process will be repeated by varying the number of neurons and filters in the layers to find the combination that gives the best accuracy results, the values of which are shown in Table4. For these tests, the kernel size has been set to 8, the dropout parameter to 0.2 and the recurrent parameter to 0.3. The results in Table4 show that the model is quite sensitive to variations in these hyperparameters, achieving a higher accuracy rate for a number of 180 in the convolution layer and 256 neurons in the LSTM layer. The second step in training the model consists of varying the dropout rates (dropout and recurrent dropout rate) of the LSTM layer in the range of 0.2 to 0.8. According to the Keras documentation (Oneiros2023), the LSTM layers have two different types of dropout rates, which are represented as a floating point number between 0 and 1. In addition, for these tests, the kernel size was set to 8, the number of filters in the convolutional layer to 180 and the number of neurons in the LSTM layer to 256. Table5 shows that varying the values generates significant changes in the accuracy of the model, and the best performance, 85.02%, is achieved with values 0.2 and 0.3, for the dropout and recurrent dropout rate parameters respectively. The next step in the model is a sweep for different values of the learning rate. This is a fundamental factor when training machine learning models, since it will determine the degree of magnitude with which adjustments to the different parameters of the model will be made, which in turn affects the convergence of the model. There are two main reasons why it is interesting to control the learning rate: to control the speed of convergence and to overcome local minima that may occur. It can be seen in Table6 that depending on the value, there is no considerable variation in the accuracy metric of the model when changing the value of this parameter, with the best value, 79.01% achieved for a learning rate of 0.01. The next training step is to reduce the total number of unique words in the corpus used in the model to increase computational efficiency. However, there is a tradeoff between the reduction of the vocabulary used and the possibility of losing important information, so it is necessary to perform the reductions in a stepwise manner (in this case in 100word jumps). Using all words, the model achieved an accuracy of FIGURE 3 | Optimisation of the number of trees. TABLE 3 | Summary of the RF model results for precision, recall and F1score metrics considering three polarity classes (P, N, NEU). Label Precision Recall F1score P74% 92% 82% N59% 37% 46% NEU 40% 8% 14% TABLE 4 | Results of the evaluation and optimisation of the hyperparameters number of neurons and filters in the hybrid CNNLSTM model. Number filters convolutional layer Number neurons LSMT layer Crossvalidation accuracy 192 256 76.97% 192 128 65.74% 192 96 76.82% 192 64 77.11% 180 256 79.15% 180 96 65.74% 160 128 77.84% 160 64 77.55% 150 256 76.24% 128 128 65.74% 128 64 65.74% 96 64 65.74% 192 150 76.8% 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 16 of 17 Expert Systems, 2025 draft, Writing – review and editing, Visualisation, Supervision, Project administration. Clara GonzálezSanguino: Conceptualization, Investigation, Data curation, Writing – original draft, Writing – review and editing, Visualisation. Alba AyusoLanchares: Conceptualization, Formal analysis, Investigation, Data curation, Writing – original draft, Writing – review and editing, Visualisation. Acknowledgements This research has been supported by the University of Valladolid. Conflicts of Interest The authors declare no conflicts of interest. Data Availability Statement The developed corpus, the labelling decalogue and the machine learning algorithms will be available on a Github repository, https:// github. com/ GCOde velop er/ Menta lHealt hDataset, for researchers to use them in contexts related to mental health in social networks. Other data and models will be made available on request. References Ahmed, A., S. Aziz, C. Toro, etal. 2022. “Machine Learning Models to Detect Anxiety and Depression Through Social Media: A Scoping Review.” Computer Methods and Programs in Biomedicine Update 2: 1–9. https:// doi. org/10.1016/j.cmpbup.2022.100066. Arco, d F M P., M. D. MolinaGonzález, L. A. U. López, and M. T. MartínValdivia. 2021. “A MultiTask Learning Approach to Hate Speech Detection Leveraging Sentiment Analysis.” IEEE Access 9: 112478–112489. https:// doi. org/ 10. 1109/ ACCESS. 2021. 3103697. Birnbaum, M., S. K. Ernala, A. Rizvi, M. Choudhury, and J. Kane. 2017. “A Collaborative Approach to Identifying Social Media Markers of Schizophrenia by Employing Machine Learning and Clinical Appraisals.” Journal of Medical Internet Research 19: e289. https:// doi. org/ 10. 2196/ jmir. 7956. Bograd, S., B. Chen, and R. Kavuluru. 2022. “Tracking Sentiments Toward Fat Acceptance Over a Decade on Twitter.” Health Informatics Journal 28, no. 1: 14604582211065702. https:// doi. org/ 10. 1177/ 14604 58221 1065702. Budenz, A., A. Klassen, J. Purtle, E. Tov, M. Yudell, and P. Massey. 2019. “Mental Illness and Bipolar Disorder on Twitter: Implications for Stigma and Social Support.” Journal of Mental Health 29: 1–9. https:// doi. org/ 10. 1080/ 09638 237. 2019. 1677878. Chollet, F. 2021. Deep Learning With Python. Shelter Island, New York: Simon and Schuster. Chrome Web Store. 2023. “One Click Comment Extractor for IG - Chrome Web Store.” https:chrome.google.com/webstore/detail/ commentexporter/cckachhlpdnncmhlhaepfcmmhadmpbgp Accessed January 26, 2023. Corrigan, P. W., and A. C. Watson. 2002. “The Paradox of SelfStigma and Mental Illness.” Clinical Psychology: Science and Practice 9, no. 1: 35–53. https:// doi. org/ 10. 1093/ clipsy.9. 1. 35. Datareportal. 2023. Accessed May 22, 2024. “Digital 2023: Global Overview Report.” https:// datar eport al. com/ repor ts/ digit al2023globa loverv iewreport. Delanys, S., F. Benamara, V. Moriceau, F. Olivier, and J. Mothe. 2020. “Psychiatry on Twitter: A Content Analysis of the Use of Psychiatric Terms in French.” JMIR Formative Research 6, no. 2: 1–13. https:// doi. org/ 10. 2196/ 18539 . European Commission. 2023. “Eurobarometer Survey 3032.” Accessed July 20, 2024. Fodeh, S., T. Li, K. Menczynski, etal. 2019. “Using Machine Learning Algorithms to Detect Suicide Risk Factors on Twitter.” In International Conference on Data Mining Workshops (ICDMW), 941–948. Fundación Mutua. 2023. “Informe de Salud Mental en España 2023.” Accessed July 12, 2024. González Sanguino, C., J. Medina, J. Redondo, E. Betegón, L. ValdiviesoLeón, and M. Irurtia. 2024. “An Exploratory CrossSectional Study on Mental Health Literacy of Spanish Adolescents.” BMC Public Health 24, no. 1: 1–10. https:// doi. org/ 10. 1186/ s1288 902418933 - 9. González Sanguino, C., A. B. SantosOlmo, S. Zamorano, I. SánchezIglesias, and L. M. Muñoz. 2023. “The Stigma of Mental Health Problems: A CrossSectional Study in a Representative Sample of Spain.” International Journal of Social Psychiatry 69, no. 8: 1928–1937. https:// doi. org/ 10. 1177/ 00207 64023 1180124. Google Brain Team. 2023. “Tensor Flow Library.” https:// www. tenso rflow. org Accessed May 10, 2024. Gronholm, P. C., and G. Thornicroft. 2022. “Impact of Celebrity Disclosure on Mental HealthRelated Stigma.” Epidemiology and Psychiatric Sciences 31, no. e62: 1–5. https:// doi. org/ 10. 1017/ S2045 79602 2000488. Guntuku, S. C., D. B. Yaden, M. L. Kern, L. H. Ungar, and J. C. Eichstaedt. 2017. “Detecting Depression and Mental Illness on Social Media: An Integrative Review.” Current Opinion in Behavioral Sciences 18: 43–49. https:// doi. org/ 10. 1016/j. cobeha. 2017. 07. 005. Hallgren, K. 2012. “Computing InterRater Reliability for Observational Data: An Overview and Tutorial.” Tutorial in Quantitative Methods for Psychology 8: 23–34. https:// doi. org/ 10. 20982/ tqmp. 08.1. p023. Statista. 2023. Accessed May 13, 2024. “Instagram: distribución mundial de usuarios por edad en 2023.” https:// es. stati sta. com/ estad istic a s/ 87525 8/ d is tr ibuc i onp ore da dde - lo su sua r ios - mund i a les - de - insta gram. Jain, P., U. Pandey, and E. Roy. 2017. “Perceived Efficacy and Intentions Regarding Seeking Mental Healthcare: Impact of Deepika Padukone, A Bollywood Celebrity's Public Announcement of Struggle With Depression.” Journal of Health Communication 22, no. 8: 713–720. https:// doi. org/ 10. 1080/ 10810 730. 2017. 1343878. Jilka, S., C. Odoi, J. Bilsen, et al. 2022. “Identifying Schizophrenia Stigma on Twitter: A Proof of Principle Model Using Service User Supervised Machine Learning.” NPJ Schizophrenia 8: 1. https:// doi. org/ 10. 1038/ s4153 702100197 - 6. Joshi, M. L., and N. Kanoongo. 2022. “Depression Detection Using Emotional Artificial Intelligence and Machine Learning: A Closer Review.” Materials Today Proceedings 58: 217–226. https:// doi. org/ 10. 1016/j. matpr. 2022. 01. 467. Larose, D. T., and C. D. Larose. 2014. Decision Treesch. Vol. 8, 165–186. Oxford, UK: John Wiley & Sons, Ltd. Lee, S. Y. 2019. “Media Coverage of Celebrity Suicide Caused by Depression and Increase in the Number of People Who Seek Depression Treatment.” Psychiatry Research 271: 598–603. https:// doi. org/ 10. 1016/j. psych res. 2018. 12. 055. Lee, Y., C. Yuan, and D. Wohn. 2021. “How Video Streamers' Mental Health Disclosures Affect Viewers' Risk Perceptions.” Health Communication 36, no. 14: 1931–1941. https:// doi. org/ 10. 1080/ 10410 236. 2020. 1808405. Lejeune, A., B. M. Robaglia, M. Walter, S. Berrouiguet, and C. Lemey. 2022. “Use of Social Media Data to Diagnose and Monitor Psychotic Disorders: Systematic Review and Perspectives (Preprint).” Journal of Medical Internet Research 24, no. 9: 1–15. https:// doi. org/ 10. 2196/ 36986 . Mäntylä, M. V., D. Graziotin, and M. Kuutila. 2018. “The Evolution of Sentiment Analysis—A Review of Research Topics, Venues, and Top Cited Papers.” Computer Science Review 27: 16–32. https:// doi. org/ 10. 1016/j. cosrev. 2017. 10. 002. 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 17 of 17 MartínezCámara, E., M. T. MartínValdivia, L. A. UreñaLópez, and A. R. MontejoRáez. 2014. “Sentiment Analysis in Twitter.” Natural Language Engineering 20, no. 1: 1–28. https:// doi. org/ 10. 1017/ S1351 32491 2000332. Mental Health America. 2023. “Mental Health America: Youth Data 2023.” Accessed July 15, 2024. Merayo, N., A. Ayuso, and C. GonzálezSanguino. 2023. “Repository of Corpus of Mental Health in Spanish Social Networks.” https:// github. com/ GCOde velop er/ Menta lHealt hDataset Accessed June 19, 2024. Mon, A. d M., C. DonatVargas, J. SantomaVilaclara, et al. 2021. “Assessment of Antipsychotic Medications on Social Media: Machine Learning Study.” Frontiers in Psychiatry 12, no. 1: 1–13. https:// doi. org/ 10. 3389/ fpsyt. 2021. 737684. NHS Digital. 2023. “Mental Health of Children and Young People in England, 2023 - Wave 4 Follow Up.” Accessed July 15, 2024. NLTK Project. 2023. “Natural Language Toolkit (NLTK) Library.” https:// www. nltk. org/ Accessed April 1, 2024. Noble, W. S. 2006. “What Is a Support Vector Machine?” Nature Biotechnology 24, no. 12: 1565–1567. https:// doi. org/ 10. 1038/ NBT12 061565. Oden, A., and L. Porter. 2023. “The Kids Are Online: Teen Social Media Use, Civic Engagement, and Affective Polarization.” Social Media + Society 9, no. 3: 20563051231186364. https:// doi. org/ 10. 1177/ 20563 05123 1186364. Oneiros. 2023. “Keras Library.” https:// keras. io/ Accessed June 26, 2024. Oscar, N., P. A. Fox, R. Croucher, R. Wernick, J. Keune, and K. Hooker. 2017. “Machine Learning, Sentiment Analysis, and Tweets: An Examination of Alzheimer's Disease Stigma on Twitter.” Journals of Gerontology: Series B 72, no. 5: 742–751. https:// doi. org/ 10. 1093/ geronb/ gbx014. Pavlova, A., and P. Berkers. 2022. ““Mental Health” as Defined by Twitter: Frames, Emotions, Stigma.” Health Communication 37, no. 5: 637–647. https:// doi. org/ 10. 1080/ 10410 236. 2020. 1862396. Pérez, J. M., D. A. Furman, L. A. Alemany, and F. M. Luque. 2021. RoBERTuito: A PreTrained Language Model for Social Media Text in Spanish. Preprint available at arXiv. https:// arxiv. org/ abs/ 2111. 09453 . Probst, P., M. N. Wright, and A. L. Boulesteix. 2019. “Hyperparameters and Tuning Strategies for Random Forest.” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 9, no. 3: e1301. https:// doi. org/ 10. 1002/ widm. 1301. Python Software Foundation. 2021. “Snowball Algorithms for Stemming.” https:// pypi. org/ proje ct/ snowb allst emmer Accessed April 15, 2024. Redondo, J., I. Fraga, I. Padrón, and M. Comesaña. 2007. “The Spanish Adaptation of ANEW (Affective Norms for English Words).” Behavior Research Methods 39: 600–605. Robinson, P., D. Turk, S. R. Jilka, and M. Cella. 2018. “Measuring Attitudes Towards Mental Health Using Social Media: Investigating Stigma and Trivialisation.” Social Psychiatry and Psychiatric Epidemiology 54: 51–58. https:// doi. org/ 10. 1007/ s0012 701815715. Shah, S., H. Ghomeshi, E. Vakaj, E. Cooper, and R. Mohammad. 2023. “An EnsembleLearningBased Technique for Bimodal Sentiment Analysis.” Big Data and Cognitive Computing 7, no. 2: 1–20. https:// doi. org/ 10. 3390/ bdcc7 020085. Srivastava, R., P. K. Bharti, and P. Verma. 2022. “Comparative Analysis of Lexicon and Machine Learning Approach for Sentiment Analysis.” International Journal of Advanced Computer Science and Applications 13, no. 3: 71–77. https:// doi. org/ 10. 14569/ IJACSA. 2022. 0130312. Taboada, M., J. Brooke, M. Tofiloski, K. D. Voll, and M. Stede. 2011. “LexiconBased Methods for Sentiment Analysis.” Computational Linguistics 37: 267–307. https:// doi. org/ 10. 1162/ COLIa 00049 . Tomar, P., K. Mathur, and U. Suman. 2022. “Unimodal Approaches for Emotion Recognition: A Systematic Review.” Cognitive Systems Research 77: 94–109. https:// doi. org/ 10. 1016/j. cogsys. 2022. 10. 012. UNICEF. 2021. “Mental Health.” Accessed July 20, 2024. Witten, I. H., E. Frank, and M. A. Hall. 2011. Data Mining: Practical Machine Learning Tools and Techniques. The Morgan Kaufmann Series in Data Management Systems. Boston: Morgan Kaufmann. World Health Organization. 2024. “Mental Health.” Accessed June 10, 2024. Xue, J., J. Chen, R. Hu, etal. 2020. “Twitter Discussions and Emotions About the COVID19 Pandemic: Machine Learning Approach.” Journal of Medical Internet Research 22, no. 11: e20550. https:// doi. org/ 10. 2196/ 20550 . Supporting Information Additional supporting information can be found online in the Supporting Information section. 14680394, 2025, 2, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13832 by Universidad De Valladolid, Wiley Online Library on [26/02/2025]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License