Full text
Spam hierarchical clustering for campaigns spotting and topic-based classification Francisco J´ a˜ nez-Martino∗†, Andr´ es Carofilis∗†, Roc´ ıo Alaiz-Rodr´ ıguez∗†, Victor Gonz´ alez-Castro∗†, Eduardo Fidalgo∗†, Enrique Alegre∗† ∗Department of Electrical, Systems and Automation, Universidad de Le´ on, Le´ on, ES †Researcher at INCIBE (Spanish National Cybersecurity Institute), Le´ on, ES Email:{francisco.janez, andres.vasco, rocio.alaiz, victor.gonzalez, eduardo.fidalgo, enrique.alegre}@unileon.es Abstract—This article focuses on the creation of multiclassification systems for spam email in cybersecurity organizations to prevent cyber-attacks and spam campaigns. We introduce two new subsets: SPEMC-15K-E and SPEMC-15K-S, comprising 14479 and 14992 spam emails, in English and Spanish, respectively. These are divided into eleven classes, defined using agglomerative hierarchical clustering. We evaluated sixteen pipelines, combining text representation techniques (TF-IDF, Bag of Words, Word2Vec, and BERT) and classifiers (Support Vector Machine, Na¨ ıve Bayes, Random Forest, and Logistic Regression). TF-IDF with Logistic Regression (LR) achieved the best results for English, with an F1-score of 0.953 and 94.6% accuracy. Similarly, TF-IDF with Na¨ ıve Bayes achieved the best results for Spanish, achieving an F1-score of 0.945 and 98.5% accuracy. Finally, it was observed that the TF-IDF with LR has the shortest processing time, completing the classification in an average of 2ms and 2.2ms per email in English and Spanish, respectively. Index Terms—Spam detection, Multi-classification, Imagebased spam, Text classification Type of contribution: Research already published [1] I. INTRODUCTION Spam is a widely studied problem [2], where mainly, researchers focus on detecting spam emails. However, spam emails are not only annoying messages, but some of them may pose a cybersecurity problem. Therefore, effort should be made to identify the topic of spam emails, in order to identify cases that supose a greater threat. This document is a summary of the article “Classifying spam emails using agglomerative hierarchical clustering and a topic-based approach” [1], in which, for the first time in the literature, the email content is analyzed for multiclass detection based on cybersecurity topics. The paper presents a preprocessing method to deal with common spammer tricks, such as embedding content in images and hiding random text in the body of the email [3]. In addition, a new dataset labeled from scratch, called Spam Email Multi-classification (SPEMC-15K), is introduced, containing two subsets, one in English and another in Spanish. Finally, we propose a framework for classifying spam into cybersecurity categories, for which the combination of multiple machine learning and natural language processing techniques is evaluated. Figure 1 depicts the complete process. II. DATASETS To create the two SPEMC-15K subsets, a total of 85k spam emails provided by INCIBE were processed and collected by honeypots from RedIRIS1. Agglomerative hierarchical 1The Spanish national research and education network (https://www.rediris. es/index.php.en) clustering [4] was used to divide each subset (i.e., English and Spanish emails) into groups to identify a hierarchy of topic-based clusters. To mitigate the computational problems associated with hierarchical clustering in large datasets, we randomly selected 15,000 emails per language. The initial partitioning was validated by cybersecurity experts from INCIBE, who oversaw the definition of the classes. The email processing (Section III-A) and text preprocessing (Section III-B) were applied, and the texts were codified using a Bag-of-Words (BOW) model. Hierarchical clustering, adopting Ward’s minimum variance entailment [5], produced 16 clusters per language, subsequently reduced to 11 classes after experts’ scrutiny. The resulting classes can be seen in Figure 1. III. METHODOLOGY A. Email processing Only the “Subject” field was considered from the email header due to the richer textual content for topic classification. Emails often use HTML formatting and plain text to enhance the layout [3]. Three scenarios are considered in the processing of the email body: (i) plain text only, (ii) HTML formatting only, and (iii) both formats together. In the dual-format cases, priority is given to the analysis of HTML content. We observed emails with hidden text, a tactic overlooked in recent spam detection efforts, so we worked under the assumption that all emails were likely to contain hidden text. Following a method derived from previous research, the HTML email bodies were first converted into greyscale images. Additionally, considering spammers’ evasion tactics, where text is embedded in images instead of the email body [6], we employed Optical Character Recognition (OCR) to extract the text from attached images. This approach ensures capturing all visible text as perceived by email clients. The text was aggregated and processed during the Text Classification stage after inferring the language from the subject, body, and attached images. The text, neither in English nor Spanish, was discarded, aligning with the interest of INCIBE. Only emails with consistent language throughout were considered. B. Text classification This stage comprises three phases: text preprocessing, representation, and classification. During preprocessing, single characters, numbers, and letters, along with characters or numbers within words, were removed. The text underwent JNIC 2024 ISBN:978-84-09-62140-8 490
Figure 1. Spam email multi-classification process: (a) extraction of around 15K random spam emails per language from resources, (b) preprocessing of emails, (c) extraction of all visible text of every email, (d) text preprocessing on each email, then encoding with Bag of Words and finally, hierarchical clustering, (e) manual review of the clusters, (f) category labelling, (g) training and evaluation of 16 pipelines of text classification. lowercase conversion, removal of stop words and duplicates, and tokenization. Stemming was omitted to mitigate ambiguity within the context of spam emails [7]. For text representation, we employed BOW and TF-IDF alongside two-word embedding techniques, i.e., Word2Vec and BERT [1]. Finally, each text representation was combined with four machine learning algorithms—Support Vector Machine (SVM), Na¨ ıve Bayes (NB), Random Forest (RF), and Logistic Regression (LR) [1]—resulting in 16 pipelines for the classification task. IV. RESULTS Among the results obtained, it was observed that the best combination for the English subset is TF-IDF and LR, achieving an accuracy of 94.6%, and TF-IDF with NB obtained the highest accuracy for Spanish, being 98.5%. The combination of TF-IDF with LR achieved the lowest execution time in both languages, with an average of 2ms and 2.2ms per email. Furthermore, the analysis included the computation of a confusion matrix, revealing that the Health, Other, and Services classes exhibited the lowest performance. In contrast, Extortion Hacking, Pharmacy, Sexual Content Dating, and Work Offer obtained the highest performance. Lastly, two data augmentation methods were assessed: (i) Random overand under-sampling and (ii) SMOTE combined with NearMiss [1]. The latter yielded an accuracy of 96.7% when applied to the English dataset with TF-IDF-LR, representing the only case where an improvement was observed from the earlier results. V. CONCLUSIONS In this paper, we analyzed email content for multiclass detection based on cybersecurity topics for the first time in the literature. Moreover, the paper introduced two new subsets, SPEMC15K-E and SPEMC-15K-S, with 14479 and 14992 spam emails, respectively, and grouped into eleven classes using hierarchical clustering. Sixteen pipelines were evaluated to identify the best combination between text representations and machine learning models, concluding that TF-IDF with LR and NB achieved the best performances. It was also concluded that TF-IDF-LR obtained the shortest inference times. ACKNOWLEDGMENTS This work has been funded by the Recovery, Transformation, and Resilience Plan, financed by the European Union (Next Generation) thanks to the LUCIA project (Fight against Cybercrime by applying Artificial Intelligence) granted by INCIBE to the University of Le´ on. REFERENCES [1] F. J´ a˜ nez-Martino, R. Alaiz-Rodr´ ıguez, V. Gonz´ alez-Castro, E. Fidalgo, and E. Alegre, “Classifying spam emails using agglomerative hierarchical clustering and a topic-based approach,” Applied Soft Computing, vol. 139, p. 110226, 2023. [2] F. J´ a˜ nez-Martino, R. Ala´ ız-Rodr´ ıguez, V. Gonz´ alez-Castro, E. Fidalgo, and E. Alegre, “A review of spam email detection: analysis of spammer strategies and the dataset shift problem,” Artificial Intelligence Review, vol. 56, no. 2, pp. 1145–1173, 2023. [3] C. Lioma, M.-F. Moens, J. C. Gomez, J. Beer, A. Bergholz, G. Paass, and P. Horkan, “Anticipating hidden text salting in emails,” in Recent Advances in Intrusion Detection, 11th International Symposium, 2008, pp. 396–397. [4] A. K. Abasi, A. T. Khader, M. A. Al-Betar, S. Naim, S. N. Makhadmeh, and Z. A. A. Alyasseri, “Link-based multi-verse optimizer for text documents clustering,” Applied Soft Computing, vol. 87, p. 106002, 2020. [5] R. Biswas, V. Gonz´ alez-Castro, E. Fidalgo, and E. Alegre, “Perceptual image hashing based on frequency dominant neighborhood structure applied to tor domains recognition,” Neurocomputing, vol. 383, pp. 24 – 38, 2020. [6] F. Naiemi, V. Ghods, and H. Khalesi, “An efficient character recognition method using enhanced hog for spam image detection,” Soft Computing 23, p. 11759–11774, 2019. [7] E. Fidalgo, E. Alegre, L. Fern´ andez-Robles, and V. Gonz´ alez-Castro, “Classifying suspicious content in tor darknet through semantic attention keypoint filtering,” Digital Investigation, vol. 30, pp. 12 – 22, 2019. Spam hierarchical clustering for campaigns spotting and topic-based classification 491