scieee AI-readable full text Open interactive document viewer

TopicMap-BN: Scalable and Explainable Framework for Cross-Source Bangla News Recommendation with BanglaBERT and BERTopic

International Journal of Computer Science, Engineering and Applications (IJCSEA)

Abstract

With the rapid growth of online Bangla news portals, thousands of articles are published daily on similar topics, resulting in an information overload for readers. Existing recommendation systems mostly focus on personalized suggestions based on user history, whereas readers frequently desire related news on a given topic across multiple sources. This challenge is amplified by the scarcity of robust Bangla Natural Language Processing (NLP) tools and the heterogeneous structure of news content. In this regard, we introduce TopicMap-BN, a scalable and explainable topic-based framework for cross-source Bangla news recommendation system. This system integrates Bangla-specific preprocessing with neural topic modeling (BERTopic with transformer embeddings), near-duplicate detection (MinHash and SimHash), and diversityaware re-ranking (MMR, xQuAD, DPP). These components facilitate coherent story grouping, interpretable topic labels, and recommendations that maintain relevance, freshness, and diversity. The effectiveness of the proposed framework is demonstrated by experiments carried out on the Potrika corpus (approximately 665,000 articles) and live crawls from five popular news portals. As a result, the system achieved topic quality, demonstrated by an NPMI score of 0.62 and a human agreement value (κ) of 0.71. In terms of story deduplication, it secured a precision of 0.91 and an F1-score of 0.88, indicating reliable clustering of nearduplicate articles. Moreover, the framework demonstrates its ability to produce precise and well-ranked recommendations by having a precision at rank five of 0.72 and an NDCG at rank five of 0.75. Compared with classical baselines such as TF–IDF with cosine similarity, TopicMap-BN achieves substantial gains across coherence, ranking, and diversity. These findings confirm the feasibility of cross-source Bangla news recommendation and emphasize the significance of domain-specific NLP frameworks in low-resource settings.

Full text

International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 DOI : 10.5121/ijcsea.2025.15501 1 TOPIC MAP-BN: SCALABLE AND EXPLAINABLE FRAMEWORK FOR CROSS-SOURCE BANGLA NEWS RECOMMENDATION WITH BANGLABERT AND BERTOPIC Md Hasan Hafizur Rahman 1 and Sumaia Afrin Sunny 2 1 Department of Computer Science and Engineering, Comilla University, Cumilla - 3506, Bangladesh 2 Department of Bangla, Comilla University, Cumilla - 3506, Bangladesh ABSTRACT With the rapid growth of online Bangla news portals, thousands of articles are published daily on similar topics, resulting in an information overload for readers. Existing recommendation systems mostly focus on personalized suggestions based on user history, whereas readers frequently desire related news on a given topic across multiple sources. This challenge is amplified by the scarcity of robust Bangla Natural Language Processing (NLP) tools and the heterogeneous structure of news content. In this regard, we introduce TopicMap-BN, a scalable and explainable topic-based framework for cross-source Bangla news recommendation system. This system integrates Bangla-specific preprocessing with neural topic modeling (BERTopic with transformer embeddings), near-duplicate detection (MinHash and SimHash), and diversityaware re-ranking (MMR, xQuAD, DPP). These components facilitate coherent story grouping, interpretable topic labels, and recommendations that maintain relevance, freshness, and diversity. The effectiveness of the proposed framework is demonstrated by experiments carried out on the Potrika corpus (approximately 665,000 articles) and live crawls from five popular news portals. As a result, the system achieved topic quality, demonstrated by an NPMI score of 0.62 and a human agreement value (κ) of 0.71. In terms of story deduplication, it secured a precision of 0.91 and an F1-score of 0.88, indicating reliable clustering of nearduplicate articles. Moreover, the framework demonstrates its ability to produce precise and well-ranked recommendations by having a precision at rank five of 0.72 and an NDCG at rank five of 0.75. Compared with classical baselines such as TF–IDF with cosine similarity, TopicMap-BN achieves substantial gains across coherence, ranking, and diversity. These findings confirm the feasibility of cross-source Bangla news recommendation and emphasize the significance of domain-specific NLP frameworks in low-resource settings. KEYWORDS Bangla News Recommendation, Topic modeling; BanglaBERT; BERTopic; News deduplication; Diversityaware re-ranking; Low-resource languages; Natural language processing 1. INTRODUCTION Bangla (Bengali) is the seventh most frequently spoken language in the world, with over 230 million native speakers and more than 300 million speakers worldwide, serving as the state language of Bangladesh and the second most widely spoken language in India (Eberhard, 2015; (UNFPA), 2025). Despite its significance to many people, Banglahas continued to remain a lowresource language within the domain of Natural Language Processing (NLP). There are fewer production-quality tools, benchmarks, and deployable systems for Bangla than for high-resource International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 2 languages like English (Alam, 2021; Sun, 2024). This shortage presents unique challenges for cross-source news recommendation tasks, since the large volume of published articles, frequently appearing media reports, and extensively edited headlines intensify the problem of information overload for audiences. Moreover, the exponential proliferation of online news portals around the world has transformed the way news is produced and consumed. Every day, thousands of news articles are published on similar topics by different news organizations, leading to information overload. Readers, on the other hand, are often interested in following particular events or topics across different sources to gain a balanced perspective and validate authenticity. Although content-based personalized news recommendation systems have been extensively examined in English and other popular languages (Li, 2011), the advancement of analogous systems for Bangla is still limited. The gap is particularly concerning due to the growing digital engagement of Bangla speakers. The Bangladesh Telecommunication Regulatory Commission (BTRC 1 ) has reported that internet subscriptions in Bangladesh have surpassed 130 million in 2023, in contrast to a just 0.3% of the population in 2007. In conjunction with this increase, over 400 Bangla news portals, such as Prothom Alo 2 , Samakal 3 , and Bdnews24 4 , disseminate thousands of news articles daily, complicating the task for users to manually identify associated information. The extensive and continuously updated news ecosystem challenges the manual identification of similar information across multiple sources, emphasizing the critical need for automated, topic-aware, and diversity-sensitive recommendation systems specifically designed for Bangla. Moreover, the diverse characteristics of online Bangla news sources provide significant challenges for users intending to navigate multiple websites manually. This manual browsing has complicated the task of retrieving pertinent articles, in particular when retrieving similar content from numerous websites. In order to resolve this matter, Bangla news information is automatically processed by evaluating the degree of similarity between documents. However, two major challenges impede this goal: (1) the structural heterogeneity of news articles across sources and (2) the restricted accessibility of NLP tools for Bangla. These issues are further complicated by the lack of uniform formatting in the representation of news on the web, since these articles are often embedded within HTML 5 pages containing noisy tags. Although a large volume of Bangla news content is available online, the tools required to process this content effectively, such as a thesaurus, stop-word lists, stemmers or morphological analyzers, and Part-of-Speech (POS) taggers, have still been under development. Although some research endeavors have attempted to address aspects of Bangla language processing, there is a scarcity of comprehensive and deployable systems for managing online Bangla text. In this work, we focus on measuring document relatedness among Bangla news articles available on the web. To this end, we have developed a dedicated web crawler to automatically retrieve articles from multiple sources. The retrieved content often contains HTML tags and formatting noise, which we remove using the jsoup 6 parser. After cleaning, we perform tokenization, stopword removal, and lemmatization to reduce the data dimensionality, representing each document as a bag-of-words in a Vector Space Model (VSM). The importance of each word is quantified using Term Frequency - Inverse Document Frequency (TF–IDF), and the relatedness between 1 https://lims.btrc.gov.bd/ 2 https://www.prothomalo.com/ 3 https://samakal.com/ 4 https://bdnews24.com/ 5 https://www.w3.org/TR/2011/WD-html5-20110405/ 6 https://jsoup.org/ International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 3 documents is computed using cosine similarity, enabling the retrieval of specific related news articles across portals. Building upon these foundations, we propose TopicMap-BN, a topic-based recommendation framework specifically designed for Bangla. The system organizes incoming articles into interpretable and stable topics and story groups and subsequently generates personalized recommendations using content-based profiles combined with diversity-aware re-ranking. Our main contribution is a Bangla-centric, topic-oriented recommendation system that integrates articles from multiple sources and classifies them into comprehensible topics using neural topic modeling. We employ BERTopic (Bhattacharjee A. a., 2021), which combines transformer-based embeddings with class-based TF–IDF (c-TF–IDF) to generate coherent and interpretable topic descriptors. In the next step, we propose a robust deduplication strategy to mitigate redundancy from syndicated reports and superficially modified headlines. This approach integrates shingling(Broder, 1997), which partitions text into overlapping word sequences, with probabilistic hashing methods such as MinHash (Broder, 1997) and SimHash (Charikar, 2002), thereby enabling efficient detection and clustering of near-duplicate articles. Next, we design a diversity-aware ranking module that balances topical relevance with diversity of sources and viewpoints. This module leverages established re-ranking methods, including Maximal Marginal Relevance (MMR) (Carbonell, 1998), xQuAD (Explicit Query Aspect Diversification (Santos R. L., 2013), and Determinantal Point Processes (DPPs) (Kulesza, 2012). By explicitly modeling diversity, these methods allow recommendations to adapt to user preferences while reducing redundancy and mitigating the risk of echo chamber formation.By integrating these components, TopicMap-BN provides a systematic and scalable solution for Bangla news recommendation that combines interpretability, personalization, and content diversity, while addressing the linguistic and infrastructural challenges of a low-resource yet globally significant language. The rest of this paper is structured as follows. Section 2 reviews existing literature on document similarity, neural topic modelling, deduplication, and diversity-aware re-ranking techniques relevant to news recommendation. The proposed methodology is explained in Section 3, which includes the pipeline for article ingestion, preprocessing, embedding generation, topic discovery, and story-level clustering. Section 4 outlines the datasets used, including live crawls from leading Bangla news portals and the Potrika corpus, in addition to summarization resources. Section 5 delineates the preprocessing and normalization tasks that are specifically designed for Bangla corpora. These procedures include morphological normalization, noise elimination, tokenization, and script standardization. Topic discovery and story clustering are incorporating BERTopic with transformer embeddings and hybrid MinHash -- SimHash deduplication in Section 6. The recommendation engine is introduced in Section 7, which also addresses cold-start and fairness issues. It emphasizes user modelling, scoring, and diversity-aware re-ranking strategies. Experimental results are presented in Section 8, which assesses the effectiveness of recommendations, deduplication accuracy, and topic quality on both live and offline datasets. Finally, Section 9 concludes with the main findings and discusses future research directions for extending the framework toward fairness-aware, multilingual news recommendation 2. RELATED WORKS The rapid growth of online news has driven substantial research into recommendation systems aimed at mitigating information overload. Early systems have predominantly utilized users’ historical reading behavior to provide personalized suggestions. For example, News Dude has been described as a content-based recommender agent that leverages TF–IDF representations and the KNearest Neighbor (KNN) algorithm to recommend articles based on prior reading preferences, with cosine similarity used to measure relationships between articles in the consumer space (Billsus, 1999). Similarly, hierarchical incremental clustering has been proposed to dynamically capture International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 4 readers’ evolving interests through tree-like structures (Godoy, 2006). Subsequent approaches have integrated user profiles with content-based features, facilitating hybrid personalization methodologies (Adomavicius, 2005). Although effective, these systems have primarily targeted resource-rich languages such as English. Large-scale benchmarks such as the Microsoft News Dataset (MIND), comprising over one million users and 160,000 articles, have accelerated the development of neural recommenders that blend content encoders with sequence models to capture user dynamics (Wu F. a.-H., 2020). For instance, self-attention-based models have been introduced for news recommendations, achieving significant performance gains (Wu C. a., 2019). However, these benchmarks remain Englishcentric, highlighting the lack of equivalent infrastructure for low-resource languages such as Bangla. Bangla has continued to be a low-resource language in NLP due to the scarcity of annotated corpora and production-quality tools. (Alam, 2021) have provided a comprehensive survey of Bangla NLP tasks, identifying persistent gaps compared to resource-rich languages despite promising transformer-based results. More recently, (Rabbi, 2024) have developed a Python-based NLP toolkit supporting annotation, tokenization, POS tagging, stemming, and sentiment analysis, emphasizing the need for open-source resources to accelerate Bangla NLP research. Representation learning for Bangla has also progressed. BanglaBERT, a monolingual BERT model trained on 27.5 GB of Bangla text, has improved a range of downstream tasks (Bhattacharjee A. a., 2021). IndicBERT and its successor IndicBERTv2, trained on the IndicCorp corpus spanning multiple Indic languages, have provided multilingual embeddings applicable to Bangla (Madanbhavi, 2024; Kakwani, 2020). In addition, multilingual models such as LaBSE (Feng, 2020) and MiniLM paraphrase models (Reimers, 2020) have supported clustering and retrieval. Recent community contributions, including BongLLaMA (Zehady, 2024) and BanglaEmbed (Kabir, 2024), have offered fine-tuned transformer models and lightweight sentence embeddings, respectively, strengthening the Bangla NLP ecosystem. Topic modeling has remained central to organizing and recommending news articles. Classical methods such as Latent Dirichlet Allocation (LDA) (Blei, 2003) and Latent Semantic Indexing (LSI) have been widely applied, though they suffer from limited coherence and interpretability. Neural approaches have shown promise. BERTopic (Grootendorst, 2022) integrates transformer embeddings with UMAP (McInnes, 2018), HDBSCAN (Campello, 2013), and class-based TF– IDF (c-TF–IDF) labeling to generate coherent, interpretable topics. In Bangla, topic modeling research is still nascent. (Yadav, 2025) have proposed BERT-LDA, a hybrid combining LDA and transformer embeddings, which has achieved superior coherence on Bangla corpora. Then, researchers have introduced GHTM (Graph-based Hybrid Topic Model) (aque, 2025), leveraging graph convolutional networks and non-negative matrix factorization, and outperforming traditional models including LDA, LSI, and BERTopic. Beyond Bangla, related studies in Hindi have also demonstrated that BERTopic outperforms classical approaches on short-text corpora (Lalitha, 2023). These findings underscore the suitability of neural topic models for low-resource languages. Redundancy in news recommendation has often arisen from syndicated articles or minimally edited headlines. Classic approaches have addressed this through shingling (Broder, 1997), which partitions text into overlapping token sequences, combined with MinHash for efficient estimation of Jaccard similarity (Broder, 1997) and SimHash (Charikar, 2002)for locality-sensitive hashing of near-duplicate documents (Charikar, 2002). These methods are widely deployed in large-scale information retrieval systems and remain directly applicable to Bangla, where duplication across outlets is pervasive. Beyond deduplication, diversity is critical for preventing echo chambers and broadening coverage. Maximal Marginal Relevance (MMR) (Carbonell, 1998) balances topical relevance and novelty by penalizing redundancy (Carbonell, 1998). xQuAD (Explicit Query International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 5 Aspect Diversification) (Santos R. L., 2013)explicitly models query sub-aspects to ensure wider coverage (Santos, 2010). Determinantal Point Processes (DPPs) (Kulesza, 2012) provide a probabilistic framework for subset selection that favors diverse items (Kulesza, 2012). Together, these re-ranking strategies have informed modern recommender systems and are central to ensuring that Bangla news recommendation balances coherence, personalization, and diversity. 3. RESEARCH METHODOLOGY Our research employs a topic-based methodology for Bangla news recommendation that integrates diverse data sources, preprocessing, neural topic modelling, story-level clustering, and personalized re-ranking into a unified pipeline. Our approach aims to address the challenges of duplicated articles, paraphrased headlines, and information overload across Bangla news portals. The pipeline initiates with getting news articles from both live sources (Prothom Alo, Bangla Tribune, bdnews24.com, and Samakal) and offline corpora such as Potrika. In order to ensure syntactic consistency, articles go through preprocessing, which encompasses Unicode normalization, Indic-aware tokenization, and stopword removal. After cleaning, articles are embedded using pretrained language models like BanglaBERT and IndicBERTv2. Then, BERTopic is used to find topics. This step uses class-based TF–IDF to produce subject descriptors that can be understood. Figure 1. Proposed methodology of the Topic-based explainable and scalable recommendation system. The pipeline begins with data sources (Prothom Alo, Potrika, and other portals), followed by preprocessing (normalization, tokenization, stopword removal). Articles are then transformed through embedding and topic modelling (BanglaBERT + BERTopic). Near-duplicate reports are clustered using story grouping with MinHash/SimHash, and finally, the recommendation engine generates personalized, topic-centric feeds. The framework utilizes a deduplication component that includes MinHash and SimHash (Charikar, 2002)to reduce redundancy that results from the cross-portal publication of close to identical stories. These methods detect and cluster paraphrased or minimally edited articles into coherent story groups. Personalized recommendations originate by calculating a composite score that weighs topic relevance, freshness, and diversity. Re-ranking methodologies, including Maximal Marginal Relevance (MMR) (Carbonell, 1998), xQuAD (Santos R. L., 2013), and Determinantal Point Output Recommendation Engine Story Grouping (MinHash/SimHash) Embedding and Topic Modelling Preprocessing Sources International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 6 Processes (DPPs) (Kulesza, 2012), are utilized to reduce redundancy and ensure multi-source representations. The final output of the methodology is a personalized and diversified recommendation feed that provides Bangla readers with coherent storylines across outlets. The methodology not only supports efficient retrieval but also ensures interpretability, fairness, and scalability for real-world deployment. 4. DATA SOURCES AND CORPORA The effectiveness of a Bangla News recommendation framework is closely tied to the breadth, diversity, and quality of data sources. Our system includes (i) live ingestion pipelines from major Bangla websites, (ii) openly accessible corpora that have been carefully chosen for Bangla Natural Language Processing (NLP), and (iii) summarizing tools that make it easier to create concise article abstracts in order to make sure that all articles are identified. 4.1. Web Crawler for Bangla News Collection To enable large-scale acquisition of Bangla news articles, we developed a breadth-first search (BFS)–based web crawler designed to systematically traverse hyperlinks starting from a given seed URL. The crawler maintains two core data structures: (i) a queue (q) for managing unvisited links in FIFO order, and (ii) a visited list to track and prevent duplicate exploration. At each iteration, the crawler dequeues a URL, retrieves its HTML content, and parses hyperlinks embedded within <a href=...> tags. For each discovered hyperlink, the crawler verifies whether the URL has not already been visited and whether its content contains Bangla text. Valid links are subsequently enqueued for further exploration, added to the visited list, and mapped to their corresponding Bangla text snippets. This approach ensures systematic coverage of relevant Bangla content while avoiding cycles and redundant processing. Algorithm 1. BFS-based Bangla News Crawler crawler(base_url) Input: base_url – the initial seed URL Output: Mapping of URLs to Bangla text 1. enqueue(base_url) into q 2. insert base_url into visited 3. while (q is NOT empty) do 4. front_url ← dequeue(q) 5. html_text ← process(front_url) 6. for each <a href="new_url"> in html_text do 7. if (new_url∉ visited) AND (contains_bangla(new_url)) then 8. enqueue(new_url) into q 9. insert new_url into visited 10. map (new_url, extract_bangla_text(new_url)) 11. end if 12. end for 13. end while Moreover, this crawler provides the foundation for constructing large, diverse, and representative Bangla news corpora. By leveraging BFS traversal, it ensures fairness in URL exploration, prevents infinite loops, and captures a balanced distribution of articles across different domains. International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 7 4.2. Live Sources and Ingestion We curated a collection of articles from leading Bangla news portals by leveraging openly accessible RSS/topic feeds and systematically crawling crawlable sections of the respective websites. The selected outlets include Prothom Alo 7 (the most widely read Bangla daily, with approximately ~10 million daily readers and 15–20 million monthly visitors), Bangla Tribune 8 (a digital-first portal attracting ~1.2 million monthly visits), bdnews24.com 9 (Bangladesh’s first webnative news service with over 2.5 million daily readers), Kaler Kantho 10 (print circulation of ~270,000), and Samakal (print circulation of ~200,000). These sources are recognized for their broad readership and credibility, thereby providing diverse coverage across politics, economy, sports, entertainment, and opinion. Table 1 presents verified entries obtained from Prothom Alo, Bangla Tribune, bdnews24.com, and Samakal, along with accurate URLs, headlines, categories, and timestamps. Table 1. Sample entries from RSS feeds with structured metadata Outlet Headline Category Timestamp (YYYY-MMDD HH:MM BST) URL Prothom Alo “চার মাস পর মূল্যস্ফীতি আবার বাড়ল্, জুল্াইয়ে মূল্যস্ফীতি ৮.৫৫%” অর্থনীতি (Economics) 2025-08-07 12:17 https://www.prothomal o.com/business/econo mics/e5rkt20xn6 Bangla Tribune “মূল্যস্ফীতি আবারও বাড়য়ল্া” অর্থ-বাতিজয → তবজয়নস তনউজ 2025-08-07 14:44 https://www.banglatrib une.com/business/news /910265/ bdnews24 .com (EN) “Inflation edges up to 8.55% in July after slight dip in June” Economy 2025-08-07 15:58 https://bdnews24.com/ economy/69432eb95cc c Samakal “জুল্াইয়ে মূল্যস্ফীতি সামানয ববয়ড়য়ে” অর্থনীতি (Economics) 2025-08-07 https://samakal.com/ec onomics/article/309348 / Each retrieved article is stored with structured metadata, including the canonical URL, publication timestamp, section label, and byline when available. 4.3. Public Bangla Corpora We use the Potrika corpus, a large Bangla news dataset with about 665,000 articles published between 2014 and 2020, for offline experimentation and model training (Ahmad, 2022). The dataset compiles articles from six prominent Bangladeshi news portals -- Jugantor 11 , Jaijaidin 12 , Ittefaq 13 , Kaler Kantho 14 , Inqilab 15 , and Somoyer Alo 16 -- and encompasses eight thematic 7 https://www.prothomalo.com/feed/ 8 https://www.banglatribune.com/feed 9 https://bdnews24.com/?getXmlFeed=true&widgetId=1150&widgetName=rssfeed 10 https://www.kalerkantho.com/rss.xml 11 https://www.jugantor.com/ 12 https://www.jaijaidinbd.com/ 13 https://www.ittefaq.com.bd/ 14 https://www.kalerkantho.com/ 15 https://dailyinqilab.com/ 16 https://www.shomoyeralo.com/ International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 8 categories: Politics, Economy, International, Sports, Entertainment, Technology, Opinion, and Lifestyle. Each entry is annotated with five structured attributes: headline, full text, category label, publication date, and source portal. Potrika is a valuable benchmark for a variety of NLP tasks, such as text classification, clustering, summarization, and topic modelling, due to its structure and scope. The structure of articles in the Potrika is demonstrated in Table 2. Table 2. Representative entries from the Potrika corpus Headline Category Source Date “টি-য় ায়েতি তবশ্বকায়প বাাংল্ায়েয়ের জে” (Bangladesh’s victory in the T20 World Cup) Sports bdnews24 .com 2019-11-05 “বাাংল্ায়েয়ে মুদ্রাস্ফীতি ববয়ড় ৯ েিাাংয়ে বপ ৌঁয়েয়ে” (Inflation in Bangladesh rises to 9%) Economy Prothom Alo 2020-07-12 “সরকার নি ুন তেক্ষা নীতি ব াষিা কয়রয়ে” (Government announces new education policy) Education Bangla Tribune 2018-03-21 “ ূতিথঝয়ড় উপকূল্ীে এল্াকাে বযাপক ক্ষতি” (Cyclone causes severe damage in coastal areas) Environment Kaler Kantho 2017-05-30 “নি ুন প্রযুতি প্রেেথনীয়ি িরুিয়ের তিড়” (Youth flock to new technology exhibition) Technology Samakal 2016-09-14 In addition to Potrika, we incorporate supplementary corpora, including IndicCorp v1/v2 (Kakwani, 2020), which provide multilingual data covering Bangla and twelve other Indic languages, and the Bangla Wikipedia dump 17 , which supports domain-general training for language modeling and entity linking. These corpora collectively ensure a balance between domain-specific news data and general encyclopedic content. 4.4. Summarization Resources Automatic text summary is crucial for ensuring effective user engagement due to the rapid pace of digital news production. We use BanglaT5, a transformer-based sequence-to-sequence (seq2seq) model that has been adapted to work with Bangla summarization and headline generation (Abhik Bhattacharjee, 2023). In addition, multilingual models such as mT5 (Raffel, 2020) can be finetuned for Bangla tasks, while frameworks like CrossSum(Bhattacharjee A. a.-F.-B., 2021) facilitate cross-lingual summarization, enabling translation-aware outputs (such as English → Bangla summaries). Table 3 presents abstractive summarization performed by BanglaT5 on entries from the Potrika corpus. Each case demonstrates the system’s ability to condense article content into concise, high-utility summaries. Table 3. Representative abstractive summaries generated by BanglaT5 on Potrika corpus entries Input Excerpt BanglaT5 Output English Gloss “টি-য় ায়েতি তবশ্বকায়প আজ বাাংল্ায়েয়ের তিয়ক েল্ পাতকস্তায়নর তবপয়ক্ষ োরুি জে অজথ ন কয়রয়ে, যা সমর্থকয়ের উচ্ছ্বায়স মাতিয়ে ি ুয়ল্য়ে” “টি-য় ায়েতি তবশ্বকায়প পাতকস্তানয়ক হাতরয়ে জে বপল্ বাাংল্ায়েে” Bangladesh secured victory over Pakistan in the T20 World Cup. “বাাংল্ায়েয়ে জুল্াই মায়স মূল্যস্ফীতি ববয়ড় ৯ েিাাংয়ে বপ ৌঁয়েয়ে, যা বিািায়ের দেনতিন জীবয়ন অতিতরি চাপ সৃতি কয়রয়ে” “জুল্াইয়ে মূল্যস্ফীতি ববয়ড় ৯ েিাাংে” Inflation in July rose to 9%. “সরকার আজ নি ুন তেক্ষা নীতি ব াষিা কয়রয়ে, বযখায়ন প্রার্তমক বর্য়ক মাধ্যতমক স্তয়র আধ্ুতনক কাতরকুল্াম অন্তি ুথ ি করা হয়েয়ে” “সরকার নি ুন তেক্ষা নীতি ব াষিা কয়রয়ে” The government announced a new education policy. 17 https://meta.wikimedia.org/wiki/Wikimedia_Bangladesh International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 9 “ ূতিথঝড় বমাখা উপকূল্ীে এল্াকাে বযাপক ক্ষেক্ষতি কয়রয়ে, হাজায়রা মানুষ গৃহহীন হয়ে পয়ড়য়ে” “ ূতিথঝয়ড় উপকূয়ল্ বযাপক ক্ষতি” Cyclone caused severe damage in coastal areas. “ঢাকাে শুরু হয়েয়ে প্রযুতি বমল্া, বযখায়ন িরুি উয়েযািারা নি ুন উদ্ভাবনী সমাধ্ান প্রেেথন করয়েন” “ঢাকাে শুরু প্রযুতি বমল্া” Technology fair begins in Dhaka. These abstractive summaries reduce textual redundancy and enable readers to rapidly scan, filter, and prioritize articles. This ability is particularly advantageous in a news ecosystem with a high volume of traffic, where individuals have restricted attention spans and immediate access to information is important. 5. PREPROCESSING AND NORMALIZATION FOR BANGLA CORPORA The heterogeneous and noisy nature of Bangla news data necessitates a robust preprocessing pipeline prior to downstream modelling. Articles collected through live ingestion (see Section 4.2), curated corpora (see Section 4.3), and summarization resources (see Section 4.4) often contain script inconsistencies, redundant boilerplate, and mixed-script artifacts that can adversely impact embedding quality and model performance. To address these challenges, we design a structured pipeline that systematically normalizes, cleans, and structures Bangla text into analysis-ready form. Figure 2. Preprocessing pipeline for Bangla corpora. The workflow applies Unicode normalization, boilerplate and noise removal, sentence segmentation, tokenization, stopword filtering, morphological normalization, and language identification to produce a clean, analysis-ready corpus with preserved metadata. The preprocessing pipeline begins with Unicode normalization and script standardization, where text is converted into Unicode Normalization Form KC (NFKC) to ensure canonical equivalence among visually similar characters. Bangla-specific digits (০–৯) and punctuation (such as““ and “,”) are standardized, and the Indic NLP Library Bengali normalizer (Kakwani, 2020)is applied to handle script features such as nukta, virama, visarga, and compound glyphs. This reduces orthographic variation across sources, enabling consistent token representation. The second step addresses boilerplate and noise removal, since web-crawled content frequently includes journalist bylines, embedded tags, advertisements, and hyperlinks. Regular expression heuristics and HTML parsing tools (e.g., jsoup 18 ) are used to isolate the main article body, ensuring only linguistically relevant text is retained. Following this, sentence segmentation and tokenization are performed using Indic-aware tokenizers from the Indic NLP Library (Kakwani, 2020) and BNLP toolkit(Sarker, 2021). Unlike whitespace-based segmentation, these tokenizers handle Bangla’s compound words and orthographic markers, producing reliable lexical units for embedding, classification, and summarization. 18 https://jsoup.org/ Crawled Articles, RSS Feeds, Potrika Corpus Summarization Data Unicode Normalization and script standardization Noise Removal Sentence Segmentation and Tokenization Stopword Removal and Normalization Morphological Consistency Clean and Analysis Ready Corpus Document Structuring and Metadata Preservation Language Identification and Filtering International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 16 (a) (b) (c) Figure 3. (a) Topic Quality on Potrika (test): NPMI=0.62, UMass=−0.23, κ=0.71 using BERTopic + BanglaBERT. (b) Story deduplication (Live, 5 news portals): Precision=0.91, Recall=0.86, F1=0.88 with MinHash + SimHash. (c) Recommendation performance (Potrika test, base model): P@5=0.72, R@5=0.64, NDCG@5=0.75. Story-level deduplication on the live crawl showed strong performance, with precision = 0.91, recall = 0.86, and F1 = 0.88. This indicates that our MinHash–SimHash hybrid was able to successfully merge near-duplicate reports across portals without excessive false positives. Qualitative inspection confirmed that duplicate inflation headlines and disaster reports were consistently clustered into single story groups, reducing redundancy. For recommendation quality, the baseline personalization model achieved P@5 = 0.72, R@5 = 0.64, and NDCG@5 = 0.75 on Potrika’s test set. Incorporating diversity-aware re-ranking yielded measurable gains: Maximal Marginal Relevance (MMR, λ=0.7) reduced redundant recommendations by 32%, while xQuAD increased source coverage by 18%. This demonstrates that diversity modules significantly improve beyond-accuracy metrics without sacrificing relevance. Ablation studies further highlight design tradeoffs. Comparing embedding backbones, BanglaBERT outperformed MiniLM by ΔNDCG@10 = +0.06, indicating that domain-specific embeddings are superior for semantic clustering in Bangla. Freshness tuning showed that λ = 12h was optimal, with P@10 = 0.74, aligning with reader preference for timely updates in breaking news while still retaining depth for feature articles. Additionally, these results show that the system balances accuracy (topic coherence, recommendation precision) with fairness and diversity (redundancy reduction, multi-source coverage). The combination of Potrika’s large-scale corpus and live crawls ensured both retrospective rigor and real-time robustness. These findings suggest that topic-centric recommendation, supported by preprocessing, deduplication, and diversity re-ranking, is a viable strategy for Bangla news personalization. 9. CONCLUSION This paper presented TopicMap-BN, a topic-centric framework for Bangla news recommendation in a low-resource setting. The framework integrates cross-source story grouping (via MinHash and SimHash), diversity-aware re-ranking (MMR, xQuAD, DPP), and interpretable labeling (BERTopic with c-TF–IDF and taxonomy mapping). Articles, clusters, and user profiles are jointly modeled as embeddings and topic histograms, supporting both personalization and reproducibility. Experiments on Potrika and live crawls demonstrated high clustering accuracy, improved International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 17 recommendation diversity, and scalable performance with ingestion-to-recommendation latency under three minutes. These contributions advance Bangla news recommendation beyond engineering practice into a methodologically principled framework that balances personalization, interpretability, scalability, and fairness. Future work will extend evaluation with larger and more diverse corpora, enhance summarization with advanced generative models, and improve user modeling by combining explicit preferences with implicit behavioral cues. Further, we plan to expand fairness auditing to capture regional and political biases and explore cross-lingual transfer between Bangla and other Indic languages. These directions aim to evolve TopicMap-BN into a comprehensive platform for fair, explainable, and multilingual news recommendation in low-resource environments. REFERENCES [1] (UNFPA), U. N. (2025). State of World Population 2025: The Real Fertility Crisis-The Pursuit of Reproductive Agency in a Changing World. Stylus Publishing, LLC. [2] Abhik Bhattacharjee, T. H. (2023). BanglaNLG and BanglaT5: Benchmarks and resources for evaluating low-resource natural language generation in Bangla [Preprint]. arXiv:2205.11081. [3] Adomavicius, G. a. (2005). Toward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions. IEEE transactions on knowledge and data engineering, 17(6), 734--749. [4] Agrawal, R. a. (2009). Diversifying search results. In Proceedings of the second ACM international conference on web search and data mining (pp. 5--14). [5] Ahmad, I. a. (2022). Potrika: Raw and balanced newspaper datasets in the bangla language with eight topics and five attributes. arXiv preprint arXiv:2210.09389. [6] Alam, F. a. (2021). A review of bangla natural language processing tasks and the utility of transformer models. arXiv preprint arXiv:2107.03844. [7] aque, F. a. (2025). GHTM: A Graph based Hybrid Topic Modeling Approach in Low-Resource Bengali Language. arXiv preprint arXiv:2508.00605. [8] Bhattacharjee, A. a. (2021). BanglaBERT: Language model pretraining and benchmarks for lowresource language understanding evaluation in Bangla. arXiv preprint arXiv:2101.00204. [9] Bhattacharjee, A. a.-F.-B. (2021). CrossSum: Beyond English-centric cross-lingual summarization for 1,500+ language pairs. arXiv preprint arXiv:2112.08804. [10] Billsus, D. a. (1999). A hybrid user model for news story classification. In UM99 User Modeling: Proceedings of the Seventh International Conference (pp. 99--108). Springer. [11] Blei, D. M. (2003). Latent dirichlet allocation. Journal of machine Learning research, 3(Jan), 993-- 1022. [12] Broder, A. Z. (1997). On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171) (pp. 21--29). IEEE. [13] Campello, R. J. (2013). Density-based clustering based on hierarchical density estimates. In PacificAsia conference on knowledge discovery and data mining (pp. 160 -- 172). Springer. [14] Carbonell, J. a. (1998). The use of MMR, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval (pp. 335--336). [15] Charikar, M. S. (2002). Similarity estimation techniques from rounding algorithms. Proceedings of the thiry-fourth annual ACM symposium on Theory of computing, (pp. 380--388). [16] Eberhard, D. M. (2015). Ethnologue: Languages of the world. Sil International, Global Publishing. [17] Feng, F. a. (2020). Language-agnostic BERT sentence embedding. arXiv preprint arXiv:2007.01852. [18] Godoy, D. a. (2006). Modeling user interests by conceptual clustering. Information Systems, 31(4-5), 247--265. [19] Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794. [20] Kabir, M. R. (2024). BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques. 2024 7th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI) (pp. 1--6). IEEE. International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 18 [21] Kakwani, D. a. (2020). IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. In Findings of the association for computational linguistics: EMNLP 2020 (pp. 4948--4961). [22] Kulesza, A. a. (2012). Determinantal point processes for machine learning. Foundations and Trends{\textregistered} in Machine Learning, 5(2--3), 123--286. [23] Lalitha, T. a. (2023). Based Topic Modeling on E-learning Web Content Titles Using BERTopic Model. In International Conference on Computing and Network Communications (pp. 559--580). Springer. [24] Li, L. a. (2011). Scene: a scalable two-stage personalized news recommendation system. In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval (pp. 125--134). [25] Madanbhavi, L. a. (2024). An Efficient Multilingual Text Classification using IndicCorp dataset. In 2024 5th IEEE Global Conference for Advancement in Technology (GCAT) (pp. 1--6). IEEE. [26] McInnes, L. a. (2018). Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. [27] Rabbi, F. a. (2024). Annotated Bangla Natural Language Processing (BNLP) Using Python and Machine Learning. Maneesha R., Annotated Bangla Natural Language Processing (BNLP) Using Python and Machine Learning (November 28, 2024). [28] Raffel, C. a. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140), 1--67. [29] Reimers, N. a. (2020). Making monolingual sentence embeddings multilingual using knowledge distillation. arXiv preprint arXiv:2004.09813. [30] Santos, R. L. (2010). Exploiting query reformulations for web search result diversification. In Proceedings of the 19th international conference on World wide web (pp. 881--890). [31] Santos, R. L. (2013). Explicit web search result diversification. University of Glasgow. [32] Sarker, S. (2021). Bnlp: Natural language processing toolkit for bengali language. arXiv preprint arXiv:2102.00405. [33] SHANAWAZ, M. (2013). Morphology and syntax: A comparative study between English and Bangla. Unpublished master’s thesis). North South University, Dhaka, Bangladesh. [34] Sun, R. a. (2024). Asian and Low-Resource Language Information Processing. ACM Transactions on, 23(4). [35] Wang, W. a. (2020). Minilm: Deep self-attention distillation for task-agnostic compression of pretrained transformers. Advances in neural information processing systems, 33, 5776--5788. [36] Wu, C. a. (2019). In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLPIJCNLP) (pp. 6389--6394). [37] Wu, F. a.-H. (2020). Mind: A large-scale dataset for news recommendation. In Proceedings of the 58th annual meeting of the association for computational linguistics (pp. 3597--3606). [38] Yadav, A. K. (2025). A Hybrid Model Integrating LDA, BERT, and Clustering for Enhanced Topic Modeling. Quality & Quantity, 1--28. [39] Zehady, A. K. (2024). Bongllama: Llama for bangla language. arXiv preprint arXiv:2410.21200. AUTHORS Md. Hasan Hafizur Rahman earned his B.Sc. in Computer Science and Engineering in 2009 and his M.S. (Engg.) in 2012, both from the University of Chittagong, Bangladesh. He is currently serving as a faculty member in the Department of Computer Science and Engineering at Comilla University, Bangladesh. His research interests include the Semantic Web, Artificial Intelligence, and Machine Learning, with a focus on developing intelligent, interoperable systems that bridge data integration, knowledge representation, and automated reasoning. He has contributed to research in geospatial knowledge bases, declarative machine learning frameworks, and ontology-driven applications, aiming to advance both theoretical foundations and practical implementations in next-generation computing. International Journal of Computer Science, Engineering and Applications (IJCSEA), Vol. 15, No. 3/4/5, October 2025 19 Dr. Sumaia Afrin Sunny is an Associate Professor in the Department of Bangla at Comilla University, Bangladesh. She obtained her BA (Hons) and MA in Bengali from Jahangirnagar University, where she consistently ranked among the top students of her class. In 2023, she was awarded a PhD from the same institution for her research on psychological realism in the works of five distinguished Bangladeshi story writers. Her academic career began as a Lecturer at Notre Dame College, Mymensingh, followed by appointments at Cambrian School and College and Barishal University. She joined Comilla University in 2015, was promoted to Assistant Professor in 2017, and advanced to Associate Professor in 2024. Dr. Sunny has published several scholarly articles and, in recognition of her academic and professional achievements, was honored in 2021 as a Joyita awardee.