scieee AI-readable full text Open interactive document viewer

A Multi-Layered India-Specific Hate Speech Detection and Censorship System for Social Media: A Hybrid Rule-ML-Graph-Fuzzy Architecture with Constitutional Alignment

Bhansari, Pawan

Abstract

This preprint presents a comprehensive, India-specific hate speech detection and censorship framework that integrates multiple AI layers—rule-based logic, machine learning classifiers, graph-based contextual analysis, and fuzzy logic reasoning—aligned with the Indian Constitution and statutory provisions such as IPC Sections 295A and 153A and the IT Rules, 2021. The system supports multilingual text and English voice inputs, handling English, Hindi, and Hinglish languages. Through extensive experimentation on Indian hate speech datasets, the proposed hybrid architecture achieves high accuracy, improved explainability, and strong legal compliance compared to traditional single-layer models. Designed with scalability and transparency in mind, this research provides a realistic blueprint for social media platforms seeking to balance freedom of expression with responsible and lawful content moderation in India.

Full text

A Multi-Layered India-Specific Hate Speech Detection and Censorship System for Social Media: A Hybrid RuleMLGraph-Fuzzy Architecture with Constitutional Alignment PAWAN BHANSARI Abstract The exponential growth of social media platforms has amplified concerns surrounding the dissemination of hate speech, especially in culturally diverse nations like India. To address this issue, we propose a multi-layered censorship framework designed specifically for the Indian context. Our architecture integrates rule-based filters, machine learning (ML) classifiers, graph-based contextual analyzers, and fuzzy logic systems to detect, flag, and moderate hate speech across English, Hindi, and Hinglish textual modalities, along with English voice inputs. The system is meticulously aligned with the Indian Constitution and statutory provisions like IPC Sections 295A and 153A, as well as the Information Technology Rules, 2021. Extensive experiments demonstrate the superiority of our model compared to traditional single-layer detection mechanisms, offering a scalable, explainable, and legally compliant solution. Keywords Hate Speech Detection; India; Social Media Censorship; Hybrid Model; Fuzzy Logic; Graph Analysis; IT Rules 2021; IPC 295A; Explainability. 1. Introduction 1.1 Background and Motivation The advent of social media platforms like Twitter, Facebook, and Instagram has democratized public discourse but also facilitated the unchecked spread of hate speech, misinformation, and abusive content [1]. While global platforms have initiated content moderation strategies, these often fail to align with regional constitutional values and legal frameworks, particularly in complex democracies like India. The Indian context presents unique challenges due to its diverse linguistic, religious, and cultural landscape, coupled with a rich but sensitive constitutional history concerning freedom of speech and its reasonable restrictions [2]. Recent incidents, such as the Karnataka High Court judgment directing strict adherence to content takedown notices [3], and controversies involving public figures like comedian Kunal Kamra [4], highlight the urgent need for a censorship system that is legally robust, explainable, and contextually sensitive. 1.2 Problem Statement Existing content moderation mechanisms often either over-censor, resulting in undue suppression of lawful free speech, or under-censor, failing to curb hate speech in compliance with Indian laws [5]. Moreover, most machine learning models act as black boxes, lacking explainability and legal auditability, critical requirements in the Indian legal context. Current solutions also inadequately handle the multilingual and code-switched nature (English, Hindi, Hinglish) of Indian social media discourse [6]. 1.3 Objective The primary objective of this research is to design and evaluate a scalable, explainable, India-specific censorship model for social media platforms that integrates multiple detection modalities — including rule-based engines, machine learning classifiers, graph-based context analyzers, and fuzzy logic systems — while maintaining strict compliance with constitutional guarantees and statutory requirements. 1.4 Contributions The principal contributions of this research are: • Development of a novel multi-layered censorship system for social media, combining rulebased logic, machine learning, graph analysis, and fuzzy reasoning. • Full alignment with Indian constitutional mandates (Articles 19(1)(a) and 19(2)) and statutory provisions (IPC 295A, IPC 153A, IT Rules 2021). • Support for multilingual inputs (English, Hindi, Hinglish text; English voice) with robust preprocessing for code-switching. • Design of a Human-in-the-Loop (HiL) framework for handling ambiguous cases to reduce bias and errors. • Extensive evaluation on Indian hate speech datasets with significant improvement over traditional models in terms of precision, recall, and explainability metrics. 2. Literature Review 2.1 Global Approaches Major social media platforms like Facebook, YouTube, and Twitter have deployed deep learning models for abusive content detection [7]. These models, based on Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformer architectures, offer reasonable precision but often suffer from lack of transparency [8]. Techniques like keyword-based filtering and community reporting have been used as additional layers [9]. However, these systems typically operate from a Western legal lens, missing nuances crucial in regional contexts such as India [10]. 2.2 India-Specific Challenges India’s socio-linguistic complexity — over 22 official languages and widespread code-switching — presents unique obstacles [11]. Hinglish (Hindi-English mixed language) is heavily used on Indian social media platforms [12]. Further, Indian society's religious, caste, and regional sensitivities mean that models must be highly context-aware [13]. Misclassifying legitimate political dissent as hate speech, or failing to censor casteist slurs, can lead to both public backlash and legal consequences [14]. 2.3 Machine Learning for Hate Speech Several machine learning-based hate speech detectors have been proposed [15]. Support Vector Machines (SVM), Logistic Regression, Random Forests, and XGBoost classifiers have been used in traditional models [16]. More recently, Transformer-based models like BERT [17] and HateXplain [18] have achieved strong performance. However, these models largely act as "black boxes," providing little insight into why a decision was made, complicating compliance in legal settings where explainability is mandatory [19]. 2.4 Hybrid Architectures Research suggests that combining rule-based and machine learning approaches can improve hate speech detection's robustness [20]. Systems like the HASOC competition winners employed hybrid methods for Indo-European languages [21]. Some studies have explored graph-based models for hate propagation detection [22], while others have incorporated fuzzy logic to handle uncertainty [23]. However, no work so far has proposed a comprehensive, India-specific, multilayered censorship model integrating all these techniques together. 3. Legal Framework and Context 3.1 Freedom of Speech under the Indian Constitution Article 19(1)(a) of the Indian Constitution guarantees the right to freedom of speech and expression [24]. However, this right is subject to reasonable restrictions outlined in Article 19(2), which include interests of sovereignty, integrity, security of the state, friendly relations with foreign states, public order, decency, morality, contempt of court, defamation, and incitement to an offense [25]. Thus, while citizens enjoy broad rights to express opinions, these rights are limited by considerations of communal harmony and public peace [26]. Any censorship system must therefore balance the right to free expression against the need to restrict content that threatens public order. 3.2 Indian Penal Code (IPC) Provisions 3.2.1 Section 295A Section 295A criminalizes deliberate and malicious acts intended to outrage religious feelings of any class by insulting its religion or religious beliefs [27]. Violation may lead to imprisonment up to three years, fine, or both. 3.2.2 Section 153A Section 153A penalizes the promotion of enmity between different groups on grounds of religion, race, place of birth, residence, language, etc., and doing acts prejudicial to the maintenance of harmony [28]. The offense is cognizable and non-bailable in serious instances. These IPC sections establish clear legal grounds for moderating religious and caste-based hate speech. 3.3 Information Technology Rules, 2021 The IT Rules (2021) provide a regulatory framework for social media intermediaries and digital media platforms [29]. • Platforms must remove unlawful content within 36 hours of notice from government agencies or courts.\nThey must appoint compliance officers and grievance officers located in India.\nThe rules require due diligence to prevent hosting or publishing information which is harmful, offensive, or unlawful.\n Thus, social media companies operating in India must implement proactive monitoring mechanisms to ensure compliance. 3.4 Recent Legal Cases 3.4.1 Karnataka High Court vs X Corp (Twitter India) In 2023, the Karnataka High Court upheld the government's authority to order blocking of certain tweets and accounts under Section 69A of the Information Technology Act [30]. Twitter was fined and criticized for delayed compliance, reinforcing the judiciary's expectation for strict adherence to takedown orders. 3.4.2 Kunal Kamra Controversy Kunal Kamra, a noted comedian, faced severe backlash for a series of tweets deemed disrespectful towards the judiciary [31]. Petitions sought contempt proceedings, sparking a national debate on where satire ends and disrespect begins under Indian law. This controversy emphasized the need for content moderation systems that can differentiate between legitimate satire and legally punishable speech. 4. Methodology 4.1 System Overview The proposed censorship system follows a multi-layered architecture specifically designed to address the complexities of the Indian socio-legal context. It processes both text and voice inputs, applies rulebased filtering grounded in Indian laws, performs machine learning-based classification, maps conversational context using graphs, aggregates outputs using fuzzy logic, and escalates ambiguous cases to human reviewers. This layered approach ensures scalable, explainable, and legally defensible censorship decisions. The major layers are: • Preprocessing Layer • Rule-Based Filtering Layer • Machine Learning Detection Layer • Graph-Based Context Mapping Layer • Fuzzy Logic Aggregation Layer • Human-in-the-Loop Review Layer Each layer incrementally refines and enriches the decision-making process. 4.2 Layer-wise Design 4.2.1 Preprocessing Layer Objective: To normalize incoming data and prepare it for subsequent analysis. Operations: • Language Detection: Classify the input as English, Hindi, Hinglish, or English voice transcript using fastText models [32]. • Tokenization: Use language-specific tokenizers to split input into meaningful tokens [33]. • Text Normalization: Lowercasing, Unicode normalization (NFKC form), slang and emoji handling [34]. • Noise Removal: Eliminate URLs, hashtags, mentions, and unnecessary punctuation. • Transcription for Voice: English voice inputs are transcribed into text using Wav2Vec 2.0 models fine-tuned on Indian accents [35]. Output: A clean, standardized text stream tagged with language metadata. 4.2.2 Rule-Based Filtering Layer Objective: To apply deterministic censorship rules based on statutory provisions. Key Components: • Keyword Dictionaries: Religious slurs, casteist insults, regional derogatory terms [36]. • Phrase Templates: Precompiled templates like "X community should be wiped out" or "Ban Y religion." • Law Mapping: Each keyword/phrase linked to relevant Indian laws like IPC 295A or 153A [37]. Processing: • Match tokens/phrases with dictionaries. • Assign law violation flags accordingly. Output: Flags indicating possible legal violations with confidence scores. 4.2.3 Machine Learning Detection Layer Objective: To predict hatefulness using probabilistic models. Models Used: • LightGBM Classifier: Trained on Indian annotated datasets [38]. • Fine-tuned BERT: Pretrained on multilingual corpora and fine-tuned on HASOC, FIRE, and custom Hinglish datasets [39]. Features Extracted: • TF-IDF vectors for traditional ML. • CLS token embeddings for BERT. Thresholding: • If hatefulness probability > 0.7, label as hate speech. • If between 0.4–0.7, pass for fuzzy aggregation. Output: Probability scores of hatefulness with attention to model explainability (feature importance, SHAP values) [40]. 4.2.4 Graph-Based Context Mapping Objective: To capture broader conversational and user network context. Graph Construction: • Nodes: Users, entities, topics extracted using spaCy + custom NER models [41]. • Edges: Relations like 'reply-to', 'mentions', 'hashtags-in-common'. Features: • Centrality Measures: Betweenness, closeness to detect influencers [42]. • Propagation Paths: To identify coordinated hate campaigns [43]. • Subgraph Clustering: Detect isolated hate communities. Trigger Conditions: • Sudden spikes in hate propagation. • High centrality hate nodes. • Clusters with high hate content density. Output: Graph anomaly scores indicating coordinated hate or isolated abuse patterns. 4.2.5 Fuzzy Logic Aggregation Layer Objective: To combine outputs from Rule, ML, and Graph layers in an explainable manner. Process: • Define fuzzy sets: LOW, MEDIUM, HIGH for each input. • Create fuzzy rules, e.g.: o IF Rule-Score is HIGH AND ML-Score is HIGH THEN HATE is VERY HIGH. o IF Rule-Score is LOW AND Graph-Score is HIGH THEN AMBIGUOUS. Fuzzy Engine: • Mamdani-type inference [44]. • Defuzzification using centroid method. Advantages: • Handles uncertainty explicitly. • Provides linguistic reasoning (“high risk due to graph cluster and partial rule match”) [45]. Output: Final hatefulness score between 0–1, with interpretability. 4.2.6 Human-in-the-Loop Review Layer Objective: To handle ambiguous or borderline cases with human judgement. Trigger Conditions: • Final hatefulness score between 0.4 and 0.6. • Content flagged as satire, political speech, or religious commentary. Process: • Cases routed to compliance officers trained in Indian legal standards. • Manual annotation whether content violates IPC/IT Rules. Benefits: • Adds explainable human oversight. • Reduces false positives in sensitive categories like satire or legitimate dissent [46]. 5. System Architecture 5.1 Detailed Flow The proposed censorship system follows a layered sequential flow where each module adds further evaluation or enrichment to the incoming content. The data journey through the system is as follows: 1. Input Capture: Text (tweets, posts, comments) or English speech (audio) is captured. 2. Preprocessing Layer: Text normalization, tokenization, language detection, and speech transcription are performed. 3. Rule-Based Filtering: Content is scanned against pre-defined dictionaries, templates, and phrase banks corresponding to legal violations. 4. Machine Learning Detection: Content is evaluated by LightGBM and fine-tuned BERT models to assign probabilistic hatefulness scores. 5. Graph-Based Context Mapping: Entities, relationships, and topic propagation are mapped. Graph-based anomalies and cluster analysis are performed. 6. Fuzzy Logic Aggregation: Inputs from Rule, ML, and Graph layers are aggregated using fuzzy rules to derive a final hatefulness score. 7. Human-in-the-Loop Review: Ambiguous cases are escalated to trained compliance officers for manual judgement. 8. Action Execution: Based on the final decision: • Content flagged for review • Content auto-deleted • User flagged or suspended • No action if content deemed legal 5.2 Token Passing between Layers Each layer appends metadata to the content token, including: • Language tag • Rule matches [8] B. Mathew, P. Saha, and A. Mukherjee, “HateXplain: Explainable Hate Speech Detection,” EMNLP, 2020. [9] J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers,” NAACL, 2019. [10] S. Kiritchenko and S. Mohammad, “Best-Worst Scaling More Reliable than Rating Scales: A Case Study on Sentiment Intensity Annotation,” ACL Anthology, 2016. [11] K. Vidgen and B. Derczynski, “Directions in Abusive Language Research: Challenges and Opportunities,” FIRE, 2020. [12] A. Wasi, M. Siddiqui, and K. Ahmed, “Islamophobia Detection on Social Media: A GraphBased Approach with Explainability,” ACL, 2024. [13] HASOC 2019 Dataset, available at: https://hasocfire.github.io/ [14] FIRE 2020 Dataset, available at: https://fire.irsi.res.in/fire/2020/home [15] Custom Dataset Annotated by Research Team, 2024. [16] G. Mishra and A. Satapathy, “Offensive Language Detection for Code-Switched Social Media Texts,” FIRE, 2021. [17] S. Basile, V. Basile, and N. Novielli, “Detecting Cyberbullying through Multimodal ContextAware Approaches,” CLiC-it, 2020. [18] A. Singh and V. Balasubramanian, “Text Normalization for Indic Languages,” COLING, 2020. [19] S. P. Saha, “Code-Switching in Indian Social Media,” Journal of South Asian Linguistics, 2022. [20] P. Mishra et al., “Trolling Detection Using Hybrid Deep Learning Approaches,” IEEE Access, 2022. [21] R. Waseem and D. Hovy, “Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter,” NAACL, 2016. [22] T. Davidson, D. Warmsley, and M. Macy, “Automated Hate Speech Detection and the Problem of Offensive Language,” ICWSM, 2017. [23] J. Yin and Y. Zubiaga, “Survey on Hate Speech Detection,” Computer Science Review, 2021. [24] Constitution of India, Article 19(1)(a) and 19(2). [25] Information Technology Rules, 2021, Ministry of Electronics and Information Technology, Government of India. [26] Indian Penal Code, 1860 (Sections 153A, 295A). [27] Karnataka High Court, “X Corp vs Union of India,” WP 13710/2022, Judgment, 2023. [28] Supreme Court of India, “Shreya Singhal v. Union of India,” AIR 2015 SC 1523. [29] D. Chhibber, “India’s New IT Rules: What They Mean for Free Speech and Censorship,” Indian Journal of Law and Technology, 2022. [30] K. Balakrishnan, “Intermediary Liability and Regulation in India,” Economic and Political Weekly, 2022. [31] Press Information Bureau, Government of India, 2022. [32] Joulin, A. Grave, P. Bojanowski, and M. Mikolov, “Bag of Tricks for Efficient Text Classification,” EACL, 2017. [33] SpaCy Documentation: Tokenization Algorithms, 2023. [34] Emoji Standardization Consortium, Unicode v14.0, 2021. [35] A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for SelfSupervised Learning of Speech Representations,” NeurIPS, 2020. [36] Compiled Hate Speech Lexicons, 2022. [37] Legal Phrase Mapping Database (Custom Research Compilation), 2024. [38] Ke, G. et al., “LightGBM: A Highly Efficient Gradient Boosting Decision Tree,” NeurIPS, 2017. [39] Transformer Fine-Tuning Practices, Huggingface Blog, 2022. [40] Lundberg, S. M., and Lee, S.-I., “A Unified Approach to Interpreting Model Predictions,” NeurIPS, 2017. [41] Explosion AI, spaCy Named Entity Recognition, 2022. [42] Freeman, L. C., “Centrality in Social Networks: Conceptual Clarifications,” Social Networks, 1979. [43] Girvan, M., and Newman, M. E. J., “Community Structure in Social and Biological Networks,” PNAS, 2002. [44] L. A. Zadeh, “Fuzzy Sets,” Information and Control, 1965. [45] Mendel, J. M., “Fuzzy Logic Systems for Engineering: A Tutorial,” Proceedings of the IEEE, 1995. [46] Human-in-the-Loop Content Moderation, Accenture Insights, 2021. [47] HASOC FIRE 2019 Overview Paper. [48] FIRE 2020 Shared Task Overview Paper. [49] NLTK Stopword Lists, 2023. [50] Urban Dictionary Internet Slang Scraping, 2022. [51] Explainability in Machine Learning: A Guide for Practitioners, Google AI, 2020. [52] Ethical Risks in Hate Speech Detection Systems, ACM Conference on Fairness, Accountability, and Transparency, 2022. [53] Bias in Hate Speech Datasets: A Survey, AAAI Workshop, 2022. [54] Social Network Analysis in Hate Propagation: IEEE Transactions on Affective Computing, 2021. [55] Multilingualism Challenges in Social Media Analysis, Journal of Language and Technology, 2021. [56] Hinglish and Its Implications for NLP, FIRE Proceedings, 2020. [57] Strategies for Code-Mixed Language Detection, EMNLP Findings, 2021. [58] Graph Neural Networks for Community Detection, ICLR 2022. [59] Using Fuzzy Logic for Uncertainty in AI Systems, Springer Journal of AI, 2021. [60] Coordinated Inauthentic Behavior and Hate Speech, Facebook Research Reports, 2020. [61] Offensive Language Detection Challenges in Multilingual Settings, LREC, 2020. [62] Twitter Transparency Report, 2022. [63] Section 69A of IT Act: Indian Censorship Laws Review, 2022. [64] UNESCO Report on Online Hate Speech, 2021. [65] Multimodal Hate Speech Detection: State of the Art, IEEE Access, 2023. [66] Detecting Hidden Hate: Implicit Offensive Language Detection, ACL Findings, 2021. [67] Waseem, Z., and Hovy, D., “Hateful Symbols or Hateful People?” NAACL 2016. [68] The Hindu, Editorial on Free Speech and Content Moderation, 2022. [69] Satirical Content vs Hate Content: Machine Learning Challenges, COLING, 2022. [70] Bias Mitigation Techniques in NLP Models, NeurIPS, 2022. [71] Transfer Learning Challenges in Multilingual Contexts, EMNLP, 2021. [72] Indian Multilingual Datasets for Hate Speech, IIT Delhi Research, 2022. [73] Speech Recognition for Indian Accents: Wav2vec2 Fine-tuning, ICASSP, 2023. [74] Human Rights Watch Report on Censorship in India, 2023. [75] Analysis of Recent Indian IT Amendments, Economic and Political Weekly, 2023. [76] Law Enforcement Use of Social Media Censorship Tools, Harvard Law Review, 2022. [77] Future Directions in AI for Hate Speech Detection, Communications of the ACM, 2022.