Bangla Racism and Body-Shaming Detection Dataset
Résumé fourni par la source
This dataset contains 5,000 manually annotated Bangla-language text instances collected and curated for the detection of racism and body-shaming content on Bangla social media. It accompanies the paper "An Optimized Multi-Transformer Framework with Explainable AI for Racism and Body-Shaming Detection in Bangla Social Media." Files Bangla_Racism_BodyShaming_Dataset.csv — the full labeled dataset (UTF-8 encoded) Format CSV, UTF-8 (with BOM for Excel compatibility) Fields Column Description Sentence The target Bangla sentence being classified Sentiment Label: "Body Shaming" or "Racism" Story with Violence A short contextual narrative embedding the sentence, framed with a violent/harsh undertone Story without Violence A short contextual narrative embedding the same sentence, framed neutrally Class distribution Body Shaming: 2,573 samples Racism: 2,427 samples Total: 5,000 samples Language Bangla (Bengali script) Intended use Training and evaluation of transformer-based models (e.g., BanglaBERT, mBERT, MuRIL, XLM-RoBERTa) for hate speech / abusive language detection in low-resource Bangla NLP, and for explainable AI (XAI) research using techniques such as SHAP. Notes Samples were manually verified prior to inclusion. One fully empty row present in the original raw export was removed during cleaning. Please cite the accompanying paper if you use this dataset.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.