A Multi-Platform Cyberbullying and Toxic Span Detection Dataset for Advanced NLP
Le résumé fourni par la source
This dataset features 75,000 carefully curated and fully balanced social media textual records designed for cyberbullying classification and sequence-labeling toxic span detection. Collected and mapped across four distinct digital ecosystems—Facebook (FB), Instagram (IG), Reddit, and online News Portals—the corpus spans seven operational categories: Racism/Xenophobia, Misogyny/Sexism, Homophobia/LGBTQphobia, Religious Hate, Harassment/Stalking, Profanity/Insult, and a baseline of Non-Toxic control samples. A standout feature of this dataset is the precise character-level localization tags (toxic_span_start and toxic_span_end), which mark the exact spans of offensive words within toxic strings. This structural architecture makes it ideal for training multi-task learning models, sequence tagging frameworks, and deep learning models (such as BERT, RoBERTa, and BiLSTM-CRF) aimed at explainable AI and targeted content moderation.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.