Arabic offensive language on twitter: Analysis and experiments

Title Arabic offensive language on twitter: Analysis and experiments
Author Mubarak, H., Rashed, Ammar, Darwish, K., Samih, Y., Abdelali, A.
Publication Date: 2021
Publication Place - Association for Computational Linguistics (ACL)
Type Document
Language English
Digital Yes
Manuscript No
Library: Özyeğin University
Library Asset ID 978-195408509-1
Record ID 4ba44e73-6907-4603-929c-e8667dbff1c9
Date 2021
Sample Text Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for building a dataset that is not biased by topic, dialect, or target. We produce the largest Arabic dataset to date with special tags for vulgarity and hate speech. We thoroughly analyze the dataset to determine which topics, dialects, and gender are most associated with offensive tweets and how Arabic speakers use offensive language. Lastly, we conduct many experiments to produce strong results (F1 = 83.2) on the dataset using SOTA techniques.
View in source Özyeğin University Özyeğin University - Historical works, archives, and periodicals search engine
Özyeğin University - Historical works, archives, and periodicals search engine Özyeğin University

Arabic offensive language on twitter: Analysis and experiments

Author Mubarak, H., Rashed, Ammar, Darwish, K., Samih, Y., Abdelali, A.
Publication Date 2021
Publication Place - Association for Computational Linguistics (ACL)
Type Document
Language English
Digital Yes
Manuscript No
Library Özyeğin University
Library Asset ID 978-195408509-1
Record ID 4ba44e73-6907-4603-929c-e8667dbff1c9
Date 2021
Sample Text Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for building a dataset that is not biased by topic, dialect, or target. We produce the largest Arabic dataset to date with special tags for vulgarity and hate speech. We thoroughly analyze the dataset to determine which topics, dialects, and gender are most associated with offensive tweets and how Arabic speakers use offensive language. Lastly, we conduct many experiments to produce strong results (F1 = 83.2) on the dataset using SOTA techniques.
Özyeğin University - Historical works, archives, and periodicals search engine
Özyeğin University You are being redirected...

Please wait