Yoruba language text-based hate detection model for social media
关键词:
classification model, confusion metrics, hate speech, support vector machine, yoruba language摘要
There is an emerging problem of hate speech sharing and spreading associated with Nigerian social media. This work was targeted to develop classification model for text based hate speech in Yoruba language. Support Vector Machine (SVM) was used on crowd sourced datasets comprising 1,420 Yoruba extracted from social media platforms. Annotated 71% Yoruba text datasets were used to train the SVM classifier. The data sets were preprocessed using the Python libraries. The technologies used for the front end are HTML version 5.0, CSS version 3.0, Java script, jQuery, and Bootstrap. The back end involved the use of Laravel version 9.0, a PHP framework, and MySQL. 1008 and 412 datasets were used to train and test the Yoruba SVM classifier, respectively. The confusion matrix indicated that the trained SVM Classifier for Yoruba was able to identify Yoruba hate speech. Accuracy, Precision, Recall, F1-Score and AUC are 75.49%, 75.21%, 0.0954, 0.8410, and 0.60, respectively for the Yoruba SVM Classifier. The model performance compared with other existing models showed that for Yoruba database; Naïve Bayes gave 69.09%, accuracy, and 72.64% precision. Logistic Regression gave 66.51% accuracy, and 68.89% precision. KNN gave 75.52% accuracy, and 86.91% precision. Random Forest gave 54.88% accuracy, and 50.65% precision.
