Main Article Content

Abstract

Detecting harmful language on online platforms is essential for creating safer and more inclusive social media environments. However, research on harmful language detection has largely focused on high-resource languages, while low-resource languages such as Pashto remain comparatively underexplored. This study compares hate speech and dehumanizing language in Pashto social media comments. To address this gap, we developed a dataset containing 1,590 Pashto comments collected from social media and manually categorized them into four classes: Hate, Dehumanization, Both, and Neither. Three traditional machine-learning algorithms—Logistic Regression, Naive Bayes, and Linear Support Vector Machine (SVM)—were evaluated using term frequency–inverse document frequency (TF–IDF) features. Among the evaluated models, Logistic Regression achieved the best overall performance, reaching an accuracy of 71.63%. The results indicated that distinguishing between hate speech and dehumanizing language remains challenging, particularly when both forms occur within the same comment. The model showed a tendency to confuse the Hate and Dehumanization categories, with the Both class presenting additional classification difficulties. Furthermore, the experiments demonstrate that appropriate text normalization and the selection of an effective feature size—approximately 3,500 features in this study—can improve classification performance.

Keywords

Hate Speech Dehumanization Insult Animalization Offensive language

Article Details

How to Cite
Nasiri, S., & Wasi , N. A. (2026). Detection of Hate and Dehumanizing Speech in Pashto Text Using Machine Learning Algorithms . Journal of Natural Sciences – Kabul University, 9(2), 117–139. https://doi.org/10.62810/jns.v9i2.580

References

  1. Abulaish, M., Wasi, N. A., & Sharma, S. (2024). The role of lifelong machine learning in bridging the gap between human and machine learning: A scientometric analysis. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(1). https://doi.org/10.1002/widm.1526
  2. Bandura, A. (1999). Moral disengagement in the perpetration of inhumanities. Personality and Social Psychology Review, 3(3), 193–209. https://doi.org/https://doi.org/10.1207/s15327957pspr0303_3
  3. Burovova, K., & Romanyshyn, M. (2024). Computational Analysis of Dehumanization of Ukrainians on Russian Social Media. 28–39.
  4. Chhabra, A., & Kumar, D. (2023). A literature survey on multimodal and multilingual automatic hate speech identification. Multimedia Systems, 29. https://doi.org/10.1007/s00530-023-01051-8
  5. Davidson, T., Warmsley, D., Macy, M., & Weber, I. (2017). Automated Hate Speech Detection and the Problem of Offensive Language ∗. Proceedings of the International AAAI Conference on Web and Social Media, 11(1), 512-515. https://doi.org/https://doi.org/10.1609/icwsm.v11i1.14955
  6. Dhanya, L. K., & Balakrishnan, K. (2021). Hate speech Detection in Asian Languages : A Survey. International Conference on Communication, Control and Infomration Sciences (ICCISc). https://doi.org/10.1109/ICCISc52257.2021.9484922
  7. Engelmann, P., & Hardmeier, C. (2024). A Dataset for the Detection of Dehumanizing Language. 14–20.
  8. Fortuna, P. (2018). A Survey on Automatic Detection of Hate Speech in Text. ACM Comput. Surv., 51(4). https://doi.org/10.1145/3232676
  9. Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Pashto offensive language detection : a benchmark dataset and monolingual. PeerJ Computer Science, 9. https://doi.org/10.7717/peerj-cs.1617
  10. Haslam, N., & Haslam, N. (2006). Personality and Social Psychology Review. https://doi.org/10.1207/s15327957pspr1003
  11. Haslam, N., & Loughnan, S. (2013). Dehumanization and Infrahumanization. June, 1–25. https://doi.org/10.1146/annurev-psych-010213-115045
  12. Jahan, S., & Oussalah, M. (2023). Neurocomputing Survey paper language processing q. Neurocomputing, 546, 126232. https://doi.org/10.1016/j.neucom.2023.126232
  13. Janisar, A. A., & Afzal, H. (2019). A Framework to Detect Hate Speech in the Pashto Language from Social Media. September.
  14. Joshi, D. (2024). Decoding Dehumanization : Leveraging NLP to Identify Dehumanizing Decoding Dehumanization : Leveraging NLP to Identify Dehumanizing Language and Its Targets. NLPIR 2024, December 2024. https://doi.org/10.1145/3711542.3711598
  15. Kteily, N., Bruneau, E., Waytz, A., & Cotterill, S. (2015). The Ascent of Man: Theoretical and Empirical Evidence for Blatant Dehumanization. Journal of Personality and Social Psychology, 109, 901–931. https://doi.org/10.1037/pspp0000048
  16. Mendelsohn, J., Tsvetkov, Y., & Jurafsky, D. (2020). A Framework for the Computational Linguistic Analysis of Dehumanization. 3(August), 1–24. https://doi.org/10.3389/frai.2020.00055
  17. Poletto, F., Basile, V., & Sanguinetti, M. (2021). Resources and benchmark corpora for hate speech detection : a systematic review. Language Resources and Evaluation, 55(2), 477–523. https://doi.org/10.1007/s10579-020-09502-8
  18. Schmidt, A., & Wiegand, M. (2017). A Survey on Hate Speech Detection using Natural Language Processing. Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, 2012, 1–10. https://doi.org/10.18653/v1/W17-1101
  19. Sharma, D. P. (2026). Hate Speech Detection Research in South Asian Languages : A Survey of Tasks , Datasets and Methods Hate Speech Detection Research in South Asian Languages : A Survey of Tasks , Datasets and Methods. 24(3). https://doi.org/10.1145/3711710
  20. Vidgen, B., & Id, L. D. (2020). Directions in abusive language training data , a systematic review : Garbage in , garbage out. https://doi.org/10.1371/journal.pone.0243300
  21. Vidgen, B., Thrush, T., Waseem, Z., & Kiela, D. (2021). Learning from the Worst : Dynamically Generated Datasets to Improve Online Hate Detection. 1667–1682.
  22. Waseem, Z., & Hovy, D. (2016). Hateful Symbols or Hateful People ? Predictive Features for Hate Speech Detection on Twitter. Proceedings of NAACL-HLT, 88–93. https://doi.org/10.18653/v1/N16-2013
  23. Wasi, N. A., & Abulaish, M. (2020). Document-Level Sentiment Analysis through Incorporating Prior Domain Knowledge into Logistic Regression. 2020 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), 969–974. https://doi.org/10.1109/WIIAT50758.2020.00148
  24. Wasi, N. A., & Abulaish, M. (2023). An Unseen Features Enhanced Text Classification Approach. 2023 International Joint Conference on Neural Networks (IJCNN), 1–8. https://doi.org/10.1109/IJCNN54540.2023.10191857
  25. Wasi, N. A., & Abulaish, M. (2024a). Leveraging Unseen Features along with their PLM-based Representation to Handle Negative Covariate Shift Problem in Text Classification. 49(4). https://doi.org/10.2478/fcds-2024-0020
  26. Wasi, N. A., & Abulaish, M. (2024b). SKEDS — An external knowledge supported logistic regression approach for document-level sentiment classification. Expert Systems with Applications, 238. https://doi.org/10.1016/j.eswa.2023.121987.