Main Article Content
Abstract
Detecting harmful language on online platforms is essential for creating safer and more inclusive social media environments. However, research on harmful language detection has largely focused on high-resource languages, while low-resource languages such as Pashto remain comparatively underexplored. This study compares hate speech and dehumanizing language in Pashto social media comments. To address this gap, we developed a dataset containing 1,590 Pashto comments collected from social media and manually categorized them into four classes: Hate, Dehumanization, Both, and Neither. Three traditional machine-learning algorithms—Logistic Regression, Naive Bayes, and Linear Support Vector Machine (SVM)—were evaluated using term frequency–inverse document frequency (TF–IDF) features. Among the evaluated models, Logistic Regression achieved the best overall performance, reaching an accuracy of 71.63%. The results indicated that distinguishing between hate speech and dehumanizing language remains challenging, particularly when both forms occur within the same comment. The model showed a tendency to confuse the Hate and Dehumanization categories, with the Both class presenting additional classification difficulties. Furthermore, the experiments demonstrate that appropriate text normalization and the selection of an effective feature size—approximately 3,500 features in this study—can improve classification performance.
Keywords
Article Details
Copyright (c) 2026 Kabul University

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
References
- Abulaish, M., Wasi, N. A., & Sharma, S. (2024). The role of lifelong machine learning in bridging the gap between human and machine learning: A scientometric analysis. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(1). https://doi.org/10.1002/widm.1526
- Bandura, A. (1999). Moral disengagement in the perpetration of inhumanities. Personality and Social Psychology Review, 3(3), 193–209. https://doi.org/https://doi.org/10.1207/s15327957pspr0303_3
- Burovova, K., & Romanyshyn, M. (2024). Computational Analysis of Dehumanization of Ukrainians on Russian Social Media. 28–39.
- Chhabra, A., & Kumar, D. (2023). A literature survey on multimodal and multilingual automatic hate speech identification. Multimedia Systems, 29. https://doi.org/10.1007/s00530-023-01051-8
- Davidson, T., Warmsley, D., Macy, M., & Weber, I. (2017). Automated Hate Speech Detection and the Problem of Offensive Language ∗. Proceedings of the International AAAI Conference on Web and Social Media, 11(1), 512-515. https://doi.org/https://doi.org/10.1609/icwsm.v11i1.14955
- Dhanya, L. K., & Balakrishnan, K. (2021). Hate speech Detection in Asian Languages : A Survey. International Conference on Communication, Control and Infomration Sciences (ICCISc). https://doi.org/10.1109/ICCISc52257.2021.9484922
- Engelmann, P., & Hardmeier, C. (2024). A Dataset for the Detection of Dehumanizing Language. 14–20.
- Fortuna, P. (2018). A Survey on Automatic Detection of Hate Speech in Text. ACM Comput. Surv., 51(4). https://doi.org/10.1145/3232676
- Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Pashto offensive language detection : a benchmark dataset and monolingual. PeerJ Computer Science, 9. https://doi.org/10.7717/peerj-cs.1617
- Haslam, N., & Haslam, N. (2006). Personality and Social Psychology Review. https://doi.org/10.1207/s15327957pspr1003
- Haslam, N., & Loughnan, S. (2013). Dehumanization and Infrahumanization. June, 1–25. https://doi.org/10.1146/annurev-psych-010213-115045
- Jahan, S., & Oussalah, M. (2023). Neurocomputing Survey paper language processing q. Neurocomputing, 546, 126232. https://doi.org/10.1016/j.neucom.2023.126232
- Janisar, A. A., & Afzal, H. (2019). A Framework to Detect Hate Speech in the Pashto Language from Social Media. September.
- Joshi, D. (2024). Decoding Dehumanization : Leveraging NLP to Identify Dehumanizing Decoding Dehumanization : Leveraging NLP to Identify Dehumanizing Language and Its Targets. NLPIR 2024, December 2024. https://doi.org/10.1145/3711542.3711598
- Kteily, N., Bruneau, E., Waytz, A., & Cotterill, S. (2015). The Ascent of Man: Theoretical and Empirical Evidence for Blatant Dehumanization. Journal of Personality and Social Psychology, 109, 901–931. https://doi.org/10.1037/pspp0000048
- Mendelsohn, J., Tsvetkov, Y., & Jurafsky, D. (2020). A Framework for the Computational Linguistic Analysis of Dehumanization. 3(August), 1–24. https://doi.org/10.3389/frai.2020.00055
- Poletto, F., Basile, V., & Sanguinetti, M. (2021). Resources and benchmark corpora for hate speech detection : a systematic review. Language Resources and Evaluation, 55(2), 477–523. https://doi.org/10.1007/s10579-020-09502-8
- Schmidt, A., & Wiegand, M. (2017). A Survey on Hate Speech Detection using Natural Language Processing. Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, 2012, 1–10. https://doi.org/10.18653/v1/W17-1101
- Sharma, D. P. (2026). Hate Speech Detection Research in South Asian Languages : A Survey of Tasks , Datasets and Methods Hate Speech Detection Research in South Asian Languages : A Survey of Tasks , Datasets and Methods. 24(3). https://doi.org/10.1145/3711710
- Vidgen, B., & Id, L. D. (2020). Directions in abusive language training data , a systematic review : Garbage in , garbage out. https://doi.org/10.1371/journal.pone.0243300
- Vidgen, B., Thrush, T., Waseem, Z., & Kiela, D. (2021). Learning from the Worst : Dynamically Generated Datasets to Improve Online Hate Detection. 1667–1682.
- Waseem, Z., & Hovy, D. (2016). Hateful Symbols or Hateful People ? Predictive Features for Hate Speech Detection on Twitter. Proceedings of NAACL-HLT, 88–93. https://doi.org/10.18653/v1/N16-2013
- Wasi, N. A., & Abulaish, M. (2020). Document-Level Sentiment Analysis through Incorporating Prior Domain Knowledge into Logistic Regression. 2020 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), 969–974. https://doi.org/10.1109/WIIAT50758.2020.00148
- Wasi, N. A., & Abulaish, M. (2023). An Unseen Features Enhanced Text Classification Approach. 2023 International Joint Conference on Neural Networks (IJCNN), 1–8. https://doi.org/10.1109/IJCNN54540.2023.10191857
- Wasi, N. A., & Abulaish, M. (2024a). Leveraging Unseen Features along with their PLM-based Representation to Handle Negative Covariate Shift Problem in Text Classification. 49(4). https://doi.org/10.2478/fcds-2024-0020
- Wasi, N. A., & Abulaish, M. (2024b). SKEDS — An external knowledge supported logistic regression approach for document-level sentiment classification. Expert Systems with Applications, 238. https://doi.org/10.1016/j.eswa.2023.121987.
References
Abulaish, M., Wasi, N. A., & Sharma, S. (2024). The role of lifelong machine learning in bridging the gap between human and machine learning: A scientometric analysis. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(1). https://doi.org/10.1002/widm.1526
Bandura, A. (1999). Moral disengagement in the perpetration of inhumanities. Personality and Social Psychology Review, 3(3), 193–209. https://doi.org/https://doi.org/10.1207/s15327957pspr0303_3
Burovova, K., & Romanyshyn, M. (2024). Computational Analysis of Dehumanization of Ukrainians on Russian Social Media. 28–39.
Chhabra, A., & Kumar, D. (2023). A literature survey on multimodal and multilingual automatic hate speech identification. Multimedia Systems, 29. https://doi.org/10.1007/s00530-023-01051-8
Davidson, T., Warmsley, D., Macy, M., & Weber, I. (2017). Automated Hate Speech Detection and the Problem of Offensive Language ∗. Proceedings of the International AAAI Conference on Web and Social Media, 11(1), 512-515. https://doi.org/https://doi.org/10.1609/icwsm.v11i1.14955
Dhanya, L. K., & Balakrishnan, K. (2021). Hate speech Detection in Asian Languages : A Survey. International Conference on Communication, Control and Infomration Sciences (ICCISc). https://doi.org/10.1109/ICCISc52257.2021.9484922
Engelmann, P., & Hardmeier, C. (2024). A Dataset for the Detection of Dehumanizing Language. 14–20.
Fortuna, P. (2018). A Survey on Automatic Detection of Hate Speech in Text. ACM Comput. Surv., 51(4). https://doi.org/10.1145/3232676
Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Pashto offensive language detection : a benchmark dataset and monolingual. PeerJ Computer Science, 9. https://doi.org/10.7717/peerj-cs.1617
Haslam, N., & Haslam, N. (2006). Personality and Social Psychology Review. https://doi.org/10.1207/s15327957pspr1003
Haslam, N., & Loughnan, S. (2013). Dehumanization and Infrahumanization. June, 1–25. https://doi.org/10.1146/annurev-psych-010213-115045
Jahan, S., & Oussalah, M. (2023). Neurocomputing Survey paper language processing q. Neurocomputing, 546, 126232. https://doi.org/10.1016/j.neucom.2023.126232
Janisar, A. A., & Afzal, H. (2019). A Framework to Detect Hate Speech in the Pashto Language from Social Media. September.
Joshi, D. (2024). Decoding Dehumanization : Leveraging NLP to Identify Dehumanizing Decoding Dehumanization : Leveraging NLP to Identify Dehumanizing Language and Its Targets. NLPIR 2024, December 2024. https://doi.org/10.1145/3711542.3711598
Kteily, N., Bruneau, E., Waytz, A., & Cotterill, S. (2015). The Ascent of Man: Theoretical and Empirical Evidence for Blatant Dehumanization. Journal of Personality and Social Psychology, 109, 901–931. https://doi.org/10.1037/pspp0000048
Mendelsohn, J., Tsvetkov, Y., & Jurafsky, D. (2020). A Framework for the Computational Linguistic Analysis of Dehumanization. 3(August), 1–24. https://doi.org/10.3389/frai.2020.00055
Poletto, F., Basile, V., & Sanguinetti, M. (2021). Resources and benchmark corpora for hate speech detection : a systematic review. Language Resources and Evaluation, 55(2), 477–523. https://doi.org/10.1007/s10579-020-09502-8
Schmidt, A., & Wiegand, M. (2017). A Survey on Hate Speech Detection using Natural Language Processing. Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, 2012, 1–10. https://doi.org/10.18653/v1/W17-1101
Sharma, D. P. (2026). Hate Speech Detection Research in South Asian Languages : A Survey of Tasks , Datasets and Methods Hate Speech Detection Research in South Asian Languages : A Survey of Tasks , Datasets and Methods. 24(3). https://doi.org/10.1145/3711710
Vidgen, B., & Id, L. D. (2020). Directions in abusive language training data , a systematic review : Garbage in , garbage out. https://doi.org/10.1371/journal.pone.0243300
Vidgen, B., Thrush, T., Waseem, Z., & Kiela, D. (2021). Learning from the Worst : Dynamically Generated Datasets to Improve Online Hate Detection. 1667–1682.
Waseem, Z., & Hovy, D. (2016). Hateful Symbols or Hateful People ? Predictive Features for Hate Speech Detection on Twitter. Proceedings of NAACL-HLT, 88–93. https://doi.org/10.18653/v1/N16-2013
Wasi, N. A., & Abulaish, M. (2020). Document-Level Sentiment Analysis through Incorporating Prior Domain Knowledge into Logistic Regression. 2020 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), 969–974. https://doi.org/10.1109/WIIAT50758.2020.00148
Wasi, N. A., & Abulaish, M. (2023). An Unseen Features Enhanced Text Classification Approach. 2023 International Joint Conference on Neural Networks (IJCNN), 1–8. https://doi.org/10.1109/IJCNN54540.2023.10191857
Wasi, N. A., & Abulaish, M. (2024a). Leveraging Unseen Features along with their PLM-based Representation to Handle Negative Covariate Shift Problem in Text Classification. 49(4). https://doi.org/10.2478/fcds-2024-0020
Wasi, N. A., & Abulaish, M. (2024b). SKEDS — An external knowledge supported logistic regression approach for document-level sentiment classification. Expert Systems with Applications, 238. https://doi.org/10.1016/j.eswa.2023.121987.