A Text Classification Approach for Detecting Cyberbullying Risk on Twitter Using Support Vector Machine with Naive Bayes and Random Forest Comparison

Authors

  • Sri Yarsasi Magister of Computer Science, Amikom Purwokerto University
  • Angga Iskoko Magister of Computer Science, Amikom Purwokerto University

DOI:

https://doi.org/10.47738/ijiis.v8i4.290

Keywords:

cyberbullying detection, Twitter, support vector machine, TF-IDF, text classification, machine learning

Abstract

The rapid development of social media as a means of digital interaction also presents serious challenges in the form of the spread of negative content, including cyberbullying. Cyberbullying is a form of verbal violence committed online and has a significant impact on mental health, especially in adolescents. This research aims to develop a text classification model to detect the risk of cyberbullying using the Support Vector Machine (SVM) algorithm. The data used comes from a collection of cyberbullying-themed tweets. The research stages include text preprocessing (normalization, cleaning, tokenization, stopword removal, and stemming), feature extraction using Term Frequency-Inverse Document Frequency (TF-IDF), data division into training and testing sets, and model training using linear kernel of SVM. The model was evaluated using accuracy, precision, recall, and F1-score metrics. The results show that this approach is able to identify risky comments quite accurately, with optimal performance on the linear kernel. This research contributes to the development of automated detection systems to create a safer and healthier digital ecosystem, and supports preventive efforts in mitigating cyberbullying online.

References

Z. Zhao, Y. Zhou, L. Chen, and L. Wang, “Hate speech detection: A solved problem? The challenging case of long tail on Twitter,” in Semantic Web Evaluation Challenges, pp. 1–16, 2016.

K. Dinakar, R. Reichart, and H. Lieberman, “Modeling the detection of textual cyberbullying,” in The Social Mobile Web, pp. 11–17, 2011.

M. Dadvar, F. de Jong, R. Ordelman, and D. Trieschnigg, “Improved cyberbullying detection using gender information,” in Proc. 12th Dutch-Belgian Inf. Retrieval Workshop, 2013.

K. Elissa, “Title of paper if known,” unpublished.

R. Nicole, “Title of paper with only first word capitalized,” J. Name Stand. Abbrev., in press.

Y. Yorozu, M. Hirano, K. Oka, and Y. Tagawa, “Electron spectroscopy studies on magneto-optical media and plastic substrate interface,” IEEE Transl. J. Magn. Japan, vol. 2, pp. 740-741, August 1987 [Digests 9th Annual Conf. Magnetics Japan, p. 301, 1982].

M. Young, The Technical Writer’s Handbook. Mill Valley, CA: University Science, 1989.

B. I. Kusuma and A. Nugroho, “Cyberbullying detection on Twitter uses the Support Vector Machine method,” J. Tek. Inform. (JUTIF), vol. 5, no. 1, pp. 11–17, 2024.

K. Lukman and S. Novianto, “Komparasi algoritma Naïve Bayes dan SVM untuk identifikasi cyberbullying selebriti di media sosial Twitter,” J. Algoritma, vol. 22, no. 1, pp. 970–981, 2025.

F. Farasalsabila, E. Utami, and H. Hanafi, “Deteksi cyberbullying menggunakan BERT dan Bi-LSTM,” J. Teknol., vol. 17, no. 1, pp. 1–6, 2024.

F. Muftie, K. Muftie, and Q. Addina, “Perbandingan performa deteksi cyberbullying dengan transformer, deep learning, dan machine learning,” J. Pendidik. Inf. dan Sains, vol. 13, no. 1, pp. 115–128, 2024.

I. A. Asqolani and E. B. Setiawan, “Hybrid deep learning approach and Word2Vec feature expansion for cyberbullying detection on Indonesian Twitter,” Indones. J. Inf. Sci., vol. 28, no. 4, pp. 123–136, 2023.

D. Krisnandi, R. N. Ambarwati, A. Y. Asih, A. Ardiansyah, and H. F. Pardede, “Analisis komentar cyberbullying terhadap kata yang mengandung toksisitas dan agresi menggunakan Bag of Words dan TF-IDF dengan klasifikasi SVM,” J. Linguistik Komput., vol. 6, no. 2, pp. 36–41, 2023.

W. A. Prabowo and F. Azizah, “Sentiment analysis for detecting cyberbullying using TF-IDF and SVM,” J. RESTI (Rekayasa Sist. dan Teknol. Inf.), vol. 4, no. 6, pp. 1250–1260, 2022.

L. Komati and K. Y. Reddy, “Cyberbullying detection on social media: Leveraging TF-IDF and LSTM for robust classification,” J. Data Acquis. Process., vol. 40, no. 1, pp. 56–71, 2025.

S. Chen, J. Wang, and K. He, “Chinese cyberbullying detection using XLNet and deep Bi-LSTM hybrid model,” Information, vol. 15, no. 2, p. 93, 2024.

R. Joshi and A. Gupta, “Performance comparison of simple Transformer and Res-CNN-BiLSTM for cyberbullying classification,” arXiv preprint, arXiv:2206.02206, 2022.

A. G. Philipo, D. S. Sarwatt, J. Ding, M. Daneshmand, and H. Ning, “Assessing text classification methods for cyberbullying detection on social media platforms,” arXiv preprint, arXiv:2412.19928, 2024.

H.-Y. Chen and C.-T. Li, “HENIN: Learning heterogeneous neural interaction networks for explainable cyberbullying detection on social media,” arXiv preprint, arXiv:2010.04576, 2020.

M. S. Akter, H. Shahriar, and A. Cuzzocrea, “A trustable LSTM-Autoencoder network for cyberbullying detection on social media using synthetic data,” arXiv preprint, arXiv:2308.09722, 2023.

Downloads

Published

2025-12-24

Issue

Section

Articles