A Text Classification Approach for Detecting Cyberbullying Risk on Twitter Using Support Vector Machine with Naive Bayes and Random Forest Comparison
DOI:
https://doi.org/10.47738/ijiis.v8i4.290Keywords:
cyberbullying detection, Twitter, support vector machine, TF-IDF, text classification, machine learningAbstract
The rapid development of social media as a means of digital interaction also presents serious challenges in the form of the spread of negative content, including cyberbullying. Cyberbullying is a form of verbal violence committed online and has a significant impact on mental health, especially in adolescents. This research aims to develop a text classification model to detect the risk of cyberbullying using the Support Vector Machine (SVM) algorithm. The data used comes from a collection of cyberbullying-themed tweets. The research stages include text preprocessing (normalization, cleaning, tokenization, stopword removal, and stemming), feature extraction using Term Frequency-Inverse Document Frequency (TF-IDF), data division into training and testing sets, and model training using linear kernel of SVM. The model was evaluated using accuracy, precision, recall, and F1-score metrics. The results show that this approach is able to identify risky comments quite accurately, with optimal performance on the linear kernel. This research contributes to the development of automated detection systems to create a safer and healthier digital ecosystem, and supports preventive efforts in mitigating cyberbullying online.References
Z. Zhao, Y. Zhou, L. Chen, and L. Wang, “Hate speech detection: A solved problem? The challenging case of long tail on Twitter,” in Semantic Web Evaluation Challenges, pp. 1–16, 2016.
K. Dinakar, R. Reichart, and H. Lieberman, “Modeling the detection of textual cyberbullying,” in The Social Mobile Web, pp. 11–17, 2011.
M. Dadvar, F. de Jong, R. Ordelman, and D. Trieschnigg, “Improved cyberbullying detection using gender information,” in Proc. 12th Dutch-Belgian Inf. Retrieval Workshop, 2013.
K. Elissa, “Title of paper if known,” unpublished.
R. Nicole, “Title of paper with only first word capitalized,” J. Name Stand. Abbrev., in press.
Y. Yorozu, M. Hirano, K. Oka, and Y. Tagawa, “Electron spectroscopy studies on magneto-optical media and plastic substrate interface,” IEEE Transl. J. Magn. Japan, vol. 2, pp. 740-741, August 1987 [Digests 9th Annual Conf. Magnetics Japan, p. 301, 1982].
M. Young, The Technical Writer’s Handbook. Mill Valley, CA: University Science, 1989.
B. I. Kusuma and A. Nugroho, “Cyberbullying detection on Twitter uses the Support Vector Machine method,” J. Tek. Inform. (JUTIF), vol. 5, no. 1, pp. 11–17, 2024.
K. Lukman and S. Novianto, “Komparasi algoritma Naïve Bayes dan SVM untuk identifikasi cyberbullying selebriti di media sosial Twitter,” J. Algoritma, vol. 22, no. 1, pp. 970–981, 2025.
F. Farasalsabila, E. Utami, and H. Hanafi, “Deteksi cyberbullying menggunakan BERT dan Bi-LSTM,” J. Teknol., vol. 17, no. 1, pp. 1–6, 2024.
F. Muftie, K. Muftie, and Q. Addina, “Perbandingan performa deteksi cyberbullying dengan transformer, deep learning, dan machine learning,” J. Pendidik. Inf. dan Sains, vol. 13, no. 1, pp. 115–128, 2024.
I. A. Asqolani and E. B. Setiawan, “Hybrid deep learning approach and Word2Vec feature expansion for cyberbullying detection on Indonesian Twitter,” Indones. J. Inf. Sci., vol. 28, no. 4, pp. 123–136, 2023.
D. Krisnandi, R. N. Ambarwati, A. Y. Asih, A. Ardiansyah, and H. F. Pardede, “Analisis komentar cyberbullying terhadap kata yang mengandung toksisitas dan agresi menggunakan Bag of Words dan TF-IDF dengan klasifikasi SVM,” J. Linguistik Komput., vol. 6, no. 2, pp. 36–41, 2023.
W. A. Prabowo and F. Azizah, “Sentiment analysis for detecting cyberbullying using TF-IDF and SVM,” J. RESTI (Rekayasa Sist. dan Teknol. Inf.), vol. 4, no. 6, pp. 1250–1260, 2022.
L. Komati and K. Y. Reddy, “Cyberbullying detection on social media: Leveraging TF-IDF and LSTM for robust classification,” J. Data Acquis. Process., vol. 40, no. 1, pp. 56–71, 2025.
S. Chen, J. Wang, and K. He, “Chinese cyberbullying detection using XLNet and deep Bi-LSTM hybrid model,” Information, vol. 15, no. 2, p. 93, 2024.
R. Joshi and A. Gupta, “Performance comparison of simple Transformer and Res-CNN-BiLSTM for cyberbullying classification,” arXiv preprint, arXiv:2206.02206, 2022.
A. G. Philipo, D. S. Sarwatt, J. Ding, M. Daneshmand, and H. Ning, “Assessing text classification methods for cyberbullying detection on social media platforms,” arXiv preprint, arXiv:2412.19928, 2024.
H.-Y. Chen and C.-T. Li, “HENIN: Learning heterogeneous neural interaction networks for explainable cyberbullying detection on social media,” arXiv preprint, arXiv:2010.04576, 2020.
M. S. Akter, H. Shahriar, and A. Cuzzocrea, “A trustable LSTM-Autoencoder network for cyberbullying detection on social media using synthetic data,” arXiv preprint, arXiv:2308.09722, 2023.
Downloads
Published
Issue
Section
License
Authors who publish with IJIIS : International Journal on Informatics and Information Systems agree to the following terms: Authors retain copyright and grant the IJIIS : International Journal on Informatics and Information Systems right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share (copy and redistribute the material in any medium or format) and adapt (remix, transform, and build upon the material) the work for any purpose, even commercially with an acknowledgement of the work's authorship and initial publication in IJIIS : International Journal on Informatics and Information Systems. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in IJIIS : International Journal on Informatics and Information Systems. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).

