Comparing SVM, Random Forest, Logistic Regression, and Decision Tree for TikTok Sentiment Analysis of Koperasi Desa Merah Putih

Main Article Content

Salman Alfarisi
Lidya Lunardi
Fauzi

Abstract

The Koperasi Desa Merah Putih program is a strategic Indonesian government initiative designed to strengthen village economies through integrated cooperative services. This study analyzes sentiment expressed in Indonesian-language TikTok comments and compares Support Vector Machine, Random Forest, Logistic Regression, and Decision Tree under identical experimental conditions. The dataset comprised 49,805 labeled comments obtained from a public Kaggle repository. Data preparation included quality screening, case folding, slang normalization, tokenization, stop-word removal while retaining negation terms, and Indonesian stemming. The processed comments were represented using Term Frequency Inverse Document Frequency features and divided through a stratified 80:20 hold-out scheme, producing 39,844 training instances and 9,961 testing instances. Because Neutral comments represented 72.89% of the test set, evaluation emphasized macro-averaged metrics alongside accuracy. Support Vector Machine achieved the best performance, with 93.35% accuracy, 91.32% macro precision, 84.78% macro recall, and 87.70% macro F1-score, providing the most balanced overall classification performance among the evaluated models.

Section
Articles

References

D. Mulyadi and A. Rahmati, “Revitalizing Merah Putih Village Cooperatives through Maqasid Sharia : A qualitative study of rural economic recovery in Indonesia,” J. Islam. Econ. Lariba, vol. 12, no. 1, pp. 673–694, 2026, doi: https://doi.org/10.20885/jielariba.vol12.iss1.art24.

“Inilah Inpres 9/2025 tentang Percepatan Pembentukan Koperasi Desa/Kelurahan Merah Putih | Sekretariat Negara.”

H. S. Ramadhan, A. S. Akbar, K. Y. Sinaga, L. Muthoharoh, A. Satria, and M. C. T. Manullang, “Sentiment Analysis of AI Adoption in Indonesian Higher Education Using Machine Learning and Transformer-Based Models,” Apr. 2026.

A. N. I. Pasha, E. F. Putri, L. Muthoharoh, A. Satria, and M. C. T. Manullang, “Hybrid TF--IDF Logistic Regression and MLP Neural Baseline for Indonesian Three-Class Sentiment Analysis on Social Media Text,” May 2026.

I. Muhandhis and A. S. Ritonga, “Public Sentiment Analysis on TikTok about Tapera Policy using Random Forest Classifier,” SISTEMASI, vol. 14, no. 1, p. 354, Jan. 2025, doi: https://doi.org/10.32520/stmsi.v14i1.4878.

Siti Rihastuti and Afnan Rosyidi, “ANALISIS SENTIMEN PENGGUNA TIKTOK TENTANG PROGRES PEMBANGUNAN IKN DENGAN METODE RANDOM FOREST,” J. Comput. Sci. Technol., vol. 5, no. 1, pp. 19–23, May 2025, doi: https://doi.org/10.54840/jcstech.v5i1.345.

S. Fide, S. Suparti, and S. Sudarno, “No Title,” J. Gaussian; Vol 10, No 3 J. GaussianDO - https://doi.org/10.14710/j.gauss.10.3.346-358 , Dec. 2021.

R. Tangke, D. Tineke Salaki, W. Widsli Kalengkongan, and E. Ketaren, “Analisis Sentimen Aplikasi TikTok Menggunakan Algoritma Support Vector Machine (SVM) dan Random Forest,” J. TIMES, vol. 13, no. 2, pp. 53–62, Dec. 2024, doi: https://doi.org/10.51351/jtm.13.2.2024762.

R. Renaldi and Y. Kurnia, “Alleged Bad Credit at Saving Cooperatives Borrow Flamboyant Assistance PPSW Jakarta With Comparasion the Algorithms Naive Bayes and C4.5,” bit-Tech, vol. 2, no. 3, pp. 141–147, Nov. 2020, doi: https://doi.org/10.32877/bt.v2i3.163.

C. F. Mariwy, L. Y. Baisa, and A. L. Sumendap, “Comparative Study of Machine Learning Methods for Sentiment Analysis of TikTok Comments Related to Cyberbullying,” Indones. J. Artif. Intell. Data Min., vol. 9, no. 1 SE-Articles, pp. 140–152, doi: https://doi.org/10.24014/ijaidm.v9i1.39183.

O. S. D. Fadhillah, J. H. Jaman, and C. Carudin, “Perbandingan Naive Bayes, Support Vector Machine, Logistic Regression dan Random Forest dalam Menganalisis Sentimen Mengenai TikTokShop,” J. Inform. dan Tek. Elektro Terap., vol. 13, no. 1, Jan. 2025, doi: https://doi.org/10.23960/jitet.v13i1.5746.

R. E. Nurfirdaus, M. A. Barata, and I. W. Prastya, “Comparison of SVM and Random Forest for TikTok E10 Fuel Sentiment Analysis,” J. Appl. Informatics Comput., vol. 10, no. 2, pp. 1426–1437, Apr. 2026, doi: https://doi.org/10.30871/jaic.v10i2.12395.

A. Hermawan, L. Lunardi, Y. Kurnia, B. Daniawan, and Junaedi, “Optimizing Convolutional Neural Networks with Particle Swarm Optimization for Enhanced Hoax News Detection,” J. Inf. Syst. Eng. Bus. Intell., vol. 11, no. 1, pp. 53–64, Mar. 2025, doi: https://doi.org/10.20473/jisebi.11.1.53-64.

N. Azzahra, A. Hermawan, Junaedi, Y. Kurnia, and Edy, “Impact of Dataset Background on Deep Learning-Based Waste Classification,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 10, no. 3, pp. 580–589, Jun. 2026, doi: https://doi.org/10.29207/resti.v10i3.6965.

Y.-W. Mak, H.-N. Goh, and A. H.-L. Lim, “Forum Text Processing and Summarization,” JOIV Int. J. Informatics Vis., vol. 8, no. 1, p. 425, Mar. 2024, doi: https://doi.org/10.62527/joiv.8.1.2279.

Y. HaCohen-Kerner, D. Miller, and Y. Yigal, “The influence of preprocessing on text classification using a bag-of-words representation,” PLoS One, vol. 15, no. 5, p. e0232525, May 2020, doi: https://doi.org/10.1371/journal.pone.0232525.

S. M. Nagarajan and U. D. Gandhi, “Classifying streaming of Twitter data based on sentiment analysis using hybridization,” Neural Comput. Appl., vol. 31, no. 5, pp. 1425–1433, 2019, doi: https://doi.org/10.1007/s00521-018-3476-3.

D. Michail, N. Kanakaris, and I. Varlamis, “Detection of fake news campaigns using graph convolutional networks,” Int. J. Inf. Manag. Data Insights, vol. 2, no. 2, p. 100104, 2022, doi: https://doi.org/10.1016/j.jjimei.2022.100104.

F. Rahutomo, I. Y. R. Pratiwi, and D. M. Ramadhani, “Eksperimen Naïve Bayes Pada Deteksi Berita Hoax Berbahasa Indonesia,” J. Penelit. Komun. Dan Opini Publik, vol. 23, no. 1, 2019, doi: https://doi.org/10.33299/jpkop.23.1.1805.

D. K. Vishwakarma, D. Varshney, and A. Yadav, “Detection and veracity analysis of fake news via scrapping and authenticating the web search,” Cogn. Syst. Res., vol. 58, pp. 217–229, 2019, doi: https://doi.org/10.1016/j.cogsys.2019.07.004.

T. Ahmed Khan, R. Sadiq, Z. Shahid, M. M. Alam, and M. Mohd Su’ud, “Sentiment Analysis using Support Vector Machine and Random Forest,” J. Informatics Web Eng., vol. 3, no. 1, p. 67, Feb. 2024, doi: https://doi.org/10.33093/jiwe.2024.3.1.5.

M. F. Ruziqiana, L. Hidayah, and M. A. Rasyidi, “Detection of Cyberbullying Using Svm, Naive Bayes, and Random Forest Algorithm,” J. Inform. dan Tek. Elektro Terap., vol. 12, no. 3, pp. 2830–7062, 2024.

R. Hidayat and F. Aminulhaq, “A Comparative Analysis of Decision Tree, Logistic Regression, and Support Vector Machine Algorithms in Sentiment Analysis of Threads App Reviews,” Intechno J. (Information Technol. Journal), vol. 7, no. 2, pp. 45–55, Dec. 2025, doi: https://doi.org/10.24076/intechnojournal.2025v7i2.2497.

A. T. P. Subandono and D. Ariatmanto, “Optimizing Feature Selection in Sentiment Analysis of Bank Saqu: A Comparative Study of SVM and Random Forest using Information Gain and Chi-Square,” SISTEMASI, vol. 14, no. 3, p. 1205, May 2025, doi: https://doi.org/10.32520/stmsi.v14i3.5106.

S. Kristianto and Y. Kurnia, “Data Mining Implementation on Choosing Potential Customers Using K-Means Algorithm on PT. Koba Metal Indonesia,” Tech-E, vol. 1, no. 2, p. 1, Feb. 2018, doi: https://doi.org/10.31253/te.v1i2.45.