Sentiment Analysis of Indonesian Tweets on the Free Lunch and Milk Program Using TF-IDF and Naïve Bayes
Main Article Content
Abstract
This study analyzes public sentiment expressed in Indonesian-language tweets concerning the Free Lunch and Milk Program using Term Frequency–Inverse Document Frequency (TF-IDF) and the Multinomial Naïve Bayes algorithm. A total of 964 relevant tweets were collected from X and classified into positive and negative sentiment categories. The research process consisted of data collection, text cleaning, normalization, tokenization, stopword removal, stemming, TF-IDF feature extraction, model training, and performance evaluation. The dataset comprised 500 positive tweets and 464 negative tweets, indicating a relatively balanced class distribution. The proposed model achieved an accuracy of 88%, precision of 85%, recall of 82%, and an F1-score of 83%. Confusion matrix analysis indicated that most tweets were correctly classified, although errors remained in tweets containing sarcasm, irony, mixed sentiment, and indirect policy references. These findings demonstrate that the combination of TF-IDF and Naïve Bayes provides an efficient and interpretable baseline for monitoring public opinion on Indonesian social policy. Nevertheless, more contextual language models are required to improve the classification of complex and implicit sentiment expressions.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
By signing this statement, I hereby assign and transfer to JCSIS all exclusive copyright rights relating to the work identified above. These rights include, without limitation, the authority to publish, reproduce, revise, adapt, distribute, transmit, market, sell, and otherwise use the work and any associated materials worldwide, either in full or in part, in any language and through any existing or future form of electronic, printed, digital, or other media. JCSIS may also authorize or license third parties to exercise any of these rights. understand that these exclusive rights will be vested in JCSIS from the date on which the article is formally accepted for publication. As the copyright owner, JCSIS will have the exclusive authority to approve, license, or permit the reproduction, distribution, and other uses of the article. Nevertheless, all proprietary rights that are separate from copyright, including patent rights and rights relating to any process, method, or procedure described in the work, will remain with the author. I further acknowledge that JCSIS permits authors to reuse and share their published articles in accordance with the terms and conditions of the applicable Creative Commons license.
References
Y. Indulkar and A. Patil, “Comparative Study of Machine Learning Algorithms for Twitter Sentiment Analysis,” in 2021 International Conference on Emerging Smart Computing and Informatics (ESCI), IEEE, Mar. 2021, pp. 295–299. doi: https://doi.org/10.1109/ESCI50559.2021.9396925.
Z. Qi, “The Text Classification of Theft Crime Based on TF-IDF and XGBoost Model,” in 2020 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA), IEEE, Jun. 2020, pp. 1241–1246. doi: https://doi.org/10.1109/ICAICA50127.2020.9182555.
M. Wongkar and A. Angdresey, “Sentiment Analysis Using Naive Bayes Algorithm Of The Data Crawler: Twitter,” in 2019 Fourth International Conference on Informatics and Computing (ICIC), IEEE, Oct. 2019, pp. 1–5. doi: https://doi.org/10.1109/ICIC47613.2019.8985884.
“Feature Extraction for Sentiment Analysis in Indonesian Twitter,” in Nusantara Science and Technology Proceedings, Galaxy Science, Apr. 2021. doi: https://doi.org/10.11594/nstp.2021.0913.
S. Shaleha, A. Saputri, and H. M. Wicaksana, “Sentiment Analysis with Supervised Topic Modelling on Twitter Data Related to Indonesian Election 2024,” in 2023 International Conference on Computer, Control, Informatics and its Applications (IC3INA), IEEE, Oct. 2023, pp. 37–42. doi: https://doi.org/10.1109/IC3INA60834.2023.10285800.
O. Oyebode and R. Orji, “Social Media and Sentiment Analysis: The Nigeria Presidential Election 2019,” in 2019 IEEE 10th Annual Information Technology, Electronics and Mobile Communication Conference (IEMCON), IEEE, Oct. 2019, pp. 0140–0146. doi: https://doi.org/10.1109/IEMCON.2019.8936139.
P. Chauhan, N. Sharma, and G. Sikka, “The emergence of social media data and sentiment analysis in election prediction,” J. Ambient Intell. Humaniz. Comput., vol. 12, no. 2, pp. 2601–2627, Feb. 2021, doi: https://doi.org/10.1007/s12652-020-02423-y.
Y. Kurnia, E. D. Kusuma, L. W. Kusuma, Suwitno, and W. Apridius, “Perbandingan Naïve Bayes dan CNN yang Dioptimasi PSO pada Identifikasi Berita Hoax Politik Indonesia,” bit-Tech, vol. 6, no. 3, pp. 340–352, Apr. 2024, doi: https://doi.org/10.32877/bt.v6i3.1225.
S. Srivastava, J. P. Singh, and D. Mangal, “Time and Domain Specific Twitter Data Mining for Plastic Ban based on Public Opinion,” in 2020 2nd International Conference on Innovative Mechanisms for Industry Applications (ICIMIA), IEEE, Mar. 2020, pp. 755–761. doi: https://doi.org/10.1109/ICIMIA48430.2020.9074935.
J. Yu, E. Domahidi, K. Benmaarouf, and N. Steinmetz, “Quantity or quality? Comparing social media data sampling strategies for government crisis communication research,” Int. J. Disaster Risk Reduct., vol. 124, p. 105531, Jun. 2025, doi: https://doi.org/10.1016/j.ijdrr.2025.105531.
L. Bozarth and C. Budak, “Keyword expansion techniques for mining social movement data on social media,” EPJ Data Sci., vol. 11, no. 1, p. 30, May 2022, doi: https://doi.org/10.1140/epjds/s13688-022-00343-9.
“Twitter Opinion Mining,” in Encyclopedia of Social Network Analysis and Mining, New York, NY: Springer New York, 2018, pp. 3241–3241. doi: https://doi.org/10.1007/978-1-4939-7131-2_101390.
A. Hermawan, L. Lunardi, Y. Kurnia, B. Daniawan, and Junaedi, “Optimizing Convolutional Neural Networks with Particle Swarm Optimization for Enhanced Hoax News Detection,” J. Inf. Syst. Eng. Bus. Intell., vol. 11, no. 1, pp. 53–64, Mar. 2025, doi: https://doi.org/10.20473/jisebi.11.1.53-64.
Nurhafnita, R. Adriman, and T. F. Abidin, “Hotel Review Sentiment Analysis Using Indonesian Language Based on Machine Learning,” in 2022 International Conference on Electrical Engineering, Computer and Information Technology (ICEECIT), IEEE, Nov. 2022, pp. 25–28. doi: https://doi.org/10.1109/ICEECIT55908.2022.10030546.
“Studying User Behavior on Twitter & Instagram Using Computer Vision & Natural Language Processing,” 2019, SAGE Publications Ltd, London United Kingdom. doi: https://doi.org/10.4135/9781526493415.
M. Alirridlo, A. F. Septiyanto, R. Sarno, and D. Sunaryono, “A Comparative Analysis of Text Normalization Techniques for Enhanced Sentiment Analysis Performance,” in 2024 Beyond Technology Summit on Informatics International Conference (BTS-I2C), IEEE, Dec. 2024, pp. 474–479. doi: https://doi.org/10.1109/BTS-I2C63534.2024.10942038.
A. W. Pradana and M. Hayaty, “The Effect of Stemming and Removal of Stopwords on the Accuracy of Sentiment Analysis on Indonesian-language Texts,” Kinet. Game Technol. Inf. Syst. Comput. Network, Comput. Electron. Control, pp. 375–380, Oct. 2019, doi: https://doi.org/10.22219/kinetik.v4i4.912.
E. Araslanov, E. Komotskiy, and E. Agbozo, “Assessing the Impact of Text Preprocessing in Sentiment Analysis of Short Social Network Messages in the Russian Language,” in 2020 International Conference on Data Analytics for Business and Industry: Way Towards a Sustainable Economy (ICDABI), IEEE, Oct. 2020, pp. 1–4. doi: https://doi.org/10.1109/ICDABI51230.2020.9325654.
B. Seref and E. Bostanci, “Sentiment Analysis using Naive Bayes and Complement Naive Bayes Classifier Algorithms on Hadoop Framework,” in 2018 2nd International Symposium on Multidisciplinary Studies and Innovative Technologies (ISMSIT), IEEE, Oct. 2018, pp. 1–7. doi: https://doi.org/10.1109/ISMSIT.2018.8567243.
F. Resyanto, Y. Sibaroni, and A. Romadhony, “Choosing The Most Optimum Text Preprocessing Method for Sentiment Analysis: Case:iPhone Tweets,” in 2019 Fourth International Conference on Informatics and Computing (ICIC), IEEE, Oct. 2019, pp. 1–5. doi: https://doi.org/10.1109/ICIC47613.2019.8985943.
B. Pang and L. Lee, “Opinion Mining and Sentiment Analysis,” Found. Trends® Inf. Retr., vol. 2, no. 1–2, pp. 1–135, Jul. 2008, doi: https://doi.org/10.1561/1500000011.
B. Liu, “Sentiment Analysis and Opinion Mining,” Synth. Lect. Hum. Lang. Technol., vol. 5, no. 1, pp. 1–167, May 2012, doi: https://doi.org/10.2200/S00416ED1V01Y201204HLT016.
L. Hong and B. D. Davison, “Empirical study of topic modeling in Twitter,” in Proceedings of the First Workshop on Social Media Analytics, New York, NY, USA: ACM, Jul. 2010, pp. 80–88. doi: https://doi.org/10.1145/1964858.1964870.
S. Zhang, X. Zhang, J. Chan, and P. Rosso, “Irony detection via sentiment-based transfer learning,” Inf. Process. Manag., vol. 56, no. 5, pp. 1633–1644, Sep. 2019, doi: https://doi.org/10.1016/j.ipm.2019.04.006.
S. Oprea and W. Magdy, “Exploring Author Context for Detecting Intended vs Perceived Sarcasm,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 2854–2859. doi: https://doi.org/10.18653/v1/P19-1275.