Klasifikasi Sentimen Berita Hiburan Menggunakan Naive Bayes Classifier Dengan TF‑IDF dan Validasi Cross 5‑fold
Kata Kunci:
Analisis Sentimen, Berita Hiburan, Cross Validation, Naive Bayes, TF IDFAbstrak
This research aims to develop a sentiment classification model for Indonesian-language entertainment news, which is increasingly consumed thru online portals and has the potential to be used to map public opinion toward entertainment figures and issues. The model was built using the Multinomial Naive Bayes algorithm with Term Frequency–Inverse Document Frequency (TF-IDF) feature representation. The data, consisting of entertainment news text from the detikHOT portal, is processed thru the preprocessing stages (case folding, cleansing, tokenization, stopword removal, and stemming), then converted into TF-IDF vectors and labeled into positive and negative sentiment classes. Evaluation was conducted using a 5-fold cross-validation scheme to obtain stable performance estimates across different data splits. The experimental results show that the model achieved an average accuracy of 0.911, precision of 0.921, recall of 0.911, and an F1 score of 0.900, indicating high and balanced performance between prediction accuracy and the ability to capture all relevant sentiment cases. Consistent metric values across each fold indicate that the combination of Naive Bayes and TF-IDF has good generalization ability on the entertainment news corpus, and is competitive compared to previous studies in similar domains. This finding implies that this approach can be utilized as a basis for developing an automated sentiment monitoring system on entertainment media and opens up opportunities for further research using other algorithms or n-gram and word embedding-based features.





