PRATAMA, REZA ADITYA (2026) PERBANDINGAN ALGORITMA NAIVE BAYES, RANDOM FOREST DAN SUPPORT VECTOR MACHINE: STUDI KASUS ANALISA SENTIMEN PENGGUNA DANA. S1 thesis, Universitas Mercu Buana Jakarta.
|
Text (HAL COVER)
COVER.pdf Download (659kB) | Preview |
|
|
Text (BAB I)
BAB I.pdf Restricted to Registered users only Download (101kB) |
||
|
Text (BAB II)
BAB II.pdf Restricted to Registered users only Download (397kB) |
||
|
Text (BAB III)
BAB III.pdf Restricted to Registered users only Download (294kB) |
||
|
Text (BAB IV)
BAB IV.pdf Restricted to Registered users only Download (437kB) |
||
|
Text (BAB V)
BAB V.pdf Restricted to Registered users only Download (28kB) |
||
|
Text (DAFTAR PUSTAKA)
DAFTAR PUSTAKA.pdf Restricted to Registered users only Download (89kB) |
||
|
Text (LAMPIRAN)
LAMPIRAN.pdf Restricted to Registered users only Download (665kB) |
Abstract
User reviews on the DANA application available on the Google Play Store can be used to understand users' opinions regarding the quality of the services provided. This study aims to compare the performance of Gaussian Naïve Bayes, Support Vector Machine (SVM), and Random Forest algorithms for sentiment analysis using the Term Frequency–Inverse Document Frequency (TF-IDF) weighting method. This study used 25,000 user reviews collected through web scraping from the Google Play Store. The research process included data preprocessing, TF-IDF weighting, splitting the dataset into training and testing sets with a 90:10 ratio, and sentiment classification using the three algorithms. Model performance was evaluated using Accuracy, Precision, Recall, F1-Score, Confusion Matrix, and Area Under the Curve (AUC). The results showed that Support Vector Machine (SVM) achieved the best performance with an accuracy of 89.00% and an AUC of 0.938, followed by Random Forest with an accuracy of 87.98% and an AUC of 0.934, while Gaussian Naïve Bayes obtained an accuracy of 85.77% and an AUC of 0.770. These findings indicate that the SVM algorithm is the most effective method for sentiment classification of DANA user reviews in this study. Keywords: Sentiment Analysis, DANA, TF-IDF, Gaussian Naïve Bayes, Support Vector Machine, Random Forest. Ulasan yang ditinggalkan pengguna pada Google Play Store menyajikan sumber data penting dalam mengukur tingkat kepuasan terhadap layanan aplikasi DANA. Kajian ini diangkat guna menguji dan membandingkan performa Gaussian Naïve Bayes, Support Vector Machine (SVM), dan Random Forest untuk analisis sentimen berpendekatan TF-IDF. Proses pengumpulan data berhasil menjaring 25.000 ulasan via teknik scraping. Keseluruhan data tersebut diproses melalui preprocessing, pembobotan teks TF-IDF, serta pembagian rasio data latih-uji sebesar 90:10 sebelum diklasifikasikan. Penilaian presisi model diukur berdasar Accuracy, Precision, Recall, F1-Score, Confusion Matrix, hingga Area Under Curve (AUC). Hasil pengujian mengindikasikan bahwa SVM menghasilkan capaian tertinggi (accuracy 89,00% dan AUC 0,938), dilanjutkan oleh Random Forest (accuracy 87,98% dan AUC 0,934), serta Gaussian Naïve Bayes (accuracy 85,77% dan AUC 0,770). Rangkaian temuan ini mempertegas posisi SVM sebagai pendekatan yang paling akurat dan efektif dalam mengklasifikasikan sentimen pengguna DANA pada eksperimen ini. Kata kunci: Analisis Sentimen, DANA, TF-IDF, Gaussian Naïve Bayes, Support Vector Machine, Random Forest.
Actions (login required)
![]() |
View Item |
