PAMUNGKAS, MAHESA HARYO (2026) PERBANDINGAN METODE PELABELAN OTOMATIS TERHADAP KINERJA SUPPORT VECTOR MACHINE DALAM ANALISIS SENTIMEN KOMENTAR POLITIK DI MEDIA SOSIAL X. S1 thesis, Universitas Mercu Buana Jakarta.
|
Text (HAL COVER)
Cover.pdf Download (706kB) | Preview |
|
|
Text (BAB I)
Bab 1.pdf Restricted to Registered users only Download (40kB) |
||
|
Text (BAB II)
Bab 2.pdf Restricted to Registered users only Download (240kB) |
||
|
Text (BAB III)
Bab 3.pdf Restricted to Registered users only Download (67kB) |
||
|
Text (BAB IV)
Bab 4.pdf Restricted to Registered users only Download (4MB) |
||
|
Text (BAB V)
Bab 5.pdf Restricted to Registered users only Download (33kB) |
||
|
Text (DAFTAR PUSTAKA)
Daftar Pustaka.pdf Restricted to Registered users only Download (92kB) |
||
|
Text (LAMPIRAN)
Lampiran.pdf Restricted to Registered users only Download (360kB) |
Abstract
X is one of the primary social media platforms where users express opinions about political figures, generating a large volume of textual data. This condition makes manual annotation inefficient and encourages the use of automatic labeling methods. This study compares the performance of three automatic labeling approaches, namely InSet Lexicon, IndoBERT, and RoBERTa, as training label sources for a Support Vector Machine (SVM) model in sentiment analysis of comments posted on Joko Widodo's X account. A total of 62,614 comments were collected using the GraphQL API, resulting in 48,546 preprocessed comments. Among them, 8,790 comments were manually annotated by three annotators and used as the ground truth. The reliability of the manual annotations was assessed using Fleiss' Kappa, yielding a value of 0.7182, which indicates Substantial Agreement. Model performance was evaluated using Accuracy, Macro F1-score, Weighted F1-score, and McNemar's Test. An additional experiment employing RandomOverSampler was conducted to examine the effect of class imbalance on SVM performance. The results show that the SVM trained with RoBERTagenerated labels achieved the best performance on the original dataset, with an Accuracy of 68.98%, Macro F1-score of 55.31%, and Weighted F1-score of 69.61%. Furthermore, McNemar's Test indicated statistically significant differences among the three labeling methods (p-value < 0.05). The RandomOverSampler experiment also demonstrated that class distribution influences SVM performance by affecting the evaluation metrics of each labeling method. These findings suggest that both automatic labeling strategies and class distribution play important roles in sentiment classification using SVM. Keywords : Sentiment Analysis, Automatic Labeling, Support Vector Machine, Transformer-based Labeling, InSet Lexicon. Media sosial X menjadi salah satu platform utama bagi masyarakat untuk menyampaikan opini terhadap tokoh politik sehingga menghasilkan data komentar dalam jumlah besar. Kondisi tersebut menyebabkan pelabelan manual menjadi kurang efisien dan mendorong penggunaan metode pelabelan otomatis. Penelitian ini bertujuan membandingkan performa tiga metode pelabelan otomatis, yaitu InSet Lexicon, IndoBERT, dan RoBERTa, sebagai sumber label pelatihan Support Vector Machine (SVM) pada analisis sentimen komentar di akun X Joko Widodo. Data diperoleh melalui proses crawling menggunakan GraphQL API sebanyak 62.614 komentar dan menghasilkan 48.546 data setelah tahap preprocessing. Sebanyak 8.790 komentar digunakan sebagai ground truth yang dianotasi oleh tiga annotator dan divalidasi menggunakan Fleiss' Kappa, dengan nilai 0,7182 (Substantial Agreement). Evaluasi dilakukan menggunakan Accuracy, Macro F1- score, Weighted F1-score, serta McNemar Test, kemudian dilengkapi dengan eksperimen RandomOverSampler untuk menganalisis pengaruh ketidakseimbangan kelas terhadap performa model. Hasil penelitian menunjukkan bahwa SVM yang dilatih menggunakan label RoBERTa memberikan performa terbaik pada dataset asli dengan Accuracy 68,98%, Macro F1-score 55,31%, dan Weighted F1-score 69,61%. Hasil McNemar Test menunjukkan seluruh perbedaan performa antar model signifikan secara statistik (p-value < 0,05). Sementara itu, eksperimen RandomOverSampler menunjukkan bahwa distribusi kelas memengaruhi performa SVM, di mana penyeimbangan data mengubah nilai metrik evaluasi pada setiap metode pelabelan. Temuan ini menunjukkan bahwa metode pelabelan otomatis dan distribusi kelas berperan penting dalam membentuk performa klasifikasi sentimen menggunakan SVM. Kata kunci: Analisis Sentimen, Pelabelan Otomatis, Support Vector Machine, Transformer-based Labeling, InSet Lexicon.
Actions (login required)
![]() |
View Item |
