PERBANDINGAN METODE PELABELAN OTOMATIS TERHADAP KINERJA SUPPORT VECTOR MACHINE DALAM ANALISIS SENTIMEN KOMENTAR POLITIK DI MEDIA SOSIAL X

PAMUNGKAS, MAHESA HARYO (2026) PERBANDINGAN METODE PELABELAN OTOMATIS TERHADAP KINERJA SUPPORT VECTOR MACHINE DALAM ANALISIS SENTIMEN KOMENTAR POLITIK DI MEDIA SOSIAL X. S1 thesis, Universitas Mercu Buana Jakarta.

[img]
Preview
Text (HAL COVER)
Cover.pdf

Download (706kB) | Preview
[img] Text (BAB I)
Bab 1.pdf
Restricted to Registered users only

Download (40kB)
[img] Text (BAB II)
Bab 2.pdf
Restricted to Registered users only

Download (240kB)
[img] Text (BAB III)
Bab 3.pdf
Restricted to Registered users only

Download (67kB)
[img] Text (BAB IV)
Bab 4.pdf
Restricted to Registered users only

Download (4MB)
[img] Text (BAB V)
Bab 5.pdf
Restricted to Registered users only

Download (33kB)
[img] Text (DAFTAR PUSTAKA)
Daftar Pustaka.pdf
Restricted to Registered users only

Download (92kB)
[img] Text (LAMPIRAN)
Lampiran.pdf
Restricted to Registered users only

Download (360kB)

Abstract

X is one of the primary social media platforms where users express opinions about political figures, generating a large volume of textual data. This condition makes manual annotation inefficient and encourages the use of automatic labeling methods. This study compares the performance of three automatic labeling approaches, namely InSet Lexicon, IndoBERT, and RoBERTa, as training label sources for a Support Vector Machine (SVM) model in sentiment analysis of comments posted on Joko Widodo's X account. A total of 62,614 comments were collected using the GraphQL API, resulting in 48,546 preprocessed comments. Among them, 8,790 comments were manually annotated by three annotators and used as the ground truth. The reliability of the manual annotations was assessed using Fleiss' Kappa, yielding a value of 0.7182, which indicates Substantial Agreement. Model performance was evaluated using Accuracy, Macro F1-score, Weighted F1-score, and McNemar's Test. An additional experiment employing RandomOverSampler was conducted to examine the effect of class imbalance on SVM performance. The results show that the SVM trained with RoBERTagenerated labels achieved the best performance on the original dataset, with an Accuracy of 68.98%, Macro F1-score of 55.31%, and Weighted F1-score of 69.61%. Furthermore, McNemar's Test indicated statistically significant differences among the three labeling methods (p-value < 0.05). The RandomOverSampler experiment also demonstrated that class distribution influences SVM performance by affecting the evaluation metrics of each labeling method. These findings suggest that both automatic labeling strategies and class distribution play important roles in sentiment classification using SVM. Keywords : Sentiment Analysis, Automatic Labeling, Support Vector Machine, Transformer-based Labeling, InSet Lexicon. Media sosial X menjadi salah satu platform utama bagi masyarakat untuk menyampaikan opini terhadap tokoh politik sehingga menghasilkan data komentar dalam jumlah besar. Kondisi tersebut menyebabkan pelabelan manual menjadi kurang efisien dan mendorong penggunaan metode pelabelan otomatis. Penelitian ini bertujuan membandingkan performa tiga metode pelabelan otomatis, yaitu InSet Lexicon, IndoBERT, dan RoBERTa, sebagai sumber label pelatihan Support Vector Machine (SVM) pada analisis sentimen komentar di akun X Joko Widodo. Data diperoleh melalui proses crawling menggunakan GraphQL API sebanyak 62.614 komentar dan menghasilkan 48.546 data setelah tahap preprocessing. Sebanyak 8.790 komentar digunakan sebagai ground truth yang dianotasi oleh tiga annotator dan divalidasi menggunakan Fleiss' Kappa, dengan nilai 0,7182 (Substantial Agreement). Evaluasi dilakukan menggunakan Accuracy, Macro F1- score, Weighted F1-score, serta McNemar Test, kemudian dilengkapi dengan eksperimen RandomOverSampler untuk menganalisis pengaruh ketidakseimbangan kelas terhadap performa model. Hasil penelitian menunjukkan bahwa SVM yang dilatih menggunakan label RoBERTa memberikan performa terbaik pada dataset asli dengan Accuracy 68,98%, Macro F1-score 55,31%, dan Weighted F1-score 69,61%. Hasil McNemar Test menunjukkan seluruh perbedaan performa antar model signifikan secara statistik (p-value < 0,05). Sementara itu, eksperimen RandomOverSampler menunjukkan bahwa distribusi kelas memengaruhi performa SVM, di mana penyeimbangan data mengubah nilai metrik evaluasi pada setiap metode pelabelan. Temuan ini menunjukkan bahwa metode pelabelan otomatis dan distribusi kelas berperan penting dalam membentuk performa klasifikasi sentimen menggunakan SVM. Kata kunci: Analisis Sentimen, Pelabelan Otomatis, Support Vector Machine, Transformer-based Labeling, InSet Lexicon.

Item Type: Thesis (S1)
NIM/NIDN Creators: 41522010017
Uncontrolled Keywords: Analisis Sentimen, Pelabelan Otomatis, Support Vector Machine, Transformer-based Labeling, InSet Lexicon.
Subjects: 000 Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 000. Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 004 Data Processing, Computer Science/Pemrosesan Data, Ilmu Komputer, Teknik Informatika
000 Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 000. Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 006 Special Computer Methods/Metode Komputer Tertentu > 006.7 Multimedia Systems/Sistem-sistem Multimedia > 006.75 Social Multimedia/Multimedia Social
500 Natural Science and Mathematics/Ilmu-ilmu Alam dan Matematika > 510 Mathematics/Matematika > 518 Numerical Analysis/Analisis Numerik, Analisa Numerik > 518.1 Algorithms/Algoritma
Divisions: Fakultas Ilmu Komputer > Informatika
Depositing User: khalimah
Date Deposited: 26 Aug 2026 04:24
Last Modified: 26 Aug 2026 04:24
URI: http://repository.mercubuana.ac.id/id/eprint/103410

Actions (login required)

View Item View Item