IKHSANUDIN, MUHAMMAD (2026) PERBANDINGAN KINERJA ALGORITMA MACHINE LEARNING DALAM DETEKSI AKUN INSTAGRAM PALSU. S1 thesis, Universitas Mercu Buana Jakarta.
|
Text (HAL COVER)
Cover.pdf Download (575kB) | Preview |
|
|
Text (BAB I)
BAB I.pdf Restricted to Registered users only Download (41kB) |
||
|
Text (BAB II)
BAB II.pdf Restricted to Registered users only Download (121kB) |
||
|
Text (BAB III)
BAB III.pdf Restricted to Registered users only Download (280kB) |
||
|
Text (BAB IV)
BAB IV.pdf Restricted to Registered users only Download (250kB) |
||
|
Text (BAB V)
BAB V.pdf Restricted to Registered users only Download (37kB) |
||
|
Text (DAFTAR PUSTAKA)
Daftar Pustaka.pdf Restricted to Registered users only Download (94kB) |
||
|
Text (LAMPIRAN)
Lampiran.pdf Restricted to Registered users only Download (413kB) |
Abstract
The presence of fake accounts on Instagram can reduce the quality of interactions, undermine user trust, and facilitate manipulative activities. This study aims to develop and compare the performance of Logistic Regression, Random Forest, and XGBoost in detecting fake Instagram accounts based on numerical profile features. The public dataset initially contained 2,594 records. After removing 35 duplicate records, 2,559 records remained, consisting of 1,660 real accounts and 899 fake accounts. The selected features included follower count, following count, biography length, media count, and the number of digits in the username. The experiments were conducted using three data-splitting scenarios: 60:40, 70:30, and 80:20. Class imbalance was addressed using the Synthetic Minority Over-sampling Technique (SMOTE), which was applied exclusively to the training data. The models were evaluated using accuracy, precision, recall, and F1-score on the internal validation data and an independently prepared test dataset containing 150 accounts, consisting of 75 real accounts and 75 fake accounts. The internal validation results showed that XGBoost with the 80:20 split achieved the highest performance, with an accuracy of 0.9570 and an F1-score of 0.9382. However, on the independent test dataset, Logistic Regression with the 60:40 split produced the best results, achieving an accuracy of 0.8867, precision of 0.9833, recall of 0.7867, and an F1-score of 0.8741. These findings indicate that the model with the highest internal validation performance does not necessarily provide the best generalization to new data. Based on the independent test evaluation, Logistic Regression was the most suitable model in this study. Keyword: fake Instagram accounts, Logistic Regression, Random Forest, XGBoost, SMOTE. Keberadaan akun palsu di Instagram dapat menurunkan kualitas interaksi, mengganggu kepercayaan pengguna, serta dimanfaatkan untuk aktivitas manipulatif. Penelitian ini bertujuan membangun dan membandingkan kinerja algoritma Logistic Regression, Random Forest, dan XGBoost dalam mendeteksi akun Instagram palsu berdasarkan fitur numerik profil akun. Data publik yang digunakan pada awalnya berjumlah 2.594 data. Setelah 35 data duplikat dihapus, diperoleh 2.559 data yang terdiri atas 1.660 akun real dan 899 akun fake. Fitur yang digunakan meliputi jumlah pengikut, jumlah akun yang diikuti, panjang biografi, jumlah unggahan, dan jumlah digit pada username. Eksperimen dilakukan menggunakan skenario pembagian data 60:40, 70:30, dan 80:20. Ketidakseimbangan kelas ditangani menggunakan Synthetic Minority Over�sampling Technique (SMOTE) yang hanya diterapkan pada data latih. Model dievaluasi menggunakan accuracy, precision, recall, dan F1-score pada data validasi internal serta 150 data uji mandiri yang terdiri atas 75 akun real dan 75 akun fake. Hasil validasi internal menunjukkan bahwa XGBoost pada rasio 80:20 memperoleh performa tertinggi dengan accuracy sebesar 0,9570 dan F1-score sebesar 0,9382. Namun, pada data uji mandiri, Logistic Regression dengan rasio 60:40 memberikan hasil terbaik dengan accuracy sebesar 0,8867, precision sebesar 0,9833, recall sebesar 0,7867, dan F1-score sebesar 0,8741. Hasil tersebut menunjukkan bahwa model dengan performa validasi internal tertinggi belum tentu memiliki kemampuan generalisasi terbaik pada data baru. Berdasarkan evaluasi data uji mandiri, Logistic Regression menjadi model yang paling sesuai dalam penelitian ini. Kata kunci: akun Instagram palsu, Logistic Regression, Random Forest, XGBoost, SMOTE.
Actions (login required)
![]() |
View Item |
