IMPLEMENTASI ALGORITMA RANDOM FOREST UNTUK KLASIFIKASI RISIKO STUNTING: ANALISIS FEATURE IMPORTANCE DAN STUDI KOMPARATIF DENGAN SVM

UMAROH, CHOIRUNNISA (2026) IMPLEMENTASI ALGORITMA RANDOM FOREST UNTUK KLASIFIKASI RISIKO STUNTING: ANALISIS FEATURE IMPORTANCE DAN STUDI KOMPARATIF DENGAN SVM. S1 thesis, Universitas Mercu Buana Jakarta.

[img]
Preview
Text (HAL COVER)
Cover.pdf

Download (472kB) | Preview
[img] Text (BAB I)
Bab 1.pdf
Restricted to Registered users only

Download (42kB)
[img] Text (BAB II)
Bab 2.pdf
Restricted to Registered users only

Download (307kB)
[img] Text (BAB III)
Bab 3.pdf
Restricted to Registered users only

Download (312kB)
[img] Text (BAB IV)
Bab 4.pdf
Restricted to Registered users only

Download (397kB)
[img] Text (BAB V)
Bab 5.pdf
Restricted to Registered users only

Download (82kB)
[img] Text (DAFTAR PUSTAKA)
Daftar Pustaka.pdf
Restricted to Registered users only

Download (131kB)
[img] Text (LAMPIRAN)
Lampiran.pdf
Restricted to Registered users only

Download (354kB)

Abstract

Stunting is one of the chronic nutritional problems among toddlers in Indonesia that affects long-term growth and development. Early detection of stunting risk is essential to enable timely nutritional intervention. This study aims to implement the Random Forest algorithm to classify stunting risk in toddlers based on secondary data from Posyandu (integrated health post) records in Kabupaten Tangerang, evaluate its performance, identify the most dominant risk factors through feature importance analysis, and compare its performance with the Support Vector Machine (SVM) algorithm to determine the best model for integration into an interactive dashboard. The study used 4,645 toddler records selected from an initial dataset of 4,839 rows, with 18 input features resulting from feature engineering. Class imbalance was handled using SMOTETomek, while optimal hyperparameter search was conducted using RandomizedSearchCV and GridSearchCV with a 10-fold cross-validation scheme. The results show that the Random Forest model achieved an accuracy of 99%, recall of 92.16% (meeting the clinical target of ≥90%), precision of 83.93%, F1-Score of 0.8785, and ROC-AUC of 0.9934 at the default threshold. Feature importance analysis identified Tinggi_Threshold_Gap (height-threshold gap) and height as the most dominant risk factors. Compared to SVM, Random Forest was superior in discriminative ability, clinical sensitivity, and model interpretability, and was therefore selected as the primary model to be integrated into a web-based stunting risk detection dashboard as a decision-support tool for Posyandu. Keywords: Stunting, Random Forest, SMOTETomek, Feature Importance, Imbalanced Data Stunting merupakan salah satu masalah gizi kronis pada balita di Indonesia yang berdampak pada gangguan pertumbuhan dan perkembangan jangka panjang. Deteksi dini risiko stunting sangat diperlukan agar intervensi gizi dapat dilakukan secara tepat waktu. Penelitian ini bertujuan mengimplementasikan algoritma Random Forest untuk mengklasifikasikan risiko stunting pada balita berdasarkan data sekunder Posyandu Kabupaten Tangerang, mengevaluasi kinerjanya, mengidentifikasi faktor risiko yang paling dominan melalui analisis feature importance, serta membandingkan performanya dengan algoritma Support Vector Machine (SVM) untuk menentukan model terbaik yang akan diintegrasikan ke dalam dashboard interaktif. Penelitian menggunakan 4.645 data balita hasil seleksi dari 4.839 baris data awal, dengan 18 fitur input hasil feature engineering. Ketidakseimbangan kelas ditangani menggunakan SMOTETomek, sedangkan pencarian hyperparameter optimal dilakukan melalui RandomizedSearchCV dan GridSearchCV dengan skema 10-fold cross-validation. Hasil penelitian menunjukkan model Random Forest mencapai akurasi 99%, recall 92,16% (memenuhi target klinis ≥90%), precision 83,93%, F1-Score 0,8785, dan ROC-AUC 0,9934 pada threshold default. Analisis feature importance mengidentifikasi Tinggi_Threshold_Gap dan tinggi badan sebagai faktor risiko paling dominan. Dibandingkan dengan SVM, Random Forest unggul dalam kemampuan diskriminasi, sensitivitas klinis, dan interpretabilitas model, sehingga dipilih sebagai model utama untuk diintegrasikan ke dalam dashboard deteksi risiko stunting berbasis web sebagai sarana pendukung keputusan di Posyandu. Kata kunci: Stunting, Random Forest, SMOTETomek, Feature Importance, Imbalanced Data

Item Type: Thesis (S1)
NIM/NIDN Creators: 41822010014
Uncontrolled Keywords: Stunting, Random Forest, SMOTETomek, Feature Importance, Imbalanced Data
Subjects: 500 Natural Science and Mathematics/Ilmu-ilmu Alam dan Matematika > 510 Mathematics/Matematika > 518 Numerical Analysis/Analisis Numerik, Analisa Numerik > 518.1 Algorithms/Algoritma
600 Technology/Teknologi > 650 Management, Public Relations, Business and Auxiliary Service/Manajemen, Hubungan Masyarakat, Bisnis dan Ilmu yang Berkaitan > 658 General Management/Manajemen Umum > 658.01-658.09 [Management of Enterprises of Specific Sizes, Scopes, Forms; Data Processing]/[Pengelolaan Usaha dengan Ukuran, Lingkup, Bentuk Tertentu; Pengolahan Data] > 658.05 Data Processing Computer Applications/Pengolahan Data Aplikasi Komputer
Divisions: Fakultas Ilmu Komputer > Sistem Informasi
Depositing User: khalimah
Date Deposited: 01 Sep 2026 02:44
Last Modified: 01 Sep 2026 02:44
URI: http://repository.mercubuana.ac.id/id/eprint/103526

Actions (login required)

View Item View Item