JUNIOR, ALBERT CHANDRA (2026) PERBANDINGAN KINERJA MODEL 3D-CNN, MOBILENETV2, DAN SWIN TRANSFORMER UNTUK KLASIFIKASI LIP READING. S1 thesis, Universitas Mercu Buana Jakarta.
|
Text (HAL COVER)
01 Cover.pdf Download (617kB) | Preview |
|
|
Text (BAB I)
02 Bab 1.pdf Restricted to Registered users only Download (35kB) |
||
|
Text (BAB II)
03 Bab 2.pdf Restricted to Registered users only Download (99kB) |
||
|
Text (BAB III)
04 Bab 3.pdf Restricted to Registered users only Download (378kB) |
||
|
Text (BAB IV)
05 Bab 4.pdf Restricted to Registered users only Download (233kB) |
||
|
Text (BAB V)
06 Bab 5.pdf Restricted to Registered users only Download (28kB) |
||
|
Text (DAFTAR PUSTAKA)
07 Daftar Pustaka.pdf Restricted to Registered users only Download (99kB) |
||
|
Text (LAMPIRAN)
08 Lampiran.pdf Restricted to Registered users only Download (1MB) |
Abstract
Visual Speech Recognition (VSR), or lip reading, is an important technology for supporting communication in acoustically noisy environments and for people with hearing impairments, although it still faces challenges such as visual ambiguity (homophenes) and variation in lip movement across individuals. This study compares the performance of three Deep Learning architectures, namely a 3D Convolutional Neural Network (3D-CNN) trained from scratch, MobileNetV2 with a transfer learning approach, and a Swin Transformer combined with a Temporal Transformer Encoder, on a 10-class lip reading classification task using the MIRACL-VC1 dataset. Each video is processed through face detection and lip region-of-interest (ROI) cropping using 68 facial landmarks, before being trained with a single-split scheme for 15 epochs. The three models are evaluated using accuracy, precision, recall, F1-score, confusion matrix, and inference speed (Frames Per Second/FPS) to assess real-time feasibility. The results show that the Swin Transformer achieves the highest test accuracy of 68.50% and a macro F1- score of 0.6791, followed by MobileNetV2 (46.30%; 0.4536) and 3D-CNN (35.31%; 0.3253). However, the Swin Transformer requires the longest training time and reaches only 3.60 FPS during inference, failing to meet the real-time criterion, unlike 3D-CNN (53.33 FPS) and MobileNetV2 (45.91 FPS), which remain real-time capable. These findings indicate a trade-off between classification accuracy and computational efficiency among the three tested architectures Keywords: Visual Speech Recognition (VSR), Lip Reading, 3D Convolutional Neural Network (3D-CNN), MobileNetV2, Swin Transformer. Visual Speech Recognition (VSR) atau lip reading merupakan teknologi penting untuk mendukung komunikasi pada lingkungan dengan gangguan akustik tinggi maupun bagi penyandang tunarungu, namun menghadapi tantangan berupa ambiguitas visual (homophenes) serta variasi karakteristik gerak bibir antarindividu. Penelitian ini membandingkan kinerja tiga arsitektur Deep Learning, yaitu 3D Convolutional Neural Network (3D-CNN) yang dilatih dari awal (from scratch), MobileNetV2 dengan pendekatan transfer learning, dan Swin Transformer yang dipadukan dengan Temporal Transformer Encoder, dalam menyelesaikan tugas klasifikasi lip reading ke dalam 10 kelas kata pada Dataset MIRACL-VC1. Setiap video diproses melalui deteksi wajah dan pemotongan region of interest (ROI) pada area bibir menggunakan 68 facial landmarks, kemudian dilatih menggunakan skema single split selama 15 epoch. Kinerja ketiga model dievaluasi menggunakan metrik akurasi, presisi, recall, F1-score, confusion matrix, serta kecepatan inferensi (Frames Per Second/FPS) untuk menilai kelayakan real-time. Hasil pengujian menunjukkan bahwa Swin Transformer memberikan akurasi uji tertinggi sebesar 68,50% dan F1-score makro 0,6791, diikuti oleh MobileNetV2 (46,30%; 0,4536) dan 3D-CNN (35,31%; 0,3253); namun Swin Transformer memerlukan waktu pelatihan paling lama dan hanya mencapai 3,60 FPS saat inferensi sehingga belum memenuhi kriteria real-time, berbeda dengan 3D-CNN (53,33 FPS) dan MobileNetV2 (45,91 FPS) yang tetap memenuhi kriteria tersebut. Temuan ini menunjukkan adanya trade-off antara akurasi klasifikasi dan efisiensi komputasi pada ketiga arsitektur yang diuji. Kata kunci: Visual Speech Recognition (VSR), Lip Reading, 3D Convolutional Neural Network (3D-CNN), MobileNetV2, Swin Transformer.
Actions (login required)
![]() |
View Item |
