PERBANDINGAN KINERJA MODEL 3D-CNN, MOBILENETV2, DAN SWIN TRANSFORMER UNTUK KLASIFIKASI LIP READING

JUNIOR, ALBERT CHANDRA (2026) PERBANDINGAN KINERJA MODEL 3D-CNN, MOBILENETV2, DAN SWIN TRANSFORMER UNTUK KLASIFIKASI LIP READING. S1 thesis, Universitas Mercu Buana Jakarta.

[img]
Preview
Text (HAL COVER)
01 Cover.pdf

Download (617kB) | Preview
[img] Text (BAB I)
02 Bab 1.pdf
Restricted to Registered users only

Download (35kB)
[img] Text (BAB II)
03 Bab 2.pdf
Restricted to Registered users only

Download (99kB)
[img] Text (BAB III)
04 Bab 3.pdf
Restricted to Registered users only

Download (378kB)
[img] Text (BAB IV)
05 Bab 4.pdf
Restricted to Registered users only

Download (233kB)
[img] Text (BAB V)
06 Bab 5.pdf
Restricted to Registered users only

Download (28kB)
[img] Text (DAFTAR PUSTAKA)
07 Daftar Pustaka.pdf
Restricted to Registered users only

Download (99kB)
[img] Text (LAMPIRAN)
08 Lampiran.pdf
Restricted to Registered users only

Download (1MB)

Abstract

Visual Speech Recognition (VSR), or lip reading, is an important technology for supporting communication in acoustically noisy environments and for people with hearing impairments, although it still faces challenges such as visual ambiguity (homophenes) and variation in lip movement across individuals. This study compares the performance of three Deep Learning architectures, namely a 3D Convolutional Neural Network (3D-CNN) trained from scratch, MobileNetV2 with a transfer learning approach, and a Swin Transformer combined with a Temporal Transformer Encoder, on a 10-class lip reading classification task using the MIRACL-VC1 dataset. Each video is processed through face detection and lip region-of-interest (ROI) cropping using 68 facial landmarks, before being trained with a single-split scheme for 15 epochs. The three models are evaluated using accuracy, precision, recall, F1-score, confusion matrix, and inference speed (Frames Per Second/FPS) to assess real-time feasibility. The results show that the Swin Transformer achieves the highest test accuracy of 68.50% and a macro F1- score of 0.6791, followed by MobileNetV2 (46.30%; 0.4536) and 3D-CNN (35.31%; 0.3253). However, the Swin Transformer requires the longest training time and reaches only 3.60 FPS during inference, failing to meet the real-time criterion, unlike 3D-CNN (53.33 FPS) and MobileNetV2 (45.91 FPS), which remain real-time capable. These findings indicate a trade-off between classification accuracy and computational efficiency among the three tested architectures Keywords: Visual Speech Recognition (VSR), Lip Reading, 3D Convolutional Neural Network (3D-CNN), MobileNetV2, Swin Transformer. Visual Speech Recognition (VSR) atau lip reading merupakan teknologi penting untuk mendukung komunikasi pada lingkungan dengan gangguan akustik tinggi maupun bagi penyandang tunarungu, namun menghadapi tantangan berupa ambiguitas visual (homophenes) serta variasi karakteristik gerak bibir antarindividu. Penelitian ini membandingkan kinerja tiga arsitektur Deep Learning, yaitu 3D Convolutional Neural Network (3D-CNN) yang dilatih dari awal (from scratch), MobileNetV2 dengan pendekatan transfer learning, dan Swin Transformer yang dipadukan dengan Temporal Transformer Encoder, dalam menyelesaikan tugas klasifikasi lip reading ke dalam 10 kelas kata pada Dataset MIRACL-VC1. Setiap video diproses melalui deteksi wajah dan pemotongan region of interest (ROI) pada area bibir menggunakan 68 facial landmarks, kemudian dilatih menggunakan skema single split selama 15 epoch. Kinerja ketiga model dievaluasi menggunakan metrik akurasi, presisi, recall, F1-score, confusion matrix, serta kecepatan inferensi (Frames Per Second/FPS) untuk menilai kelayakan real-time. Hasil pengujian menunjukkan bahwa Swin Transformer memberikan akurasi uji tertinggi sebesar 68,50% dan F1-score makro 0,6791, diikuti oleh MobileNetV2 (46,30%; 0,4536) dan 3D-CNN (35,31%; 0,3253); namun Swin Transformer memerlukan waktu pelatihan paling lama dan hanya mencapai 3,60 FPS saat inferensi sehingga belum memenuhi kriteria real-time, berbeda dengan 3D-CNN (53,33 FPS) dan MobileNetV2 (45,91 FPS) yang tetap memenuhi kriteria tersebut. Temuan ini menunjukkan adanya trade-off antara akurasi klasifikasi dan efisiensi komputasi pada ketiga arsitektur yang diuji. Kata kunci: Visual Speech Recognition (VSR), Lip Reading, 3D Convolutional Neural Network (3D-CNN), MobileNetV2, Swin Transformer.

Item Type: Thesis (S1)
NIM/NIDN Creators: 41522110044
Uncontrolled Keywords: Visual Speech Recognition (VSR), Lip Reading, 3D Convolutional Neural Network (3D-CNN), MobileNetV2, Swin Transformer.
Subjects: 000 Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 000. Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 004 Data Processing, Computer Science/Pemrosesan Data, Ilmu Komputer, Teknik Informatika
000 Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 000. Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 006 Special Computer Methods/Metode Komputer Tertentu > 006.3 Artificial Intelligence/Kecerdasan Buatan > 006.32 Neural Nets (Neural Network)/Jaringan Saraf Buatan
000 Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 000. Computer Science, Information and General Works/Ilmu Komputer, Informasi, dan Karya Umum > 006 Special Computer Methods/Metode Komputer Tertentu > 006.4 Computer Pattern Recognition/Pola Pengenalan Komputer > 006.45 Acoustical Pattern Recognition/Pengenalan Pola Akustik > 006.454 Speech Recognition/Pengenalan Suara
Divisions: Fakultas Ilmu Komputer > Informatika
Depositing User: khalimah
Date Deposited: 14 Sep 2026 09:42
Last Modified: 14 Sep 2026 09:42
URI: http://repository.mercubuana.ac.id/id/eprint/103865

Actions (login required)

View Item View Item