A Review of Recent Advances on Deep Learning Methods for Audio-Visual Speech Recognition

This article provides a detailed review of recent advances in audio-visual speech recognition (AVSR) methods that have been developed over the last decade (2013–2023). Despite the recent success of audio speech recognition systems, the problem of audio-visual (AV) speech decoding remains challenging...

Full description

Bibliographic Details
Main Authors:	Denis Ivanko, Dmitry Ryumin, Alexey Karpov
Format:	Article
Language:	English
Published:	MDPI AG 2023-06-01
Series:	Mathematics
Subjects:	audio-visual speech recognition lip-reading audio-visual fusion computer vision deep learning review
Online Access:	https://www.mdpi.com/2227-7390/11/12/2665

Description
Summary:	This article provides a detailed review of recent advances in audio-visual speech recognition (AVSR) methods that have been developed over the last decade (2013–2023). Despite the recent success of audio speech recognition systems, the problem of audio-visual (AV) speech decoding remains challenging. In comparison to the previous surveys, we mainly focus on the important progress brought with the introduction of deep learning (DL) to the field and skip the description of long-known traditional “hand-crafted” methods. In addition, we also discuss the recent application of DL toward AV speech fusion and recognition. We first discuss the main AV datasets used in the literature for AVSR experiments since we consider it a data-driven machine learning (ML) task. We then consider the methodology used for visual speech recognition (VSR). Subsequently, we also consider recent AV methodology advances. We then separately discuss the evolution of the core AVSR methods, pre-processing and augmentation techniques, and modality fusion strategies. We conclude the article with a discussion on the current state of AVSR and provide our vision for future research.
ISSN:	2227-7390

A Review of Recent Advances on Deep Learning Methods for Audio-Visual Speech Recognition

Similar Items