Exploring deep learning approaches for video captioning: A comprehensive review

While humans can easily describe visual data at varying levels of detail, the same task presents a significant challenge for machines. This challenge becomes even more complex when dealing with video data. The process of understanding a video and generating descriptive text for it is known as video...

Full description

Bibliographic Details
Main Authors:	Adel Jalal Yousif, Mohammed H. Al-Jammas
Format:	Article
Language:	English
Published:	Elsevier 2023-12-01
Series:	e-Prime: Advances in Electrical Engineering, Electronics and Energy
Subjects:	Evaluation metrics Video captioning Video description Computer vision Deep learning
Online Access:	http://www.sciencedirect.com/science/article/pii/S277267112300267X

_version_	1797388591301656576
author	Adel Jalal Yousif Mohammed H. Al-Jammas
author_facet	Adel Jalal Yousif Mohammed H. Al-Jammas
author_sort	Adel Jalal Yousif
collection	DOAJ
description	While humans can easily describe visual data at varying levels of detail, the same task presents a significant challenge for machines. This challenge becomes even more complex when dealing with video data. The process of understanding a video and generating descriptive text for it is known as video captioning. Video captioning requires not only understanding the visual content but also producing human-like descriptions that accurately capture its semantics. Achieving this level of understanding requires the collaborative efforts of both the computer vision and natural language processing research communities. The captions produced through video captioning serve as valuable resources that can be further leveraged for various applications such as video search, accessibility for visually impaired people, and human-robot interaction. Deep learning strategies have emerged as powerful tools in addressing the complexities of video captioning. By leveraging large scale annotated video caption datasets and sophisticated neural network architectures, deep learning approaches have made significant advances in this challenging task. In the existing literature, numerous techniques, benchmark datasets, and evaluation metrics have been developed, emphasizing the necessity for a comprehensive examination to concentrate research efforts in this rapidly evolving field. This paper provides a survey of deep learning based methods for video captioning, highlighting their key components, challenges, and recent advancements.
first_indexed	2024-03-08T22:42:58Z
format	Article
id	doaj.art-13deabbc7cab49c08a5b532d65f6cf70
institution	Directory Open Access Journal
issn	2772-6711
language	English
last_indexed	2024-03-08T22:42:58Z
publishDate	2023-12-01
publisher	Elsevier
record_format	Article
series	e-Prime: Advances in Electrical Engineering, Electronics and Energy
spelling	doaj.art-13deabbc7cab49c08a5b532d65f6cf702023-12-17T06:43:35ZengElseviere-Prime: Advances in Electrical Engineering, Electronics and Energy2772-67112023-12-016100372Exploring deep learning approaches for video captioning: A comprehensive reviewAdel Jalal Yousif0Mohammed H. Al-Jammas1University of Mosul, Mosul, Iraq; Corresponding author.Ninevah University, Mosul, IraqWhile humans can easily describe visual data at varying levels of detail, the same task presents a significant challenge for machines. This challenge becomes even more complex when dealing with video data. The process of understanding a video and generating descriptive text for it is known as video captioning. Video captioning requires not only understanding the visual content but also producing human-like descriptions that accurately capture its semantics. Achieving this level of understanding requires the collaborative efforts of both the computer vision and natural language processing research communities. The captions produced through video captioning serve as valuable resources that can be further leveraged for various applications such as video search, accessibility for visually impaired people, and human-robot interaction. Deep learning strategies have emerged as powerful tools in addressing the complexities of video captioning. By leveraging large scale annotated video caption datasets and sophisticated neural network architectures, deep learning approaches have made significant advances in this challenging task. In the existing literature, numerous techniques, benchmark datasets, and evaluation metrics have been developed, emphasizing the necessity for a comprehensive examination to concentrate research efforts in this rapidly evolving field. This paper provides a survey of deep learning based methods for video captioning, highlighting their key components, challenges, and recent advancements.http://www.sciencedirect.com/science/article/pii/S277267112300267XEvaluation metricsVideo captioningVideo descriptionComputer visionDeep learning
spellingShingle	Adel Jalal Yousif Mohammed H. Al-Jammas Exploring deep learning approaches for video captioning: A comprehensive review e-Prime: Advances in Electrical Engineering, Electronics and Energy Evaluation metrics Video captioning Video description Computer vision Deep learning
title	Exploring deep learning approaches for video captioning: A comprehensive review
title_full	Exploring deep learning approaches for video captioning: A comprehensive review
title_fullStr	Exploring deep learning approaches for video captioning: A comprehensive review
title_full_unstemmed	Exploring deep learning approaches for video captioning: A comprehensive review
title_short	Exploring deep learning approaches for video captioning: A comprehensive review
title_sort	exploring deep learning approaches for video captioning a comprehensive review
topic	Evaluation metrics Video captioning Video description Computer vision Deep learning
url	http://www.sciencedirect.com/science/article/pii/S277267112300267X
work_keys_str_mv	AT adeljalalyousif exploringdeeplearningapproachesforvideocaptioningacomprehensivereview AT mohammedhaljammas exploringdeeplearningapproachesforvideocaptioningacomprehensivereview

Exploring deep learning approaches for video captioning: A comprehensive review

Similar Items