Self-Supervised Learning for Audio-Visual Relationships of Videos With Stereo Sounds

Learning cross-modal features is an essential task for many multimedia applications such as sound localization, audio-visual alignment, and image/audio retrieval. Most existing methods mainly focus on the semantic correspondence between videos and monaural sounds, and spatial information of sound so...

Full description

Bibliographic Details
Main Authors:	Tomoya Sato, Yusuke Sugano, Yoichi Sato
Format:	Article
Language:	English
Published:	IEEE 2022-01-01
Series:	IEEE Access
Subjects:	Computer vision feature extraction machine learning self-supervised learning audio-visual learning cross-modal retrieval
Online Access:	https://ieeexplore.ieee.org/document/9877882/

Internet

https://ieeexplore.ieee.org/document/9877882/

Self-Supervised Learning for Audio-Visual Relationships of Videos With Stereo Sounds

Internet

Similar Items