Self-Supervised Learning for Audio-Visual Relationships of Videos With Stereo Sounds

Learning cross-modal features is an essential task for many multimedia applications such as sound localization, audio-visual alignment, and image/audio retrieval. Most existing methods mainly focus on the semantic correspondence between videos and monaural sounds, and spatial information of sound so...

Full description

Bibliographic Details
Main Authors: Tomoya Sato, Yusuke Sugano, Yoichi Sato
Format: Article
Language:English
Published: IEEE 2022-01-01
Series:IEEE Access
Subjects:
Online Access:https://ieeexplore.ieee.org/document/9877882/