Self-Supervised Learning for Audio-Visual Relationships of Videos With Stereo Sounds

Learning cross-modal features is an essential task for many multimedia applications such as sound localization, audio-visual alignment, and image/audio retrieval. Most existing methods mainly focus on the semantic correspondence between videos and monaural sounds, and spatial information of sound so...

詳細記述

書誌詳細
主要な著者: Tomoya Sato, Yusuke Sugano, Yoichi Sato
フォーマット: 論文
言語:English
出版事項: IEEE 2022-01-01
シリーズ:IEEE Access
主題:
オンライン・アクセス:https://ieeexplore.ieee.org/document/9877882/