Self-Supervised Learning for Audio-Visual Relationships of Videos With Stereo Sounds
Learning cross-modal features is an essential task for many multimedia applications such as sound localization, audio-visual alignment, and image/audio retrieval. Most existing methods mainly focus on the semantic correspondence between videos and monaural sounds, and spatial information of sound so...
主要な著者: | , , |
---|---|
フォーマット: | 論文 |
言語: | English |
出版事項: |
IEEE
2022-01-01
|
シリーズ: | IEEE Access |
主題: | |
オンライン・アクセス: | https://ieeexplore.ieee.org/document/9877882/ |