Audio captioning and retrieval with improved cross-modal objectives

Automated Audio Captioning (AAC) is the task of generating descriptive captions from an input audio clip, while Language-Based Audio Retrieval (LBAR) is the task of retrieving the most relevant audio clip based on an input text query. AAC requires a model that is not only able to comprehend the acou...

ver descrição completa

Detalhes bibliográficos
Autor principal: Koh, Andrew Jin Jie
Outros Autores: Chng Eng Siong
Formato: Thesis-Doctor of Philosophy
Idioma:English
Publicado em: Nanyang Technological University 2023
Assuntos:
Acesso em linha:https://hdl.handle.net/10356/172437

Registros relacionados