Improving automatic speech recognition through head pose driven visual grounding

In this paper, we present a multimodal speech recognition system for real world scene description tasks. Given a visual scene, the system dynamically biases its language model based on the content of the visual scene and visual attention of the speaker. Visual attention is used to focus on likely ob...

Full description

Bibliographic Details
Main Author: Vosoughi, Soroush
Other Authors: Massachusetts Institute of Technology. Media Laboratory
Format: Article
Language:en_US
Published: Association for Computing Machinery 2014
Online Access:http://hdl.handle.net/1721.1/86943
https://orcid.org/0000-0002-2564-8909