Grounding referring expressions in images by variational context

We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., 'largest elephant standing behind baby elephant'. This is a general yet challenging vision-language task since it does not only require the localization of objects, but also the multimodal comprehens...

Full description

Bibliographic Details
Main Authors:	Zhang, Hanwang, Niu, Yulei, Chang, Shih-Fu
Other Authors:	School of Computer Science and Engineering
Format:	Conference Paper
Language:	English
Published:	2020
Subjects:	Engineering::Computer science and engineering Grounding Context Modeling
Online Access:	https://hdl.handle.net/10356/143054

Internet

https://hdl.handle.net/10356/143054

Grounding referring expressions in images by variational context

Internet

Similar Items