Grounding referring expressions in images by variational context
We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., 'largest elephant standing behind baby elephant'. This is a general yet challenging vision-language task since it does not only require the localization of objects, but also the multimodal comprehens...
Main Authors: | , , |
---|---|
Other Authors: | |
Format: | Conference Paper |
Language: | English |
Published: |
2020
|
Subjects: | |
Online Access: | https://hdl.handle.net/10356/143054 |