all AI news
When and why vision-language models behave like bags-of-words, and what to do about it?. (arXiv:2210.01936v2 [cs.CV] UPDATED)
Oct. 7, 2022, 1:16 a.m. | Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, James Zou
cs.CV updates on arXiv.org arxiv.org
Despite the success of large vision and language models (VLMs) in many
downstream applications, it is unclear how well they encode compositional
information. Here, we create the Attribution, Relation, and Order (ARO)
benchmark to systematically evaluate the ability of VLMs to understand
different types of relationships, attributes, and order. ARO consists of Visual
Genome Attribution, to test the understanding of objects' properties; Visual
Genome Relation, to test for relational understanding; and COCO &
Flickr30k-Order, to test for order sensitivity. ARO …
More from arxiv.org / cs.CV updates on arXiv.org
Jobs in AI, ML, Big Data
Lead GNSS Data Scientist
@ Lurra Systems | Melbourne
Senior Machine Learning Engineer (MLOps)
@ Promaton | Remote, Europe
Data Engineer
@ Contact Government Services | Trenton, NJ
Data Engineer
@ Comply365 | Bristol, UK
Masterarbeit: Deep learning-basierte Fehler Detektion bei Montageaufgaben
@ Fraunhofer-Gesellschaft | Karlsruhe, DE, 76131
Assistant Manager ETL testing 1
@ KPMG India | Bengaluru, Karnataka, India