Compositional Chain-of-Thought Prompting for Large Multimodal Models | allainews.com

April 1, 2024, 4:43 a.m. | Chancharik Mitra, Brandon Huang, Trevor Darrell, Roei Herzig

cs.LG updates on arXiv.org arxiv.org

arXiv:2311.17076v2 Announce Type: replace-cross
Abstract: The combination of strong visual backbones and Large Language Model (LLM) reasoning has led to Large Multimodal Models (LMMs) becoming the current standard for a wide range of vision and language (VL) tasks. However, recent research has shown that even the most advanced LMMs still struggle to capture aspects of compositional visual reasoning, such as attributes and relationships between objects. One solution is to utilize scene graphs (SGs)--a formalization of objects and their relations and …

arxiv cs.ai cs.cl cs.cv cs.lg large multimodal models multimodal multimodal models prompting thought type

More from arxiv.org / cs.LG updates on arXiv.org

Red-Teaming for Generative AI: Silver Bullet or Security Theater? an hour ago | arxiv.org

abstract arxiv concerns cs.cy +15

Efficient Data-Driven MPC for Demand Response of Commercial Buildings an hour ago | arxiv.org

abstract arxiv buildings commercial +20

BrepGen: A B-rep Generative Diffusion Model with Structured Latent Geometry an hour ago | arxiv.org

arxiv cs.cv cs.lg diffusion +5

Data-Driven Physics-Informed Neural Networks: A Digital Twin Perspective an hour ago | arxiv.org

abstract arxiv automated construction +26

Testing the Segment Anything Model on radiology data an hour ago | arxiv.org

abstract applications arxiv become +20

Robust Point Matching with Distance Profiles an hour ago | arxiv.org

abstract analyze arxiv cs.lg +13

Cell Maps Representation For Lung Adenocarcinoma Growth Patterns Classification In Whole Slide Images an hour ago | arxiv.org

abstract arxiv behavior classification +18

Improved Baselines with Visual Instruction Tuning an hour ago | arxiv.org

abstract academic arxiv clip +25

Calorimeter shower superresolution an hour ago | arxiv.org

abstract arxiv challenge computational +16

Software Engineer for AI Training Data (School Specific)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Python)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Tier 2)

@ G2i Inc | Remote

View on ai-jobs.net

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net