Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects | allainews.com

April 15, 2024, 4:47 a.m. | Junyu Lu, Dixiang Zhang, Songxin Zhang, Zejian Xie, Zhuoyang Song, Cong Lin, Jiaxing Zhang, Bingyi Jing, Pingjian Zhang

cs.CL updates on arXiv.org arxiv.org

arXiv:2312.05278v2 Announce Type: replace
Abstract: Large Vision Language Models (LVLMs) have demonstrated impressive zero-shot capabilities in various vision-language dialogue scenarios. However, the absence of fine-grained visual object detection hinders the model from understanding the details of images, leading to irreparable visual hallucinations and factual errors. In this paper, we propose Lyrics, a novel multi-modal pre-training and instruction fine-tuning paradigm that bootstraps vision-language alignment from fine-grained cross-modal collaboration. Building on the foundation of BLIP-2, Lyrics infuses local visual features extracted from …

abstract alignment arxiv boosting capabilities cs.cl detection dialogue errors fine-grained hallucinations however images language language models lyrics object objects paper semantic type understanding via vision visual zero-shot

More from arxiv.org / cs.CL updates on arXiv.org

Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults 22 hours ago | arxiv.org

abstract arxiv asr automatic speech recognition +19

Beyond Turing: A Comparative Analysis of Approaches for Detecting Machine-Generated Text 22 hours ago | arxiv.org

abstract analysis arxiv beyond +18

Tackling Fake News in Bengali: Unraveling the Impact of Summarization vs. Augmentation on Pre-trained Language … 22 hours ago | arxiv.org

arxiv augmentation cs.cl fake +7

Matching domain experts by training from scratch on domain knowledge 22 hours ago | arxiv.org

abstract arxiv cs.ai cs.cl +24

QueryNER: Segmentation of E-commerce Queries 22 hours ago | arxiv.org

abstract arxiv commerce cs.ai +13

ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata 22 hours ago | arxiv.org

abstract arxiv cs.ai cs.cl +8

Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts 22 hours ago | arxiv.org

abstract adapt applications arxiv +20

Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models 22 hours ago | arxiv.org

abstract analysis arxiv cs.cl +15

Facilitating Opinion Diversity through Hybrid NLP Approaches 22 hours ago | arxiv.org

abstract arxiv challenges cs.ai +22

Software Engineer for AI Training Data (School Specific)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Python)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Tier 2)

@ G2i Inc | Remote

View on ai-jobs.net

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net