Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models | allainews.com

Feb. 27, 2024, 5:50 a.m. | Chaoya Jiang, Wei Ye, Mengfan Dong, Hongrui Jia, Haiyang Xu, Ming Yan, Ji Zhang, Shikun Zhang

cs.CL updates on arXiv.org arxiv.org

arXiv:2402.15721v1 Announce Type: cross
Abstract: Large Vision Language Models exhibit remarkable capabilities but struggle with hallucinations inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine grained hallucinatory data …

abstract arxiv capabilities cs.ai cs.cl evaluation fine-grained framework hal hallucination hallucinations images language language models narrative objects relations struggle studies terms type universal vision

More from arxiv.org / cs.CL updates on arXiv.org

Sketch-Guided Constrained Decoding for Boosting Blackbox Large Language Models without Logit Access 1 day, 2 hours ago | arxiv.org

abstract access application arxiv +21

LLaMA Pro: Progressive LLaMA with Block Expansion 1 day, 2 hours ago | arxiv.org

abstract arxiv block codellama +15

Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning 1 day, 2 hours ago | arxiv.org

arxiv captioning chart charts +4

Sibyl: Sensible Empathetic Dialogue Generation with Visionary Commonsense Knowledge 1 day, 2 hours ago | arxiv.org

abstract access arxiv building +19

PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models 1 day, 2 hours ago | arxiv.org

abstract accessibility art arxiv +17

ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models 1 day, 2 hours ago | arxiv.org

abstract arxiv challenges core +23

Cross-Lingual Knowledge Editing in Large Language Models 1 day, 2 hours ago | arxiv.org

arxiv cross-lingual cs.ai cs.cl +8

Hi Model, generating 'nice' instead of 'good' is not as bad as generating 'rice'! Towards … 1 day, 2 hours ago | arxiv.org

abstract arxiv context cs.cl +16

Chatlaw: A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model 1 day, 2 hours ago | arxiv.org

abstract agent ai legal arxiv +28

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

View on ai-jobs.net

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

View on ai-jobs.net

Seeking Developers and Engineers for AI T-Shirt Generator Project

@ Chevon Hicks | Remote

View on ai-jobs.net

Security Data Engineer

@ ASML | Veldhoven, Building 08, Netherlands

View on ai-jobs.net

Data Engineer

@ Parsons Corporation | Pune - Business Bay

View on ai-jobs.net

Data Engineer

@ Parsons Corporation | Bengaluru, Velankani Tech Park

View on ai-jobs.net