Beyond Performance: Quantifying and Mitigating Label Bias in LLMs | allainews.com

May 7, 2024, 4:50 a.m. | Yuval Reif, Roy Schwartz

cs.CL updates on arXiv.org arxiv.org

arXiv:2405.02743v1 Announce Type: new
Abstract: Large language models (LLMs) have shown remarkable adaptability to diverse tasks, by leveraging context prompts containing instructions, or minimal input-output examples. However, recent work revealed they also exhibit label bias -- an undesirable preference toward predicting certain answers over others. Still, detecting and measuring this bias reliably and at scale has remained relatively unexplored. In this study, we evaluate different approaches to quantifying label bias in a model's predictions, conducting a comprehensive investigation across 279 …

abstract adaptability arxiv beyond bias context cs.cl diverse examples however input-output language language models large language large language models llms measuring performance prompts tasks type work

More from arxiv.org / cs.CL updates on arXiv.org

Sketch-Guided Constrained Decoding for Boosting Blackbox Large Language Models without Logit Access 1 day, 17 hours ago | arxiv.org

abstract access application arxiv +21

LLaMA Pro: Progressive LLaMA with Block Expansion 1 day, 17 hours ago | arxiv.org

abstract arxiv block codellama +15

Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning 1 day, 17 hours ago | arxiv.org

arxiv captioning chart charts +4

Sibyl: Sensible Empathetic Dialogue Generation with Visionary Commonsense Knowledge 1 day, 17 hours ago | arxiv.org

abstract access arxiv building +19

PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models 1 day, 17 hours ago | arxiv.org

abstract accessibility art arxiv +17

ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models 1 day, 17 hours ago | arxiv.org

abstract arxiv challenges core +23

Cross-Lingual Knowledge Editing in Large Language Models 1 day, 17 hours ago | arxiv.org

arxiv cross-lingual cs.ai cs.cl +8

Hi Model, generating 'nice' instead of 'good' is not as bad as generating 'rice'! Towards … 1 day, 17 hours ago | arxiv.org

abstract arxiv context cs.cl +16

Chatlaw: A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model 1 day, 17 hours ago | arxiv.org

abstract agent ai legal arxiv +28

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

View on ai-jobs.net

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

View on ai-jobs.net

Seeking Developers and Engineers for AI T-Shirt Generator Project

@ Chevon Hicks | Remote

View on ai-jobs.net

Cloud Data Platform Engineer

@ First Central | Home Office (Remote)

View on ai-jobs.net

Associate Director, Data Science

@ MSD | USA - New Jersey - Rahway

View on ai-jobs.net

Data Scientist Sr.

@ MSD | CHL - Santiago - Santiago (Calle Mariano)

View on ai-jobs.net