Leveraging Weakly Annotated Data for Hate Speech Detection in Code-Mixed Hinglish: A Feasibility-Driven Transfer Learning Approach with Large Language Models | allainews.com

March 5, 2024, 2:52 p.m. | Sargam YadavDundalk Institute of Technology, Dundalk, Abhishek KaushikDundalk Institute of Technology, Dundalk, Kevin McDaidDundalk Institute of Techn

cs.CL updates on arXiv.org arxiv.org

arXiv:2403.02121v1 Announce Type: new
Abstract: The advent of Large Language Models (LLMs) has advanced the benchmark in various Natural Language Processing (NLP) tasks. However, large amounts of labelled training data are required to train LLMs. Furthermore, data annotation and training are computationally expensive and time-consuming. Zero and few-shot learning have recently emerged as viable options for labelling data using large pre-trained models. Hate speech detection in mix-code low-resource languages is an active problem area where the use of LLMs has …

abstract advanced annotated data annotation arxiv benchmark code cs.ai cs.cl data data annotation detection hate speech hate speech detection language language models language processing large language large language models llms mixed natural natural language natural language processing nlp processing speech tasks train training training data transfer transfer learning type

More from arxiv.org / cs.CL updates on arXiv.org

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models 15 hours ago | arxiv.org

abstract arxiv become contents +17

Temporal Knowledge Question Answering via Abstract Reasoning Induction 15 hours ago | arxiv.org

abstract arxiv cs.ai cs.cl +8

Large Language Models can Contrastively Refine their Generation for Better Sentence Representation Learning 15 hours ago | arxiv.org

abstract application arxiv capabilities +19

ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base 15 hours ago | arxiv.org

abstract arxiv cognitive cs.ai +23

FOLIO: Natural Language Reasoning with First-Order Logic 15 hours ago | arxiv.org

abstract arxiv benchmarks capabilities +21

Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks 15 hours ago | arxiv.org

arxiv attention attention mechanisms cs.cl +6

SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks 15 hours ago | arxiv.org

abstract arxiv capabilities communities +17

A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers 15 hours ago | arxiv.org

abstract academia accessibility advances +28

COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain 15 hours ago | arxiv.org

abstract advanced art artificial +25

Software Engineer for AI Training Data (School Specific)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Python)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Tier 2)

@ G2i Inc | Remote

View on ai-jobs.net

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net