CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data | allainews.com

April 4, 2024, 4:47 a.m. | Rian Touchent, Laurent Romary, Eric de la Clergerie

cs.CL updates on arXiv.org arxiv.org

arXiv:2306.15550v3 Announce Type: replace
Abstract: Clinical data in hospitals are increasingly accessible for research through clinical data warehouses. However these documents are unstructured and it is therefore necessary to extract information from medical reports to conduct clinical studies. Transfer learning with BERT-like models such as CamemBERT has allowed major advances for French, especially for named entity recognition. However, these models are trained for plain language and are less efficient on biomedical data. Addressing this gap, we introduce CamemBERT-bio, a dedicated …

abstract arxiv bert bio biomedical clinical continual cost cs.ai cs.cl data data warehouses documents extract french hospitals however information major medical pre-training reports research studies through training transfer transfer learning type unstructured warehouses

More from arxiv.org / cs.CL updates on arXiv.org

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models an hour ago | arxiv.org

abstract arxiv become contents +17

Temporal Knowledge Question Answering via Abstract Reasoning Induction an hour ago | arxiv.org

abstract arxiv cs.ai cs.cl +8

Large Language Models can Contrastively Refine their Generation for Better Sentence Representation Learning an hour ago | arxiv.org

abstract application arxiv capabilities +19

ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base an hour ago | arxiv.org

abstract arxiv cognitive cs.ai +23

FOLIO: Natural Language Reasoning with First-Order Logic an hour ago | arxiv.org

abstract arxiv benchmarks capabilities +21

Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks an hour ago | arxiv.org

arxiv attention attention mechanisms cs.cl +6

SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks an hour ago | arxiv.org

abstract arxiv capabilities communities +17

A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers an hour ago | arxiv.org

abstract academia accessibility advances +28

COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain an hour ago | arxiv.org

abstract advanced art artificial +25

Software Engineer for AI Training Data (School Specific)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Python)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Tier 2)

@ G2i Inc | Remote

View on ai-jobs.net

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net