The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis | allainews.com

April 2, 2024, 7:52 p.m. | Chen Yang, Junzhuo Li, Xinyao Niu, Xinrun Du, Songyang Gao, Haoran Zhang, Zhaoliang Chen, Xingwei Qu, Ruibin Yuan, Yizhi Li, Jiaheng Liu, Stephen W. H

cs.CL updates on arXiv.org arxiv.org

arXiv:2404.01204v1 Announce Type: new
Abstract: Uncovering early-stage metrics that reflect final model performance is one core principle for large-scale pretraining. The existing scaling law demonstrates the power-law correlation between pretraining loss and training flops, which serves as an important indicator of the current training state for large language models. However, this principle only focuses on the model's compression properties on the training data, resulting in an inconsistency with the ability improvements on the downstream tasks. Some follow-up works attempted to …

abstract analysis arxiv capability core correlation cs.cl current language language model large language large language model law line loss metrics performance power power-law pretraining scale scaling scaling law stage state streaming training type

More from arxiv.org / cs.CL updates on arXiv.org

Sketch-Guided Constrained Decoding for Boosting Blackbox Large Language Models without Logit Access 2 days, 9 hours ago | arxiv.org

abstract access application arxiv +21

LLaMA Pro: Progressive LLaMA with Block Expansion 2 days, 9 hours ago | arxiv.org

abstract arxiv block codellama +15

Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning 2 days, 9 hours ago | arxiv.org

arxiv captioning chart charts +4

Sibyl: Sensible Empathetic Dialogue Generation with Visionary Commonsense Knowledge 2 days, 9 hours ago | arxiv.org

abstract access arxiv building +19

PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models 2 days, 9 hours ago | arxiv.org

abstract accessibility art arxiv +17

ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models 2 days, 9 hours ago | arxiv.org

abstract arxiv challenges core +23

Cross-Lingual Knowledge Editing in Large Language Models 2 days, 9 hours ago | arxiv.org

arxiv cross-lingual cs.ai cs.cl +8

Hi Model, generating 'nice' instead of 'good' is not as bad as generating 'rice'! Towards … 2 days, 9 hours ago | arxiv.org

abstract arxiv context cs.cl +16

Chatlaw: A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model 2 days, 9 hours ago | arxiv.org

abstract agent ai legal arxiv +28

Senior Machine Learning Engineer

@ GPTZero | Toronto, Canada

View on ai-jobs.net

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

View on ai-jobs.net

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

View on ai-jobs.net

Seeking Developers and Engineers for AI T-Shirt Generator Project

@ Chevon Hicks | Remote

View on ai-jobs.net

Data Scientist, Mid

@ Booz Allen Hamilton | DEU, Stuttgart (Kurmaecker St)

View on ai-jobs.net

Tech Excellence Data Scientist

@ Booz Allen Hamilton | Undisclosed Location - USA, VA, Mclean

View on ai-jobs.net