Advancing LLM Reasoning Generalists with Preference Trees | allainews.com

April 3, 2024, 4:42 a.m. | Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao

cs.LG updates on arXiv.org arxiv.org

arXiv:2404.02078v1 Announce Type: cross
Abstract: We introduce Eurus, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B and CodeLlama-70B, Eurus models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, Eurus-70B beats GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 tests covering five tasks, and achieves a 33.3% pass@1 accuracy on LeetCode and 32.6% on TheoremQA, two challenging benchmarks, substantially outperforming existing …

70b abstract art arxiv benchmarks code code generation codellama cs.ai cs.cl cs.lg diverse gpt gpt-3 gpt-3.5 language language models large language large language models llm llm reasoning llms mathematics mistral open-source models reasoning results set state through trees turbo type

More from arxiv.org / cs.LG updates on arXiv.org

LangProp: A code optimization framework using Large Language Models applied to driving 19 hours ago | arxiv.org

arxiv code cs.ai cs.lg +10

MRI Scan Synthesis Methods based on Clustering and Pix2Pix 19 hours ago | arxiv.org

abstract arxiv automated brain +16

Continual Diffusion with STAMINA: STack-And-Mask INcremental Adapters 19 hours ago | arxiv.org

abstract arxiv concept concepts +21

Improving Interpretation Faithfulness for Vision Transformers 19 hours ago | arxiv.org

abstract adversarial adversarial attacks architectures +21

Training robust and generalizable quantum models 19 hours ago | arxiv.org

abstract adversarial arxiv context +15

Causal Discovery Under Local Privacy 19 hours ago | arxiv.org

abstract application arxiv causal +19

From Neural Activations to Concepts: A Survey on Explaining Concepts in Neural Networks 19 hours ago | arxiv.org

abstract act arxiv concepts +13

It's About Time: Temporal References in Emergent Communication 19 hours ago | arxiv.org

abstract agents arxiv autonomous +21

Learning Risk-Aware Quadrupedal Locomotion using Distributional Reinforcement Learning 19 hours ago | arxiv.org

arxiv cs.lg cs.ro reinforcement +3

Founding AI Engineer, Agents

@ Occam AI | New York

View on ai-jobs.net

AI Engineer Intern, Agents

@ Occam AI | US

View on ai-jobs.net

AI Research Scientist

@ Vara | Berlin, Germany and Remote

View on ai-jobs.net

Data Architect

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Data ETL Engineer

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Alternance DATA/AI Engineer (H/F)

@ SQLI | Le Grand-Quevilly, France

View on ai-jobs.net