[R] The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | allainews.com

Feb. 28, 2024, 10:03 a.m. | /u/Civil_Collection7267

Machine Learning www.reddit.com

[https://arxiv.org/abs/2402.17764](https://arxiv.org/abs/2402.17764)

**Abstract**

>Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}. It matches the full-precision (i.e., FP16 or BF16) Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in …

abstract every fp16 language language models large language large language models llm llms machinelearning precision research the way transformer work

More from www.reddit.com / Machine Learning

[D] Kolmogorov-Arnold Network is just an MLP 57 minutes ago | www.reddit.com

machinelearning mlp network relu +1

[D] Why Gemma has such crazy big MLP hidden dim size? an hour ago | www.reddit.com

big gemma hidden machinelearning +1

[R] Why can Llama-3 work with 32K context if it only had 8K context length? 2 hours ago | www.reddit.com

32k context config context dynamic +7

[D] Is there a formal name for "dialogue classification?" 8 hours ago | www.reddit.com

agents classification customer customer service +11

How Large Language Models play video games [D] 9 hours ago | www.reddit.com

agents case engineering explore +15

[Project] An LLM-Powered Web App for SEC Filing Insights 9 hours ago | www.reddit.com

apis app financial future +18

[Research] Understanding The Attention Mechanism In Transformers: A 5-minute visual guide. 🧠 13 hours ago | www.reddit.com

architectures attention dictionary guide +12

[D] Is there a more systematic way of choosing the layers or how deep the … 17 hours ago | www.reddit.com

architecture deep learning least machinelearning +6

[D] Where does the real value of a data scientist come from? 22 hours ago | www.reddit.com

code companies data data scientist +11

Founding AI Engineer, Agents

@ Occam AI | New York

View on ai-jobs.net

AI Engineer Intern, Agents

@ Occam AI | US

View on ai-jobs.net

AI Research Scientist

@ Vara | Berlin, Germany and Remote

View on ai-jobs.net

Data Architect

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Data ETL Engineer

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Codec Avatars Research Engineer

@ Meta | Pittsburgh, PA

View on ai-jobs.net