May 13, 2024, 4:46 a.m. | Jean Mercat, Igor Vasiljevic, Sedrick Keh, Kushal Arora, Achal Dave, Adrien Gaidon, Thomas Kollar

cs.CL updates on arXiv.org arxiv.org

arXiv:2405.06640v1 Announce Type: new
Abstract: Linear transformers have emerged as a subquadratic-time alternative to softmax attention and have garnered significant interest due to their fixed-size recurrent state that lowers inference cost. However, their original formulation suffers from poor scaling and underperforms compute-matched transformers. Recent linear models such as RWKV and Mamba have attempted to address these shortcomings by proposing novel time-mixing and gating architectures, but pre-training large language models requires significant data and compute investments. Thus, the search for subquadratic …

arxiv cs.cl language language models large language large language models type

Senior Machine Learning Engineer

@ GPTZero | Toronto, Canada

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

Director, Global Success Business Intelligence

@ Salesforce | Texas - Austin

Deep Learning Compiler Engineer - MLIR

@ NVIDIA | US, CA, Santa Clara

Commerce Data Engineer (Remote)

@ CrowdStrike | USA TX Remote