Latent Attention for Linear Time Transformers | allainews.com

Feb. 28, 2024, 5:44 a.m. | Rares Dolga, Marius Cobzarenco, David Barber

stat.ML updates on arXiv.org arxiv.org

arXiv:2402.17512v1 Announce Type: cross
Abstract: The time complexity of the standard attention mechanism in a transformer scales quadratically with the length of the sequence. We introduce a method to reduce this to linear scaling with time, based on defining attention via latent vectors. The method is readily usable as a drop-in replacement for the standard attention mechanism. Our "Latte Transformer" model can be implemented for both bidirectional and unidirectional tasks, with the causal version allowing a recurrent implementation which is …

abstract arxiv attention complexity cs.cl linear reduce replacement scaling standard stat.ml transformer transformers type vectors via

More from arxiv.org / stat.ML updates on arXiv.org

Nuisance Function Tuning for Optimal Doubly Robust Estimation 2 days, 23 hours ago | arxiv.org

abstract arxiv convergence function +12

Fast Topological Signal Identification and Persistent Cohomological Cycle Matching 2 days, 23 hours ago | arxiv.org

abstract analysis applications art +20

Neural Networks for Extreme Quantile Regression with an Application to Forecasting of Flood Risk 2 days, 23 hours ago | arxiv.org

abstract application arxiv assessment +17

The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms 2 days, 23 hours ago | arxiv.org

abstract algorithms arxiv call +15

Comparison of Point Process Learning and its special case Takacs-Fiksel estimation 2 days, 23 hours ago | arxiv.org

abstract arxiv case comparison +14

Algorithmically Designed Artificial Neural Networks (ADANNs): Higher order deep operator learning for parametric partial differential … 3 days, 23 hours ago | arxiv.org

abstract ann architectures article +18

Adaptive posterior concentration rates for sparse high-dimensional linear regression with random design and unknown error … 3 days, 23 hours ago | arxiv.org

abstract analyze arxiv design +13

CHANI: Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration 3 days, 23 hours ago | arxiv.org

abstract aggregation arxiv bio +14

Principled Probabilistic Imaging using Diffusion Models as Plug-and-Play Priors 3 days, 23 hours ago | arxiv.org

abstract arxiv bayesian capability +15

Senior Machine Learning Engineer

@ GPTZero | Toronto, Canada

View on ai-jobs.net

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

View on ai-jobs.net

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

View on ai-jobs.net

Seeking Developers and Engineers for AI T-Shirt Generator Project

@ Chevon Hicks | Remote

View on ai-jobs.net

Senior Applied Data Scientist

@ dunnhumby | London

View on ai-jobs.net

Principal Data Architect - Azure & Big Data

@ MGM Resorts International | Home Office - US, NV

View on ai-jobs.net