[D] Why is Vaswani et al still the SOTA when the attention mechanism is O(n²)?

Aug. 11, 2022, 6:28 a.m. | /u/tororo-in

There are so many attention mechanisms (guided attention, apple's aft, nystromformer, etc) that alleviate the O(n²) to something like O(n). Why don't recent LMs use these techniques to speed up training and matrix multiplication in the SA layer?

attention machinelearning sota

Visit resource

More from www.reddit.com / Machine Learning

Do you think Reinforcement Learning still got it? [D] an hour ago | www.reddit.com

alphago architectures big computer +15

[P] TorchFix - a linter for PyTorch-using code with autofix support 3 hours ago | www.reddit.com

machinelearning

[P] AI-based Language Teacher that can run locally on a 12GB graphics card (RTX 4070) 7 hours ago | www.reddit.com

application card fun graphics +7

[D] Embeddings search "drowning" in a sea of noise! Can you solve this riddle? 8 hours ago | www.reddit.com

application concept dimensions embeddings +15

Any ways to improve TabNet..??? [D] 13 hours ago | www.reddit.com

machinelearning

[Discussion] Are there specific technical/scientific breakthroughs that have allowed the significant jump in maximum context … 15 hours ago | www.reddit.com

claude context gpt gpt-4 +14

[D] How to evaluate RAG - both retrieval and generation, when all I have is … 17 hours ago | www.reddit.com

data documents embedding embedding models +7

[R] Unifying Bias and Unfairness in Information Retrieval: A Survey of Challenges and Opportunities with … 17 hours ago | www.reddit.com

abstract advancement biases challenges +20

[D] Has anyone tried distilling large language models the old way? 21 hours ago | www.reddit.com

distillation however language language model +9

Senior Machine Learning Engineer (MLOps)

@ Promaton | Remote, Europe

View on ai-jobs.net

Applied Scientist, Control Stack, AWS Center for Quantum Computing

@ Amazon.com | Pasadena, California, USA

View on ai-jobs.net

Specialist Marketing with focus on ADAS/AD f/m/d

@ AVL | Graz, AT

View on ai-jobs.net

Machine Learning Engineer, PhD Intern

@ Instacart | United States - Remote

View on ai-jobs.net

Supervisor, Breast Imaging, Prostate Center, Ultrasound

@ University Health Network | Toronto, ON, Canada

View on ai-jobs.net

Senior Manager of Data Science (Recommendation Science)

@ NBCUniversal | New York, NEW YORK, United States

View on ai-jobs.net

View more jobs

all AI news

[D] Why is Vaswani et al still the SOTA when the attention mechanism is O(n²)?

More from www.reddit.com / Machine Learning

Jobs in AI, ML, Big Data

Senior Machine Learning Engineer (MLOps)

Applied Scientist, Control Stack, AWS Center for Quantum Computing

Specialist Marketing with focus on ADAS/AD f/m/d

Machine Learning Engineer, PhD Intern

Supervisor, Breast Imaging, Prostate Center, Ultrasound

Senior Manager of Data Science (Recommendation Science)