[R] Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding - Carnegie Mellon University 2024 - Allows running an unquantized Llama2-70B on an RTX4090 with half-second per token latency! | allainews.com

March 13, 2024, 4:35 p.m. | /u/Singularian2501

Machine Learning www.reddit.com

Paper: [https://arxiv.org/abs/2402.12374](https://arxiv.org/abs/2402.12374)

Github: [https://github.com/Infini-AI-Lab/Sequoia/tree/main](https://github.com/Infini-AI-Lab/Sequoia/tree/main)

Abstract:

>As the usage of large language models (LLMs) grows, performing efficient inference with these models becomes increasingly important. While speculative decoding has recently emerged as a promising direction for speeding up inference, existing methods are limited in their ability to scale to larger speculation budgets, and adapt to different hyperparameters and hardware. This paper introduces Sequoia, a scalable, robust, and hardware-aware algorithm for speculative decoding. To attain better scalability, Sequoia introduces a dynamic programming algorithm …

abstract adapt budgets decoding hardware inference language language models large language large language models llms machinelearning paper robust scalable scale sequoia speculation usage

More from www.reddit.com / Machine Learning

[D]What Are Your Favorite Tools That You Use For Research? 42 minutes ago | www.reddit.com

become found latest machinelearning +8

[P] Automated LoRA Discovery 3 hours ago | www.reddit.com

adapter automated discovery explore +11

[D] Can other areas researches such as the recent mapping of a cubic millimeter of … 10 hours ago | www.reddit.com

algorithms brain build human +7

[D] Is sequence packing common for training transformers? 19 hours ago | www.reddit.com

machinelearning training transformers

[D] ML Conferences and Organization Metrics 21 hours ago | www.reddit.com

machinelearning

[D] KAN == multi-layer GAM ? 1 day ago | www.reddit.com

function functions generalized kan +11

[R] Lipreading with LipNet: End-to-End Sentence-level Lipreading 1 day, 3 hours ago | www.reddit.com

complexity features gru hey +5

[R] Machine learning introspection 1 day, 11 hours ago | www.reddit.com

machinelearning urls videos

[R] Research Collaboration in CV /Structured Light/ 3D Reconstruction 1 day, 11 hours ago | www.reddit.com

3d reconstruction 3d scanning analysis challenges +8

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

View on ai-jobs.net

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

View on ai-jobs.net

Seeking Developers and Engineers for AI T-Shirt Generator Project

@ Chevon Hicks | Remote

View on ai-jobs.net

Security Data Engineer

@ ASML | Veldhoven, Building 08, Netherlands

View on ai-jobs.net

Data Engineer

@ Parsons Corporation | Pune - Business Bay

View on ai-jobs.net

Data Engineer

@ Parsons Corporation | Bengaluru, Velankani Tech Park

View on ai-jobs.net