all AI news
[D] How to make AI model infernce faster?
Nov. 27, 2023, 12:34 p.m. | /u/HughLee_1999
Machine Learning www.reddit.com
I am currently trying to serve some LM models like GPT2 and some LLMs like llama or bloomz. I tried converting GPT2 to TensorRT (based on Nvidia's documentation) and serving with Triton server. The output is also quite good. However, if you want to continue to optimize, will there be any more methods to perform optimization? And I'm also quite curious how to make the LLM model able to serve many users at the same time. Is there …
ai model documentation faster good llama llms machinelearning nvidia serve server tensorrt triton will
More from www.reddit.com / Machine Learning
[D] software to design figures
10 hours ago |
www.reddit.com
[Discussion] Should I go to ICML and present my paper?
1 day, 4 hours ago |
www.reddit.com
Jobs in AI, ML, Big Data
AI Engineer Intern, Agents
@ Occam AI | US
AI Research Scientist
@ Vara | Berlin, Germany and Remote
Data Architect
@ University of Texas at Austin | Austin, TX
Data ETL Engineer
@ University of Texas at Austin | Austin, TX
Lead GNSS Data Scientist
@ Lurra Systems | Melbourne
Lead Data Modeler
@ Sherwin-Williams | Cleveland, OH, United States