Nov. 27, 2023, 12:34 p.m. | /u/HughLee_1999

Machine Learning www.reddit.com

Hi everyone,

I am currently trying to serve some LM models like GPT2 and some LLMs like llama or bloomz. I tried converting GPT2 to TensorRT (based on Nvidia's documentation) and serving with Triton server. The output is also quite good. However, if you want to continue to optimize, will there be any more methods to perform optimization? And I'm also quite curious how to make the LLM model able to serve many users at the same time. Is there …

ai model documentation faster good llama llms machinelearning nvidia serve server tensorrt triton will

Senior Machine Learning Engineer

@ GPTZero | Toronto, Canada

ML/AI Engineer / NLP Expert - Custom LLM Development (x/f/m)

@ HelloBetter | Remote

Doctoral Researcher (m/f/div) in Automated Processing of Bioimages

@ Leibniz Institute for Natural Product Research and Infection Biology (Leibniz-HKI) | Jena

Seeking Developers and Engineers for AI T-Shirt Generator Project

@ Chevon Hicks | Remote

Principal Data Architect - Azure & Big Data

@ MGM Resorts International | Home Office - US, NV

GN SONG MT Market Research Data Analyst 11

@ Accenture | Bengaluru, BDC7A