Web: https://www.reddit.com/r/MachineLearning/comments/xjg9bp/p_gpt_inference_on_the_cpu_in_cc/

Sept. 20, 2022, 6:19 p.m. | /u/ggerganov

Machine Learning reddit.com

I wanted to learn a bit more about the GPT models and understand how they work, so I decided to try and implement the inference from scratch. My programming language of choice is C/C++.

This weekend I got it working, and I can now run GPT-J on my MacBook. The inference runs on the CPU and I think the performance is quite reasonable - around 125 ms per token.

Here is a short write up and instructions how you can …

cpu gpt inference machinelearning

Postdoctoral Fellow: ML for autonomous materials discovery

@ Lawrence Berkeley National Lab | Berkeley, CA

Research Scientists

@ ODU Research Foundation | Norfolk, Virginia

Embedded Systems Engineer (Robotics)

@ Neo Cybernetica | Bedford, New Hampshire

2023 Luis J. Alvarez and Admiral Grace M. Hopper Postdoc Fellowship in Computing Sciences

@ Lawrence Berkeley National Lab | San Francisco, CA

Senior Manager Data Scientist

@ NAV | Remote, US

Senior AI Research Scientist

@ Earth Species Project | Remote anywhere

Research Fellow- Center for Security and Emerging Technology (Multiple Opportunities)

@ University of California Davis | Washington, DC

Staff Fellow - Data Scientist

@ U.S. FDA/Center for Devices and Radiological Health | Silver Spring, Maryland

Staff Fellow - Senior Data Engineer

@ U.S. FDA/Center for Devices and Radiological Health | Silver Spring, Maryland

Research Engineer - VFX, Neural Compositing

@ Flawless | Los Angeles, California, United States

[Job-TB] Senior Data Engineer

@ CI&T | Brazil

Data Analytics Engineer

@ The Fork | Paris, France