The Local LLM Index / Inference Engines / #206
xaskasdf/ntransformer
by xaskasdf · Inference Engines · updated 6mo ago
High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.
30
momentum
465
stars
19
forks
#206
rank
High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.