The Local LLM Index / Inference Engines / #85
jmaczan/tiny-vllm
by jmaczan · Inference Engines · updated 19d ago
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
63
momentum
1,106
stars
89
forks
#85
rank
aiattentionbatchingcoursecppcudahpcinferencellmllm-inferencepagedattentiontiny-vllm
View on GitHub →