The Local LLM Index / Quantization & Formats / #90

localai-org/vllm.cpp

by localai-org · Quantization & Formats · updated today

a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)

64
momentum
447
stars
54
forks
#90
rank
continous-batchingcpp20llmmlxpaged-attentionradixattentionvllm
View on GitHub →