The Local LLM Index / Quantization & Formats / #90
localai-org/vllm.cpp
by localai-org · Quantization & Formats · updated today
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
64
momentum
447
stars
54
forks
#90
rank
continous-batchingcpp20llmmlxpaged-attentionradixattentionvllm
View on GitHub →