The Local LLM Index / Quantization & Formats / #209
SqueezeAILab/KVQuant
by SqueezeAILab · Quantization & Formats · updated 2y ago
[NeurIPS 2024] KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
30
momentum
434
stars
48
forks
#209
rank
compressionefficient-inferenceefficient-modellarge-language-modelsllamallmlocalllamalocalllmmistralmodel-compressionnatural-language-processingquantization
View on GitHub →