The Local LLM Index / Quantization & Formats / #178
RahulSChand/gpu_poor
by RahulSChand · Quantization & Formats · updated 1y ago
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
35
momentum
1,403
stars
88
forks
#178
rank
ggmlgpuhuggingfacelanguage-modelllamallama2llamacppllmpytorchquantization
View on GitHub →