The Local LLM Index / Quantization & Formats / #187
RahulSChand/gpu_poor
by RahulSChand · Quantization & Formats · updated 1y ago
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
35
momentum
1,403
stars
88
forks
#187
rank
ggmlgpuhuggingfacelanguage-modelllamallama2llamacppllmpytorchquantization
View on GitHub →