The Local LLM Index / Quantization & Formats / #27

Michael-A-Kuykendall/shimmy

by Michael-A-Kuykendall · Quantization & Formats · updated 12d ago

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

72
momentum
5,866
stars
569
forks
#27
rank
api-servercommand-line-tooldeveloper-toolsggufhuggingfacehuggingface-modelshuggingface-transformersinference-serverllamallamacppllm-inferencelocal-ai
View on GitHub →