The Local LLM Index / Quantization & Formats / #101
raketenkater/ggrun
by raketenkater · Quantization & Formats · updated today
Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.
59
momentum
260
stars
14
forks
#101
rank
cudaggufgolanginference-serverllama-cppllamacppllmlocal-llmlocalllamametalmoemulti-gpu
View on GitHub →