The Local LLM Index / Quantization & Formats / #48
sudoingX/qwen38-mtp
by sudoingX · Quantization & Formats · updated 2d ago
One llama.cpp flag unlocks +33-39% decode speed for Qwen3.8-27B on consumer GPUs. The MTP head already ships inside your GGUF. Recipe, paired benchmarks, probe tool.
69
momentum
292
stars
86
forks
#48
rank