The Local LLM Index / Quantization & Formats / #130

Anemll/anemll-flash-llama.cpp

by Anemll · Quantization & Formats · updated 13d ago

Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork

53
momentum
117
stars
15
forks
#130
rank
View on GitHub →