The Local LLM Index / Inference Engines / #43

youssofal/MTPLX

by youssofal · Inference Engines · updated today

The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

70
momentum
2,344
stars
174
forks
#43
rank
anthropic-compatibleapple-siliconclaude-codeflash-nextinference-enginellm-inferencelocal-ailocal-llmmacosmetalmlxmtp
View on GitHub →