The Local LLM Index / Inference Engines / #60
waybarrios/vllm-mlx
by waybarrios · Inference Engines · updated 6d ago
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
67
momentum
1,570
stars
221
forks
#60
rank
anthropicanthropic-apiapple-siliconclaude-codecontinuous-batchinginference-serverllmlocal-llmmacosmcpmlxmultimodal-ai
View on GitHub →