The Local LLM Index / Quantization & Formats / #12
syv-ai/HyperQwen
by syv-ai · Quantization & Formats · updated today
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
77
momentum
1,450
stars
200
forks
#12
rank
consumer-gpucudainference-optimizationkv-cachellm-inferencelocal-llmquantizationqwenqwen3rtx-3090speculative-decodingvllm
View on GitHub →