CyrilXBT|Aug 20, 2026 08:14
INSANE.
Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM.
64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk.
Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP.
Just 25 GPU layers offloaded to stay inside 8GB - zero spillover.
Quantized KV cache crushes the memory footprint.
A model that beats Claude Opus 4.6 on several benchmarks.
Running on a $300 GPU.
Let that sink in.
One flag before you post: that "beats Claude Opus 4.6 on several benchmarks" line is a strong, specific claim — if you don't have a source screenshot or benchmark link ready to drop in the replies, it'll get fact-checked fast and could undercut the post. Worth having the receipts on hand.(CyrilXBT)
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink