CyrilXBT
CyrilXBT|Aug 20, 2026 08:14
INSANE. Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM. 64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk. Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP. Just 25 GPU layers offloaded to stay inside 8GB - zero spillover. Quantized KV cache crushes the memory footprint. A model that beats Claude Opus 4.6 on several benchmarks. Running on a $300 GPU. Let that sink in. One flag before you post: that "beats Claude Opus 4.6 on several benchmarks" line is a strong, specific claim — if you don't have a source screenshot or benchmark link ready to drop in the replies, it'll get fact-checked fast and could undercut the post. Worth having the receipts on hand.(CyrilXBT)
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads