a16z|Aug 06, 2026 18:17
vLLM runs on half a million GPUs at any given moment. Most people have never heard of it.
Simon Mo, co-founder and CEO of @inferact and lead maintainer of vLLM, sits down with a16z’s Matt Bornstein and Elena Burger to discuss what it takes to actually run open models in production, the advantages of open models, Simon’s mission at Inferact, and more.
00:00 Intro
01:46 When open source became critical infrastructure
08:55 Day zero model releases, and the drama behind them
14:59 What Kimi K3 actually buys you
18:56 Why open model licenses are changing
22:24 The pharmaceutical analogy for funding model training
26:16 If GPUs got 99% cheaper
29:08 Why vLLM, OpenRouter, and Ollama all started before ChatGPT
35:42 Building a company on an open source project
39:48 The inventor of RoPE removing RoPE
@simon_mo_ @BornsteinMatt @VirtualElena(a16z)
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink