Reasoning
Frontier model inference costs fell ~12x from GPT-4 (2023) to o1-preview (2024) per token; at 2025-2026 rates, continued 4-6x annual hardware and algorithmic efficiency gains imply cumulative 20-40x cost reduction by end-2028, exceeding 95% decline from 2025 baselines. Open-weight releases (Llama-3-405B, DeepSeek-V3) already show 70-85% cheaper inference than proprietary equivalents, accelerating price compression. Energy and capex constraints could slow deployment, but recent TSMC 2 nm and Blackwell ramp data indicate supply growth outpacing demand through 2027.Key uncertainty
Whether US export controls on advanced GPUs to China materially restrict global chip supply and slow the observed efficiency curve.