Reasoning
Current 1M+ context pricing sits at $0.60–$3.00 per million tokens (Gemini 1.5 Pro 1M, Claude 3 200K, GPT-4o 128K), while the 2023–2025 compound annual decline has averaged ~45% driven by HBM3e supply ramp and speculative decoding optimizations. 2026 supply schedules show TSMC 2 nm risk production in H2-2025 plus Samsung HBM4 sampling, projecting another 50–70% reduction in inference FLOPs/$ that historically translates to 55–65% price cuts when passed through to list rates.Key uncertainty
Whether hyperscale providers maintain 2025–26 gross-margin targets above 70% or accelerate price cuts to capture share.