Two labs cut inference prices on the same night, and the cache line did more work than the headline rate.
- OpenAI halved GPT-6 Sol to 2 dollars per million input tokens and 10 dollars per million output. GPT-6 Luna fell to 10 cents and 50 cents.
- Anthropic’s Claude Opus 5.5 landed the same day at 4 dollars and 20 dollars per million, with cache reads down 60 percent to 20 cents.
- OpenAI prices cached input at a tenth of the uncached rate, so a repeated prefix is the cheapest token on the card.
- Artificial Analysis measured the result per task. GPT-6 Sol at maximum effort costs 1.06 dollars against 1.99 dollars for GPT-5.6 Sol, and Luna drops to 7 cents from 18.
- vLLM’s tiered KV cache offloading explains the floor. Past 128 concurrent conversations, a storage-backed cache more than doubles throughput over a full recompute.
Read the full analysis. Why the bill now lives in the cache.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
