Photo by Shubham Dhage on Unsplash. Source: https://unsplash.com/photos/a-group-of-cubes-that-are-on-a-black-surface-T9rKvI3N0NM (Unsplash License).

Two labs cut inference prices on the same night, and the cache line did more work than the headline rate.

  • OpenAI halved GPT-6 Sol to 2 dollars per million input tokens and 10 dollars per million output. GPT-6 Luna fell to 10 cents and 50 cents.
  • Anthropic’s Claude Opus 5.5 landed the same day at 4 dollars and 20 dollars per million, with cache reads down 60 percent to 20 cents.
  • OpenAI prices cached input at a tenth of the uncached rate, so a repeated prefix is the cheapest token on the card.
  • Artificial Analysis measured the result per task. GPT-6 Sol at maximum effort costs 1.06 dollars against 1.99 dollars for GPT-5.6 Sol, and Luna drops to 7 cents from 18.
  • vLLM’s tiered KV cache offloading explains the floor. Past 128 concurrent conversations, a storage-backed cache more than doubles throughput over a full recompute.

Read the full analysis. Why the bill now lives in the cache.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.