Photo by Jakub Żerdzicki on Unsplash. Source: https://unsplash.com/photos/a-person-pointing-at-a-calculator-on-a-desk-eGI0aGwuE-A (Unsplash License).

Executive Summary

Price and quality in model API pricing have stopped moving together. LiveBench publishes a cost per successful task alongside its capability scores, and on that measure DeepSeek V4.1 Flash costs 2.9 cents against 1.21 dollars for Claude Fable 5.1, for an overall score difference of 2.3 points. The cheapest model scoring above 80 overall is DeepSeek Flash on off-peak rates at 2.9 cents. The most expensive strong model is 42 times that.

The top performer is a different answer, and it now costs less than it used to. Claude Opus 5.5 leads the Artificial Analysis Intelligence Index at 58, and at 4 dollars per million input tokens and 20 dollars output it undercuts the previous Opus 5 at 5 and 25 while scoring higher. Two traps sit in the cheap end. Several budget tiers carry published expiry dates, and the single cheapest line in the market trades a discount for the right to train on your prompts.

Every price below was read from the vendor’s own pricing page on 5 October 2026. Prices in this market change monthly, so the date matters as much as the number.

Table comparing monthly cost and LiveBench score across DeepSeek Flash, Gemini 3.8 Flash, Grok 4.6, Claude Sonnet 5.5 and Claude Opus 5 for a fixed token workload, showing a 42 times cost spread.
A 42 times spread in price for a few points of measured capability.

Best value is DeepSeek Flash off-peak, and the arithmetic is not close

Take a common agent workload using 10 million input tokens and 2.5 million output tokens a month, which is roughly a four to one ratio. At DeepSeek’s published off-peak rates of 0.15 and 0.60 dollars per million, that workload costs 3 dollars. On peak rates for the same model it costs 6. GPT-6.1 Sol at 2 dollars input and 10 dollars output costs 45 dollars for the same tokens, roughly 15 times more.

Cost per token is only half the picture, because a weaker model needs more retries and produces more tokens to reach the same answer. LiveBench’s cost per successful task already bakes that in, and it is where the case gets strong. DeepSeek V4.1 Flash scores 81.1 overall at 2.9 cents per successful task, the cheapest figure of any model above 80 on that board. GPT-6.1 Sol scores 81.6 at 14.2 cents, nearly five times the cost for half a point. It is also the strongest model on the board for agentic coding at 77.3, ahead of every frontier model listed.

Two honest caveats. The off-peak discount is genuinely a scheduling lever, since peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, but third-party analysis found the off-peak rates sat above the flat rate DeepSeek replaced. And Meta’s nominally cheapest tier, Muse Spark Contributor at 0.10 and 0.20 dollars, pays for that discount with permission to train on your data, which disqualifies it for many buyers.

Best raw performance is Claude Opus 5.5, and the leader depends on the task

On raw capability the answer is Claude Opus 5.5, which tops the Intelligence Index at 58, two clear of Claude Sonnet 5.5 at 56. It also undercuts its own predecessor, coming in at 4 and 20 dollars against 5 and 25 for Opus 5, while beating it on the index. Paying frontier prices no longer means paying the highest price.

Beyond that, the leader changes with the job. Claude 5.5 Opus leads coding at 89.3 on LiveBench and mathematics at 97.1. Gemini 4 Argon leads human preference on the LMArena leaderboard at 1525 Elo. Gemini 3.8 Flash leads instruction following at 81.4. DeepSeek Flash leads agentic coding. Long context is effectively tied, with Google, OpenAI, Alibaba and Qwen all publishing windows at or near a million tokens.

The pattern worth internalising is that the capability gap at the top has narrowed to a few points while the price gap has widened to a factor of 42. The choice is less about which model is smartest and more about which failure mode you can tolerate at which price.

Three questions for a build decision. Does your workload run in DeepSeek’s off-peak window, and if not, what is the peak surcharge? Which of your tiers carries an expiry date, since Google’s Flash rates double on 1 January 2027? And if the strongest model is only a few points ahead, what does your evaluation say about your own task rather than a public leaderboard?

For the wider cost picture, see why GPU rental prices are rising and why one lab bought inference instead of building it.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.