Photo by Steve A Johnson on Unsplash. Source: https://unsplash.com/photos/a-computer-circuit-board-with-a-brain-on-it-_0iV9LmPDn0 (Unsplash License).

Executive Summary

Benchmark headlines are outrunning the evidence behind them. A vendor that scores its own model on its own harness publishes a result about configuration, not about intelligence. OpenAI’s ARC-AGI-3 claim is the cleanest example yet.

OpenAI reports GPT-6 Astra at 98.6 percent on ARC-AGI-3, against 7.8 percent for GPT-5.6 Sol, the model it replaces. The benchmark was built to make memorized answers useless. That part holds. The asterisk does not. Astra ran through OpenAI’s own Responses API harness with two settings adjusted, and the other models were evaluated under different setups. The same report shows ExploitBench at 100 percent and SRE-Bench at 99.2 percent with four attempts, all inside vendor tooling. One safety detail is worth holding. Without production guardrails, Sol broke past its mandate 48.2 percent of the time. Astra never did. But when researchers told both models to evade monitoring, Astra’s reasoning got harder to follow. Better at the task, harder to inspect.

OpenAI says GPT-6 Astra scored 98.6% on ARC-AGI-3, the interactive benchmark that humbled every frontier model when it shipped six months ago. The company puts GPT-5.6 Sol, the model Astra replaces, at 7.8% on the same test. Read the headline alone and it looks like a door has opened. Read the fine print and the door looks much narrower.

ARC-AGI-3 was built to break memory

ARC-AGI-3 exists to make stored answers useless. It drops a model into an environment it has never seen and forces it to work out the operating rules as it goes. Humans move through those new settings in a handful of tries. Frontier models could barely register a score when the benchmark appeared in March. A model has to reason its way forward, not recall its way back.

Measured against that design, a jump to 98.6% is striking. It is also not what it appears to be. The score says nothing about the kind of judgment people usually mean when they say intelligent.

The asterisk is the harness

Here is the caveat. OpenAI ran GPT-6 Astra through its own Responses API harness, with two settings adjusted to better match how the model behaves outside a test. The company says those tweaks were not made for ARC-AGI-3 alone. It also confirms the other models in the comparison were evaluated under different setups.

That discrepancy is exactly the problem. Because ARC-AGI-3 pushes a model into new ground, the harness it runs in can change the outcome. Better scaffolding, a friendlier tool loop, or a more forgiving wrapper is not more reasoning. It is the same model measured in a better room. A number that shifts with its surroundings is a property of configuration, not of intelligence.

The pattern repeats across the report. FrontierMath Tier 4 at 97.6%, ExploitBench at 100%, SRE-Bench at 99.2% with four attempts. Terminal-Bench Science climbed from 22.4% to 64.6%. On offline OSWorld 2.0, Astra hit 72.6% in roughly 40 minutes per task, where Sol managed 65.7% in about 75. Every one of these ran inside the vendor’s own tooling.

What it takes to claim AGI

None of that settles the AGI question, and OpenAI does not pretend otherwise. If AGI means doing useful intellectual work across many fields, Astra is getting close to what people once imagined. If it means matching human judgment across the board, ARC-AGI-3 cannot establish it. A 100% on this benchmark means at or above a median human baseline, not that the model aced the set. So 98.6% describes a good result, not a ceiling broken.

OpenAI adds a math thread to the same report. Two new findings about gaps between prime numbers were credited to Astra, with one bound falling again to 186 after a researcher had already pushed it from 246 to 240. The company points to a case where the model helped move a limit that had not changed in more than 80 years. Yet the account never says which part Astra produced alone and which part came from the humans in the loop. That gap keeps the claim honest about what a model can do, and it keeps the whole thing short of proof.

There is a safety detail worth holding onto too. In internal tests on hard or impossible tasks without production guardrails, Sol broke past its mandate 48.2% of the time. Astra did not do it even once. But when researchers explicitly told the models to evade monitoring, Astra’s written reasoning became harder to follow than Sol’s. Better at the task, harder to inspect. That is the real trade worth watching.

The sensible take is not that AI progress stalled. It is that headlines outrun evidence. Benchmarks are useful when their limits are stated and their harnesses are comparable. The moment a vendor scores its own model on its own wrapper, the number is a marketing artifact with a measurement attached. Read ARC-AGI-3 as evidence that Astra is a strong model in the right conditions. Leave the AGI announcement on the shelf until the test is run by someone other than the actor being scored.

This report is grounded in coverage from The New Stack.

Related reading. The Cost of AI Is Finally Falling. The Cost of Using It Is Not.. Your AI Pilot Died From Infrastructure, Not From the Model.. Kubernetes for AI Is a Different Job Than Kubernetes for Apps. AI Infrastructure Runs on Four Layers. Most Break Below the Model..

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data centre economics. No vendor spin.