GPT-6 Astra Coding Benchmarks Show Gains, but Anthropic Lead Is Not Clear-Cut
OpenAI reported strong agentic coding scores for GPT-6 Astra on DeepSWE and Terminal-Bench, but overlapping uncertainty ranges and rival Anthropic results complicate claims of a definitive lead.
OpenAI's headline numbers
OpenAI reported that GPT-6 Astra scored 74.1% on DeepSWE v1.1, a 113-task agentic coding benchmark, compared with 70.8% for GPT-5.6 Sol. On Terminal-Bench 4.0, which tests command-line research tasks across scientific fields, Astra reached 57.7% versus 37.3% for Sol. OpenAI also cited 64.6% on Terminal-Bench Science against Anthropic's reported 52.6% for Claude Fable 5.1.
The company highlighted these results as evidence that Astra is its best model for software engineering, with stronger performance on complex tasks in real codebases. FrontierCode 1.1 Extended and an internal database-migration benchmark also showed double-digit percentage gains over the prior generation.
Where the leaderboard tells a different story
The public DeepSWE leaderboard currently places Gemini 3.8 Flash and Claude Opus 5 at 74%, with Sol at 73%. Reported uncertainty ranges overlap across these results, meaning Astra's 74.1% does not establish a clear leader. OpenAI's comparison chart excluded Meta's Muse and used a 67.4% Fable 5.1 result, making Astra's advantage appear larger than the broader set of published scores suggests.
Independent benchmarking firm Artificial Analysis assigned Astra an Intelligence Index of 61, matching GPT-5.6 Sol exactly and placing it five points behind Claude Fable 5.1 at 66. In the Coding Agent Index, Astra reached 67 points in the Codex harness, roughly on par with Claude Opus 5, but Fable 5.1 in Claude Code still leads at 70 points.
OSWorld and harness differences matter
On OSWorld V2-Offline, OpenAI reported 72.6% for Astra against 65.7% for Sol, with task time falling from about 75 to 40 minutes. Anthropic has reported 77.9% for Fable 5.1, but says that result used a different OSWorld release and should not be compared with previously published scores. Benchmark comparisons in agentic coding increasingly depend as much on the agent harness and system configuration as on the underlying model.
OpenAI's 98.6% score on ARC-AGI-3 similarly reflects both the model and the Responses API harness that retains reasoning between turns. The company has previously demonstrated that system choices can substantially raise ARC-AGI-3 scores without changing the underlying model weights.
Enterprise buyers weigh benchmarks against cost
For developers and companies evaluating frontier models, the picture is nuanced. In agentic coding, GPT-6 Astra delivers competitive performance, but Anthropic retains an edge on broader intelligence indices and in Claude Code harness results. OpenAI trained Astra in its largest run to date, using more than 100,000 GPUs at the Stargate data centre in Texas.
Both companies price their flagship models at $10 and $50 per million input and output tokens, but Anthropic's 75% cut to cache read pricing on Fable 5.1 can reduce effective cost by 25% on typical workloads and up to 45% on heavily agentic tasks. Benchmark leadership and bill size are therefore related but not identical questions.
Sources & References
Editorial Team
Editorial
In-house writers and editors producing original explainers, guides, and analysis. Articles cite authoritative public sources where helpful.