Skip to main content

GPT-6 Astra is out with 72.6% on OSWorld 2.0 — now the EU AI Act sets the real test

Ilustrační obrázek
GPT-6 Astra is here. OpenAI’s next flagship launched on September 3, 2026 with agent-ready skills for computer use and coding. The vendor-reported numbers are strong — 72.6% on OSWorld 2.0 and a lower internal computer-use overstep rate than its predecessor. The harder question for European teams is whether the pricing and the shifting EU AI Act timetable make it a safe production choice.

On September 3, 2026, OpenAI released GPT-6 Astra as a limited preview, the successor to the GPT-5.6 generation that included its “Sol” flagship. The consumer rollout targets ChatGPT Plus, Pro, Business and Enterprise subscribers. Developer access is through the OpenAI API and Microsoft Azure, as the Deccan Herald report details. Some reports also mention AWS Bedrock availability, but we could not verify that at publication.

OpenAI has framed the release in AGI-era language, but that is vendor framing, not an independent assessment. Impressive benchmark scores are not the same as general intelligence.

What GPT-6 Astra was actually built for

Astra is not a chat model with a new coat of paint. Its design targets the messy jobs where previous GPT models struggled:

  • Autonomous computer use — navigating interfaces, filling forms, operating software across windows.
  • Software engineering — long, multi-file coding tasks that require planning, not just snippet generation.
  • Multi-step workflows — chaining dozens of actions with verification between steps.
  • Scientific analysis — research reasoning across chemistry, physics and mathematics.
  • Cybersecurity — defensive analysis and controlled security testing.

OpenAI has not published full training details for Astra at the time of writing. The widely circulated 100,000-GPU training figure is unverified, so we are not treating it as a confirmed specification.

The benchmark table that matters

The figures reported by OpenAI put Astra ahead of its own predecessor on several agentic tasks. They are vendor-reported, not independently reproduced by us or by an external evaluator, so treat them as provisional.

BenchmarkGPT-6 AstraGPT-5.6 Sol
OSWorld 2.0 (computer use)72.6% in ~40 min65.7% in ~75 min
ScreenSpot-Pro (UI grounding)92.7%not disclosed
Agents’ Last Exam (long-horizon agent tasks)59.3%not disclosed
AutomationBench (process automation)41.4%not disclosed
Internal computer-use safety overstep2.4%22.0%

Read OSWorld 2.0 carefully: Astra did not simply score higher; it finished in about 40 minutes versus Sol’s 75 minutes, approximately 47% less time. That is a meaningful efficiency gain, but it is still a vendor-reported result.

Safety: a lower overstep rate, not zero

The biggest reported change is in safety monitoring. In OpenAI’s internal computer-use evaluation, Astra overstepped its authorized target in 2.4% of cases, compared with 22.0% for Sol. That is a substantial reduction, but it is not the zero-overstep result that was initially circulating.

That is closer to a real production safety mechanism than “we added a system prompt.” For a model designed to click around a computer on its own, the difference between 22.0% and 2.4% may matter operationally, but it is not a guarantee of safety — and the updated figure still needs external replication.

We have not yet put Astra through our own pipelines at AI Arena, where we benchmark local models on an RTX 5060 Ti 16 GB rig. Astra is a closed cloud model, so it will never run on our hardware — cloud flagship claims have to be judged by published benchmarks and independent replication.

Pricing: flagship tokens, flagship bills

Astra is priced at $10 per million input tokens and $50 per million output tokens. Audio input costs $0.06 per minute and audio output $0.24 per minute.

Let’s do the arithmetic a developer will actually meet. A typical agentic code-review task consuming 60,000 input tokens and 8,000 output tokens costs:

$10 × 0.06 + $50 × 0.008 = $0.60 + $0.40 = $1.00 per task before volume discounts. A team running 100 such tasks a day spends $100 — and that is a realistic workload, not an exaggerated one.

OpenAI also describes efficiency gains in complex workflows, but those figures were not independently verified at publication. Token efficiency may partially offset the high price — but only partially, and the output-token cost remains the budget line to watch.

The European reality check

For EU companies, the cost question is not only about dollars per token. Astra lands in a regulatory environment with obligations already in force and more phasing in:

  • General-Purpose AI obligations under the EU AI Act have been enforceable since August 2025.
  • High-risk AI system rules (Annex III): the 2026 Digital Omnibus was adopted and moved the standalone high-risk deadline from 2 August 2026 to 2 December 2027. This deadline is separate from the GPAI obligations that already apply.

In practice, European deployers need legal review: using Astra for autonomous computer use in areas such as recruitment, credit decisions or critical infrastructure can trigger high-risk obligations once the relevant Annex III deadlines apply — technical documentation, transparency to affected people, human oversight and conformity assessments. The Annex III standalone high-risk deadline is now 2 December 2027, but the general-purpose AI obligations are already in force; the old comfort of “we are still in a grace period” is no longer safe.

There is also a genuine procurement angle. Astra is available through Microsoft Azure. Azure regional availability may support an EU data-residency arrangement, but it does not automatically guarantee that every processing operation remains in the EU. Buyers should verify the exact service, region, retention and data-processing terms before treating the deployment as EU-data-resident. But the model itself remains closed and proprietary: no local weights, no on-premise instance, no way to run it inside your own VPC on your own GPUs.

What European teams should actually decide

GPT-6 Astra is technically impressive, and the vendor-reported 2.4% internal safety overstep rate is a more credible engineering signal than the earlier zero-overstep claim. But European buyers should evaluate it the same way they would evaluate any expensive new hire: check references, run it on their own workflows, and calculate what happens when a task goes off the rails. The EU AI Act timetable is in motion, and compliance planning cannot wait for the final legal date.

Is GPT-6 Astra actually AGI?

No. OpenAI has used AGI-era language around the release, but that is vendor framing, not an independent technical assessment. The model still operates within defined benchmarks and uses automated monitoring to prevent unauthorized actions. Those are properties of a very capable, carefully constrained system — not general intelligence.

Can we run GPT-6 Astra on our own servers in the EU?

No. Astra is a proprietary cloud model offered through ChatGPT and the OpenAI API, with Microsoft Azure integration. Reported AWS Bedrock availability is not confirmed at publication. There are no open weights for local or on-premise deployment. Organisations that need full data control should look at open-weight alternatives, which typically cost far less per token but do not match Astra’s reported agentic benchmark scores.

How much does a real Astra task cost in euros?

In a realistic 60,000-input / 8,000-output token task, the API cost is $1.00. As an illustrative exchange-rate assumption, that is around €0.85–0.95 before VAT; a different exchange rate changes the euro figure directly. VAT is a separate addition and varies by EU member state. A company running 100 such tasks per day should budget around $100 per day — roughly €2,500 per month at that illustrative rate — before any volume discounts or tax considerations.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.