OpenAI President Greg Brockman described the release on X as the beginning of the “AGI era” (source: Greg Brockman’s post on X). That claim deserves the usual skepticism: what matters more is what Astra actually does in production, what it costs, and how it behaves under a European regulatory framework that is no longer waiting around.
The launch was not smooth. OpenAI’s rollout updates refer to a July 2026 safety incident that prompted a delay and additional production-alignment controls (source: OpenAI rollout updates); OpenAI has not publicly confirmed details of the incident beyond that acknowledgement. The model then went through the company’s Preparedness Framework. In its Preparedness Framework summary (source: OpenAI Preparedness Framework), OpenAI assigned Astra its first “Critical” cyber-risk tier classification. The framework text OpenAI cites scopes that tier to models that can “autonomously perform sophisticated cyber operations, including discovering and exploiting zero-day vulnerabilities in real-world targets” within the framework’s controlled evaluation environment. OpenAI has not stated that Astra can do this outside that supervised evaluation scope.
Not a chatbot: a computer-use model with a 1.05-million-token context
GPT-6 Astra succeeds GPT-5.6 Sol as OpenAI’s flagship reasoning model, but its design goal is different. Astra operates software through UI elements — pixels, mouse, keyboard — rather than only through text output. OpenAI positions it for software engineering, complex math, research, and multi-step workflows that run autonomously inside a browser or a virtual machine.
The numbers behind that positioning, as cited by OpenAI unless otherwise noted:
- Context window: 1.05 million tokens, according to OpenAI’s model specifications.
- Maximum output: 128,000 tokens, according to OpenAI’s model specifications.
- Average task execution: about 40 minutes on computer-use workflows, a vendor-reported estimate from OpenAI.
- Speed claim: roughly 47% faster than GPT-5.6 Sol, which OpenAI says averaged about 75 minutes per task; both figures are vendor-reported estimates.
That execution-time figure is worth pausing on. A single agent run can take the better part of an hour. That is not a search query; it is a digital coworker working a shift. OpenAI is betting that companies will pay flagship prices for that endurance.
Benchmarks: impressive, with caveats
OpenAI published benchmark results that, on paper, look like a clean sweep. The figures below are reported scores from OpenAI and Artificial Analysis, with the source identified for each row; treat them as directional until independent evaluation catches up.
| Benchmark | Reported score | Source |
|---|---|---|
| ExploitBench | 100% | OpenAI |
| ARC-AGI-3 (via provider harness) | 99.9% | OpenAI |
| FrontierMath Tier 4 (v2) | 97.6% – 98% | OpenAI |
| OSWorld 2.0 (computer use) | 72.6% | OpenAI |
| Terminal-Bench Science 0.1 | 64.6% | OpenAI |
| Humanity’s Last Exam (with tools) | 57.2% | OpenAI |
| Artificial Analysis Intelligence Index | 61.2 | Artificial Analysis |
The disparity between the scores is itself informative. Astra records a 100% ExploitBench result in OpenAI’s reported evaluation — a benchmark result that should not be treated as proof of reliable real-world vulnerability discovery — but scores 72.6% on OSWorld 2.0, the computer-use benchmark that measures whether an agent can actually operate a desktop environment. The model appears exceptionally strong at narrow, high-difficulty reasoning and still fallible at ordinary computer chores. The “Critical” risk label and the imperfect OSWorld score together tell the real story: this is a powerful tool that needs supervision, not an autonomous coworker to be left alone with your production systems.
Pricing: flagship reasoning still costs flagship money
GPT-6 Astra API pricing is $10 per 1 million input tokens and $50 per 1 million output tokens. OpenAI’s published pricing lists GPT-5.6 Sol at $4 per 1 million input tokens and $20 per 1 million output tokens; Astra is therefore 2.5× more expensive than its predecessor on both rates, rather than Sol costing 2.5× less. Anthropic’s published pricing should be checked separately; the available information does not establish that the $50 output rate is exactly equal to Anthropic’s rate, nor is the launch-date comparison needed for this price analysis.
In euros, at roughly $1.10 per euro, that is about €9.10 per million input tokens and €45.50 per million output tokens. An illustrative token-only calculation — exactly one million input tokens plus the maximum 128,000 output tokens — costs $16.40, or roughly €14.90, before any retries. That is not a typical run: it assumes the full one-million-token context is used and the model produces the maximum 128,000 output tokens. More importantly, this calculation does not establish the typical cost of a 40-minute agent run; real computer-use sessions may involve repeated calls and variable token consumption.
The market has split into two speeds. Public vendor pricing varies widely across model families and service tiers, while OpenAI’s flagship agentic pricing reaches $50 per million output tokens. European teams building agentic workflows should compare current vendor pricing directly rather than assume that competitor launch dates or rates are equivalent. They should also ask whether they genuinely need frontier-grade autonomy or just reliable automation — the price difference can be substantial.
The European angle: arriving just as the AI Act starts to bite
Timing matters for European adopters. The EU AI Act’s GPAI obligations have applied since 2 August 2025. The Act’s broader application and the Commission’s enforcement powers for GPAI provisions are tied to 2 August 2026, although some provisions have separate transition periods. Annex III high-risk obligations are not scheduled for 2 December 2027: for high-risk systems classified under Annex III, the relevant obligations generally apply from 2 August 2026. The separate 2 August 2027 date concerns certain high-risk AI systems embedded in regulated products under Article 6(1), not Annex III systems generally.
GPT-6 Astra therefore lands after the August 2025 GPAI obligations began and around the August 2026 GPAI enforcement milestone. For an application-level system, whether Annex III high-risk obligations apply depends on its use and deployment context; the model’s GPAI enforcement timetable is a separate legal question from the obligations applying to a high-risk application.
A general computer-use model is not automatically “high risk” under Annex III — the Act largely classifies based on deployment context. But Astra’s classification profile makes the question practical rather than academic. A model that OpenAI places in the Critical cyber-risk tier, deployed in finance, critical infrastructure, or public administration, changes the risk calculation for European companies.
The availability picture for European users is more reassuring. According to OpenAI’s rollout announcement, Astra is beginning a staged rollout across ChatGPT plan tiers, the OpenAI API, Microsoft Azure, and AWS Bedrock. OpenAI has not announced a blanket EU-specific restriction, but actual access will depend on the specific OpenAI product, Azure region, or Bedrock region used.
The Azure and Bedrock routes matter for European enterprises because both clouds offer EU data-residency regions — but only if the deployment is configured in those regions. This is a relevant GDPR consideration for companies that cannot send data to US-hosted endpoints. As with previous OpenAI launches, regional availability should be checked before building production pipelines, and nothing in the announcement suggests a blanket European geo-block.
Bottom line for European teams
Astra is a serious engineering release wrapped in an even bigger narrative. If you need state-of-the-art autonomous computer use, the API price is defensible only when measured against the hourly cost of a human doing the same multi-step work. But European companies should do the math in euros before committing: €45 per million output tokens makes every long agentic session a budget line, not an experiment.
OpenAI itself says this model needs safeguards at the “Critical” risk level. European buyers should treat Astra the same way — with clear human oversight, restricted permissions, and a hard look at whether their deployment falls into AI Act high-risk territory. For many internal workflows, an open-weight model hosted on European infrastructure remains a cheaper and more predictable alternative. We are already running those comparisons on our AI Arena rig at ai-jarvis.eu/ai-arena, and this release gives us a much higher bar to measure against.
Is GPT-6 Astra available in the EU?
OpenAI is rolling Astra out through ChatGPT plan tiers, the OpenAI API, Microsoft Azure, and AWS Bedrock, and has not announced a blanket EU-specific restriction in its rollout announcement. Actual availability depends on the OpenAI product, Azure region, or Bedrock region used. European companies should verify regional availability before building production workflows, as is standard for OpenAI launches.
How much does an agentic GPT-6 Astra run actually cost?
At $10 per million input and $50 per million output tokens, an illustrative token-only calculation that assumes exactly one million input tokens and 128,000 output tokens costs $16.40 before retries — roughly €14.90 at current exchange rates. That is not a typical run; real computer-use sessions may involve repeated calls. Because OpenAI reports Astra averages about 40 minutes per computer-use task, a complex job with multiple attempts will cost substantially more than that illustrative figure.
What does the “Critical” cybersecurity risk classification mean in practice?
OpenAI’s Preparedness Framework classified Astra as Critical because of its demonstrated behaviour in controlled evaluations. For European businesses this means strict access controls, human oversight, and careful evaluation of the deployment context, which also determines whether EU AI Act high-risk obligations apply.