OpenAI’s new model is built to use a computer, not to chat

The benchmark that says most about this release is not a maths score. It is 72.6% on a computer-use test that the previous model passed at 65.7% — while taking about 47% less time per task, roughly forty minutes instead of seventy-five. That is a procurement number, not a research one.
What shipped, and where
OpenAI released GPT-6 Astra on 3 September 2026, first to a limited set of organisations and then to ChatGPT Plus, Pro, Business and Enterprise users. For developers it is gpt-6-astra in the OpenAI API, and it is also available through Microsoft Azure and Amazon Bedrock.
One line matters for anyone running a workspace: enterprise access is off by default at launch, and an administrator has to turn it on.
What the new model costs to run
$10 per million input tokens and $50 per million output tokens at standard rates, with separate rates for cache reads and writes. A Fast mode in the API runs up to twice the speed at twice the price. Usage is included in existing subscription allowances, with credits available on top.
What it is actually for
Long tasks in real software, rather than answers in a chat window. OpenAI reports 59.3% on Agents’ Last Exam — professional work in real applications, from financial modelling to media production — against 55.5% for Claude Opus 5 and 53.6% for its own previous model, while using about 65% fewer output tokens. On BenchCAD, which asks a model to rebuild a 3D object as CAD code, it reports 95.9% at an estimated API cost roughly 43% below its predecessor.
The company also claims a 1.9× faster completion of browser tasks when the model is paired with an updated Codex harness.
Where it is not first
Worth reading before anyone quotes a clean sweep. On Humanity’s Last Exam with tools, Astra scores 57.2% — below Claude Fable 5.1 at 65.0% and below two other Claude models in the same table. On the Artificial Analysis Intelligence Index it scores 61.2 against Fable 5.1’s 65.7. The gains are concentrated in computer use, coding and science, and OpenAI’s own tables say so.
The security clause
Astra is the first model OpenAI has classified as meeting the Critical threshold in cybersecurity under its Preparedness Framework. It scored 100% on ExploitBench, and during an internal evaluation it found and used two previously unknown vulnerabilities, which OpenAI says it is disclosing to the maintainers. The shipped model refuses the more advanced offensive tasks, and extra safety checks can pause or stop a job mid-run — in the API, the task simply stops.
For a business the useful frame is not «is it smarter» but «what does an hour of it cost, and who signs off». Three things follow from this release. First, the price is a specialist’s: at $50 per million output tokens nobody should be routing a support inbox through it, and the efficiency claims — fewer output tokens, less wall-clock time per task — are the argument for using it on long agentic work where a cheaper model burns the saving in retries. Second, the availability path runs through Azure and Bedrock, which for most regulated buyers is the only path that clears procurement at all; the API listing is the easy part. Third, plan for the interruptions: OpenAI is explicit that the new safety checks can pause or stop legitimate work, and that in the API a stopped task does not resume. Any workflow you build on this needs to survive being halted halfway, which is a design constraint rather than a footnote.
Source: OpenAI checked against the source
More Dubai news

Six clocks, one AI-designed drug
Six proteomic aging clocks read a lower biological age in patients given rentosertib, an AI-designed drug. The authors say the reading is not yet clean.

Marked text, private detector
Anthropic began watermarking Claude's output on 2 August. The mark is in the model's word choices, and the tool that reads it is in private preview.

2,100 jobs, and no denominator
The Real Estate Empowerment Programme created more than 2,100 jobs for Emiratis in Dubai property between 2023 and 2026 — about fifty a month. The release gives no baseline to measure them against.