GPT-6 Astra arrived on 7 September 2026, and OpenAI has sold it on something other than better prose. The company calls it state of the art on computer use, browsing and professional work, and the demonstrations show a model filling in forms, updating records in a customer system and drafting inside a document editor rather than answering questions in a chat window. For a firm that has spent two years learning to check what a model writes, the question has moved.
The number to hold on to is 59.3 per cent. That is what Astra scored on Agents' Last Exam, a benchmark of complex professional tasks carried out in real software, against 55.5 per cent for Claude Opus 5 and 53.6 per cent for the OpenAI model it replaces. It leads the comparison OpenAI published, and it still means that four attempts in ten at a whole professional task fall short. Something that finishes six jobs in ten on its own is worth having. It is not something you point at a completion and walk away from.
What improved, and what cuts the other way
The gains are real and they are measured. OpenAI's own hallucination benchmark puts Astra at 4.2 per cent against 12.2 per cent for the previous flagship, which is a marked fall in the rate at which the model asserts things that are not so. The alignment result reads better still. OpenAI built a test, informed by the Hugging Face incident, of whether a model given a hard or impossible task will step outside the scope it was set. Run without production safeguards, the older model went beyond the authorised target in 48 per cent of cases. Astra did so in none.
Two findings in the same announcement pull the other way, and OpenAI states both plainly. Astra meets the Critical threshold for cybersecurity under the company's own preparedness framework, having scored 100 per cent on a benchmark that turns known software vulnerabilities into working exploits. Its written reasoning is also harder to monitor than its predecessor's, because it now solves problems in fewer written steps. A model that shows less of its working is one you have to judge on output alone, which puts more weight on whoever reads it.
Computer use changes what you are approving
Every model release so far has put the same question to your firm, which is whether the words it produced are right. A model built to operate software puts a different one, which is what it is allowed to reach. Filling in a form, updating a matter in your case management system and sending a draft out of your document editor are actions inside your systems rather than text on a screen. An action taken in error does not sit quietly on a page waiting for a fee earner to notice it.
The National Cyber Security Centre made this point in its agentic AI advice in August and the practical answer has not changed since. Work out what the tool may read, what it may write to, and what waits for a person to press the button. Your client ledger, your case management system and anything that leaves the building in the firm's name belong in the last of those. Nothing in the benchmark tables alters that, because a score of 59.3 per cent tells you the tool will act wrongly on a real task often enough to matter.
What to settle before anyone turns it on
Astra reaches ChatGPT Plus, Pro, Business and Enterprise accounts over the coming days, and the OpenAI API, Microsoft Azure and Amazon Bedrock alongside them. If your firm runs a Business or Enterprise workspace, access is off by default at launch and an administrator has to switch it on. That gives you a short window in which to set the terms before anyone in the firm uses it, and windows of that kind close quietly.
Start with whether the connectors that let a model act inside your systems are switched on at all, since a tool is only as contained as the access you granted it. Settle who may use the model on client work and on which matters. Then say what gets recorded, because a supervisor who signs off work produced by an agent has to be able to say what the agent did and what they checked themselves. Neither the SRA warning notice on the misuse of AI nor the disciplinary tribunal in its first case on AI citations will accept that the software looked dependable.
The firms that handle this well will be the ones that decide access before the capability lands, rather than after a fee earner has connected a model to the practice management system to see what it does.
OpenAI has set out the release, the full benchmark tables and its safety findings on its own page for GPT-6 Astra, which is open to read without registration, and Artificial Lawyer collected the early views from Harvey and Legora on 7 September 2026.
If you want the access question settled in writing before anybody in your firm switches the model on, that is an afternoon: talk it through with us.
