What happened. OpenAI reported on 6 October that GPT-6 Astra averaged 55% across the requirements in 11 Ironclad research tasks, compared with 41.6% for GPT-5.6 Sol.

What it actually means. That is a rubric score, not a claim that 55% of contracts were handled correctly. Procurement should test the finished workflow: the approval above the threshold, the exception below it and the record left behind. A convincing sequence of clicks can still produce the wrong business process.

What to watch. Whether evaluations start reporting critical control failures separately from average scores. These results cover research tasks, and the reported time estimates are simulated, not measured customer savings.

Put this to work.