On October 6, 2026, OpenAI and Ironclad published the results of a research evaluation of AI agents working with contract processes in enterprise software. Across 11 tasks, GPT-6 Astra earned an average score of 55.0% against the completion criteria, while GPT-5.6 Sol scored 41.6%.

The tasks covered legal, commercial, and procurement work. Examples included configuring a nondisclosure agreement, creating procurement approvals, and modifying a standard legal clause for a selected jurisdiction. OpenAI describes the study’s goal as testing whether an agent can complete a multistep process while preserving business rules and checking its output against the original requirements.
How performance was measured
Ironclad employees and people who use Ironclad at OpenAI helped define the 11 tasks. Each task was assessed against 8–50 criteria, depending on its complexity. Ironclad provided hosted environments of its software for model practice; OpenAI says it improved the models using synthetic tasks and reinforcement learning.
The comparison used the settings at which each model achieved its highest score: Max reasoning for Astra and High reasoning for Sol. According to OpenAI’s published table, the average estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. These are simulated estimates based on assumed processing and generation speeds, not measured reductions in customer time.
What the results show
The evaluation shows one way to test a computer agent by examining the final state of a complex workflow: a set of criteria can reveal which requirements the system fulfilled and which it missed. For companies evaluating similar tools, a practical guide is to check not just individual interface actions but also how a configured process handles different conditions, such as approval thresholds and exceptions.
The figures relate to 11 research tasks and do not describe all Ironclad workflows. In a separate example in the publication, Astra completed about 94% of the criteria in an estimated 20 minutes, while Sol completed about 85% in 32 minutes; this is an example of one task, not an average. OpenAI also says the synthetic tasks were created from public contracts in the SEC EDGAR database after personal information was filtered out; according to the company, confidential customer contracts were not used for training or evaluation.
Sunita Verma, Ironclad’s chief technology officer, emphasized the importance of the end-to-end workflow. In translation, her quote from OpenAI’s publication reads: “Agents need to understand the entire contract lifecycle, including how workflows connect, while preserving the control teams rely on.”
OpenAI describes the collaboration as research into training and evaluating models. The published scores offer a picture of performance on selected tasks; the publication does not report an independent audit of the evaluation or measurements across a broader set of customer workflows.
Translated quote: “Agents need to understand the entire contract lifecycle, including how workflows connect, while preserving the control teams rely on.”