AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How OpenAI’s Software-Based Agent Training Raises Questions For Ironclad on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra in hosted copies of contract-management platform Ironclad, using synthetic tasks based on public contracts. Astra met an average 55% of rubric criteria across 11 workflows; the reported time estimates were simulated, and the results do not establish customer productivity gains or readiness for unsupervised work.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra in hosted copies of contract-management software company Ironclad’s product, testing it on 11 legal, commercial and procurement workflows. The results show progress on complex software tasks, but Astra met an average of 55% of evaluation criteria, and OpenAI says its time figures are simulated estimates—not measured customer savings.

Ironclad staff and OpenAI employees who use the product selected 11 tasks, including creating nondisclosure agreements, configuring procurement approvals and revising a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes per task. The tasks were scored against rubrics containing 8 to 50 criteria, depending on complexity.

OpenAI reported that GPT-6 Astra, run at its maximum setting, met an average 55.0% of the criteria. GPT-5.6 Sol, run at high effort, averaged 41.6%, while an internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria. These are rubric scores across the research tasks, not percentages of tasks fully completed.

For training, Ironclad supplied hosted product copies where models could practise. OpenAI said it generated synthetic tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. The company stated that it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. OpenAI estimated Astra’s time per attempt at 19.2 minutes, compared with 37 minutes for Sol, but described those figures as simulated estimates based on assumed processing and generation speeds.

At a glance
reportWhen: Published October 6; reported results c…
The developmentOpenAI described a partnership with Ironclad to train and evaluate a frontier model on contract-work tasks inside the company’s software, and invited other software firms to participate in similar research.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Work Needs More Than a Score

The test addresses a practical question for businesses: can an AI agent carry out multi-step work inside specialized software while preserving company rules and controls? A model may complete many actions correctly and still make a consequential error if it misses one required approval or applies the wrong contract provision. In the procurement example described by OpenAI, a workflow could require Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing one condition could undermine the whole process.

That makes the reported 55% average criteria score a measure of progress, not evidence that the system is ready to run contract workflows without review. OpenAI’s account itself says human oversight remains important when an agent may lose track of a business rule. For buyers, the relevant question is not only how many criteria an agent meets on average, but which criteria it misses, how serious those failures are and whether a person can detect them before work is finalized.

The initiative also points to a possible change in how AI systems are developed: software companies could serve as training and testing environments for agents learning professional work. That may make products more useful, but it could also shift the value of software away from its screens and toward the rules, records, audit trails and controls behind them. That is an implication of the approach, not a result established by this test.

Amazon

AI contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Became a Test Environment

The October 6 post was titled “Advancing computer use with Ironclad.” The name refers to Ironclad, the contract-management software company, rather than a new AI framework. OpenAI described the project as training models to understand business rules, complete multi-step workflows in specialized software and check results against the original requirements.

OpenAI said it is inviting a small number of software companies to work on tasks that current agents cannot reliably complete. It asked prospective partners to bring a concrete example of an agent failure, people with expertise in the relevant work, a secure test environment and data that can safely be used for research. The stated aim is to study difficult real-world workflows, rather than only general computer-use tasks.

The source account attributes to Ironclad CTO Sunita Verma the point that agents must preserve “the controls teams rely on.” OpenAI’s post also presents the work as evidence that a full contracting platform remains important. Those statements describe the companies’ framing; they do not establish how the partnership will affect Ironclad’s product or its customers.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Cannot Yet Show

The published results cover 11 selected research tasks; they do not show how Astra performs across Ironclad workflows generally, across different customers or under live production conditions. The account does not provide task-by-task details of which rubric criteria were missed, how severe those failures were, or how often a human reviewer would catch them. An average score can conceal a failure in a mandatory control.

The 19.2-minute and 37-minute figures are simulated estimates, according to OpenAI, not observed completion times or measured savings for customers. The comparison therefore does not establish that Astra currently makes contract work faster. It also remains unclear how the reported results would change with different workflows, data or software configurations, and what review procedures would be required before any deployment.

OpenAI says it did not use non-public Ironclad customer data, but the source material does not specify all technical or contractual arrangements for the hosted test environments. Nor does it provide a deployment schedule, commercial terms or details about which other software companies may take part. These points should not be inferred from the research announcement.

Amazon

AI-powered NDA generator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Buyers and Partners Should Check

OpenAI’s stated next step is to invite a small group of software companies to propose difficult tasks, provide domain experts and supply secure environments and research-eligible data. The company has not named additional partners or announced a timetable in the source material. Further results would be needed to show whether agent performance improves on a wider range of workflows and whether those gains translate into dependable use.

Organizations considering agents in contract, finance or customer-record systems can ask vendors for criterion-level results, rather than a single average score; details on failed requirements; and evidence from testing under conditions similar to their own. They can also ask how human review, approval thresholds, audit records and access controls work. Until such information is available, this evaluation is best read as a research demonstration of training in specialized software—not proof of safe, hands-off automation or realized productivity gains.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s October 6 post describes training and testing a model inside hosted copies of Ironclad’s product; it is not announcing a framework called Ironclad.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the 11 tasks. It does not mean Astra completed 55% of the tasks, and it does not show that each workflow was safe or usable without review.

Did OpenAI show that the agent saves customers time?

No. OpenAI described the reported per-attempt times as simulated estimates based on assumed processing and generation speeds. They are not measured time savings from customer use.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Is Astra ready to run contract workflows on its own?

The reported results do not establish that. OpenAI’s account says human oversight remains important, and the average score does not reveal which individual requirements were missed. The evaluation is a limited research test, not evidence of readiness for unsupervised deployment.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Caseload Pipeline Can Help Restorative Justice Programs Manage Their Work

A proposed case pipeline would track restorative justice referrals, consent, scheduling and agreements; a small live-case test is suggested.

Data retention cleanup assistant for small law firms

A new data retention cleanup assistant aimed at small law firms is entering testing, focusing on managing legacy files and improving operational workflows.

EFF To Courts: Don’t Rewrite Copyright Over AI Hype

Electronic Frontier Foundation warns courts against broad copyright rulings fueled by AI hype, emphasizing the need for clear legal boundaries.

Musk’s xAI Challenges Minnesota’s AI “Nudification” Ban

A Bloomberg Law headline says xAI won a block on Minnesota’s AI “nudification” ban, but the court order and its scope are not available.