What a run does
1
Measure the baseline
The Optimizer evaluates the current version against the dataset. This is iteration 0. Everything after is judged against it.
2
Propose
The Optimizer studies where the previous results went wrong and writes improved field descriptions.
3
Evaluate
The Optimizer runs the proposed descriptions over the dataset, like a regular evaluation.
4
Keep only improvements
An iteration is accepted only if its accuracy beats the best result so far. Propose → evaluate repeats up to the number of iterations you chose.
“No improvement found” is a valid outcome. Your current descriptions already perform as well as the proposals. The promoted version matches your baseline, and nothing changes.
What it needs
- An Extract node. The Optimizer improves extraction field descriptions, so parse-only and split workflows are excluded.
- A dataset with ground truth. The run scores every proposal against your documents’ known-correct values, so a dataset without ground truth cannot start one. See Building a dataset.
Starting a run
Choose Run optimization and set:- Scope: the full dataset.
- Metric: overall accuracy.
- Iterations: how many propose-and-evaluate rounds to attempt, from 1 to 10. The default is 3.
A workflow allows one active optimization at a time. A second start is rejected until the current run finishes.
Following a run
Runs appear in the Optimizer table. Status, accuracy, and rounds update live as the run progresses:
Open a run for the full picture: an accuracy-per-iteration chart, the timeline of iterations, and each iteration’s proposed description changes. Open a single iteration to see its per-field accuracy, which fields improved, which still mismatch, and the exact description edits it tried.
The per-field view doubles as a work list. A field that stays inaccurate after an optimization needs more dataset examples, or more varied ones.
Versions and rolling back
The promoted result goes live automatically. It becomes the workflow’s latest version, with no separate review step. Every change is a normal version-history entry. If you disagree with a rewrite, restore an earlier version.Asking annie
annie drives the Optimizer conversationally, which helps when you are already in a chat about the workflow’s results. “improve this workflow” starts a run. “how did the optimization go?” reports accuracy movement and the changes she kept. “list the optimizations” shows the history.Current limits
- A running optimization cannot be cancelled. It runs to completion. Size your first runs small.
- The Optimizer changes extraction field descriptions only. It never touches Classify, Split, or Validate nodes.
- Scope and metric are fixed for now: full dataset, overall accuracy.
What’s next?
Building a dataset
Add documents and author the ground truth an optimization measures against.
Running evaluations
Score a version by hand and compare accuracy across versions.

