Skip to main content
POST
Create workflow
POST /v2/workflows/ creates a workflow from a strongly-typed graph. Use it to:
  • Configure parse-node settings: standard or agentic mode, prompt hints, and figure enhancement.
  • Build a parse-only workflow, with no extract node, that returns markdown only.
  • Route documents through a classifier or a splitter to several extract nodes.
  • Run a linear parse → extract workflow.
Build the workflow in the anyformat platform instead, and iterate faster: configure fields visually and test against sample documents. Then copy the workflow id and run workflows through the API.

Request Body

Parse Node

The entry-point node. Exactly one per workflow.

Extract Node

Pulls structured fields from upstream parsed content. Field Types gives the field shape. Any field may set "source": "smart_lookup" to become a smart-lookup field, resolved from a reference file rather than read from the document. "source": "lookup_if_missing" extracts the field too, and lets the lookup fill only the empty slots. The default is "extraction".

Classify Node

Routes the document to one of categories[] based on an LLM verdict. Outgoing edges must set branch to a category id. Each Category:

Splitter Node

Partitions a multi-document file into per-rule sub-documents. Each rule fires an outgoing edge whose branch matches the rule id. The downstream extract node runs once per resulting partition. Each Rule:

Validate Node

Runs after an extract node and emits per-rule validation results. The results endpoint returns them as extractions[].validations[], beside the fields they judged. See Response Formats. A rule is either AI-evaluated or deterministic. An AI rule sets kind="ai" and carries a natural-language description. A deterministic rule sets kind="deterministic" and carries a structured check, evaluated in pure Python with no model call. Each Rule: The deterministic check shapes are discriminated on type. Every shape except expression names its field operands by persistent_id. expression addresses fields by extracted name instead. An arithmetic check for subtotal + tax == total, within ±1¢:

Edges

Examples

Linear parse → extract

Parse-only with agentic mode

A workflow with one parse node and no edges. It produces markdown only, and extracts nothing.
Agentic mode routes each block through a typed text, table or figure strategy, using Gemini to route. The mode is binary: mode: "agentic" turns on the full agentic pipeline. See the Agentic Parse to Markdown recipe for an end-to-end walkthrough.

Classify-then-extract (branched)

Each outgoing edge from a classify or splitter node sets branch to a category or rule .id on the source node. .name is the display label shown to the LLM, and is not a valid branch value: using it returns 400. In the example below, cat_invoice routes the edge, and Invoice is the label the classifier sees.

Smart lookup

Enrich extracted values by matching them against a reference file instead of reading them off the page. Flag the looked-up field with "source": "smart_lookup". Attach the reference file inline as base64 through lookup_file_uploads. Describe the join with lookup_suggestion if it helps. See Smart Lookup for how matching works and for the Studio equivalent of these fields. In the example below, the model reads vendor_name from the document, then resolves the canonical vendor_id from a vendor catalog.
lookup_file_uploads[].content is the base64-encoded bytes of the reference CSV or text file. The results payload returns the lookup field, vendor_id, alongside the extraction fields. It has no separate response section. An extract node carrying a source: smart_lookup or lookup_if_missing field but no reference file is rejected with 400.

Topology Rules

The endpoint enforces graph correctness. A violated rule returns 400, and the API persists nothing.

Response

201 Created returns the workflow resource:
Save the id. You reuse it when uploading documents and fetching results.

Next Steps

Agentic Parse to Markdown

End-to-end recipe: create, upload, results, using agentic parse mode

Parse-Only Workflow

Convert documents to markdown without extraction

Field Types

All supported data_type values for extraction fields

Response Formats

The shape of parse, extractions, splits and classifications on the results endpoint

Authorizations

Authorization
string
header
required

API key issued from app.anyformat.ai/api-key. Send as Authorization: Bearer <key>.

Body

application/json

Public-surface workflow create body — typed graph of parse / classify / splitter / extract / validate nodes.

A validate node carries rules; each rule is either AI-evaluated (kind="ai" + a natural-language description) or deterministic (kind="deterministic" + a structured check, evaluated in pure Python with no model call).

Uses PublicNode (public request models with only the fields callers may set); the domain AnyNode also carries filter (worker-only) and executor-injected / staff-only fields, so the public bodies pin a stricter node schema. Call :meth:to_domain before forwarding to the backend.

name
string
required
Minimum string length: 1
Example:

"Invoice or receipt"

nodes
(PublicParseNode · object | ClassifyNode · object | SplitterNode · object | PublicExtractNode · object | PublicValidateNode · object | PublicIfElseNode · object | PublicSlackAlertNode · object | PublicEditNode · object | KnowledgeNode · object)[]
required
Minimum array length: 1
description
string | null
default:""
edges
Edge · object[]

Response

Successful Response

A workflow defines the extraction template — what fields to extract from documents, their types, and validation rules.

id
string
required

Unique identifier of the workflow (UUID).

Example:

"0686bb97-8c30-70f0-8000-97669e000eb8"

name
string
required

Human-readable name of the workflow.

Example:

"Invoice Processing"

description
string | null

Optional description of what this workflow extracts.

Example:

"A workflow for processing invoices and retrieving invoice details."

created_at
string<date-time> | null

Timestamp when the workflow was created (ISO 8601).

updated_at
string<date-time> | null

Timestamp when the workflow was last modified (ISO 8601).

fields
Fields · object[] | null

List of extraction field definitions configured for this workflow. null if not yet configured.