Updated Apr 27, 2026
core concepts

Templates #

Reusable extraction schemas. Fields, validation, OCR settings; defined once, referenced by ID forever.

A template is a saved field schema you can reference by ID across extractions. Define the shape once (field keys, data types, validation rules, OCR settings) and pass template_id on every POST /extract to use it. The alternative is an inline extraction_schema per request, which is fine for one-offs but expensive at scale.

Template or inline schema?

Use a template_id whenUse extraction_schema when
The same shape repeats: invoices, receipts, IDs, contracts.One-off extraction; no reuse.
You want versioning and an audit trail of changes.You're prototyping or A/B testing schemas.
Field validation matters.Validation isn't part of the contract.
You want OCR / confidence settings centralised.You don't care about settings.

Passing both template_id and extraction_schema in the same request returns 400. Pick one.

Anatomy of a template

A template owns:

  • Identity: name, slug, description, category (invoice, receipt, id_card, contract, custom), tags.
  • Schema: an ordered list of TemplateFieldDef, each with a key, label, data_type, required flag, and optional validation.
  • Settings: auto_save, confidence_threshold, ocr_engine.
  • Lifecycle: status (draft / active / archived), version, usage_count.

POST /templates:

json
{
  "name":     "Acme invoice v3",
  "category": "invoice",
  "tags":     ["acme", "us-invoice"],
  "fields": [
    { "key": "invoice_number", "label": "Invoice #", "data_type": "string",
      "required": true },
    { "key": "issued_at",      "label": "Issue date", "data_type": "date" },
    { "key": "vendor_name",    "label": "Vendor",     "data_type": "string" },
    { "key": "subtotal",       "label": "Subtotal",   "data_type": "currency" },
    { "key": "tax",            "label": "Tax",        "data_type": "currency" },
    { "key": "total",          "label": "Total",      "data_type": "currency",
      "required": true,
      "validation": { "min_value": 0 } }
  ],
  "settings": { "confidence_threshold": 0.8, "ocr_engine": "default" }
}

The response includes the template's id (e.g. tpl_x8y9z…). Use it in every subsequent extraction:

json
{
  "filename":      "invoice.pdf",
  "artifact_type": "application/pdf",
  "artifact_size": 184320,
  "template_id":   "tpl_x8y9z…"
}

Field data types

Every field declares one of six data types:

data_typeReturned shape
string"INV-2026-001"
number42 or 3.14
dateISO-8601 "2026-04-29"
currency{ "value": 1240.00, "currency": "USD" }
booleantrue / false
tableArray of row objects with the field's nested schema

required: true causes extraction to fail (or flag for review, depending on confidence_threshold) if the field can't be located.

validation (regex, min, max) is enforced after extraction; failures appear in the result and may flip needs_review to true.

Versioning and patches

Templates are versioned. Every PATCH /templates/{id}/schema increments version and preserves the prior schema for any in-flight job that already referenced the old version. Existing artifacts and results are unaffected by schema changes.

Three operations on the schema:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X PATCH https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/schema \
  -H 'Content-Type: application/json' \
  -d '{
    "add_fields":   [{ "key": "po_number", "label": "PO #", "data_type": "string" }],
    "update_fields":[{ "key": "total", "label": "Grand total", "data_type": "currency", "required": true }],
    "remove_fields":["legacy_notes"]
  }'
  • add_fields: adds new fields. Fails with 400 if a key already exists.
  • update_fields: overwrites by key. Fails with 404 if the key is missing.
  • remove_fields: deletes by key.

For larger restructures, clone the template into a draft, edit freely, and switch your template_id once the new version is ready.

Cloning

POST /templates/{id}/clone produces a fresh draft template with a new id and the prior schema copied in. Useful for:

  • Forking a stable production template before risky edits.
  • Branching one shape into multiple variants (e.g. US vs EU invoice).
  • Sharing a starting point across teams.
bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/clone \
  -H 'Content-Type: application/json' \
  -d '{ "new_name": "Acme invoice v4 (draft)", "include_settings": true }'

The clone starts in draft status. Promote it to active with PUT /templates/{id} once you're ready to point production traffic at it.

Discovery

GET /templates/?category=invoice&search=acme filters the templates owned by your account. Pair with limit and offset for pagination.

GET /templates/categories returns the canonical list of category strings. Useful for populating a UI selector; never hard-code the list, since new categories ship without an API version bump.

Soft vs hard delete

DELETE /templates/{id} is a soft delete by default. The template can no longer be referenced by new jobs, but existing jobs and the audit trail remain intact. To permanently remove the record:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X DELETE 'https://api.datadistillers.com/api/v1/templates/tpl_x8y9z?hard_delete=true'

Use hard delete sparingly. It's irreversible and breaks any historical job that referenced the template.

Esc
↑↓Navigate↵OpenEscClose