Updated Apr 27, 2026
guides

Working with templates #

Create, version, clone, and reference reusable extraction templates.

This guide walks through the full lifecycle of a template, from designing fields, through creating, patching, and cloning, to retiring an old version. For the conceptual overview see Templates; for the per-endpoint reference see Templates API.

Step 1 · Design the field schema

Before creating the template, decide:

  • Which fields are mandatory? Mark them required: true. Required fields that can't be located fail extraction (or trigger review, depending on threshold).
  • Which need validation? Currency totals should usually have min_value: 0. Tax IDs need a regex. Dates rarely need anything beyond the date type.
  • What's your confidence floor? 0.7 is permissive (will accept noisy OCR); 0.9 flags more for review but produces cleaner output. Default is 0.7.

A starting schema for invoices:

json
{
  "name":        "Acme invoice v1",
  "category":    "invoice",
  "description": "Standard US-format vendor invoices.",
  "tags":        ["acme", "us-invoice"],
  "fields": [
    { "key": "invoice_number", "label": "Invoice #",     "data_type": "string",   "required": true,
      "validation": { "regex": "^[A-Z0-9-]+$" } },
    { "key": "issued_at",      "label": "Issue date",    "data_type": "date",     "required": true },
    { "key": "due_at",         "label": "Due date",      "data_type": "date" },
    { "key": "vendor_name",    "label": "Vendor",        "data_type": "string",   "required": true },
    { "key": "subtotal",       "label": "Subtotal",      "data_type": "currency", "validation": { "min_value": 0 } },
    { "key": "tax",            "label": "Tax",           "data_type": "currency", "validation": { "min_value": 0 } },
    { "key": "total",          "label": "Total",         "data_type": "currency", "required": true,
      "validation": { "min_value": 0 } }
  ],
  "settings": { "confidence_threshold": 0.8, "ocr_engine": "default" }
}

key must be ^[a-z0-9_]+$; they end up as JSON keys in the result, so choose names that are stable database column names.

Step 2 · Create the template

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/templates/ \
  -H 'Content-Type: application/json' \
  -d @template.json

Response (201 Created):

json
{
  "id":          "tpl_x8y9z",
  "user_id":     "usr_…",
  "name":        "Acme invoice v1",
  "slug":        "acme-invoice-v1",
  "category":    "invoice",
  "status":      "active",
  "version":     1,
  "usage_count": 0,
  "schema_definition": { "fields": [...] }
}

Reference the new template in POST /extract via template_id: "tpl_x8y9z". You're ready to extract.

Step 3 · Iterate on the schema

Real templates need to evolve. Add a field, tighten validation, drop a field no consumer reads. Use PATCH /templates/{id}/schema for any of these:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X PATCH https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/schema \
  -H 'Content-Type: application/json' \
  -d '{
    "add_fields": [
      { "key": "po_number", "label": "PO #", "data_type": "string" }
    ],
    "update_fields": [
      { "key": "total", "label": "Grand total", "data_type": "currency",
        "required": true, "validation": { "min_value": 0 } }
    ],
    "remove_fields": ["due_at"]
  }'

The patch returns the updated template with version incremented. New jobs reference the new version automatically; in-flight jobs continue against the version they started under.

Patches are partial

update_fields overwrites the entire field by key. To change just one attribute (say, required), include the full field def with the new attribute applied. Anything you omit is reset, not preserved.

Step 4 · Clone before risky changes

For a major restructure (renaming many fields, splitting the template by region, etc.), clone the production template first and edit the draft:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/clone \
  -H 'Content-Type: application/json' \
  -d '{ "new_name": "Acme invoice v2 (draft)", "include_settings": true }'

The clone:

  • Has a fresh id and slug.
  • Starts in status: "draft".
  • Carries the full schema and (optionally) settings of the source.

Iterate on the draft, run a few test extractions against it (dk_test_… keys are useful here), then either:

  • Promote the draft with PUT /templates/{draft_id} setting status: "active", then point production traffic at the new ID, or
  • Discard with DELETE /templates/{draft_id} if the experiment doesn't pan out.

Listing and searching

GET /templates/ lists every template owned by your account. Filter with:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  'https://api.datadistillers.com/api/v1/templates/?category=invoice&search=acme&limit=25&offset=0'
Query paramNotes
categoryOne of invoice, receipt, id_card, contract, custom.
searchCase-insensitive prefix match against name and description.
limit1–100 (default 50).
offsetStandard offset pagination.

For a UI selector that lets users pick a template, paginate at limit=25 and hit GET /templates/categories once at startup to populate filters.

Step 5 · Retire old versions

When you've migrated all traffic to a new template, archive the old one to clear it from the active list without losing the audit trail:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X PUT https://api.datadistillers.com/api/v1/templates/tpl_x8y9z \
  -H 'Content-Type: application/json' \
  -d '{ "status": "archived" }'

Archived templates can still be referenced by historical jobs (the audit trail is preserved) but won't appear in default GET /templates/ results and can't be selected for new extractions in the dashboard. To bring one back, set status: "active" again.

To permanently delete:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X DELETE 'https://api.datadistillers.com/api/v1/templates/tpl_x8y9z?hard_delete=true'

hard_delete=true is destructive; historical jobs will lose their template reference. Soft-delete (the default) is almost always what you want.

Esc
↑↓Navigate↵OpenEscClose