Working with templates #
Create, version, clone, and reference reusable extraction templates.
This guide walks through the full lifecycle of a template, from designing fields, through creating, patching, and cloning, to retiring an old version. For the conceptual overview see Templates; for the per-endpoint reference see Templates API.
Step 1 · Design the field schema
Before creating the template, decide:
- Which fields are mandatory? Mark them
required: true. Required fields that can't be located fail extraction (or trigger review, depending on threshold). - Which need validation? Currency totals should usually have
min_value: 0. Tax IDs need aregex. Dates rarely need anything beyond thedatetype. - What's your confidence floor?
0.7is permissive (will accept noisy OCR);0.9flags more for review but produces cleaner output. Default is0.7.
A starting schema for invoices:
{
"name": "Acme invoice v1",
"category": "invoice",
"description": "Standard US-format vendor invoices.",
"tags": ["acme", "us-invoice"],
"fields": [
{ "key": "invoice_number", "label": "Invoice #", "data_type": "string", "required": true,
"validation": { "regex": "^[A-Z0-9-]+$" } },
{ "key": "issued_at", "label": "Issue date", "data_type": "date", "required": true },
{ "key": "due_at", "label": "Due date", "data_type": "date" },
{ "key": "vendor_name", "label": "Vendor", "data_type": "string", "required": true },
{ "key": "subtotal", "label": "Subtotal", "data_type": "currency", "validation": { "min_value": 0 } },
{ "key": "tax", "label": "Tax", "data_type": "currency", "validation": { "min_value": 0 } },
{ "key": "total", "label": "Total", "data_type": "currency", "required": true,
"validation": { "min_value": 0 } }
],
"settings": { "confidence_threshold": 0.8, "ocr_engine": "default" }
}
key must be ^[a-z0-9_]+$; they end up as JSON keys in the result, so
choose names that are stable database column names.
Step 2 · Create the template
curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/templates/ \ -H 'Content-Type: application/json' \ -d @template.json
Response (201 Created):
{
"id": "tpl_x8y9z",
"user_id": "usr_…",
"name": "Acme invoice v1",
"slug": "acme-invoice-v1",
"category": "invoice",
"status": "active",
"version": 1,
"usage_count": 0,
"schema_definition": { "fields": [...] }
}
Reference the new template in POST /extract via template_id: "tpl_x8y9z".
You're ready to extract.
Step 3 · Iterate on the schema
Real templates need to evolve. Add a field, tighten validation, drop a field
no consumer reads. Use PATCH /templates/{id}/schema for any of these:
curl -u "$DD_KEY:$DD_SECRET" \
-X PATCH https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/schema \
-H 'Content-Type: application/json' \
-d '{
"add_fields": [
{ "key": "po_number", "label": "PO #", "data_type": "string" }
],
"update_fields": [
{ "key": "total", "label": "Grand total", "data_type": "currency",
"required": true, "validation": { "min_value": 0 } }
],
"remove_fields": ["due_at"]
}'
The patch returns the updated template with version incremented. New jobs
reference the new version automatically; in-flight jobs continue against the
version they started under.
update_fields overwrites the entire field by key. To change just one
attribute (say, required), include the full field def with the new
attribute applied. Anything you omit is reset, not preserved.
Step 4 · Clone before risky changes
For a major restructure (renaming many fields, splitting the template by region, etc.), clone the production template first and edit the draft:
curl -u "$DD_KEY:$DD_SECRET" \
-X POST https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/clone \
-H 'Content-Type: application/json' \
-d '{ "new_name": "Acme invoice v2 (draft)", "include_settings": true }'
The clone:
- Has a fresh
idandslug. - Starts in
status: "draft". - Carries the full schema and (optionally) settings of the source.
Iterate on the draft, run a few test extractions against it (dk_test_…
keys are useful here), then either:
- Promote the draft with
PUT /templates/{draft_id}settingstatus: "active", then point production traffic at the new ID, or - Discard with
DELETE /templates/{draft_id}if the experiment doesn't pan out.
Listing and searching
GET /templates/ lists every template owned by your account. Filter with:
curl -u "$DD_KEY:$DD_SECRET" \ 'https://api.datadistillers.com/api/v1/templates/?category=invoice&search=acme&limit=25&offset=0'
| Query param | Notes |
|---|---|
category | One of invoice, receipt, id_card, contract, custom. |
search | Case-insensitive prefix match against name and description. |
limit | 1–100 (default 50). |
offset | Standard offset pagination. |
For a UI selector that lets users pick a template, paginate at limit=25
and hit GET /templates/categories once at startup to populate filters.
Step 5 · Retire old versions
When you've migrated all traffic to a new template, archive the old one to clear it from the active list without losing the audit trail:
curl -u "$DD_KEY:$DD_SECRET" \
-X PUT https://api.datadistillers.com/api/v1/templates/tpl_x8y9z \
-H 'Content-Type: application/json' \
-d '{ "status": "archived" }'
Archived templates can still be referenced by historical jobs (the audit
trail is preserved) but won't appear in default GET /templates/ results
and can't be selected for new extractions in the dashboard. To bring one
back, set status: "active" again.
To permanently delete:
curl -u "$DD_KEY:$DD_SECRET" \ -X DELETE 'https://api.datadistillers.com/api/v1/templates/tpl_x8y9z?hard_delete=true'
hard_delete=true is destructive; historical jobs will lose their template
reference. Soft-delete (the default) is almost always what you want.