Templates #
Reusable extraction schemas. Fields, validation, OCR settings; defined once, referenced by ID forever.
A template is a saved field schema you can reference by ID across
extractions. Define the shape once (field keys, data types, validation
rules, OCR settings) and pass template_id on every POST /extract to use
it. The alternative is an inline extraction_schema per request, which is
fine for one-offs but expensive at scale.
Template or inline schema?
Use a template_id when | Use extraction_schema when |
|---|---|
| The same shape repeats: invoices, receipts, IDs, contracts. | One-off extraction; no reuse. |
| You want versioning and an audit trail of changes. | You're prototyping or A/B testing schemas. |
| Field validation matters. | Validation isn't part of the contract. |
| You want OCR / confidence settings centralised. | You don't care about settings. |
Passing both template_id and extraction_schema in the same request returns
400. Pick one.
Anatomy of a template
A template owns:
- Identity:
name,slug,description,category(invoice,receipt,id_card,contract,custom),tags. - Schema: an ordered list of
TemplateFieldDef, each with akey,label,data_type,requiredflag, and optionalvalidation. - Settings:
auto_save,confidence_threshold,ocr_engine. - Lifecycle:
status(draft/active/archived),version,usage_count.
POST /templates:
{
"name": "Acme invoice v3",
"category": "invoice",
"tags": ["acme", "us-invoice"],
"fields": [
{ "key": "invoice_number", "label": "Invoice #", "data_type": "string",
"required": true },
{ "key": "issued_at", "label": "Issue date", "data_type": "date" },
{ "key": "vendor_name", "label": "Vendor", "data_type": "string" },
{ "key": "subtotal", "label": "Subtotal", "data_type": "currency" },
{ "key": "tax", "label": "Tax", "data_type": "currency" },
{ "key": "total", "label": "Total", "data_type": "currency",
"required": true,
"validation": { "min_value": 0 } }
],
"settings": { "confidence_threshold": 0.8, "ocr_engine": "default" }
}
The response includes the template's id (e.g. tpl_x8y9z…). Use it in
every subsequent extraction:
{
"filename": "invoice.pdf",
"artifact_type": "application/pdf",
"artifact_size": 184320,
"template_id": "tpl_x8y9z…"
}
Field data types
Every field declares one of six data types:
data_type | Returned shape |
|---|---|
string | "INV-2026-001" |
number | 42 or 3.14 |
date | ISO-8601 "2026-04-29" |
currency | { "value": 1240.00, "currency": "USD" } |
boolean | true / false |
table | Array of row objects with the field's nested schema |
required: true causes extraction to fail (or flag for review, depending on
confidence_threshold) if the field can't be located.
validation (regex, min, max) is enforced after extraction; failures appear
in the result and may flip needs_review to true.
Versioning and patches
Templates are versioned. Every PATCH /templates/{id}/schema increments
version and preserves the prior schema for any in-flight job that already
referenced the old version. Existing artifacts and results are unaffected by
schema changes.
Three operations on the schema:
curl -u "$DD_KEY:$DD_SECRET" \
-X PATCH https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/schema \
-H 'Content-Type: application/json' \
-d '{
"add_fields": [{ "key": "po_number", "label": "PO #", "data_type": "string" }],
"update_fields":[{ "key": "total", "label": "Grand total", "data_type": "currency", "required": true }],
"remove_fields":["legacy_notes"]
}'
add_fields: adds new fields. Fails with400if akeyalready exists.update_fields: overwrites bykey. Fails with404if the key is missing.remove_fields: deletes bykey.
For larger restructures, clone the template into a draft, edit
freely, and switch your template_id once the new version is ready.
Cloning
POST /templates/{id}/clone produces a fresh draft template with a new
id and the prior schema copied in. Useful for:
- Forking a stable production template before risky edits.
- Branching one shape into multiple variants (e.g. US vs EU invoice).
- Sharing a starting point across teams.
curl -u "$DD_KEY:$DD_SECRET" \
-X POST https://api.datadistillers.com/api/v1/templates/tpl_x8y9z/clone \
-H 'Content-Type: application/json' \
-d '{ "new_name": "Acme invoice v4 (draft)", "include_settings": true }'
The clone starts in draft status. Promote it to active with PUT /templates/{id} once you're ready to point production traffic at it.
Discovery
GET /templates/?category=invoice&search=acme filters the templates owned
by your account. Pair with limit and offset for pagination.
GET /templates/categories returns the canonical list of category strings.
Useful for populating a UI selector; never hard-code the list, since new
categories ship without an API version bump.
Soft vs hard delete
DELETE /templates/{id} is a soft delete by default. The template can no
longer be referenced by new jobs, but existing jobs and the audit trail
remain intact. To permanently remove the record:
curl -u "$DD_KEY:$DD_SECRET" \ -X DELETE 'https://api.datadistillers.com/api/v1/templates/tpl_x8y9z?hard_delete=true'
Use hard delete sparingly. It's irreversible and breaks any historical job that referenced the template.