Extract a single document #
The recommended /extract flow for one file at a time. Submit, upload, confirm, poll or webhook.
This guide walks through the most common DataDistillers integration: extract fields from one document, get a result back. It assumes you have a working API key (see Authentication) and an optional template ID (see Templates).
Step 1 · Pick template or inline schema
You need to tell the extractor what fields to look for. Two ways:
template_id: reference a saved template. Best for repeat shapes.extraction_schema: declare fields inline. Best for one-offs.
Send exactly one. Sending both returns 400 Bad Request.
Step 2 · Submit the extraction
POST /extract returns a presigned upload URL plus the IDs you'll need to
track the job.
The response (202 Accepted):
{
"job_id": "job_8f3c2e1a",
"artifact_id": "art_b4d8c0f1",
"upload_url": "https://s3.amazonaws.com/dd-uploads/…?X-Amz-Signature=…",
"fields": {},
"expires_in": 3600
}
Step 3 · PUT the file to the upload URL
The bytes go directly to S3; your API request never carries the file.
The presigned URL is signed against the artifact_type from step 2. If
you submitted application/pdf but PUT with a different MIME type, S3
returns 403 SignatureDoesNotMatch.
S3's response is a 200 with no body. The bucket has the file; the API doesn't know yet. The backend has no S3 event listener — you have to tell it the upload completed.
Step 4 · Confirm the upload
POST /artifacts/{artifact_id}/confirm-upload is the signal that flips the
artifact to uploaded, sets the job to queued_for_processing, and
dispatches the worker. Without it, the artifact stays in pending_upload
forever.
The endpoint is idempotent — calling it twice on the same artifact is
safe. It returns 409 if the artifact is already past uploaded (e.g.
already processing or completed).
Step 5 · Wait for the result
Two ways to know the extraction is done:
Polling (GET /job/{job_id}):
curl -u "$DD_KEY:$DD_SECRET" \ https://api.datadistillers.com/api/v1/job/job_8f3c2e1a
A backoff-friendly polling loop:
import time
backoff = 1
while True:
job = requests.get(
f'https://api.datadistillers.com/api/v1/job/{submit["job_id"]}',
auth=(KEY, SECRET),
).json()
if job['status'] in ('success', 'failed', 'cancelled', 'expired'):
break
time.sleep(backoff)
backoff = min(backoff * 1.5, 10)
Webhook: register one upfront, attach webhook_id on submit, and skip
the polling entirely. See Webhook setup.
Step 6 · Read the result
When status == "success":
{
"job_id": "job_8f3c2e1a",
"status": "success",
"completed_at": "2026-05-03T09:12:11Z",
"result_download_url": "https://s3.amazonaws.com/…/result.json?X-Amz-…",
"result": {
"invoice_number": "INV-2026-001",
"total": { "value": 1240.00, "currency": "USD" },
"issued_at": "2026-04-29"
}
}
The result field is the inline JSON. For larger results (tables with
thousands of rows), the inline payload may be omitted and only
result_download_url set; fetch the JSON from S3 in that case. Treat
download URLs as short-lived and don't cache them.
Handling failures
A failed job carries an error string that explains the cause:
{
"job_id": "job_8f3c2e1a",
"status": "failed",
"error": "schema_mismatch: required field 'total' not found"
}
Common failure categories:
| Error | Cause | Action |
|---|---|---|
schema_mismatch | A required field couldn't be located. | Loosen required, lower confidence_threshold, or fix the schema. |
unsupported_type | MIME type isn't supported. | Convert the file or use a supported format. |
corrupted_file | File failed to open / decode. | Re-upload from the source. |
timeout | Pipeline exceeded the per-job time budget. | Split the document or contact support if recurring. |
To re-attempt against the same artifact (e.g. after fixing a template),
call POST /artifacts/{id}/rerun; that
creates a new job with fresh ID and preserves billing history.