Updated Apr 27, 2026
guides

Extract a single document #

The recommended /extract flow for one file at a time. Submit, upload, confirm, poll or webhook.

This guide walks through the most common DataDistillers integration: extract fields from one document, get a result back. It assumes you have a working API key (see Authentication) and an optional template ID (see Templates).

Step 1 · Pick template or inline schema

You need to tell the extractor what fields to look for. Two ways:

  • template_id: reference a saved template. Best for repeat shapes.
  • extraction_schema: declare fields inline. Best for one-offs.

Send exactly one. Sending both returns 400 Bad Request.

Step 2 · Submit the extraction

POST /extract returns a presigned upload URL plus the IDs you'll need to track the job.

The response (202 Accepted):

json
{
  "job_id":      "job_8f3c2e1a",
  "artifact_id": "art_b4d8c0f1",
  "upload_url":  "https://s3.amazonaws.com/dd-uploads/…?X-Amz-Signature=…",
  "fields":      {},
  "expires_in":  3600
}

Step 3 · PUT the file to the upload URL

The bytes go directly to S3; your API request never carries the file.

Content-Type must match

The presigned URL is signed against the artifact_type from step 2. If you submitted application/pdf but PUT with a different MIME type, S3 returns 403 SignatureDoesNotMatch.

S3's response is a 200 with no body. The bucket has the file; the API doesn't know yet. The backend has no S3 event listener — you have to tell it the upload completed.

Step 4 · Confirm the upload

POST /artifacts/{artifact_id}/confirm-upload is the signal that flips the artifact to uploaded, sets the job to queued_for_processing, and dispatches the worker. Without it, the artifact stays in pending_upload forever.

The endpoint is idempotent — calling it twice on the same artifact is safe. It returns 409 if the artifact is already past uploaded (e.g. already processing or completed).

Step 5 · Wait for the result

Two ways to know the extraction is done:

Polling (GET /job/{job_id}):

bash
curl -u "$DD_KEY:$DD_SECRET" \
  https://api.datadistillers.com/api/v1/job/job_8f3c2e1a

A backoff-friendly polling loop:

py
import time

backoff = 1
while True:
    job = requests.get(
        f'https://api.datadistillers.com/api/v1/job/{submit["job_id"]}',
        auth=(KEY, SECRET),
    ).json()
    if job['status'] in ('success', 'failed', 'cancelled', 'expired'):
        break
    time.sleep(backoff)
    backoff = min(backoff * 1.5, 10)

Webhook: register one upfront, attach webhook_id on submit, and skip the polling entirely. See Webhook setup.

Step 6 · Read the result

When status == "success":

json
{
  "job_id": "job_8f3c2e1a",
  "status": "success",
  "completed_at": "2026-05-03T09:12:11Z",
  "result_download_url": "https://s3.amazonaws.com/…/result.json?X-Amz-…",
  "result": {
    "invoice_number": "INV-2026-001",
    "total": { "value": 1240.00, "currency": "USD" },
    "issued_at": "2026-04-29"
  }
}

The result field is the inline JSON. For larger results (tables with thousands of rows), the inline payload may be omitted and only result_download_url set; fetch the JSON from S3 in that case. Treat download URLs as short-lived and don't cache them.

Handling failures

A failed job carries an error string that explains the cause:

json
{
  "job_id": "job_8f3c2e1a",
  "status": "failed",
  "error":  "schema_mismatch: required field 'total' not found"
}

Common failure categories:

ErrorCauseAction
schema_mismatchA required field couldn't be located.Loosen required, lower confidence_threshold, or fix the schema.
unsupported_typeMIME type isn't supported.Convert the file or use a supported format.
corrupted_fileFile failed to open / decode.Re-upload from the source.
timeoutPipeline exceeded the per-job time budget.Split the document or contact support if recurring.

To re-attempt against the same artifact (e.g. after fixing a template), call POST /artifacts/{id}/rerun; that creates a new job with fresh ID and preserves billing history.

Esc
↑↓Navigate↵OpenEscClose