Updated Apr 27, 2026
getting started

Quickstart #

Submit a document, upload it, confirm the upload, poll for the result. End-to-end extraction in a handful of HTTP calls.

This is the path from "I have an API key" to "I have JSON" in four HTTP calls: submit, upload, confirm, poll. The same flow scales from a one-off receipt to ten thousand invoices a day.

You'll need an API key and its secret. Both are issued from the dashboard and shown exactly once. If you don't have them yet, see Authentication.

Base URL and authentication

EnvironmentBase URL
Productionhttps://api.datadistillers.com/api/v1

Every request authenticates with HTTP Basic. Username = API key, password = secret.

bash
export DD_KEY='dk_live_…'
export DD_SECRET='sk_…'

Verify your credentials work before going further:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  https://api.datadistillers.com/api/v1/wallet

A 200 with a JSON balance means you're set. A 401 means the key/secret pair is wrong; double-check the header isn't being mangled by your shell.

Step 1 · Submit the extraction

POST /extract describes the file you're about to upload. You don't send bytes here; you send filename, MIME type, size, and either a saved template_id or an inline extraction_schema.

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/extract \
  -H 'Content-Type: application/json' \
  -d '{
    "filename": "invoice-2026-001.pdf",
    "artifact_type": "application/pdf",
    "artifact_size": 184320,
    "extraction_schema": {
      "fields": [
        { "key": "invoice_number", "data_type": "string", "required": true },
        { "key": "total",          "data_type": "currency" },
        { "key": "issued_at",      "data_type": "date" }
      ]
    }
  }'

Response (202 Accepted):

json
{
  "job_id":     "job_8f3c2e1a",
  "artifact_id": "art_b4d8c0f1",
  "upload_url":  "https://s3.amazonaws.com/dd-uploads/…?X-Amz-Signature=…",
  "fields":      {},
  "expires_in":  3600
}

Save job_id (you'll poll on it) and upload_url (you have one hour to use it).

Step 2 · Upload the file

PUT the raw bytes to the upload_url. The upload goes directly to S3; no auth header, no JSON envelope, no API server in the path.

bash
curl -X PUT "$UPLOAD_URL" \
  -H 'Content-Type: application/pdf' \
  --upload-file ./invoice-2026-001.pdf

A 200 with an empty body means the upload landed in S3. Nothing happens on the API side yet — the backend doesn't watch S3 for new objects.

Match the Content-Type

The presigned URL is signed against the artifact_type you sent in step 1. If you submitted application/pdf but PUT with image/png, S3 returns a signature mismatch (403 SignatureDoesNotMatch).

Step 3 · Confirm the upload

POST /artifacts/{artifact_id}/confirm-upload tells the backend the bytes have landed. The platform flips the artifact to uploaded, the job to queued_for_processing, and dispatches the worker.

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/confirm-upload

Skip this call and the artifact stays in pending_upload forever; the job will never run.

Step 4 · Poll for the result

GET /job/{job_id} returns the current status. Status transitions are:

pending → queued_for_processing → running → success (or failed, cancelled, expired)

bash
curl -u "$DD_KEY:$DD_SECRET" \
  https://api.datadistillers.com/api/v1/job/job_8f3c2e1a

While processing:

json
{
  "job_id": "job_8f3c2e1a",
  "status": "running",
  "created_at": "2026-05-03T09:12:04Z"
}

When complete:

json
{
  "job_id": "job_8f3c2e1a",
  "status": "success",
  "created_at":   "2026-05-03T09:12:04Z",
  "completed_at": "2026-05-03T09:12:11Z",
  "result_download_url": "https://s3.amazonaws.com/…/result.json?X-Amz-…",
  "result": {
    "invoice_number": "INV-2026-001",
    "total":          { "value": 1240.00, "currency": "USD" },
    "issued_at":      "2026-04-29"
  }
}

A simple polling loop, with backoff, is two lines of bash. Production code should prefer a webhook so you're not paying latency on a fixed interval.

End-to-end script

Where to go next

  • Extraction flow. The whole lifecycle in one diagram.
  • Webhook setup. Stop polling, start receiving signed pushes.
  • Templates. Reuse a field schema across thousands of jobs.
  • Batch processing. Submit up to 50 files in one call.
  • Errors. What every status code means and how to retry.
  • Going to production. The checklist for swapping the dk_test_… key for dk_live_… and leaving it running unattended.
Esc
↑↓Navigate↵OpenEscClose