Updated Apr 27, 2026
core concepts

Extraction flow #

The submit → upload → confirm → process → receive lifecycle that every extraction follows.

Every extraction goes through the same five-stage lifecycle, regardless of whether you submit one file or fifty, use a saved template or an inline schema, poll for the result or receive it via webhook. Understanding the state machine — including the confirm-upload step that tells the backend the bytes have landed — makes the rest of the API obvious.

The five stages

text
┌────────┐  POST /extract  ┌────────┐  PUT (S3)  ┌────────┐  POST .../confirm-upload  ┌─────────┐  worker  ┌─────────┐
│ client │ ──────────────▶ │ submit │ ─────────▶ │ upload │ ────────────────────────▶ │ confirm │ ───────▶ │ process │
└────────┘                 └────────┘            └────────┘                           └─────────┘          └────┬────┘
                                                                                                                │
                                            GET /job/{id}        ┌─────────┐                                    │
                                       ◀─────────────────────────│ receive │ ◀──────────────────────────────────┘
                                            webhook push         └─────────┘
StageWhat happensWho triggers it
SubmitAPI issues a presigned upload_url and creates a pending_upload artifact + job.You. POST /extract.
UploadFile bytes go directly to S3 via the presigned URL. The artifact stays in pending_upload until you confirm.You. PUT to upload_url.
ConfirmYou tell the backend the upload landed. The artifact flips to uploaded, the job to queued_for_processing, and the worker is dispatched.You. POST /artifacts/{artifact_id}/confirm-upload.
ProcessA worker pulls the file, runs the pipeline against your template/schema, writes the result to S3, finalises billing.The platform, automatically.
ReceiveYou either poll GET /job/{job_id} until status == success, or receive an extraction.completed webhook.You. Pull or push.
The backend never finds out on its own

There is no S3 event listener watching the upload bucket. The platform only learns that an upload completed when you call the confirm-upload endpoint. Skip the confirm step and the artifact stays in pending_upload forever; the job is never enqueued.

Stage 1 · Submit

POST /extract is the recommended entry point; it auto-binds the API key's default pipeline and returns one ID per resource you'll need.

Request:

json
{
  "filename":      "invoice-2026-001.pdf",
  "artifact_type": "application/pdf",
  "artifact_size": 184320,
  "template_id":   "tpl_invoice_v3"
}

Response (202 Accepted):

json
{
  "job_id":      "job_8f3c2e1a",
  "artifact_id": "art_b4d8c0f1",
  "upload_url":  "https://s3.amazonaws.com/dd-uploads/…?X-Amz-Signature=…",
  "fields":      {},
  "expires_in":  3600
}

The upload_url is single-use and expires in one hour. If you don't upload within that window, the artifact stays in pending_upload and you'll need to submit again. (Alternatively, see POST /artifacts/{id}/rerun once the artifact has at least been uploaded once.)

Stage 2 · Upload

Bytes go directly to S3. Your API server is not in the data path. This is why extraction can scale to gigabyte-scale PDFs without proxying.

bash
curl -X PUT "$UPLOAD_URL" \
  -H 'Content-Type: application/pdf' \
  --upload-file ./invoice.pdf

The Content-Type you PUT must match the artifact_type you submitted; the URL is signed against it. A mismatch returns 403 SignatureDoesNotMatch from S3, not from the API.

A 200 from S3 only tells the bucket; it does not notify the API. The artifact stays in pending_upload until you call confirm-upload (see the next stage).

Stage 3 · Confirm

Once the PUT returns 200, tell the backend the bytes have landed:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/confirm-upload

The platform verifies the object exists in S3, flips the artifact to uploaded, sets the job to queued_for_processing, commits, then dispatches the processing worker. Until you call this endpoint, nothing happens.

Batch submissions confirm in bulk via POST /artifacts/confirm-batch-upload/{batch_id} — one round-trip enqueues every file that finished uploading. Direct multipart uploads to POST /artifacts/upload are the one exception: the bytes and the queue dispatch happen inline as part of the same request, so no separate confirm is needed.

Stage 4 · Process

The worker:

  1. Fetches the file from S3.
  2. Runs OCR if needed (controlled by template settings).
  3. Extracts fields against your template or extraction_schema.
  4. Computes per-field confidence scores.
  5. Writes the structured result + a markdown rendering back to S3.
  6. Finalises billing: actual GPU compute time, not file size.

While processing, the artifact status walks uploaded → processing → completed (or failed / cancelled / quarantined). The job status walks queued_for_processing → running → success (or failed / cancelled / expired).

Two status fields, not one

Artifacts and jobs each have their own status. The artifact tracks the lifecycle of the file; the job tracks the lifecycle of the extraction attempt. Reruns create a new job against the same artifact.

Stage 5 · Receive

Two ways to find out the job is done:

Pull (polling): GET /job/{job_id} returns the live status. Cheap, simple, fine for prototypes and low-volume integrations. Use the lightweight /artifacts/{id}/status endpoint if you're polling at high frequency.

Push (webhook): register a webhook once, attach webhook_id (or pass a one-shot webhook_url) on submit, get a signed extraction.completed POST the moment the job finishes. See Webhook setup.

Most integrations end up with both: webhook for happy-path latency, polling as a safety net in case a delivery fails after all retries.

Terminal states

Job statusMeaningResult available?
successExtraction completed; result in result and result_download_url.Yes
failedPipeline raised. See error field for the reason.No
cancelledYou called POST /artifacts/{id}/cancel while in flight.No
expiredThe artifact never uploaded within the window, or was archived.No

A terminal state is, well, terminal. The job will never transition out of it on its own. To re-process the same artifact, call POST /artifacts/{id}/rerun; that creates a new job with a new job_id and preserves the prior billing history.

Cancelling in flight

POST /artifacts/{artifact_id}/cancel stops a running job, releases the held balance back to the wallet, and best-effort kills the GPU worker. It's a no-op (returns 409) if the job is already in a terminal state.

Use this when a user retracts a submission, when you've detected the wrong file was uploaded, or as a circuit breaker if you're being throttled upstream.

Esc
↑↓Navigate↵OpenEscClose