Extraction flow #
The submit → upload → confirm → process → receive lifecycle that every extraction follows.
Every extraction goes through the same five-stage lifecycle, regardless of whether you submit one file or fifty, use a saved template or an inline schema, poll for the result or receive it via webhook. Understanding the state machine — including the confirm-upload step that tells the backend the bytes have landed — makes the rest of the API obvious.
The five stages
┌────────┐ POST /extract ┌────────┐ PUT (S3) ┌────────┐ POST .../confirm-upload ┌─────────┐ worker ┌─────────┐
│ client │ ──────────────▶ │ submit │ ─────────▶ │ upload │ ────────────────────────▶ │ confirm │ ───────▶ │ process │
└────────┘ └────────┘ └────────┘ └─────────┘ └────┬────┘
│
GET /job/{id} ┌─────────┐ │
◀─────────────────────────│ receive │ ◀──────────────────────────────────┘
webhook push └─────────┘
| Stage | What happens | Who triggers it |
|---|---|---|
| Submit | API issues a presigned upload_url and creates a pending_upload artifact + job. | You. POST /extract. |
| Upload | File bytes go directly to S3 via the presigned URL. The artifact stays in pending_upload until you confirm. | You. PUT to upload_url. |
| Confirm | You tell the backend the upload landed. The artifact flips to uploaded, the job to queued_for_processing, and the worker is dispatched. | You. POST /artifacts/{artifact_id}/confirm-upload. |
| Process | A worker pulls the file, runs the pipeline against your template/schema, writes the result to S3, finalises billing. | The platform, automatically. |
| Receive | You either poll GET /job/{job_id} until status == success, or receive an extraction.completed webhook. | You. Pull or push. |
There is no S3 event listener watching the upload bucket. The platform
only learns that an upload completed when you call the
confirm-upload endpoint. Skip the confirm step and the artifact stays
in pending_upload forever; the job is never enqueued.
Stage 1 · Submit
POST /extract is the recommended entry point; it auto-binds the API key's
default pipeline and returns one ID per resource you'll need.
Request:
{
"filename": "invoice-2026-001.pdf",
"artifact_type": "application/pdf",
"artifact_size": 184320,
"template_id": "tpl_invoice_v3"
}
Response (202 Accepted):
{
"job_id": "job_8f3c2e1a",
"artifact_id": "art_b4d8c0f1",
"upload_url": "https://s3.amazonaws.com/dd-uploads/…?X-Amz-Signature=…",
"fields": {},
"expires_in": 3600
}
The upload_url is single-use and expires in one hour. If you don't upload
within that window, the artifact stays in pending_upload and you'll need to
submit again. (Alternatively, see POST /artifacts/{id}/rerun
once the artifact has at least been uploaded once.)
Stage 2 · Upload
Bytes go directly to S3. Your API server is not in the data path. This is why extraction can scale to gigabyte-scale PDFs without proxying.
curl -X PUT "$UPLOAD_URL" \ -H 'Content-Type: application/pdf' \ --upload-file ./invoice.pdf
The Content-Type you PUT must match the artifact_type you submitted;
the URL is signed against it. A mismatch returns 403 SignatureDoesNotMatch
from S3, not from the API.
A 200 from S3 only tells the bucket; it does not notify the API. The
artifact stays in pending_upload until you call confirm-upload (see the
next stage).
Stage 3 · Confirm
Once the PUT returns 200, tell the backend the bytes have landed:
curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/confirm-upload
The platform verifies the object exists in S3, flips the artifact to
uploaded, sets the job to queued_for_processing, commits, then dispatches
the processing worker. Until you call this endpoint, nothing happens.
Batch submissions confirm in bulk via
POST /artifacts/confirm-batch-upload/{batch_id}
— one round-trip enqueues every file that finished uploading. Direct
multipart uploads to POST /artifacts/upload
are the one exception: the bytes and the queue dispatch happen inline as
part of the same request, so no separate confirm is needed.
Stage 4 · Process
The worker:
- Fetches the file from S3.
- Runs OCR if needed (controlled by template settings).
- Extracts fields against your template or
extraction_schema. - Computes per-field confidence scores.
- Writes the structured result + a markdown rendering back to S3.
- Finalises billing: actual GPU compute time, not file size.
While processing, the artifact status walks uploaded → processing →
completed (or failed / cancelled / quarantined). The job status walks
queued_for_processing → running → success (or failed / cancelled /
expired).
Artifacts and jobs each have their own status. The artifact tracks the lifecycle of the file; the job tracks the lifecycle of the extraction attempt. Reruns create a new job against the same artifact.
Stage 5 · Receive
Two ways to find out the job is done:
Pull (polling): GET /job/{job_id} returns the live status. Cheap, simple, fine for prototypes and low-volume integrations. Use the lightweight /artifacts/{id}/status endpoint if you're polling at high frequency.
Push (webhook): register a webhook once, attach webhook_id (or pass a one-shot webhook_url) on submit, get a signed extraction.completed POST the moment the job finishes. See Webhook setup.
Most integrations end up with both: webhook for happy-path latency, polling as a safety net in case a delivery fails after all retries.
Terminal states
Job status | Meaning | Result available? |
|---|---|---|
success | Extraction completed; result in result and result_download_url. | Yes |
failed | Pipeline raised. See error field for the reason. | No |
cancelled | You called POST /artifacts/{id}/cancel while in flight. | No |
expired | The artifact never uploaded within the window, or was archived. | No |
A terminal state is, well, terminal. The job will never transition out of it
on its own. To re-process the same artifact, call
POST /artifacts/{id}/rerun; that creates a
new job with a new job_id and preserves the prior billing history.
Cancelling in flight
POST /artifacts/{artifact_id}/cancel stops a running job, releases the held
balance back to the wallet, and best-effort kills the GPU worker. It's a
no-op (returns 409) if the job is already in a terminal state.
Use this when a user retracts a submission, when you've detected the wrong file was uploaded, or as a circuit breaker if you're being throttled upstream.