Updated Apr 27, 2026
core concepts

Artifacts #

What an artifact is, the statuses it walks through, and how it relates to jobs and results.

An artifact is the API's record of a single file you've uploaded: its metadata, its storage location, its processing state, and any extraction results derived from it. Every POST /extract creates one artifact and one initial job. Reruns create new jobs against the same artifact.

Anatomy of an artifact

GET /artifacts/{id} returns the full record:

json
{
  "id":             "art_b4d8c0f1",
  "user_id":        "usr_…",
  "filename":       "invoice-2026-001.pdf",
  "artifact_type":  "application/pdf",
  "artifact_size":  184320,
  "artifact_hash":  "sha256:9c1f…",
  "status":         "completed",
  "uploaded_at":    "2026-05-03T09:12:04Z",
  "processing_started_at":   "2026-05-03T09:12:06Z",
  "processing_completed_at": "2026-05-03T09:12:11Z",
  "processing_duration_seconds": 5,
  "s3_path_raw":    "s3://dd-uploads/usr_…/art_b4d8c0f1.pdf",
  "tags":           ["fy26", "us-east"],
  "description":    "Acme Corp invoice batch, May",
  "extracted_data": { "confidence": 0.94, "needs_review": false, … },
  "extraction_jobs": [{ "id": "job_…", "status": "completed", … }]
}

The fields you'll touch most often:

FieldNotes
idUse for download, rerun, cancel, status, metadata patch.
statusLifecycle state. See below.
artifact_hashSHA-256 of the uploaded bytes. Useful for client-side dedupe.
tags, descriptionMutable metadata. Patch with PATCH /artifacts/{id}.
extracted_dataPresent only after a job succeeds.
processing_duration_secondsDrives the bill. Visible after completion.

Status lifecycle

text
                  PUT to S3          POST /artifacts/{id}/confirm-upload         worker pulls
pending_upload  ─────────────▶  pending_upload  ────────────────────────▶  uploaded  ───────────▶  processing
                (bytes in S3,                                                                          │
                 backend not                                  ┌────────────────────────────────────────┤
                 yet notified)                                ▼                                        ▼
                                                          completed                       failed / quarantined / cancelled

The artifact does not advance to uploaded when the PUT to S3 finishes. The platform only learns about the upload when the client calls POST /artifacts/{id}/confirm-upload. Skip that call and the artifact stays in pending_upload indefinitely.

ArtifactStatusMeaning
pending_uploadSubmitted. Either the bytes haven't arrived in S3 yet, or they have but no one has called confirm-upload to tell the backend.
uploadedconfirm-upload succeeded; the job is queued.
processingA worker is actively running the pipeline.
completedExtraction finished successfully. extracted_data is populated.
failedPipeline raised. See error_details_json.
quarantinedFile flagged (malware, banned MIME, signed corruption). Cannot rerun.
cancelledYou stopped the in-flight job. No charge for compute that didn't happen.

For high-frequency polling, prefer GET /artifacts/{id}/status over the full artifact response; it returns just status, processed_at, and error, which is much cheaper to serve and parse.

Artifacts vs jobs

An artifact is the file. A job is one attempt to extract from it.

  • One artifact, one job: the common case. Submit, upload, success.
  • One artifact, many jobs: you reran the extraction (template changed, transient failure, schema iteration).

Every job is reflected in extraction_jobs[] on the artifact. Each job has its own id, status, created_at, and bills independently. Old jobs are preserved for audit. The most recent successful job's output is what extracted_data reflects.

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/rerun

A rerun returns 202 and a new job_id. Useful when the artifact is fine but the extraction logic changed.

Downloading results

GET /artifacts/{id}/download?file_type={raw|json|md} returns a short-lived presigned URL for the requested file:

file_typeContents
rawThe original uploaded file.
json (default)Structured extraction result.
mdMarkdown rendering of the extracted fields.
bash
URL=$(curl -s -u "$DD_KEY:$DD_SECRET" \
  "https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/download?file_type=json" \
  | jq -r .download_url)

curl -s "$URL" > result.json

The download URL is signed against your account and expires in a few minutes. Don't cache it; re-request on each download.

Mutable metadata

Three fields on an artifact are user-mutable after upload:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X PATCH https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1 \
  -H 'Content-Type: application/json' \
  -d '{
    "filename":    "renamed.pdf",
    "description": "FY26 Q2, Acme batch",
    "tags":        ["fy26", "q2", "acme"]
  }'

Use tags for downstream filtering (e.g. tax year, business unit, source system). Tags are searchable in the dashboard and surface in webhook payloads.

Deletion and retention

DELETE /artifacts/{id} removes the artifact record and its stored files immediately. There's no soft delete, no recycle bin.

For automated cleanup, set a retention policy at submission time (retention_policy: "7d", "30d", etc.); the platform will delete the file and its results when the policy expires. The artifact record (metadata + extraction result) can be retained longer than the source file.

Esc
↑↓Navigate↵OpenEscClose