Artifacts #
What an artifact is, the statuses it walks through, and how it relates to jobs and results.
An artifact is the API's record of a single file you've uploaded: its
metadata, its storage location, its processing state, and any extraction
results derived from it. Every POST /extract creates one artifact and one
initial job. Reruns create new jobs against the same artifact.
Anatomy of an artifact
GET /artifacts/{id} returns the full record:
{
"id": "art_b4d8c0f1",
"user_id": "usr_…",
"filename": "invoice-2026-001.pdf",
"artifact_type": "application/pdf",
"artifact_size": 184320,
"artifact_hash": "sha256:9c1f…",
"status": "completed",
"uploaded_at": "2026-05-03T09:12:04Z",
"processing_started_at": "2026-05-03T09:12:06Z",
"processing_completed_at": "2026-05-03T09:12:11Z",
"processing_duration_seconds": 5,
"s3_path_raw": "s3://dd-uploads/usr_…/art_b4d8c0f1.pdf",
"tags": ["fy26", "us-east"],
"description": "Acme Corp invoice batch, May",
"extracted_data": { "confidence": 0.94, "needs_review": false, … },
"extraction_jobs": [{ "id": "job_…", "status": "completed", … }]
}
The fields you'll touch most often:
| Field | Notes |
|---|---|
id | Use for download, rerun, cancel, status, metadata patch. |
status | Lifecycle state. See below. |
artifact_hash | SHA-256 of the uploaded bytes. Useful for client-side dedupe. |
tags, description | Mutable metadata. Patch with PATCH /artifacts/{id}. |
extracted_data | Present only after a job succeeds. |
processing_duration_seconds | Drives the bill. Visible after completion. |
Status lifecycle
PUT to S3 POST /artifacts/{id}/confirm-upload worker pulls
pending_upload ─────────────▶ pending_upload ────────────────────────▶ uploaded ───────────▶ processing
(bytes in S3, │
backend not ┌────────────────────────────────────────┤
yet notified) ▼ ▼
completed failed / quarantined / cancelled
The artifact does not advance to uploaded when the PUT to S3
finishes. The platform only learns about the upload when the client calls
POST /artifacts/{id}/confirm-upload. Skip that call and the artifact
stays in pending_upload indefinitely.
ArtifactStatus | Meaning |
|---|---|
pending_upload | Submitted. Either the bytes haven't arrived in S3 yet, or they have but no one has called confirm-upload to tell the backend. |
uploaded | confirm-upload succeeded; the job is queued. |
processing | A worker is actively running the pipeline. |
completed | Extraction finished successfully. extracted_data is populated. |
failed | Pipeline raised. See error_details_json. |
quarantined | File flagged (malware, banned MIME, signed corruption). Cannot rerun. |
cancelled | You stopped the in-flight job. No charge for compute that didn't happen. |
For high-frequency polling, prefer GET /artifacts/{id}/status
over the full artifact response; it returns just status, processed_at,
and error, which is much cheaper to serve and parse.
Artifacts vs jobs
An artifact is the file. A job is one attempt to extract from it.
- One artifact, one job: the common case. Submit, upload, success.
- One artifact, many jobs: you reran the extraction (template changed, transient failure, schema iteration).
Every job is reflected in extraction_jobs[] on the artifact. Each job has
its own id, status, created_at, and bills independently. Old jobs are
preserved for audit. The most recent successful job's output is what
extracted_data reflects.
curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/rerun
A rerun returns 202 and a new job_id. Useful when the artifact is fine
but the extraction logic changed.
Downloading results
GET /artifacts/{id}/download?file_type={raw|json|md} returns a short-lived
presigned URL for the requested file:
file_type | Contents |
|---|---|
raw | The original uploaded file. |
json (default) | Structured extraction result. |
md | Markdown rendering of the extracted fields. |
URL=$(curl -s -u "$DD_KEY:$DD_SECRET" \ "https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/download?file_type=json" \ | jq -r .download_url) curl -s "$URL" > result.json
The download URL is signed against your account and expires in a few minutes. Don't cache it; re-request on each download.
Mutable metadata
Three fields on an artifact are user-mutable after upload:
curl -u "$DD_KEY:$DD_SECRET" \
-X PATCH https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1 \
-H 'Content-Type: application/json' \
-d '{
"filename": "renamed.pdf",
"description": "FY26 Q2, Acme batch",
"tags": ["fy26", "q2", "acme"]
}'
Use tags for downstream filtering (e.g. tax year, business unit, source
system). Tags are searchable in the dashboard and surface in webhook payloads.
Deletion and retention
DELETE /artifacts/{id} removes the artifact record and its stored files
immediately. There's no soft delete, no recycle bin.
For automated cleanup, set a retention policy
at submission time (retention_policy: "7d", "30d", etc.); the platform
will delete the file and its results when the policy expires. The artifact
record (metadata + extraction result) can be retained longer than the source
file.