Monitoring jobs #
Polling strategies, lightweight status checks, cancellation, and rerun semantics.
This guide covers everything that happens between "I submitted a job" and "I have a result" without using webhooks: how to poll efficiently, when to use the lightweight status endpoint, how to cancel an in-flight job, and how reruns work.
Two status endpoints, two purposes
| Endpoint | Returns | Use it for |
|---|---|---|
GET /job/{job_id} | Full job: status, timestamps, result, result_download_url, error. | Final status check + result fetch. |
GET /artifacts/{id}/status | Just status, processed_at, error. | High-frequency polling. |
If you're polling once at the end to fetch the result, /job/{id} is fine.
If you're polling every second to surface a progress bar, use the
artifact-status endpoint; it's purpose-built for that.
A backoff-friendly polling loop
Linear polling at 1Hz is fine for prototypes. For anything with traffic, back off; most jobs complete in a few seconds, the long tail is rarer.
import time, requests
def wait_for_job(job_id: str, max_wait: int = 600) -> dict:
start, delay = time.monotonic(), 1.0
while time.monotonic() - start < max_wait:
r = requests.get(f'{BASE}/job/{job_id}', auth=AUTH)
r.raise_for_status()
job = r.json()
if job['status'] in ('success', 'failed', 'cancelled', 'expired'):
return job
time.sleep(delay)
delay = min(delay * 1.5, 15) # 1, 1.5, 2.25, … cap 15s
raise TimeoutError(f'job {job_id} did not complete in {max_wait}s')
For job collections (e.g. an entire batch), poll the artifact-status endpoint and only fetch the full job once it's terminal:
import asyncio, httpx
async def wait_for_batch(artifact_ids, base, auth):
pending = set(artifact_ids)
results = {}
async with httpx.AsyncClient(auth=auth) as client:
while pending:
for aid in list(pending):
r = await client.get(f'{base}/artifacts/{aid}/status')
s = r.json()['status']
if s in ('completed', 'failed', 'cancelled', 'quarantined'):
job = await client.get(f'{base}/artifacts/{aid}')
results[aid] = job.json()
pending.discard(aid)
await asyncio.sleep(1)
return results
Status transitions
Job statuses (/job/{id}):
pending ─ file not yet uploaded │ queued_for_processing ─ confirm-upload was called; waiting for a worker │ running ─ actively processing │ ├─▶ success (result + result_download_url populated) ├─▶ failed (error populated) ├─▶ cancelled (you called /cancel) └─▶ expired (timed out or archived)
Artifact statuses (/artifacts/{id} and /artifacts/{id}/status):
pending_upload ──confirm-upload──▶ uploaded ─▶ processing ─▶ completed | failed | quarantined | cancelled
The artifact only leaves pending_upload when the client calls
POST /artifacts/{id}/confirm-upload. A job stuck on pending usually
means the confirm call was skipped.
Only the four terminal states matter for "is it done?": completed,
failed, quarantined, cancelled. Any other status means keep polling.
Cancelling an in-flight job
POST /artifacts/{id}/cancel stops a running job, releases the held
balance back to the wallet, and best-effort kills the GPU worker.
curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/cancel
Returns:
200 OKif the cancellation took effect.409 Conflictif the artifact is already in a terminal state. Idempotent; a second cancel after success is a 409 not a 5xx.
Use cancellation when:
- The user retracts their submission.
- You detect the wrong file was uploaded (hash mismatch).
- A circuit breaker fires upstream and you want to drop in-flight work.
Cancellation is not a refund mechanism for already-completed work. If a
job has reached success, the compute is already paid for; cancel returns 409.
Re-running a completed or failed artifact
POST /artifacts/{id}/rerun queues a fresh job against an existing artifact.
Works for any terminal state: failed (try again with a fixed schema),
completed (re-extract with a new template), or cancelled (resume work
you abandoned).
curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/rerun
A new Job row is created with a new job_id. The original job stays in
the artifact's extraction_jobs[] array; billing history is preserved,
not overwritten.
If the artifact's retention policy has
already purged the source file, rerun returns 400. Choose a longer
retention policy upfront if you anticipate reruns.
Debugging a failure
When a job lands in failed, the error field on the job response is the
first thing to look at; it's a short string identifying the failure
category and any specific detail.
For the rich payload (stack-shaped diagnostics, partial results, etc.),
fetch the artifact and look at error_details_json:
curl -u "$DD_KEY:$DD_SECRET" \ https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1 | \ jq '.error_details_json'
If the failure was during webhook delivery (job succeeded, callback
didn't), the data lives elsewhere; see
GET /job/{id}/webhook-logs.