Updated Apr 27, 2026
guides

Monitoring jobs #

Polling strategies, lightweight status checks, cancellation, and rerun semantics.

This guide covers everything that happens between "I submitted a job" and "I have a result" without using webhooks: how to poll efficiently, when to use the lightweight status endpoint, how to cancel an in-flight job, and how reruns work.

Two status endpoints, two purposes

EndpointReturnsUse it for
GET /job/{job_id}Full job: status, timestamps, result, result_download_url, error.Final status check + result fetch.
GET /artifacts/{id}/statusJust status, processed_at, error.High-frequency polling.

If you're polling once at the end to fetch the result, /job/{id} is fine. If you're polling every second to surface a progress bar, use the artifact-status endpoint; it's purpose-built for that.

A backoff-friendly polling loop

Linear polling at 1Hz is fine for prototypes. For anything with traffic, back off; most jobs complete in a few seconds, the long tail is rarer.

py
import time, requests

def wait_for_job(job_id: str, max_wait: int = 600) -> dict:
    start, delay = time.monotonic(), 1.0
    while time.monotonic() - start < max_wait:
        r = requests.get(f'{BASE}/job/{job_id}', auth=AUTH)
        r.raise_for_status()
        job = r.json()
        if job['status'] in ('success', 'failed', 'cancelled', 'expired'):
            return job
        time.sleep(delay)
        delay = min(delay * 1.5, 15)              # 1, 1.5, 2.25, … cap 15s
    raise TimeoutError(f'job {job_id} did not complete in {max_wait}s')

For job collections (e.g. an entire batch), poll the artifact-status endpoint and only fetch the full job once it's terminal:

py
import asyncio, httpx

async def wait_for_batch(artifact_ids, base, auth):
    pending = set(artifact_ids)
    results = {}
    async with httpx.AsyncClient(auth=auth) as client:
        while pending:
            for aid in list(pending):
                r = await client.get(f'{base}/artifacts/{aid}/status')
                s = r.json()['status']
                if s in ('completed', 'failed', 'cancelled', 'quarantined'):
                    job = await client.get(f'{base}/artifacts/{aid}')
                    results[aid] = job.json()
                    pending.discard(aid)
            await asyncio.sleep(1)
    return results

Status transitions

Job statuses (/job/{id}):

text
pending                ─ file not yet uploaded
   │
queued_for_processing  ─ confirm-upload was called; waiting for a worker
   │
running                ─ actively processing
   │
   ├─▶ success    (result + result_download_url populated)
   ├─▶ failed     (error populated)
   ├─▶ cancelled  (you called /cancel)
   └─▶ expired    (timed out or archived)

Artifact statuses (/artifacts/{id} and /artifacts/{id}/status):

text
pending_upload ──confirm-upload──▶ uploaded ─▶ processing ─▶ completed | failed | quarantined | cancelled

The artifact only leaves pending_upload when the client calls POST /artifacts/{id}/confirm-upload. A job stuck on pending usually means the confirm call was skipped.

Only the four terminal states matter for "is it done?": completed, failed, quarantined, cancelled. Any other status means keep polling.

Cancelling an in-flight job

POST /artifacts/{id}/cancel stops a running job, releases the held balance back to the wallet, and best-effort kills the GPU worker.

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/cancel

Returns:

  • 200 OK if the cancellation took effect.
  • 409 Conflict if the artifact is already in a terminal state. Idempotent; a second cancel after success is a 409 not a 5xx.

Use cancellation when:

  • The user retracts their submission.
  • You detect the wrong file was uploaded (hash mismatch).
  • A circuit breaker fires upstream and you want to drop in-flight work.

Cancellation is not a refund mechanism for already-completed work. If a job has reached success, the compute is already paid for; cancel returns 409.

Re-running a completed or failed artifact

POST /artifacts/{id}/rerun queues a fresh job against an existing artifact. Works for any terminal state: failed (try again with a fixed schema), completed (re-extract with a new template), or cancelled (resume work you abandoned).

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1/rerun

A new Job row is created with a new job_id. The original job stays in the artifact's extraction_jobs[] array; billing history is preserved, not overwritten.

Rerun needs the source file

If the artifact's retention policy has already purged the source file, rerun returns 400. Choose a longer retention policy upfront if you anticipate reruns.

Debugging a failure

When a job lands in failed, the error field on the job response is the first thing to look at; it's a short string identifying the failure category and any specific detail.

For the rich payload (stack-shaped diagnostics, partial results, etc.), fetch the artifact and look at error_details_json:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  https://api.datadistillers.com/api/v1/artifacts/art_b4d8c0f1 | \
  jq '.error_details_json'

If the failure was during webhook delivery (job succeeded, callback didn't), the data lives elsewhere; see GET /job/{id}/webhook-logs.

Esc
↑↓Navigate↵OpenEscClose