Updated Apr 27, 2026
reference

Errors #

Status codes, error envelopes, retry guidance, and common failure modes across every endpoint.

The DataDistillers API uses standard HTTP status codes and a consistent error envelope. This page covers the canonical error shapes, the status-code matrix, and how to retry safely.

Error envelope

All non-2xx responses share the same envelope:

json
{
  "error": {
    "code":       "template_not_found",
    "message":    "No template with id tpl_unknown",
    "request_id": "req_8f3c2e1a"
  }
}
FieldNotes
codeStable, machine-readable identifier. Switch on this in code.
messageHuman-readable. Subject to change; never parse it.
request_idPass to support to correlate with server-side logs. Always include it.

Validation errors (HTTP 422) use FastAPI's HTTPValidationError shape instead: an array of { loc, msg, type } entries pointing at the offending field path.

json
{
  "detail": [
    {
      "loc":  ["body", "artifact_size"],
      "msg":  "ensure this value is greater than 0",
      "type": "value_error.number.not_gt"
    }
  ]
}

Status-code matrix

StatusClassRetryable?Common causes
400ClientnoValidation failure, both template_id and extraction_schema, invalid webhook URL.
401ClientnoMissing or invalid Authorization: Basic.
402ClientconditionalWallet frozen or spendable insufficient. Retry after adding funds.
403ClientnoResource belongs to a different account, or key lacks scope.
404Clientnojob_id, artifact_id, template_id, webhook_id not found.
409ClientnoCancel against a terminal job; duplicate creation.
422ClientnoBody shape didn't pass schema validation.
429ClientyesRate limit. See Retry-After.
5xxServeryesTransient service issue. Retry with backoff.
When to retry

Only 429 and 5xx are safe to retry automatically. Use exponential backoff with jitter, starting at 1s and capped at 60s. Cap total attempts at 5; beyond that you're papering over a real outage.

Common error codes

codeStatusMeaning
unauthorized401Missing or malformed credentials.
invalid_credentials401Key/secret pair rejected.
wallet_frozen402New jobs blocked until the freeze clears.
insufficient_balance402spendable < estimated job cost.
template_not_found404template_id doesn't exist or was hard-deleted.
webhook_not_found404webhook_id doesn't exist.
artifact_not_found404artifact_id doesn't exist.
job_not_found404job_id doesn't exist.
schema_template_conflict400Both template_id and extraction_schema were sent.
invalid_webhook_url400Non-HTTPS, private IP, or > 2083 chars.
duplicate_filename400Two files in a batch share a filename.
terminal_state409Cancel attempted against a job in a terminal status.
rate_limited429Too many requests. See Retry-After.
internal_error500Server-side bug. Include request_id if reporting.

Job failure errors (status: failed)

A job that finishes with status: "failed" carries an error string with the failure category and a short detail:

json
{
  "job_id": "job_8f3c2e1a",
  "status": "failed",
  "error":  "schema_mismatch: required field 'total' not found"
}
PrefixCauseFix
schema_mismatchA required field couldn't be located.Loosen required, lower confidence_threshold, or fix the schema.
unsupported_typeMIME type not supported by the pipeline.Convert the file or use a supported format.
corrupted_fileFile failed to open / decode.Re-upload from the source.
timeoutPipeline exceeded the per-job time budget.Split the document or contact support if recurring.
quarantinedFile flagged by safety checks.Review the file; cannot rerun.

For richer diagnostics, fetch the artifact and look at error_details_json; it contains the structured error payload, partial extraction (if any), and worker-side stack info when applicable.

Implementing retry-with-backoff

A correct retry loop respects status code, includes jitter, and caps attempts:

py
import random, time, requests

RETRYABLE = {429, 500, 502, 503, 504}

def retrying(call, max_attempts: int = 5):
    for attempt in range(1, max_attempts + 1):
        r = call()
        if r.status_code not in RETRYABLE:
            return r                                      # success or fatal
        if attempt == max_attempts:
            return r                                      # give up

        # Honour Retry-After if the server set it
        wait = float(r.headers.get('Retry-After', 0)) or \
               min(60, (2 ** attempt) + random.random())
        time.sleep(wait)

For 429s, prefer the server-supplied Retry-After header over your own backoff schedule; the server has more context about when to come back.

Correlating with support

Every error response includes a request_id. Capture it in your logs:

py
try:
    r = requests.post(url, auth=AUTH, json=payload)
    r.raise_for_status()
except requests.HTTPError:
    rid = r.json().get('error', {}).get('request_id')
    log.error(f'DD API {r.status_code}: request_id={rid}')
    raise

When opening a support ticket, include the request_id of the failing call plus the rough timestamp. This is the fastest path to a server-side trace.

Esc
↑↓Navigate↵OpenEscClose