Errors #
Status codes, error envelopes, retry guidance, and common failure modes across every endpoint.
The DataDistillers API uses standard HTTP status codes and a consistent error envelope. This page covers the canonical error shapes, the status-code matrix, and how to retry safely.
Error envelope
All non-2xx responses share the same envelope:
{
"error": {
"code": "template_not_found",
"message": "No template with id tpl_unknown",
"request_id": "req_8f3c2e1a"
}
}
| Field | Notes |
|---|---|
code | Stable, machine-readable identifier. Switch on this in code. |
message | Human-readable. Subject to change; never parse it. |
request_id | Pass to support to correlate with server-side logs. Always include it. |
Validation errors (HTTP 422) use FastAPI's HTTPValidationError shape
instead: an array of { loc, msg, type } entries pointing at the offending
field path.
{
"detail": [
{
"loc": ["body", "artifact_size"],
"msg": "ensure this value is greater than 0",
"type": "value_error.number.not_gt"
}
]
}
Status-code matrix
| Status | Class | Retryable? | Common causes |
|---|---|---|---|
400 | Client | no | Validation failure, both template_id and extraction_schema, invalid webhook URL. |
401 | Client | no | Missing or invalid Authorization: Basic. |
402 | Client | conditional | Wallet frozen or spendable insufficient. Retry after adding funds. |
403 | Client | no | Resource belongs to a different account, or key lacks scope. |
404 | Client | no | job_id, artifact_id, template_id, webhook_id not found. |
409 | Client | no | Cancel against a terminal job; duplicate creation. |
422 | Client | no | Body shape didn't pass schema validation. |
429 | Client | yes | Rate limit. See Retry-After. |
5xx | Server | yes | Transient service issue. Retry with backoff. |
Only 429 and 5xx are safe to retry automatically. Use exponential
backoff with jitter, starting at 1s and capped at 60s. Cap total attempts
at 5; beyond that you're papering over a real outage.
Common error codes
code | Status | Meaning |
|---|---|---|
unauthorized | 401 | Missing or malformed credentials. |
invalid_credentials | 401 | Key/secret pair rejected. |
wallet_frozen | 402 | New jobs blocked until the freeze clears. |
insufficient_balance | 402 | spendable < estimated job cost. |
template_not_found | 404 | template_id doesn't exist or was hard-deleted. |
webhook_not_found | 404 | webhook_id doesn't exist. |
artifact_not_found | 404 | artifact_id doesn't exist. |
job_not_found | 404 | job_id doesn't exist. |
schema_template_conflict | 400 | Both template_id and extraction_schema were sent. |
invalid_webhook_url | 400 | Non-HTTPS, private IP, or > 2083 chars. |
duplicate_filename | 400 | Two files in a batch share a filename. |
terminal_state | 409 | Cancel attempted against a job in a terminal status. |
rate_limited | 429 | Too many requests. See Retry-After. |
internal_error | 500 | Server-side bug. Include request_id if reporting. |
Job failure errors (status: failed)
status: failed)A job that finishes with status: "failed" carries an error string with
the failure category and a short detail:
{
"job_id": "job_8f3c2e1a",
"status": "failed",
"error": "schema_mismatch: required field 'total' not found"
}
| Prefix | Cause | Fix |
|---|---|---|
schema_mismatch | A required field couldn't be located. | Loosen required, lower confidence_threshold, or fix the schema. |
unsupported_type | MIME type not supported by the pipeline. | Convert the file or use a supported format. |
corrupted_file | File failed to open / decode. | Re-upload from the source. |
timeout | Pipeline exceeded the per-job time budget. | Split the document or contact support if recurring. |
quarantined | File flagged by safety checks. | Review the file; cannot rerun. |
For richer diagnostics, fetch the artifact and look at error_details_json;
it contains the structured error payload, partial extraction (if any),
and worker-side stack info when applicable.
Implementing retry-with-backoff
A correct retry loop respects status code, includes jitter, and caps attempts:
import random, time, requests
RETRYABLE = {429, 500, 502, 503, 504}
def retrying(call, max_attempts: int = 5):
for attempt in range(1, max_attempts + 1):
r = call()
if r.status_code not in RETRYABLE:
return r # success or fatal
if attempt == max_attempts:
return r # give up
# Honour Retry-After if the server set it
wait = float(r.headers.get('Retry-After', 0)) or \
min(60, (2 ** attempt) + random.random())
time.sleep(wait)
For 429s, prefer the server-supplied Retry-After header over your own
backoff schedule; the server has more context about when to come back.
Correlating with support
Every error response includes a request_id. Capture it in your logs:
try:
r = requests.post(url, auth=AUTH, json=payload)
r.raise_for_status()
except requests.HTTPError:
rid = r.json().get('error', {}).get('request_id')
log.error(f'DD API {r.status_code}: request_id={rid}')
raise
When opening a support ticket, include the request_id of the failing call
plus the rough timestamp. This is the fastest path to a server-side trace.