Updated Apr 27, 2026
reference

Pipelines #

List pipelines and submit jobs to a specific pipeline (legacy /pipeline_async endpoint).

A pipeline is a saved configuration that bundles a template, default webhook, and processing settings under one ID. Most accounts have exactly one pipeline auto-created with the API key; the unified POST /extract endpoint resolves it automatically, so you rarely need to touch pipelines directly.

The endpoints below are kept for legacy integrations and for accounts that maintain multiple named pipelines.

GET /pipelines

Lists all pipelines belonging to the authenticated account.

bash
curl -u "$DD_KEY:$DD_SECRET" \
  https://api.datadistillers.com/api/v1/pipelines

Response:

json
[
  {
    "id":              "pl_default",
    "name":            "default",
    "is_active":       true,
    "template_id":     "tpl_invoice_v3",
    "execution_count": 12482,
    "created_at":      "2026-01-12T00:00:00Z"
  }
]
FieldNotes
idReference value for POST /pipeline_async and usage filtering.
nameHuman-readable.
is_activeIf false, no new jobs accepted under this pipeline.
template_idThe default template the pipeline applies. May be null.
execution_countLifetime count of jobs executed under this pipeline.
created_atWhen the pipeline was created.

POST /pipeline_async (legacy)

Submit a job against a specific pipeline. Same upload-then-poll flow as POST /extract, but requires you to pass pipeline_id explicitly.

Prefer /extract for new integrations

This endpoint exists for legacy clients. New integrations should use POST /extract, which auto-resolves the pipeline from the API key. The shape and behavior are otherwise identical.

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/pipeline_async \
  -H 'Content-Type: application/json' \
  -d '{
    "pipeline_id":   "pl_default",
    "filename":      "invoice.pdf",
    "artifact_type": "application/pdf",
    "artifact_size": 184320,
    "template_id":   "tpl_invoice_v3"
  }'
FieldRequiredNotes
pipeline_idyesThe pipeline to run under.
filenameyesOriginal filename.
artifact_typeyesMIME type.
artifact_sizeyesFile size in bytes.
retention_policynoDefaults to 30d.
template_idnoOverride the pipeline default.
extraction_schemanoInline schema. Mutually exclusive with template_id.
webhook_idnoOverride pipeline-default webhook.
webhook_urlnoOne-shot webhook URL.

Returns 202 Accepted + AsyncJobStartResponse:

json
{
  "job_id":     "job_8f3c2e1a",
  "upload_url": "https://s3.amazonaws.com/…?X-Amz-Signature=…",
  "fields":     {},
  "expires_in": 3600,
  "webhook_signing_secret": null
}

The shape mirrors ExtractResponse minus the artifact_id (legacy); fetch the artifact ID from GET /job/{id} if you need it.

When to use multiple pipelines

Most teams run on the auto-created default pipeline. Maintain multiple pipelines when you have:

  • Distinct templates per business unit, with per-pipeline webhooks routing to different services.
  • Separate dev/staging/prod pipelines under one account (cleaner than spinning up separate accounts).
  • Tiered SLAs: one pipeline for high-priority customer-facing work, another for batch reconciliation.

If none of those apply, stick with POST /extract and a single pipeline.

Errors

StatusCause
400Invalid request shape, both template_id and extraction_schema, etc.
401Auth missing or invalid.
403Pipeline belongs to a different account.
404pipeline_id doesn't exist or is inactive.
422Body validation failed.
Esc
↑↓Navigate↵OpenEscClose