Pipelines #
List pipelines and submit jobs to a specific pipeline (legacy /pipeline_async endpoint).
A pipeline is a saved configuration that bundles a template, default
webhook, and processing settings under one ID. Most accounts have exactly
one pipeline auto-created with the API key; the unified POST /extract
endpoint resolves it automatically, so you rarely need to touch pipelines
directly.
The endpoints below are kept for legacy integrations and for accounts that maintain multiple named pipelines.
GET /pipelines
GET /pipelinesLists all pipelines belonging to the authenticated account.
curl -u "$DD_KEY:$DD_SECRET" \ https://api.datadistillers.com/api/v1/pipelines
Response:
[
{
"id": "pl_default",
"name": "default",
"is_active": true,
"template_id": "tpl_invoice_v3",
"execution_count": 12482,
"created_at": "2026-01-12T00:00:00Z"
}
]
| Field | Notes |
|---|---|
id | Reference value for POST /pipeline_async and usage filtering. |
name | Human-readable. |
is_active | If false, no new jobs accepted under this pipeline. |
template_id | The default template the pipeline applies. May be null. |
execution_count | Lifetime count of jobs executed under this pipeline. |
created_at | When the pipeline was created. |
POST /pipeline_async (legacy)
POST /pipeline_async (legacy)Submit a job against a specific pipeline. Same upload-then-poll flow as
POST /extract, but requires you to pass
pipeline_id explicitly.
This endpoint exists for legacy clients. New integrations should use
POST /extract, which auto-resolves the pipeline from the API key. The
shape and behavior are otherwise identical.
curl -u "$DD_KEY:$DD_SECRET" \
-X POST https://api.datadistillers.com/api/v1/pipeline_async \
-H 'Content-Type: application/json' \
-d '{
"pipeline_id": "pl_default",
"filename": "invoice.pdf",
"artifact_type": "application/pdf",
"artifact_size": 184320,
"template_id": "tpl_invoice_v3"
}'
| Field | Required | Notes |
|---|---|---|
pipeline_id | yes | The pipeline to run under. |
filename | yes | Original filename. |
artifact_type | yes | MIME type. |
artifact_size | yes | File size in bytes. |
retention_policy | no | Defaults to 30d. |
template_id | no | Override the pipeline default. |
extraction_schema | no | Inline schema. Mutually exclusive with template_id. |
webhook_id | no | Override pipeline-default webhook. |
webhook_url | no | One-shot webhook URL. |
Returns 202 Accepted + AsyncJobStartResponse:
{
"job_id": "job_8f3c2e1a",
"upload_url": "https://s3.amazonaws.com/…?X-Amz-Signature=…",
"fields": {},
"expires_in": 3600,
"webhook_signing_secret": null
}
The shape mirrors ExtractResponse minus the artifact_id (legacy); fetch
the artifact ID from GET /job/{id} if you need it.
When to use multiple pipelines
Most teams run on the auto-created default pipeline. Maintain multiple pipelines when you have:
- Distinct templates per business unit, with per-pipeline webhooks routing to different services.
- Separate dev/staging/prod pipelines under one account (cleaner than spinning up separate accounts).
- Tiered SLAs: one pipeline for high-priority customer-facing work, another for batch reconciliation.
If none of those apply, stick with POST /extract and a single pipeline.
Errors
| Status | Cause |
|---|---|
400 | Invalid request shape, both template_id and extraction_schema, etc. |
401 | Auth missing or invalid. |
403 | Pipeline belongs to a different account. |
404 | pipeline_id doesn't exist or is inactive. |
422 | Body validation failed. |