Batch processing #
Submit up to 50 documents in one call. One template, one webhook, parallel uploads.
The POST /batch endpoint takes up to 50 file descriptors in one request,
returns one presigned upload URL per file, and queues every job under a
shared batch_id. All files share the same template, schema, and webhook
configuration. It's the right shape for nightly imports, bulk backfills, and
"my user dragged 30 invoices into the upload widget" flows.
When to use batch instead of /extract
Use /batch when | Use /extract per file when |
|---|---|
| ≥ 5 files at a time | One file from a user-driven action |
| All files share template/schema | Each file needs different settings |
| One webhook covers them all | You want per-job webhook overrides |
| You can upload concurrently | You're constrained to serial uploads |
50 is a hard cap per request. For more than 50, submit multiple batches.
Step 1 · Submit the batch
curl -u "$DD_KEY:$DD_SECRET" \
-X POST https://api.datadistillers.com/api/v1/batch \
-H 'Content-Type: application/json' \
-d '{
"template_id": "tpl_invoice_v3",
"files": [
{ "filename": "invoice-001.pdf", "artifact_type": "application/pdf", "artifact_size": 184320 },
{ "filename": "invoice-002.pdf", "artifact_type": "application/pdf", "artifact_size": 192011 },
{ "filename": "invoice-003.pdf", "artifact_type": "application/pdf", "artifact_size": 175920 }
]
}'
Response (202 Accepted):
{
"batch_id": "bat_z9y8x7",
"total": 3,
"files": [
{
"filename": "invoice-001.pdf",
"job_id": "job_aaaa",
"artifact_id": "art_aaaa",
"upload_url": "https://s3.amazonaws.com/…?X-Amz-Signature=…",
"fields": {}
},
{
"filename": "invoice-002.pdf",
"job_id": "job_bbbb",
"artifact_id": "art_bbbb",
"upload_url": "https://s3.amazonaws.com/…?X-Amz-Signature=…",
"fields": {}
},
{ "filename": "invoice-003.pdf", "job_id": "job_cccc", "artifact_id": "art_cccc",
"upload_url": "https://…", "fields": {} }
]
}
Each file has its own job_id and upload_url, but they share the
batch_id for grouping in the dashboard and webhook payloads.
The API rejects duplicate filenames within a single batch with 400. If
you have two invoice.pdf from different sources, prefix or rename
before submitting.
Step 2 · Upload concurrently
Each upload URL is independent. Push them all in parallel rather than serially; that's the whole point of batch.
For 50 files, eight parallel uploads will saturate most home connections. Tune up if you're on a wide office or datacenter pipe.
Step 3 · Confirm the batch
Just like the single-file flow, the backend does not watch S3 for the
uploaded objects. Once every (or as many as you can) presigned PUT has
returned, call POST /artifacts/confirm-batch-upload/{batch_id} so the
platform knows the batch is ready to process:
curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/confirm-batch-upload/bat_z9y8x7
The endpoint walks every artifact in the batch:
- Files that uploaded successfully are flipped to
uploaded, their jobs toqueued_for_processing, and dispatched. - Files that never uploaded are left in
pending_upload. They are not enqueued and not held against your wallet. - Already-confirmed files are skipped — confirm is idempotent, so calling it again after a partial upload is safe.
Until you call this endpoint, every job in the batch sits in pending.
Step 4 · Track completions
Two patterns, depending on whether you set up a webhook:
With a webhook: register once on the API key, set the events to
extraction.completed and extraction.failed, and you'll receive one POST
per job as it finishes. The payload includes batch_id for grouping.
Without a webhook: poll each job_id independently:
remaining = {f['job_id']: f for f in batch['files']}
results = {}
while remaining:
for jid in list(remaining):
job = requests.get(f'{BASE}/job/{jid}', auth=AUTH).json()
if job['status'] in ('success', 'failed', 'cancelled', 'expired'):
results[jid] = job
del remaining[jid]
time.sleep(2)
For high-volume polling, prefer the lighter
GET /artifacts/{id}/status over /job/{id}.
Lower-level variant: request-then-confirm
POST /batch is the convenient front-door; under the hood it's the same
shape as POST /artifacts/request-batch-upload followed by the same
confirm endpoint described above. The lower-level pair is useful if you
want to assemble the manifest separately from triggering the batch.
# 1. Request URLs (same response shape as POST /batch). curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/request-batch-upload \ -H 'Content-Type: application/json' \ -d @manifest.json # 2. Upload all files to the returned upload_url values (parallel, as above). # 3. Confirm; queues all uploaded files in one call. curl -u "$DD_KEY:$DD_SECRET" \ -X POST https://api.datadistillers.com/api/v1/artifacts/confirm-batch-upload/bat_z9y8x7
Why not just loop /extract?
Compared with calling /extract 50 times in a loop, batch saves:
- 49 round-trips to the API server (one batch call vs 50 single calls).
- Per-request overhead for auth and submission validation.
- Webhook setup: one config covers the whole batch.
- Dashboard noise: the dashboard groups batches into one card.
It does not speed up the actual extraction; jobs queue at the same priority as single submissions. If you want to lower latency, parallelise the uploads (see step 2); that's where the wall-clock time lives.