Updated Apr 27, 2026
guides

Batch processing #

Submit up to 50 documents in one call. One template, one webhook, parallel uploads.

The POST /batch endpoint takes up to 50 file descriptors in one request, returns one presigned upload URL per file, and queues every job under a shared batch_id. All files share the same template, schema, and webhook configuration. It's the right shape for nightly imports, bulk backfills, and "my user dragged 30 invoices into the upload widget" flows.

When to use batch instead of /extract

Use /batch whenUse /extract per file when
≥ 5 files at a timeOne file from a user-driven action
All files share template/schemaEach file needs different settings
One webhook covers them allYou want per-job webhook overrides
You can upload concurrentlyYou're constrained to serial uploads

50 is a hard cap per request. For more than 50, submit multiple batches.

Step 1 · Submit the batch

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/batch \
  -H 'Content-Type: application/json' \
  -d '{
    "template_id": "tpl_invoice_v3",
    "files": [
      { "filename": "invoice-001.pdf", "artifact_type": "application/pdf", "artifact_size": 184320 },
      { "filename": "invoice-002.pdf", "artifact_type": "application/pdf", "artifact_size": 192011 },
      { "filename": "invoice-003.pdf", "artifact_type": "application/pdf", "artifact_size": 175920 }
    ]
  }'

Response (202 Accepted):

json
{
  "batch_id": "bat_z9y8x7",
  "total":    3,
  "files": [
    {
      "filename":    "invoice-001.pdf",
      "job_id":      "job_aaaa",
      "artifact_id": "art_aaaa",
      "upload_url":  "https://s3.amazonaws.com/…?X-Amz-Signature=…",
      "fields":      {}
    },
    {
      "filename":    "invoice-002.pdf",
      "job_id":      "job_bbbb",
      "artifact_id": "art_bbbb",
      "upload_url":  "https://s3.amazonaws.com/…?X-Amz-Signature=…",
      "fields":      {}
    },
    { "filename": "invoice-003.pdf", "job_id": "job_cccc", "artifact_id": "art_cccc",
      "upload_url": "https://…", "fields": {} }
  ]
}

Each file has its own job_id and upload_url, but they share the batch_id for grouping in the dashboard and webhook payloads.

Duplicate filenames are rejected

The API rejects duplicate filenames within a single batch with 400. If you have two invoice.pdf from different sources, prefix or rename before submitting.

Step 2 · Upload concurrently

Each upload URL is independent. Push them all in parallel rather than serially; that's the whole point of batch.

For 50 files, eight parallel uploads will saturate most home connections. Tune up if you're on a wide office or datacenter pipe.

Step 3 · Confirm the batch

Just like the single-file flow, the backend does not watch S3 for the uploaded objects. Once every (or as many as you can) presigned PUT has returned, call POST /artifacts/confirm-batch-upload/{batch_id} so the platform knows the batch is ready to process:

bash
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/confirm-batch-upload/bat_z9y8x7

The endpoint walks every artifact in the batch:

  • Files that uploaded successfully are flipped to uploaded, their jobs to queued_for_processing, and dispatched.
  • Files that never uploaded are left in pending_upload. They are not enqueued and not held against your wallet.
  • Already-confirmed files are skipped — confirm is idempotent, so calling it again after a partial upload is safe.

Until you call this endpoint, every job in the batch sits in pending.

Step 4 · Track completions

Two patterns, depending on whether you set up a webhook:

With a webhook: register once on the API key, set the events to extraction.completed and extraction.failed, and you'll receive one POST per job as it finishes. The payload includes batch_id for grouping.

Without a webhook: poll each job_id independently:

py
remaining = {f['job_id']: f for f in batch['files']}
results = {}

while remaining:
    for jid in list(remaining):
        job = requests.get(f'{BASE}/job/{jid}', auth=AUTH).json()
        if job['status'] in ('success', 'failed', 'cancelled', 'expired'):
            results[jid] = job
            del remaining[jid]
    time.sleep(2)

For high-volume polling, prefer the lighter GET /artifacts/{id}/status over /job/{id}.

Lower-level variant: request-then-confirm

POST /batch is the convenient front-door; under the hood it's the same shape as POST /artifacts/request-batch-upload followed by the same confirm endpoint described above. The lower-level pair is useful if you want to assemble the manifest separately from triggering the batch.

bash
# 1. Request URLs (same response shape as POST /batch).
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/request-batch-upload \
  -H 'Content-Type: application/json' \
  -d @manifest.json

# 2. Upload all files to the returned upload_url values (parallel, as above).

# 3. Confirm; queues all uploaded files in one call.
curl -u "$DD_KEY:$DD_SECRET" \
  -X POST https://api.datadistillers.com/api/v1/artifacts/confirm-batch-upload/bat_z9y8x7

Why not just loop /extract?

Compared with calling /extract 50 times in a loop, batch saves:

  • 49 round-trips to the API server (one batch call vs 50 single calls).
  • Per-request overhead for auth and submission validation.
  • Webhook setup: one config covers the whole batch.
  • Dashboard noise: the dashboard groups batches into one card.

It does not speed up the actual extraction; jobs queue at the same priority as single submissions. If you want to lower latency, parallelise the uploads (see step 2); that's where the wall-clock time lives.

Esc
↑↓Navigate↵OpenEscClose