Going to production #
Checklist for moving an integration from a sandbox key on a laptop to a hardened live deployment: keys, retries, webhooks, rate limits, retention, and observability.
This is the path from "it works on my laptop with a dk_test_… key" to a
production integration you can leave running unattended. Each section is a
short checklist; most of the deep detail lives in the linked pages.
Switch from test to live credentials
Test keys (dk_test_…) return synthetic results and never bill the wallet.
The day you cut over to production:
- Issue a fresh
dk_live_…key from the dashboard. Don't reuse a key that's been pasted into a dev shell. - Move both halves into your secret manager, not into git, CI logs, or container images. See the table in Installation for the per-platform options.
- Add
GET /walletas a deploy-time smoke test. A200confirms the live key resolved correctly inside the prod runtime; a401catches misconfigured secret injection before any user traffic hits it. - Schedule a rotation cadence (most teams: every 90 days) per Authentication → Rotating a secret. The key stays stable; only the secret rotates.
Resist the urge to share one live key across staging and prod. Separate keys let you revoke a leak in one environment without grounding the other, and keep usage analytics clean.
Make every call retry-safe
Network blips, transient 5xxs, and rate-limit 429s are not "if". They
are "when". The full retry contract is in
Errors → Implementing retry-with-backoff;
the production-shape summary:
- Only retry
429and5xx. Everything else is a real failure; don't mask it. - Honour
Retry-Afteron429before falling back to your own exponential schedule. - Use jitter. A fleet of retrying clients without jitter is a self-inflicted thundering herd.
- Cap attempts at ~5. Past that, page a human; the API is probably down and your retry loop is hiding it.
- On
POST /extract, treat the response as side-effecting: if you timed out before reading it, the job may have been accepted. Don't blindly resubmit; query by your own client-side dedup key (e.g.tags) or accept the duplicate and clean up after.
Harden webhook delivery
Polling is fine for prototypes. Production should be webhook-driven so you aren't paying latency on a fixed interval. The full setup is in Webhook setup; the production checklist:
- Verify the HMAC signature on every delivery using the raw request body. A receiver that doesn't verify is an open relay.
- Reject signatures with a timestamp skew over 5 minutes to close the replay window.
- Use
X-Webhook-Idas an idempotency key; the same value is sent across all retries. Persist it before processing. - Ack with 2xx in under a few seconds. Push the actual work to a queue; slow handlers cause retries, retries cause duplicates.
- Plan for the 24h overlap window when rotating the signing secret; verifiers should accept either the current or previous secret during rollouts. See Webhook setup → Rotate.
- Subscribe to
extraction.failedin addition toextraction.completedso failures don't fall on the floor. The endpoint already receives both; make sure your handler routes them differently.
Respect rate limits at the edge of your fleet
Rate limits apply per API key, not per host or per process. A fleet of workers sharing one live key shares one bucket.
- Centralise submit traffic through a shared queue or limiter rather
than letting every worker
POST /extractdirectly. One queue worker enforcing your own concurrency cap is easier to reason about than N workers all hitting429and backing off independently. - Alert on sustained
429rate (not single events;429is fine occasionally). Sustained rate-limiting usually means a workload grew past what's provisioned; talk to support before bolting on heavier retry logic. - Status polling counts. If you're polling
/job/{id}at 1Hz across thousands of in-flight jobs, switch to the artifact-status endpoint or webhooks per Monitoring jobs.
Choose retention deliberately
Retention controls how long the platform stores the uploaded file and the extracted result. The default is fine for development; production should be an explicit choice driven by compliance, not defaults.
- Pick a retention policy per workload, not per account. PII-heavy
extractions should run
immediateor a short TTL; archival workflows that need rerun support need a longer TTL. - Confirm
rerunrequirements upfront. Rerun needs the source file. If retention has purged it,POST /artifacts/{id}/rerunreturns 400. See Monitoring jobs → Rerun. - Document the choice somewhere your security review can find it. See Retention for the policy matrix.
Wire up observability
The platform exposes enough telemetry to drive your existing dashboards. You don't need a custom stack. You need the right four signals.
- Log
request_idon every non-2xx response. It's the only thing that lets support correlate to a server-side trace. See Errors → Correlating with support. - Alert on wallet state, not just balance.
is_frozen: truerejects new jobs with402, so your pipeline will stop end-to-end. A 5-minute cron againstGET /walletcovers it; the script is in Billing → Building a balance monitor. - Track job failure rate as a SLI. A sudden jump usually means an upstream change (new document shape, OCR config drift), not a platform problem. Surface it before users notice.
- Track webhook delivery health via
GET /webhooks/{id}/deliveries?status_filter=failed. A spike in failed deliveries means your endpoint is the problem, not the platform. - Pull weekly spend from
GET /usage?period=7dinto whatever your org uses for budget tracking. The wallet'sbalancelags finalisation;total_spendis authoritative.
Deployment hygiene
A few mechanics that bite people on first prod deploys:
- Run
GET /walletas a post-deploy smoke test gated on the live key. Fail the deploy on401/5xxrather than discovering it from user reports. - Pin the API base URL in one place
(
https://api.datadistillers.com/api/v1). Don't sprinkle it across services; when the version changes, you want one PR. - If you ingest webhooks behind a CDN or WAF, confirm the raw body makes it through unmodified. Body rewrites (gzip strip, JSON reformat) break HMAC verification silently.
- If your platform sets
NODE_ENVautomatically (Vercel, Cloudflare, Netlify) leave it alone. If you're on bare metal or Docker, setNODE_ENV=productionexplicitly. It's a signal to your own runtime, not to this API.
Cutover checklist
The last hour before flipping production traffic:
- Live key in the secret manager, test key removed from the same path.
-
GET /walletreturns200from the prod runtime. - Wallet topped up; auto-recharge configured (or a low-balance pager).
- Webhook registered against the prod URL; signature verifier deployed; one test event ack'd 2xx.
- Retry-with-backoff, idempotency dedup, and
request_idlogging shipped on the same release, not a follow-up. - Retention policy documented and set on the template (or in the per-call config).
- Dashboards for failure rate,
429rate, and wallet state wired into the existing on-call rotation.
If every line is checked, the integration is production-shaped. You can leave it running unattended without an "and one day we should fix X" backlog.