Batch PDF Generation: Patterns for Rendering Thousands of Documents

Generating one PDF is easy. Generating ten thousand on a deadline is a systems problem. Here are the patterns that hold up in production.

pdfbatch processingnode.jsqueuesperformancebackend

Batch PDF Generation: Patterns for Rendering Thousands of Documents

Generating a single PDF is a solved problem. You pick a library, render a template, return a buffer. Done. But the moment you need to generate ten thousand invoices on the first of the month, or render a million statements overnight, the naive approach falls apart fast. Memory balloons, headless browsers crash, downstream services time out, and your retry logic turns into a footgun.

This post covers the patterns that actually hold up when you need to generate PDFs in bulk.

The naive approach (and why it fails)

Most batch jobs start like this:

js
for (const user of users) { const pdf = await renderPdf(user); await s3.putObject({ Key: `${user.id}.pdf`, Body: pdf }); }

This works for 50 users. At 50,000, you'll hit at least one of these:

  • Memory pressure: each PDF buffer sits in memory until uploaded.
  • Headless browser leaks: Chromium processes accumulate if you don't manage them carefully.
  • No backpressure: a slow S3 upload blocks the next render.
  • No isolation: one bad template input crashes the entire job.

Sequential is too slow. Unbounded Promise.all is a denial-of-service attack on your own infrastructure. You need something in between.

Pattern 1: Bounded concurrency with a worker pool

The simplest improvement: process N items in parallel, never more. Libraries like p-limit make this trivial.

js
import pLimit from 'p-limit'; const limit = pLimit(10); // tune based on CPU + memory await Promise.all( users.map(user => limit(async () => { try { const pdf = await renderPdf(user); await uploadToS3(user.id, pdf); } catch (err) { await markFailed(user.id, err); } }) ) );

A few things to note:

  • The concurrency limit is the most important number in your job. Profile it. CPU-bound rendering benefits from numCores; I/O-heavy workloads can go higher.
  • Wrap each task in try/catch. One failure must not abort the batch.
  • Stream the upload if your storage client supports it — don't hold the buffer longer than necessary.

Pattern 2: Queue-backed jobs

Inline batch processing breaks when the job takes longer than your request timeout, your deploy window, or the patience of whoever clicked the button. Move the work to a queue.

js
// enqueue for (const user of users) { await queue.add('render-pdf', { userId: user.id }, { attempts: 3, backoff: { type: 'exponential', delay: 2000 }, }); } // worker new Worker('render-pdf', async (job) => { const user = await db.users.findById(job.data.userId); const pdf = await renderPdf(user); await uploadToS3(user.id, pdf); }, { concurrency: 8 });

Benefits you get for free:

  • Retries with backoff for transient failures (network blips, cold starts).
  • Horizontal scaling by adding more worker instances.
  • Visibility: dashboards like Bull Board show what's stuck.
  • Idempotency: jobs can be safely re-run if you key uploads by a deterministic ID.

BullMQ, SQS, and Cloud Tasks all work well here. Pick whatever fits your stack.

Pattern 3: Offload rendering entirely

Running headless Chromium at scale is its own job. You're managing process lifecycles, font caches, sandbox flags, and memory leaks that only show up after 6 hours. For many teams, this isn't core engineering work.

A managed PDF API like Kamy takes the rendering off your plate — you POST HTML or a template ID, get a PDF (or a URL) back. Your worker becomes:

js
new Worker('render-pdf', async (job) => { const user = await db.users.findById(job.data.userId); const res = await fetch('https://api.kamy.dev/v1/pdf', { method: 'POST', headers: { 'Authorization': `Bearer ${process.env.KAMY_KEY}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ template: 'invoice', data: user }), }); const pdf = Buffer.from(await res.arrayBuffer()); await uploadToS3(user.id, pdf); }, { concurrency: 25 });

Concurrency goes up because workers no longer hold a Chromium process. Failure modes shrink to "API call failed," which your retry logic already handles.

Pattern 4: Aggregate the result

Sometimes the deliverable isn't 10,000 PDFs — it's one zip, or a single merged file. Don't merge in memory.

js
import archiver from 'archiver'; const zip = archiver('zip', { zlib: { level: 6 } }); zip.pipe(fs.createWriteStream('batch.zip')); for await (const { id, stream } of pdfStreams) { zip.append(stream, { name: `${id}.pdf` }); } await zip.finalize();

Stream PDFs directly into the archive, never landing them on disk if you can avoid it. For merged single PDFs, pdf-lib or a server-side merge endpoint is the way.

Operational guardrails

A few things worth building before you need them:

  • Per-batch IDs: tag every job with a batchId so you can query progress, cancel, or retry the whole set.
  • Dead letter queues: failed jobs go somewhere you can inspect, not /dev/null.