DEV Community

Vitalii Holben
Vitalii Holben

Posted on

I Built a Screenshot Pipeline That Handles 10K URLs Daily


title: I Built a Screenshot Pipeline That Handles 10K URLs Daily (Here's What Broke First) published: true tags: webdev, api, automation, javascript

Last quarter I needed to generate preview thumbnails for a link aggregation tool. The requirement seemed simple: given a URL, produce a 1200x630 screenshot. At first it was maybe 50 URLs a day. Then the product grew.

At around 2,000 URLs per day things started falling apart.

The first bottleneck: memory

Each Puppeteer page instance uses somewhere between 50-150MB of RAM depending on the page complexity. Running 10 concurrent pages meant 1.5GB just for Chrome, plus Node overhead. Our 4GB server started swapping.

// the naive approach that doesn't scale
const pages = urls.map(async (url) => {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'networkidle0' });
  await page.screenshot({ path: `${hash(url)}.png` });
  await page.close();
});
await Promise.all(pages); // boom, 10K pages at once

The fix was a simple semaphore — limit to 4 concurrent pages and queue the rest. But even with that, Chrome would occasionally leak memory and need a restart every ~500 screenshots.

The second bottleneck: timeouts

Some URLs just don't load. Or they load but never reach networkidle0 because of persistent WebSocket connections or infinite polling. I was losing 15-20% of my daily runs to timeout-related failures.

What worked:

await page.goto(url, {
  waitUntil: 'domcontentloaded', // not networkidle0
  timeout: 15000
});
// then wait for specific readiness signal
await page.waitForFunction(
  () => document.readyState === 'complete',
  { timeout: 5000 }
).catch(() => {}); // screenshot whatever we have

Switching from networkidle0 to domcontentloaded + a short secondary wait cut my failure rate from 18% to under 3%.

The third bottleneck: me

At some point I was spending more time maintaining the screenshot infrastructure than building the actual product. Chrome updates breaking things, memory leak debugging, managing the screenshot queue, handling retries, storage rotation...

I switched the whole thing to ScreenshotRun — send a URL via API, get back an image. No browser management, no server scaling, no 3am Chrome crashes.

// entire screenshot pipeline, after
const response = await fetch('https://api.screenshotrun.com/capture', {
  method: 'POST',
  headers: { 'Authorization': `Bearer ${API_KEY}` },
  body: JSON.stringify({ url, width: 1200, height: 630 })
});
const image = await response.buffer();

My screenshot server went from a dedicated 8GB instance to zero infrastructure. The API handles the browser pool, the timeouts, the retries, the scaling.

What I'd do differently

If I was starting over and knew I'd hit scale:

  1. Don't self-host screenshots past ~500/day — the operational overhead isn't worth it unless you need very specific browser customization
  2. Always use domcontentloaded over networkidle0 — the few pages that need network idle aren't worth the timeout penalty on everything else
  3. Store WebP, not PNG — 60-70% smaller files, virtually identical quality for thumbnails
  4. Build the retry logic from day one — transient failures are guaranteed at scale, not edge cases

Has anyone else hit similar scaling issues with browser-based tools? Curious what concurrency limits others settled on.

Top comments (0)