title: I Built a Screenshot Pipeline That Handles 10K URLs Daily (Here's What Broke First) published: true tags: webdev, api, automation, javascript
Last quarter I needed to generate preview thumbnails for a link aggregation tool. The requirement seemed simple: given a URL, produce a 1200x630 screenshot. At first it was maybe 50 URLs a day. Then the product grew.
At around 2,000 URLs per day things started falling apart.
The first bottleneck: memory
Each Puppeteer page instance uses somewhere between 50-150MB of RAM depending on the page complexity. Running 10 concurrent pages meant 1.5GB just for Chrome, plus Node overhead. Our 4GB server started swapping.
// the naive approach that doesn't scale
const pages = urls.map(async (url) => {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle0' });
await page.screenshot({ path: `${hash(url)}.png` });
await page.close();
});
await Promise.all(pages); // boom, 10K pages at once
The fix was a simple semaphore — limit to 4 concurrent pages and queue the rest. But even with that, Chrome would occasionally leak memory and need a restart every ~500 screenshots.
The second bottleneck: timeouts
Some URLs just don't load. Or they load but never reach networkidle0 because of persistent WebSocket connections or infinite polling. I was losing 15-20% of my daily runs to timeout-related failures.
What worked:
await page.goto(url, {
waitUntil: 'domcontentloaded', // not networkidle0
timeout: 15000
});
// then wait for specific readiness signal
await page.waitForFunction(
() => document.readyState === 'complete',
{ timeout: 5000 }
).catch(() => {}); // screenshot whatever we have
Switching from networkidle0 to domcontentloaded + a short secondary wait cut my failure rate from 18% to under 3%.
The third bottleneck: me
At some point I was spending more time maintaining the screenshot infrastructure than building the actual product. Chrome updates breaking things, memory leak debugging, managing the screenshot queue, handling retries, storage rotation...
I switched the whole thing to ScreenshotRun — send a URL via API, get back an image. No browser management, no server scaling, no 3am Chrome crashes.
// entire screenshot pipeline, after
const response = await fetch('https://api.screenshotrun.com/capture', {
method: 'POST',
headers: { 'Authorization': `Bearer ${API_KEY}` },
body: JSON.stringify({ url, width: 1200, height: 630 })
});
const image = await response.buffer();
My screenshot server went from a dedicated 8GB instance to zero infrastructure. The API handles the browser pool, the timeouts, the retries, the scaling.
What I'd do differently
If I was starting over and knew I'd hit scale:
- Don't self-host screenshots past ~500/day — the operational overhead isn't worth it unless you need very specific browser customization
-
Always use
domcontentloadedovernetworkidle0— the few pages that need network idle aren't worth the timeout penalty on everything else - Store WebP, not PNG — 60-70% smaller files, virtually identical quality for thumbnails
- Build the retry logic from day one — transient failures are guaranteed at scale, not edge cases
Has anyone else hit similar scaling issues with browser-based tools? Curious what concurrency limits others settled on.
Top comments (0)