I build Filewhisk, a set of file tools that run entirely in the browser, and this week I added "resize PDF pages": change every page to A4, US Letter or any size. The obvious way to do that with pdf-lib takes about ten lines. It also quietly deletes every link and bookmark in the document.
Here is the shortcut, what it costs, and the in-place approach I ended up with.
The ten-line version
const src = await PDFDocument.load(bytes);
const out = await PDFDocument.create();
const embedded = await out.embedPdf(bytes, src.getPageIndices());
for (const ep of embedded) {
const page = out.addPage([595.28, 841.89]); // A4
const s = Math.min(595.28 / ep.width, 841.89 / ep.height);
page.drawPage(ep, {
x: (595.28 - ep.width * s) / 2,
y: (841.89 - ep.height * s) / 2,
xScale: s, yScale: s,
});
}
embedPdf turns each old page into a form XObject, a reusable drawing, and drawPage stamps it onto a fresh blank page. The text stays vector and selectable. It looks perfect.
What it loses
I ran both versions on the same test files and read the results back with pdf.js, independently of the code that produced them:
| File | Links | Bookmarks | Text items | |
|---|---|---|---|---|
| 5-page report | original | 4 | 3 | 137 |
| embedPage | 0 | 0 | 137 | |
| in place | 4 | 3 | 137 | |
| 3 pages with links | original | 4 | 0 | 3 |
| embedPage | 0 | 0 | 3 | |
| in place | 4 | 0 | 3 |
The reason is structural. Links, comments and form fields are not part of a page's drawing. They are annotations, separate objects listed in the page's /Annots array, with their own rectangles in page coordinates. A form XObject only carries the content stream, so the annotations stay behind on the old page, which is never written out. Bookmarks point at page objects that no longer exist, so they go too. Fillable form fields are annotations as well; on a form whose pages had no content stream at all, embedPdf simply threw Can't embed page with missing Contents.
Changing the page in place instead
The alternative is to keep every page object and change three things on it: the drawing, the page box, and every coordinate that refers to the page.
1. Wrap the existing content in a transform. A content stream can start with q (save state) and a cm matrix, and end with Q. Everything in between is scaled and moved. I also clip to the old visible box, so anything that was hidden outside the crop box stays hidden:
const pre = `q ${a} 0 0 ${d} ${e} ${f} cm ${x} ${y} ${w} ${h} re W n\n`;
const preRef = ctx.register(ctx.stream(pre));
const postRef = ctx.register(ctx.stream('\nQ'));
node.set(PDFName.of('Contents'),
ctx.obj([preRef, ...existingContentRefs, postRef]));
node.set(PDFName.of('MediaBox'), ctx.obj([0, 0, W, H]));
['CropBox', 'BleedBox', 'TrimBox', 'ArtBox'].forEach(k => node.delete(PDFName.of(k)));
Prepending and appending separate streams means the original content stream is never decoded or rewritten, so compressed streams stay byte-identical.
For "fit", the matrix is a uniform scale s = min(W / w, H / h) plus a translation that centres the old box (x, y, w, h) on the new page: e = (W - s*w) / 2 - s*x, and the same for f. The - s*x term matters: crop boxes don't always start at the origin.
2. Move every annotation with the same matrix. /Rect is the obvious one. The less obvious ones are /QuadPoints (text highlights), /Vertices (polygons), /L (lines), /CL (callouts) and /InkList (freehand drawings, an array of arrays). Miss /QuadPoints and highlights stay where the text used to be.
3. Fix every destination. A link or bookmark that jumps to "page 3, at this position" stores [pageRef /XYZ left top zoom]. The page reference is still valid, because the page object survived, but left and top are coordinates on the old page. Destinations live in link annotations, in the outline, in named destination trees and in actions, so instead of chasing each structure I walk every indirect object, find arrays shaped like a destination, look up the transform of the page they point to, and apply it:
if (kind === 'XYZ') {
if (left !== null) set(2, t.a * left + t.e);
if (top !== null) set(3, t.d * top + t.f);
} else if (kind === 'FitH' || kind === 'FitBH') {
if (top !== null) set(2, t.d * top + t.f);
} // FitV, FitR likewise
Rotated pages
/Rotate is easy to get wrong. A page with /Rotate 90 is displayed turned clockwise, so a "portrait A4" request for a page that looks portrait means a landscape box in the page's own, unrotated coordinates. I compute the target size in display orientation, swap width and height when the rotation is 90 or 270, keep /Rotate as it is, and do all the matrix work in unrotated space.
Telling the user when something gets cut
"Fill" and "keep size" modes can push content off the new page. Rather than guess from the page box, I render the strip that will be cut with pdf.js at about 50 dpi and look for dark pixels. If there is ink there, the page says which pages lose something; if it is only blank margin, it says nothing is cut.
One gotcha in my own helper
My shared PDF loader removes page references from annotations when it opens a file, to avoid a different bloat problem with copyPages. For resizing that was exactly wrong: it would have broken the very links I was trying to keep. The resize tool loads the document with plain PDFDocument.load instead.
Results
On the 5-page report, the in-place version kept 4 of 4 links and 3 of 3 bookmarks, and the file came out smaller than the original (32.6 KB against 37.4 KB), because nothing was duplicated. The embedPage version kept 0 of 4 and 0 of 3. On the link test file, every link rectangle and every destination landed within 0.6 pt of where the matrix puts it.
If you only need pages as pictures of themselves, embedPage is fine. If anyone will click anything in the result, change the page in place.
The tool is at filewhisk.com/resize-pdf-pages. It runs locally in the browser; nothing is uploaded.
Top comments (0)