Every PDF-to-Markdown tool claims it "preserves the structure". I build one of them (CleanMD), and I started it because a DeepSeek paper I wanted in my notes came out of every tool I tried as soup. So I had an obvious bias and an obvious question: how would I know if that claim were false about mine?
I ended up writing a small benchmark. Five public PDFs from five genres, three tools with default settings, and a ground truth that comes from the documents themselves, never from any tool's output. Here is what I found, losses first.
The setup
Documents (all public, none picked after seeing results):
| Document | Why it is hard |
|---|---|
| RFC 9110 HTTP Semantics (194 pp.) | 291 numbered sections, ABNF grammar, ASCII diagrams, running footer |
| BERT (arXiv 1810.04805, 16 pp.) | two-column paper with result tables and captions right next to them |
| Think Python 2e (292 pp.) | technical book: table of contents, 500+ code listings |
| Loper Bright v. Raimondo (SCOTUS, 114 pp.) | legal opinion: parts I/II and sub-parts A/B are single centered letters at body size |
| NIST CSF 2.0 (32 pp.) | standard with appendices and tables with multi-line cells |
Tools: MarkItDown 0.1.7 (Microsoft), pymupdf4llm 1.28.2 (Artifex), and CleanMD 0.93.0 (mine, the same engine that runs in the browser, executed in Node). Pandoc is not in the table because it cannot read PDF at all (PDF is an output format for it): return code 21 on all five files.
Ground truth: RFC sections from the table of contents of the official rfc9110.txt; Think Python's 126 entries from the book's own ToC; BERT's 29 numbered headings from the lines set in the Medium weight; NIST's 8 entries from its ToC. For the court opinion there is no ToC, so the 31 part markers were defined geometrically (a single token, centered, at body size), and that definition is close to the heuristic my own engine uses, which I flag on the page. Weigh that document accordingly.
Metrics: sections recognised as headings (and at a depth consistent with the numbering), fenced code blocks, tables and "prose cells" (a table cell holding a sentence = a fake or contaminated table), plus a few document-specific checks like "is the collected ABNF one block?".
Where CleanMD loses
- RFC 9110 section recall, first run: pymupdf4llm 290/291, CleanMD 280/291. Eleven bold, body-size headings (8.8.2 Last-Modified, 13.1.3 If-Modified-Since, 15.5.10 409 Conflict…) came out of CleanMD as plain paragraphs. The cause turned out to be embarrassing and specific: pdf.js splits a line into two font ids when a glyph (the hyphen, here) comes from a second subset of the same bold font, and my detector only looked at lines with one uniform font. Fixed the same day; the re-run is 291/291. I am keeping the first number in the article and on the page, because a benchmark that only shows the after is marketing.
-
Fence count on Think Python: pymupdf4llm 657, CleanMD 569. With a caveat: 328 of pymupdf4llm's fences are single-line fences wrapping inline code words. The operator listing that opens §5.2 (
x != y … x >= y) is a fence in CleanMD and not there. - Prose cells on NIST: MarkItDown 0, CleanMD 23. Most of CleanMD's are the Tier table in Appendix B, whose cells are paragraphs by design (it now comes out as one table per page instead of one block per Tier, so the metric counts every wrapped line); MarkItDown's tables are cleaner on that document. It also emits zero headings on it, so I wouldn't take that trade, but the row is there.
Where it wins
- Code survives. RFC 9110: 158 fenced blocks vs 0 and 0. The ABNF grammar, the message diagrams and the examples come out of the other two as prose; the collected ABNF appendix is one block only in CleanMD.
- Running footers are gone. "Fielding, et al. Standards Track Page N" left in the text: 0 vs 187 (MarkItDown) vs 194 (pymupdf4llm).
- Hierarchy has depth. Think Python: 122 of 126 sections at the right depth for CleanMD; pymupdf4llm finds 113 but emits all of them as H4, so the outline is flat. On the court opinion, 28 of 31 part markers become headings (I → H2, A → H3, 1 → H4); pymupdf4llm gets 18 and promotes the running page header to an H3 on 80 pages.
- No fake tables. MarkItDown turns BERT into 134 pseudo-tables (the title becomes a 9-column table) and Think Python's code listings into 683 tables.
The numbers
| CleanMD | MarkItDown | pymupdf4llm | |
|---|---|---|---|
| RFC 9110: sections / 291 (right depth) | 291 (291) · first run 280 | 0 | 290 (275) |
| RFC 9110: fences / ABNF one block | 158 / yes | 0 / no | 0 / no |
| RFC 9110: footer lines in text | 0 | 187 | 194 |
| BERT: sections / 29 | 26 | 0 | 26 |
| BERT: tables (prose cells) | 11 (0) | 134 (0) | 9 (0) |
| Think Python: sections / 126 (right depth) | 122 (122) | 0 | 113 (0) |
| Think Python: fences (single-line) | 569 (52) | 0 | 657 (328) |
| SCOTUS: part markers as headings / 31 | 28 | 0 | 18 |
| SCOTUS: running headers promoted | 0 | 0 | 80 |
| NIST: sections / 8 | 8 | 0 | 8 |
| Seconds, RFC 9110 (same laptop) | 1.3 | 5.5 | 10.9 |
What this does not tell you
Five documents are five documents. Nothing here is scanned (all five have a text layer), nothing measures reading order inside dense three-column layouts (where all three tools struggle), nothing measures math or prose quality. It shows failure modes, not a universal ranking.
Reproduce it
The page has the full tables with the best value per row highlighted and a kit you can download: run.sh (downloads the PDFs, installs the two open-source tools in a venv, builds the ground truth, scores everything in about two minutes), the ground truth, the results, and the raw Markdown CleanMD produced, so every sentence above can be checked against the actual output.
→ https://cleanmd.dev/benchmarks/pdf-to-markdown
If you know a public PDF that breaks all three, I want it.


Top comments (6)
I read your PDF-to-Markdown benchmark page after your comment, and it changed what I built next, so I wanted to say so here.
Two things stuck with me. First, using the document's own numbering (1 → H1, 1.1 → H2) as the ground truth for heading depth, instead of anything the tools emit. Second, scoring "contamination" (prose turned into tables) as a failure rather than ignoring it. Both are the opposite of "more structure is always better", which is what I would have assumed.
So the next update of my export (v1.15.0, submitting now) adds headings, and only headings. A line becomes H1 to H3 when it is short, clearly larger or bolder than the page's body text, and, if it carries a numbering pattern, the number's depth wins over the font size. Numbered list items, emphasized words, and amounts stay as body text. Anything uncertain stays as body text too; I would rather miss a heading than invent one.
I left tables out on purpose. Ruled tables in Japanese business forms might be doable from the drawing objects, but unruled tables are exactly the contamination case your scorer penalizes, and I don't have your benchmark ground truth to validate that against.
One question, if you have time: how does CleanMD decide heading depth for documents with no numbering at all (reports, letters), where only font size and spacing are available? That's the case where my heuristic is weakest, and I'm curious whether you found a signal I missed.
Thanks for the pointer. It was a better design review than most I've had.
Glad the page was useful, and v1.15.0 sounds like the right cut: headings only, and silence when unsure. That is also the choice I would make in your place.
On your question. For documents with no numbering I stopped treating font size as an absolute and started treating it as a rank. The body size is the most common size in the document. Anything at least 15% larger is a heading candidate, and the three largest candidate sizes become H1, H2, H3 in that order. So a 14pt line is not "an H2" by itself; it is an H2 only if 14 is the second largest recurring size in that document. In a report with 18/14/12 you get three levels, in a letter with only 12 on top of 10 you get one. That alone fixed more than anything else I tried.
The word "recurring" is doing real work there. A size only earns a rung on the ladder if it appears on at least three lines across at least three pages. That rule came from a TI datasheet where the block-diagram labels lived on two pages at a big font and were being promoted to H3. It also handles the title page: sizes that appear only on page one, on one or two lines, in a document of three pages or more, are the title, the subtitle and the author. The largest of them is the H1, the others are not structure at all. Before that rule "Kenneth Reitz" at 17pt was a section.
Then there are two ladders below the size ladder, for the cases where the size does not move. Lines set entirely in the bold face at roughly body size get their own levels under the size levels, but only if bold is a minority at body size, otherwise the document is simply set in bold and weight means nothing. And lines in small caps get levels under those, with a density guard: a page of a paper carries two or three section titles at most, a bibliography page carries fifteen to forty small-caps author names, so above five small-caps lines on one page that size is not structure. The bibliography case cost me a whole afternoon.
Spacing I use less than you would think. It decides paragraph breaks, not depth. What I do use is a text guard: over 120 characters, or ending in a period, comma or colon, is never a heading whatever the font says.
Where it still fails is honest to say. A document whose title is small and whose part headings are huge but live on fewer than three pages loses those parts as a level. And a figure label at 19.5pt in the Attention paper once sat above the 17pt title, took the top rung, and shifted the whole hierarchy down by one; the fix was that nothing non-recurring may sit above the title. I expect Japanese business forms have their own version of that trap.
If you ever want to test against the same five PDFs, the ground truth is going up as a public repo, with the SCOTUS one rebuilt from the Cornell HTML first, as promised in the comment above.
Thank you, this is exactly the kind of answer I was hoping for, and more concrete than anything I had found written down.
Where I am now: after v1.15.0 shipped, I built a small benchmark (synthetic documents with known headings), and it caught a bug from the same family as your Attention example. I was deciding levels page by page, so from page 2 onward everything sat one rung too high. v1.15.1, now live, carries the title size and heading ladder across the whole document. My levels are already relative in the way you describe: the title is H1, numbered headings map to their numbering depth plus one, and the remaining size steps become H2/H3. But whether a line is a candidate is still based on a fixed ratio to the body size (1.15 / 1.3 / 1.6), rather than its rank. Your rank ladder is the cleaner version of what I have, and I will try it against my test set.
Two of your rules I don't have yet and want to test: "nothing non-recurring may sit above the title," and "bold only counts when bold is a minority at body size." Right now I trust bold only when numbering or a gap above backs it up, which is weaker.
On the Japanese trap you expected: the biggest one is that your recurrence rule can rarely fire. Japanese business documents (estimates, invoices, notices) are usually one or two pages, so "three lines on three pages" almost never happens, and I have to lean on size ratio, numbering, and spacing instead.
The traps I actually hit:
No small caps, at least.
I would be glad to run mine on your five PDFs once the ground truth is up, and I'll report where it loses, the same way you did.
How do you separate heading accuracy from heuristic agreement when the SCOTUS ground truth uses a geometric rule close to CleanMD's detector? That choice could reward matching the implementation rather than identifying headings.
Fair point, and it is the reason that row carries a caveat on the page. For the other four documents the ground truth comes from something none of the tools ever sees (a table of contents, the official .txt). The opinion has no ToC, so I defined the 31 markers by geometry: one token, centered, body size. That is close to what my detector does, and on that document a perfect score would prove little. I agree with you there.
Two things keep it from being fully circular, though. The rule generated the list, but the list can be checked without the rule. The 31 markers form a clean outline with no gaps (I A B C II A B C III A B 1 2 3 IV for the majority, then the concurrence, then I II III IV for the dissent), and a centered "I" picked up by accident would break the sequence somewhere. Today I also compared the majority opinion against the Cornell LII HTML, where every marker is its own paragraph: 15 for 15. I have not done the concurrence and the dissent yet.
The other thing is that the score is 28 out of 31. CleanMD misses A, B and C of Part I of the Gorsuch concurrence, two of them fused into the paragraph and one lost entirely. Sharing the idea with the ground truth did not hand it the points. And pymupdf4llm reaches 18 of the same 31 by a completely different route, so those lines are not something only my heuristic can see.
Still, the cleaner fix is the one you are hinting at: build the list from the LII HTML and keep the geometry only as a cross-check. It is a small change to gt_build.py in the kit and I will make it.
Useful benchmark, and the honest caveats are what make it credible — most people would have let the SCOTUS row stand without a word. The single-line fence finding is a good reminder that raw counts need normalization before they mean anything.