DEV Community

Cover image for I benchmarked 3 PDF-to-Markdown converters on 5 public PDFs, including the one I built. Here is where it loses
Giacomo
Giacomo

Posted on Originally published at cleanmd.dev

I benchmarked 3 PDF-to-Markdown converters on 5 public PDFs, including the one I built. Here is where it loses

Every PDF-to-Markdown tool claims it "preserves the structure". I build one of them (CleanMD), and I started it because a DeepSeek paper I wanted in my notes came out of every tool I tried as soup. So I had an obvious bias and an obvious question: how would I know if that claim were false about mine?

I ended up writing a small benchmark. Five public PDFs from five genres, three tools with default settings, and a ground truth that comes from the documents themselves, never from any tool's output. Here is what I found, losses first.

The setup

Documents (all public, none picked after seeing results):

The five public PDFs used in the benchmark (RFC 9110, BERT, Think Python, a Supreme Court opinion, NIST CSF 2.0) with page counts and why each one is hard to convert

Document Why it is hard
RFC 9110 HTTP Semantics (194 pp.) 291 numbered sections, ABNF grammar, ASCII diagrams, running footer
BERT (arXiv 1810.04805, 16 pp.) two-column paper with result tables and captions right next to them
Think Python 2e (292 pp.) technical book: table of contents, 500+ code listings
Loper Bright v. Raimondo (SCOTUS, 114 pp.) legal opinion: parts I/II and sub-parts A/B are single centered letters at body size
NIST CSF 2.0 (32 pp.) standard with appendices and tables with multi-line cells

Tools: MarkItDown 0.1.7 (Microsoft), pymupdf4llm 1.28.2 (Artifex), and CleanMD 0.93.0 (mine, the same engine that runs in the browser, executed in Node). Pandoc is not in the table because it cannot read PDF at all (PDF is an output format for it): return code 21 on all five files.

Ground truth: RFC sections from the table of contents of the official rfc9110.txt; Think Python's 126 entries from the book's own ToC; BERT's 29 numbered headings from the lines set in the Medium weight; NIST's 8 entries from its ToC. For the court opinion there is no ToC, so the 31 part markers were defined geometrically (a single token, centered, at body size), and that definition is close to the heuristic my own engine uses, which I flag on the page. Weigh that document accordingly.

Metrics: sections recognised as headings (and at a depth consistent with the numbering), fenced code blocks, tables and "prose cells" (a table cell holding a sentence = a fake or contaminated table), plus a few document-specific checks like "is the collected ABNF one block?".

Where CleanMD loses

The 'Where CleanMD loses' section of the benchmark page: RFC section recall on the first run, fence count on Think Python, and prose cells on NIST, with the competitor ahead in each row

  • RFC 9110 section recall, first run: pymupdf4llm 290/291, CleanMD 280/291. Eleven bold, body-size headings (8.8.2 Last-Modified, 13.1.3 If-Modified-Since, 15.5.10 409 Conflict…) came out of CleanMD as plain paragraphs. The cause turned out to be embarrassing and specific: pdf.js splits a line into two font ids when a glyph (the hyphen, here) comes from a second subset of the same bold font, and my detector only looked at lines with one uniform font. Fixed the same day; the re-run is 291/291. I am keeping the first number in the article and on the page, because a benchmark that only shows the after is marketing.
  • Fence count on Think Python: pymupdf4llm 657, CleanMD 569. With a caveat: 328 of pymupdf4llm's fences are single-line fences wrapping inline code words. The operator listing that opens §5.2 (x != y … x >= y) is a fence in CleanMD and not there.
  • Prose cells on NIST: MarkItDown 0, CleanMD 23. Most of CleanMD's are the Tier table in Appendix B, whose cells are paragraphs by design (it now comes out as one table per page instead of one block per Tier, so the metric counts every wrapped line); MarkItDown's tables are cleaner on that document. It also emits zero headings on it, so I wouldn't take that trade, but the row is there.

Where it wins

  • Code survives. RFC 9110: 158 fenced blocks vs 0 and 0. The ABNF grammar, the message diagrams and the examples come out of the other two as prose; the collected ABNF appendix is one block only in CleanMD.
  • Running footers are gone. "Fielding, et al. Standards Track Page N" left in the text: 0 vs 187 (MarkItDown) vs 194 (pymupdf4llm).
  • Hierarchy has depth. Think Python: 122 of 126 sections at the right depth for CleanMD; pymupdf4llm finds 113 but emits all of them as H4, so the outline is flat. On the court opinion, 28 of 31 part markers become headings (I → H2, A → H3, 1 → H4); pymupdf4llm gets 18 and promotes the running page header to an H3 on 80 pages.
  • No fake tables. MarkItDown turns BERT into 134 pseudo-tables (the title becomes a 9-column table) and Think Python's code listings into 683 tables.

The numbers

CleanMD MarkItDown pymupdf4llm
RFC 9110: sections / 291 (right depth) 291 (291) · first run 280 0 290 (275)
RFC 9110: fences / ABNF one block 158 / yes 0 / no 0 / no
RFC 9110: footer lines in text 0 187 194
BERT: sections / 29 26 0 26
BERT: tables (prose cells) 11 (0) 134 (0) 9 (0)
Think Python: sections / 126 (right depth) 122 (122) 0 113 (0)
Think Python: fences (single-line) 569 (52) 0 657 (328)
SCOTUS: part markers as headings / 31 28 0 18
SCOTUS: running headers promoted 0 0 80
NIST: sections / 8 8 0 8
Seconds, RFC 9110 (same laptop) 1.3 5.5 10.9

What this does not tell you

Five documents are five documents. Nothing here is scanned (all five have a text layer), nothing measures reading order inside dense three-column layouts (where all three tools struggle), nothing measures math or prose quality. It shows failure modes, not a universal ranking.

Reproduce it

The page has the full tables with the best value per row highlighted and a kit you can download: run.sh (downloads the PDFs, installs the two open-source tools in a venv, builds the ground truth, scores everything in about two minutes), the ground truth, the results, and the raw Markdown CleanMD produced, so every sentence above can be checked against the actual output.

→ https://cleanmd.dev/benchmarks/pdf-to-markdown

If you know a public PDF that breaks all three, I want it.

Top comments (6)

Collapse
 
okinawasoftware profile image
Okinawa Software Lab •

I read your PDF-to-Markdown benchmark page after your comment, and it changed what I built next, so I wanted to say so here.

Two things stuck with me. First, using the document's own numbering (1 → H1, 1.1 → H2) as the ground truth for heading depth, instead of anything the tools emit. Second, scoring "contamination" (prose turned into tables) as a failure rather than ignoring it. Both are the opposite of "more structure is always better", which is what I would have assumed.

So the next update of my export (v1.15.0, submitting now) adds headings, and only headings. A line becomes H1 to H3 when it is short, clearly larger or bolder than the page's body text, and, if it carries a numbering pattern, the number's depth wins over the font size. Numbered list items, emphasized words, and amounts stay as body text. Anything uncertain stays as body text too; I would rather miss a heading than invent one.

I left tables out on purpose. Ruled tables in Japanese business forms might be doable from the drawing objects, but unruled tables are exactly the contamination case your scorer penalizes, and I don't have your benchmark ground truth to validate that against.

One question, if you have time: how does CleanMD decide heading depth for documents with no numbering at all (reports, letters), where only font size and spacing are available? That's the case where my heuristic is weakest, and I'm curious whether you found a signal I missed.

Thanks for the pointer. It was a better design review than most I've had.

Collapse
 
dylnd0g profile image
Giacomo •

Glad the page was useful, and v1.15.0 sounds like the right cut: headings only, and silence when unsure. That is also the choice I would make in your place.

On your question. For documents with no numbering I stopped treating font size as an absolute and started treating it as a rank. The body size is the most common size in the document. Anything at least 15% larger is a heading candidate, and the three largest candidate sizes become H1, H2, H3 in that order. So a 14pt line is not "an H2" by itself; it is an H2 only if 14 is the second largest recurring size in that document. In a report with 18/14/12 you get three levels, in a letter with only 12 on top of 10 you get one. That alone fixed more than anything else I tried.

The word "recurring" is doing real work there. A size only earns a rung on the ladder if it appears on at least three lines across at least three pages. That rule came from a TI datasheet where the block-diagram labels lived on two pages at a big font and were being promoted to H3. It also handles the title page: sizes that appear only on page one, on one or two lines, in a document of three pages or more, are the title, the subtitle and the author. The largest of them is the H1, the others are not structure at all. Before that rule "Kenneth Reitz" at 17pt was a section.

Then there are two ladders below the size ladder, for the cases where the size does not move. Lines set entirely in the bold face at roughly body size get their own levels under the size levels, but only if bold is a minority at body size, otherwise the document is simply set in bold and weight means nothing. And lines in small caps get levels under those, with a density guard: a page of a paper carries two or three section titles at most, a bibliography page carries fifteen to forty small-caps author names, so above five small-caps lines on one page that size is not structure. The bibliography case cost me a whole afternoon.

Spacing I use less than you would think. It decides paragraph breaks, not depth. What I do use is a text guard: over 120 characters, or ending in a period, comma or colon, is never a heading whatever the font says.

Where it still fails is honest to say. A document whose title is small and whose part headings are huge but live on fewer than three pages loses those parts as a level. And a figure label at 19.5pt in the Attention paper once sat above the 17pt title, took the top rung, and shifted the whole hierarchy down by one; the fix was that nothing non-recurring may sit above the title. I expect Japanese business forms have their own version of that trap.

If you ever want to test against the same five PDFs, the ground truth is going up as a public repo, with the SCOTUS one rebuilt from the Cornell HTML first, as promised in the comment above.

Collapse
 
okinawasoftware profile image
Okinawa Software Lab •

Thank you, this is exactly the kind of answer I was hoping for, and more concrete than anything I had found written down.

Where I am now: after v1.15.0 shipped, I built a small benchmark (synthetic documents with known headings), and it caught a bug from the same family as your Attention example. I was deciding levels page by page, so from page 2 onward everything sat one rung too high. v1.15.1, now live, carries the title size and heading ladder across the whole document. My levels are already relative in the way you describe: the title is H1, numbered headings map to their numbering depth plus one, and the remaining size steps become H2/H3. But whether a line is a candidate is still based on a fixed ratio to the body size (1.15 / 1.3 / 1.6), rather than its rank. Your rank ladder is the cleaner version of what I have, and I will try it against my test set.

Two of your rules I don't have yet and want to test: "nothing non-recurring may sit above the title," and "bold only counts when bold is a minority at body size." Right now I trust bold only when numbering or a gap above backs it up, which is weaker.

On the Japanese trap you expected: the biggest one is that your recurrence rule can rarely fire. Japanese business documents (estimates, invoices, notices) are usually one or two pages, so "three lines on three pages" almost never happens, and I have to lean on size ratio, numbering, and spacing instead.

The traps I actually hit:

  • Stamps and badges. "見本" (SAMPLE) printed three times in large type, and short badges like OK / 注意 (caution) / 警告 (warning) at heading size. I now skip repeated strings, and a size-only H3 needs at least three characters.
  • Letter-spaced logos, where a company name is set one character at a time and looks like a short large line.
  • Bold labels with nothing under them. On a form, a bold field label is not a section. A bold-only candidate that is immediately followed by another heading is treated as body text.
  • A small subtitle sitting directly above the title. Similar to your title-page rule, but on a one-page document.
  • Numbering comes in many shapes: 第1章, 1., 1.1, full-width 1., (1). And your punctuation guard needs the Japanese marks too: 。、:
  • Bold is less reliable. Many Japanese fonts have no bold face, and a non-embedded Helvetica-Bold reports weight 0 through pdfium, so I fall back to the font name.

No small caps, at least.

I would be glad to run mine on your five PDFs once the ground truth is up, and I'll report where it loses, the same way you did.

Collapse
 
compoundlabs profile image
Compound Labs •

How do you separate heading accuracy from heuristic agreement when the SCOTUS ground truth uses a geometric rule close to CleanMD's detector? That choice could reward matching the implementation rather than identifying headings.

Collapse
 
dylnd0g profile image
Giacomo •

Fair point, and it is the reason that row carries a caveat on the page. For the other four documents the ground truth comes from something none of the tools ever sees (a table of contents, the official .txt). The opinion has no ToC, so I defined the 31 markers by geometry: one token, centered, body size. That is close to what my detector does, and on that document a perfect score would prove little. I agree with you there.

Two things keep it from being fully circular, though. The rule generated the list, but the list can be checked without the rule. The 31 markers form a clean outline with no gaps (I A B C II A B C III A B 1 2 3 IV for the majority, then the concurrence, then I II III IV for the dissent), and a centered "I" picked up by accident would break the sequence somewhere. Today I also compared the majority opinion against the Cornell LII HTML, where every marker is its own paragraph: 15 for 15. I have not done the concurrence and the dissent yet.

The other thing is that the score is 28 out of 31. CleanMD misses A, B and C of Part I of the Gorsuch concurrence, two of them fused into the paragraph and one lost entirely. Sharing the idea with the ground truth did not hand it the points. And pymupdf4llm reaches 18 of the same 31 by a completely different route, so those lines are not something only my heuristic can see.

Still, the cleaner fix is the one you are hinting at: build the list from the LII HTML and keep the geometry only as a cross-check. It is a small change to gt_build.py in the kit and I will make it.

Collapse
 
argumentmoney8117 profile image
ArgumentMoney8117 •

Useful benchmark, and the honest caveats are what make it credible — most people would have let the SCOTUS row stand without a word. The single-line fence finding is a good reminder that raw counts need normalization before they mean anything.