DEV Community

Johmisking
Johmisking

Posted on Fully Autonomous

Pretty JSON costs 3 CSV in tokens. Columnar JSON costs the same. (I measured 9 formats)

Most of us paste data into LLM prompts as JSON, often straight from JSON.stringify(data, null, 2). I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and counted.

Short version: row-by-row JSON, pretty-printed, uses about 3× the tokens of CSV. The same data as columnar JSON costs about the same as CSV. It's the repeated keys, not JSON itself.

Update: a reader pointed out I'd left out columnar JSON (one object, one array per field). I measured it on the same table and tokenizer and added it below.

The setup

A product table: 20 rows, 5 fields (id, name, price, in_stock, category). One row in CSV:

1001,Wireless Mouse,9.99,false,electronics
Enter fullscreen mode Exit fullscreen mode

Same data, seven formats, counted with o200k_base, the tokenizer behind GPT-4o and later OpenAI models:

import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base");
const count = (s) => enc.encode(s).length;

count(csv);                            // 300
count(JSON.stringify(rows));           // 525
count(JSON.stringify(rows, null, 2));  // 884
Enter fullscreen mode Exit fullscreen mode

Results

Format Tokens vs CSV
TSV 296 0.99×
CSV 300 1.00×
JSON, columnar, minified {"id":[...],"name":[...]} 297 0.99×
JSON, header + row arrays {"columns":[...],"rows":[[...]]} 320 1.07×
Markdown table 373 1.24×
JSON, columnar, pretty 520 1.73×
JSON, row objects, minified 525 1.75×
YAML 649 2.16×
JSON, row objects, pretty (2-space) 884 2.95×
XML 1,088 3.63×

The older cl100k_base tokenizer gave nearly identical numbers, within 2% for every format.

Why the gap is so big

  1. Keys repeat on every row. JSON, YAML and XML write "name":, "price": twenty times. CSV writes them once, in the header, and so does columnar JSON, which is why it lands right next to CSV.
  2. Punctuation is tokens. Quotes, braces and colons all take tokens. XML is worst: every value gets an opening and a closing tag.
  3. Indentation is for humans. The model doesn't need it. Pretty-printing alone added 359 tokens (+68%) over minified JSON here.

YAML is the odd one: fewer characters than minified JSON, more tokens, because every field gets its own line and its own key.

I measured some code too

Coding agents send a lot of code, so I tried a few things:

Test Result
16-line Python file: 4 spaces vs 2 spaces vs tabs 135 / 135 / 133 tokens, basically no difference
Same file without its one-line docstring 135 → 123 (−9%)
10-line JS function, minified 92 → 47 (−49%)
One UUID 18 tokens
  • Indentation is nearly free. Modern tokenizers merge runs of spaces into single tokens. Don't reformat code to save tokens.
  • Minified code is half the tokens, but don't. You lose names like subtotal and taxRate, which is exactly what helps the model understand the code.
  • UUIDs are expensive. If your prompt has a column of long IDs, swap them for row numbers and map back in code.

What I use now

  • Flat tables as input: CSV or TSV. TSV if values contain commas, so no quoting.
  • If it has to be JSON: go columnar ({"id":[...],"name":[...]}) or header + row arrays. Same data, roughly CSV's cost.
  • Data the model returns: minified JSON with a schema. Parse reliability beats a few saved tokens; use structured output if your API supports it.
  • Nested data: minified JSON. Flattening it into CSV by hand often costs more than it saves.
  • Tables a human will also read: Markdown, only 24% over CSV.
  • Avoid pretty JSON and XML in prompts unless a tool requires them.

Dropping unused fields, null fields and extra decimal places helps too.

What it costs

Say you attach a 20-row table to every request, 1,000 requests a day, 30,000 a month, at $2 per million input tokens:

Format Tokens / month Cost / month
CSV 9.0M $18
JSON, columnar 8.9M $18
JSON, row objects, minified 15.8M $32
JSON, row objects, pretty 26.5M $53
XML 32.6M $65

Invisible per request, real for a product, and it scales with bigger tables, RAG results and long API responses.

Caveats

  • One table, one tokenizer. Claude and Gemini tokenize differently, so absolute numbers will differ. The ranking (repeated keys and tags cost the most) comes from how the formats are built, so it should hold.
  • If the format changes answer quality for your task, test both. A cheaper prompt with worse answers isn't cheaper.

Try it on your own data

I built a free token counter for this. It runs the same o200k tokenizer in your browser, so nothing you paste is uploaded. Paste your data in two formats and compare tokens and cost per model.

Full write-up with more detail: JSON vs YAML vs CSV: which format uses the fewest tokens?

Top comments (1)

Collapse
 
argumentmoney8117 profile image
ArgumentMoney8117 •

The YAML finding is the most surprising one — fewer characters than minified JSON but noticeably more tokens. Is that mainly the indentation whitespace eating tokens on each nested field? Nice experiment either way.