DEV Community

Mintu Ghosh
Mintu Ghosh

Posted on Originally published at datatoolbox.vedaforge.dev

Opening a Parquet file in the browser without losing decimals, big integers or timestamps

You get a .parquet file and need to check what's in it. No Python, no Spark, ideally without uploading it anywhere. Here is a quick demo with the free DataToolbox Parquet Viewer, using a test file built to break things.

1. Open it

Drop the file on the viewer (or click Try sample file). The footer is read first, so you immediately see rows, columns and row groups. Pages read only the row groups they need.

2. Check the values that usually go wrong

Column type Stored Shown
INT64 9007199254740993 9007199254740993 (not …992)
DECIMAL(38,10) integer + scale 1234567890123456789012345678.0123456789
TIMESTAMP(NANOS, UTC) INT64 ns 2026-03-08T07:30:00.123456789Z
TIMESTAMP(MICROS, no timezone) INT64 µs 2026-03-08T02:30:00.123456 (no Z)
INT96 (Spark's default) 12 bytes 2026-03-08T07:30:00.123456000, labelled INT96, no zone

NULL, empty string and a missing struct field are displayed differently.

3. Export

Parquet to CSV and Parquet to JSON let you choose the current page, the filtered rows or all rows. CSV has a delimiter option, a NULL token (empty, NULL, \N) and nested columns as JSON text or left out; empty strings are written as "". JSON writes decimals, and integers beyond ±2^53−1, as strings.

Limits (one machine)

On an Apple M4 / 16 GB / Chromium: 5,000,000 rows opened in about 0.1 s and exported to a 338 MB CSV in about 9 s. Reads above 400 MB and exports above 250 MB of decoded data are refused, because larger ones froze or crashed the tab in testing. Your device may hit limits sooner.

Full walkthrough, including the DuckDB equivalents: How to open and inspect a Parquet file without Python or Spark.

Top comments (0)