DEV Community

William Rodriguez
William Rodriguez

Posted on

Row-by-row is the death of OLAP: High-speed bulk ingestion in ClickHouse

ClickHouse is a columnar OLAP beast, not an OLTP database. Inserting row-by-row will kill it. WClickHouse insert_many() delivers maximum batch speed.

Day 02 of the WClickHouse Open-Source Engineering Series.

The Pain Points We Faced

  • Triggering 'Too many parts' exceptions by inserting individual rows
  • CPU starvation from constant small disk merges on the ClickHouse server
  • Slow analytical ingestion pipelines taking hours instead of seconds

The Implementation

from wclickhouse import WClickHouse

# Prepare 50,000 validated events
events = [
    AnalyticsEvent(event_id=i, event_name="metric_ping", properties=["v1"])
    for i in range(50000)
]

# Bulk insert executed in a single atomic network call
db = WClickHouse(AnalyticsEvent, db_config)
db.insert_many(events)
Enter fullscreen mode Exit fullscreen mode

Why This Architecture Wins

  • 100k+ Rows/Sec: Direct columnar batch insertion with insert_many().
  • No 'Too Many Parts': Large compressed parts written cleanly in single operations.
  • Memory Efficient: Serializes batches directly into ClickHouse native wire format.

Verification & Status

Tested and verified against live ClickHouse server instances with 95%+ test coverage. Built for Python 3.9 through 3.14 with Apache Arrow and Pydantic v2.

Author: William Steve Rodríguez Villamizar (Wisrovi)

Top comments (1)

Collapse
 
william_rodriguez_65a5898 profile image
William Rodriguez •

When designing high-throughput analytical ingestion pipelines for ClickHouse, balancing memory pressure against CPU serialization is critical. Lazy chunking and vectorization prevent memory spikes while keeping client network throughput saturated.

What chunk sizing and compression thresholds have you found most optimal in production OLAP workloads?