ClickHouse is a columnar OLAP beast, not an OLTP database. Inserting row-by-row will kill it. WClickHouse insert_many() delivers maximum batch speed.
Day 02 of the WClickHouse Open-Source Engineering Series.
The Pain Points We Faced
- Triggering 'Too many parts' exceptions by inserting individual rows
- CPU starvation from constant small disk merges on the ClickHouse server
- Slow analytical ingestion pipelines taking hours instead of seconds
The Implementation
from wclickhouse import WClickHouse
# Prepare 50,000 validated events
events = [
AnalyticsEvent(event_id=i, event_name="metric_ping", properties=["v1"])
for i in range(50000)
]
# Bulk insert executed in a single atomic network call
db = WClickHouse(AnalyticsEvent, db_config)
db.insert_many(events)
Why This Architecture Wins
- 100k+ Rows/Sec: Direct columnar batch insertion with insert_many().
- No 'Too Many Parts': Large compressed parts written cleanly in single operations.
- Memory Efficient: Serializes batches directly into ClickHouse native wire format.
Verification & Status
Tested and verified against live ClickHouse server instances with 95%+ test coverage. Built for Python 3.9 through 3.14 with Apache Arrow and Pydantic v2.
Author: William Steve Rodríguez Villamizar (Wisrovi)
Top comments (1)
When designing high-throughput analytical ingestion pipelines for ClickHouse, balancing memory pressure against CPU serialization is critical. Lazy chunking and vectorization prevent memory spikes while keeping client network throughput saturated.
What chunk sizing and compression thresholds have you found most optimal in production OLAP workloads?