Our air-quality heatmaps took about 5 seconds per image, on an hourly Python cron that used 4 GB of RAM and 2 CPUs and still wasn't enough. The rebuild is a Go service that makes them on request: 20 images in 30 ms, for any past time range.
I lead the software team at Oizom, an air-quality monitoring company (the platform is Envizom). This is how the heatmap backend was rebuilt there, including the bottleneck I didn't expect: PNG encoding.
How the old version worked
We used the IDW (inverse distance weighting) algorithm with wind speed and direction to figure out the value for each spot.
- We set up a heatmap config in our database.
- A cron job ran every hour.
- We grabbed all the configs.
- For each config, we had boundaries, known devices, limits, and color codes.
- We pre-computed a grid of small 1km x 1km boxes and stored it in the CDN.
- We used the center of each box in the pre-computed grid as the unknown value.
- We knew each gas parameter for all devices, plus real-time geolocation, wind speeds, and directions.
- Using IDW with wind effects, we calculated the value for each box in the grid.
- Colors were added to the grid based on user-defined limits.
- A Base64 PNG image was created and stored in a time-series database.
- This image was served to the frontend when requested.
This was written in Python and used up a ton of resources (4GB RAM, 2 CPUs, 1-2 instances depending on the load), and it still wasn't enough. Each image took 10-20 seconds to process. It just wasn't fast. Back then, heatmap was a brand new feature, and we only had 3 configs, so it clearly wasn't scalable. Plus, on closer look, the wind effects weren't even used correctly.
A weekend proof of concept in the browser
When I found out about this, I suggested generating the heatmap on the frontend. My boss thought the idea was nuts. He didn't believe it was even possible. My argument was simple: today, there are tools like Canva and Figma, and many others that let you edit photos and videos right in the browser. There are even proper 2D games and simulations being played in the browser, and some resume websites are like 3D games. Technology has improved, and coding patterns and styles have changed a lot in the past 2 years, so I thought it was doable. He was convinced and decided to hire an intern to do the R&D. Typical corporate.
This really bugged me because I've always wanted to dive into image generation projects, and this seemed like my golden opportunity. So, I spent a weekend putting together a POC for a heatmap in the browser. It worked.
This was just the start. I still needed to add wind effects, map (long, lat) to (x, y) coordinates, include user-defined boundaries (from a GeoJson polygon), and finally, add gradient colors.
Boss saw it, boss liked it, and everyone was happy.
He still had a good point. We don't want to do this on the frontend because our frontend runs on old phones, corporate laptops, and sometimes even TVs—in other words, places without much computing power. So, maybe we can set up a proxy server to handle all the heatmaps and generate them on the fly.
Node.js, then Go, then the real bottleneck
So, I went ahead and set up a Node server and built the whole thing. It worked great. Turns out, when you use gradient colors, there's not much difference between a 150px resolution and a 1000px resolution, except for having sharper edges.
Test case: each request with 100px resolution, 10 known points with 40 images. Each load ran for 5 seconds.
In Node.js the results were already so much better than anything we had before. After this, I decided that to make it even faster, we should use a lower-level language. We also wanted to ensure the language is easy for other developers to understand. So, I chose GoLang. Go was faster and throttled less.
Something still felt off; it shouldn't take this long. After adding time logs for each function, I discovered the bottleneck was Base64 PNG encoding. Each image took 1-3ms, which was 80% of the time. Switching to JPEG boosted performance dramatically.
| Load (5 s) | Node.js | Go, Base64 PNG | Go, JPEG |
|---|---|---|---|
| 5 req/sec | 108ms-874ms (average 419ms) | 80ms-175ms (average 114ms) | 18ms-56ms (average 29ms) |
| 50 req/sec | 97ms-26,588ms (average 14,687ms, with 2 req/sec failing) | 157ms-265ms (average 458ms) | 29ms-132ms (average 88ms) |
| 500 req/sec | 172ms-37048ms (average 22,152ms, with 432 requests failing) | 241ms-10,891ms (average 4,405ms, 296 req/sec failing) | 44ms-929ms (average 496ms, 80 req/sec failing) |
The only downside to JPEG is that it cannot create transparent images, as it lacks an alpha channel. However, we have a straightforward solution: let the backend handle all the computational tasks, while the frontend focuses on one task—masking the image and selecting only the required portions. This can be accomplished in the frontend using JavaScript, within 10-30ms. If the client can render GeoJSON, maps, and images, it can also perform image masking.
This approach also offers an advantage: the difference between a 150px and a 1000px image is that the 1000px image has sharper edges. To achieve sharper edges for a 150px image, we generate an image with an additional 2px around the boundary and crop the necessary section in the frontend. This results in sharper edges.
What the new pipeline does
- The frontend creates a config from backend APIs and stores it in the database.
- The frontend requests images for a config, gas parameters, and given time bounds.
- The backend fetches all the resources and sets up a payload to send to the GoLang heatmap server.
- In the GoLang server: It takes all the boundary points in (long, lat), maps them to (x, y) within 0-1, creates a grid of boxes based on resolution, and finds the ones within the polygon.
- It converts all the known points/devices' locations from (long, lat) to (x, y) based on grid transformation.
- This grid context is used to precompute the weight of each known point on each target pixel.
- Based on the given colors, we create a gradient of values ranging from 0-256 integers.
- We set up the encoder and precompute the (x, y) to (pixel position) in the data image array.
- For each snapshot of values, we quickly compute the grid's value using precomputed weights.
- We magnify each value to bring it between 0-256.
- We fill the RGB values in the data image array.
- We use the encoder for the base64 image.
- These images are sent back to the core server.
- The core server sends them back to the client.
- The frontend masks the required portion using geo boundaries. Done!
Before and after
Images are generated on the fly, so we can now create past images too. We use fewer resources, which makes everything cheaper overall. Plus, it's finally scalable and versatile.
- Heatmaps were only available after they were created → Now we can get heatmaps for any time range.
- For a fixed 1-hour average → You can completely customize it.
- 5 seconds per image → Now it's 30 ms for 20 images.
- Python → GoLang.
- Complex scheduler-based solution → Simple request/response-based solution.
- Fixed options → Custom options for resolution, distance power, wind power, and wind effect.
Finally:
Of course, the data used to create this heatmap is fake and random. But this is how it would look.
The fix I didn't see coming wasn't the language or the algorithm. It was a time log on every function, which showed the image encoder taking 80% of the time. What was the last bottleneck you found that way, somewhere you weren't looking?




Top comments (9)
Thanks for the correction on math.Pow — good point, precomputing the weights per request makes the power cost one-time, and every snapshot after that is multiply-and-add.
On WebP encode time: I haven't measured it in Go, so treat this as a hypothesis rather than a number. If you're on the pure-Go encoder (chai2010/webp — golang.org/x/image only decodes), you're paying several-fold over cgo libwebp, so the codec wrapper may matter more than the format choice. One structural saving: encode RGB-only and let the frontend mask own the transparency — the alpha channel is where lossy WebP spends extra time on smooth gradients, and you keep the polygon crop either way. Might be worth a quick A/B before committing to a format change.
@contentclips_st fair point on the wrapper. A pure-Go encoder against cgo libwebp could swamp the format difference, so an A/B has to name the library on both sides, not just the format.
The PNG cost only showed up because of a time log on every function, so the same harness would give the WebP number: same snapshot, JPEG vs WebP, encode time and bytes. RGB-only makes sense there, since the mask owns transparency anyway.
Would you judge it on encode time alone, or on bytes over the wire too?
Both, but in different columns. Encode time first — if the WebP encoder eats the same share of the pipeline the JPEG one did, that's the fix to make before anything else. Bytes over the wire is what the end user actually pays for on a heatmap endpoint, so once encoding isn't the bottleneck I'd report ms and KB side by side from the same harness (same snapshot, both formats, encoder library named on each side). One caveat: WebP only wins if the smaller file actually ships — with the mask forcing RGB anyway, keep an eye on whether the alpha handling erodes the gain. At 30 ms total you're already fast, so I'd treat wire size as the next metric to optimize and keep encode time as a regression guard.
@contentclips_st makes sense. Encode time answers whether the change is worth making at all, and bytes answer whether the user notices. One table from the same snapshot, ms and KB side by side, encoder named on each side.
Agreed on where the next cost is, too. At 30 ms for 20 images the server isn't the slow part any more. The images still have to reach old phones and corporate laptops, so wire size is the better next target. Thanks for staying with this thread.
Nice rebuild - the PNG encoding being the bottleneck is a classic trap. A few things that could push it further:
Your precomputed weight matrix only holds if wind speed/direction are baked into the weights. IDW with wind anisotropy usually means each known point's weight changes per snapshot - if wind varies per timestamp, keep the base distance weights precomputed and apply the wind term as a cheap per-snapshot correction instead of recomputing the full matrix.
Since you already quantize values to the 0-256 gradient integers, you can drop math.Pow: with p=2, IDW is 1/squared distance and weights normalize by their sum - one reciprocal per known point per pixel, no transcendentals.
For transparency you don't strictly need the frontend mask: WebP supports alpha and encodes these gradients in the same ballpark as JPEG. Frontend masking is still the right call for polygon cropping though.
JPEG at 100-150px will band on smooth gas gradients - encoding at 2x resolution and letting the frontend downsample gives smoother ramps for negligible cost.
@contentclips_st thanks, this is a useful list.
On math.Pow: distance power is a request option, not fixed at 2, so the power can't go away entirely. But the weights are precomputed once per request (step 6), so whatever the power costs is paid once per known point per pixel. Every snapshot after that is multiply-and-add. A p=2 fast path would still be cheap to add.
On wind: the post is thin on how wind enters the weights, that's a fair call. Keeping the distance weights fixed and applying wind as a per-snapshot correction is a clean way to split it.
On WebP: the frontend mask stays either way, because we crop to the GeoJSON polygon, so alpha in the image wouldn't save us a step.
Have you measured WebP encode time against JPEG in Go? The encoder was 80% of our time once, so that's the number I'd check first.
Glad the list was useful, Panth — and thanks for pushing on the details.
On math.Pow: agreed, distance power being a request option means it cannot disappear entirely — but since the weights are precomputed once per request anyway, a p=2 fast path should be nearly free. Worth adding.
On wind: per-snapshot correction on top of fixed distance weights is the clean split — we will model it that way.
On WebP: fair point on the GeoJSON mask making alpha moot. Honest answer on encode time: we have not measured WebP vs JPEG in Go yet, it is on the bench list. Your "encoder was 80% of our time" data point is exactly the number that decides it — we will profile before claiming anything.
Do not follow any external links! DEV.to uses Sloan for automated messages, this is likely phishing.