Memory failures at scale rarely begin with an obvious leak; they usually begin with a process that owns enough free memory but cannot use it efficiently.
That distinction matters. A service can report gigabytes of free heap space while its resident set size keeps climbing.
A JVM can pass heap pressure checks while native memory grows outside -Xmx.
A Go service can show stable live objects while containers are killed by the Linux OOM killer. These are not classic leaks.
They are fragmentation, allocator retention, GC pacing failures, and mismatches between runtime assumptions and production reality.
Small systems hide these problems. Large systems expose them.
Fragmentation Is Not Just “Wasted Memory”
Memory fragmentation means available memory is split into pieces that do not match future allocation needs.
There are two major forms:
1. External fragmentation
External fragmentation happens when free memory exists, but it is divided into chunks too small or poorly placed to satisfy larger allocations.
A simplified example:
[ used 64 KB ][ free 8 KB ][ used 128 KB ][ free 12 KB ][ used 32 KB ][ free 16 KB ]
The process has 36 KB free, but a 32 KB contiguous request may fail depending on allocator structure, page layout, and span ownership.
Modern allocators reduce this problem with size classes, arenas, slabs, and per-thread caches, but they do not eliminate it. They often trade external fragmentation for speed.
2. Internal fragmentation
Internal fragmentation happens when the allocator reserves more memory than the application requested.
If an allocator rounds a 33-byte request to a 48-byte size class, 15 bytes are unused. That sounds small until a service creates hundreds of millions of short-lived objects per hour.
Common examples:
Requested size Allocated size class Waste
17 bytes 32 bytes 15 bytes
65 bytes 80 bytes 15 bytes
513 bytes 640 bytes 127 bytes
4097 bytes 8192 bytes 4095 bytes
At scale, these rounding effects are not background noise. They become capacity planning variables.
Why Fragmentation Gets Worse in Large Systems
Fragmentation grows with allocation diversity, concurrency, and uneven object lifetimes.
A service that allocates objects of similar sizes and frees them in roughly the same order is allocator-friendly. A service that mixes request buffers, JSON parse trees, TLS records, compression buffers, cache entries, protobuf messages, and tracing spans creates a much harder pattern.
Three production traits make the problem worse.
A. Mixed object lifetimes
Short-lived request objects often sit beside medium-lived cache entries and long-lived connection structures. The short-lived objects are freed quickly, but the surrounding long-lived allocations keep whole pages, spans, or regions from being returned to the operating system.
A page may be 95% empty and still not releasable because one object remains live.
This is common in services with:
• Large in-memory caches
• Long-lived HTTP/2 or gRPC connections
• Per-tenant metadata
• Streaming workloads
• Background queues
• Metrics label maps with high cardinality
B. Per-thread and per-core allocation caches
Allocators such as jemalloc, tcmalloc, Go’s allocator, and many JVM native paths use thread-local or per-CPU caches to reduce lock contention.
That improves throughput. It can also inflate memory use.
A service with 96 worker threads may hold idle memory in dozens of local caches.
The memory is “free” from the thread’s point of view but not always visible to other threads or returnable to the OS. During traffic spikes, those caches grow. After traffic falls, RSS may remain elevated for minutes or hours.
C. Container limits
Many runtimes were designed before containers became the default deployment unit. A process running in a 4 GB cgroup does not have the same safety margin as a process on a 256 GB host.
A runtime may believe it can grow, compact, or defer reclamation based on host memory signals. The kernel enforces the container limit with less sympathy.
The result is familiar:
Killed process 18472 (service-api) total-vm:8123456kB, anon-rss:3920012kB
The service may have had reclaimable memory inside its runtime. The kernel did not wait.
Garbage Collection Does Not Solve Fragmentation Automatically
Garbage collectors reclaim unreachable objects. That is not the same as returning memory to the operating system, compacting all memory, or making future allocations cheap.
Different runtimes handle this differently.
1. JVM: Heap Looks Fine, RSS Does Not
The Java Virtual Machine has mature garbage collectors, including G1, ZGC, and Shenandoah. They can handle huge heaps with low pause targets.
Still, JVM processes often surprise teams because process memory includes more than Java heap.
A typical JVM process may contain:
• Java heap
• Metaspace
• Code cache
• Thread stacks
• Direct ByteBuffer memory
• JNI allocations
• GC bookkeeping structures
• Allocator arenas from native libraries
• Memory-mapped files
A service configured with:
-Xmx6g
can easily consume 8 GB or more RSS if direct buffers, thread stacks, Netty arenas, compression libraries, and TLS native allocations are active.
2. G1 humongous allocations
G1 divides the heap into regions. Objects larger than 50% of a region are treated as humongous. These allocations can fragment the heap because they require contiguous regions.
If the region size is 4 MB, an object larger than 2 MB becomes humongous. Large byte arrays, serialized payloads, and oversized buffers can trigger this path.
Symptoms often include:
• Increased mixed GC frequency
• Allocation stalls
• Heap occupancy below maximum but allocation failures still occurring
• Logs showing humongous region pressure
GC logs tell the story:
[gc,heap] Humongous regions: 428->391
[gc] Pause Young (Concurrent Start) (G1 Humongous Allocation)
The fix is rarely “increase heap” as a first move.
Better options include capping payload size, chunking large buffers, tuning G1 region size, or removing avoidable large object allocation.
3. Direct buffers and Netty arenas
Netty often uses pooled direct memory for performance. That memory lives outside the Java heap, so heap metrics can look healthy while RSS grows.
Useful checks:
jcmd VM.native_memory summary
jcmd GC.heap_info
cat /proc//smaps_rollup
If Native Memory Tracking is enabled, jcmd can separate Java heap from class metadata, threads, code, GC, compiler, and internal allocations.
For Netty, inspect:
-XX:MaxDirectMemorySize
io.netty.maxDirectMemory
io.netty.allocator.numDirectArenas
Too many arenas can retain memory after burst traffic.
Go: Stable Heap, Rising RSS
Go’s garbage collector is concurrent, non-generational, and tuned around a target heap growth ratio controlled by GOGC. Since Go 1.19, GOMEMLIMIT gives operators a soft memory limit.
Still, Go services can show confusing memory behavior:
heap_alloc = 1.2 GB
heap_idle = 2.8 GB
heap_released = 300 MB
RSS = 4.4 GB
heap_alloc is live heap. heap_idle includes spans not currently holding live objects.
Some idle memory may not be returned to the OS immediately. The scavenger releases memory gradually, and its behaviour depends on allocation pressure, CPU availability, runtime version, and memory limit settings.
The slice retention trap
A small reference can keep a large backing array alive:
func firstKB(buf []byte) []byte {
return buf[:1024]
}
If buf is a 50 MB buffer, the returned 1 KB slice still references the 50 MB backing array. The GC correctly keeps the whole array alive.
The fix is to copy:
func firstKB(buf []byte) []byte {
out := make([]byte, 1024)
copy(out, buf[:1024])
return out
}
This is not fragmentation in the allocator sense, but it creates the same operational symptom: memory remains resident long after the team expects it to fall.
Allocation shape matters
Go’s allocator groups objects by size class. A workload that allocates many objects just over a size boundary can waste significant memory.
For high-volume paths, measure object sizes. A struct that grows from 128 bytes to 136 bytes may move into a larger size class. Multiplied across tens of millions of objects, that small field addition becomes measurable RSS.
Tools that help:
go test -bench=. -benchmem
go tool pprof http://host/debug/pprof/heap
go tool pprof http://host/debug/pprof/allocs
GODEBUG=gctrace=1
Look at allocation rate, not only live heap. A service allocating 8 GB/s will create GC pressure even if live memory stays flat.
Node.js and V8: Old Space Is Only Part of the Story
V8 separates memory into spaces, including new space, old space, large object space, code space, and external memory. Node.js adds buffers, native addons, OpenSSL, libuv, and memory used by C++ bindings.
A Node service may hit container limits while V8 heap statistics look acceptable:
process.memoryUsage()
returns:
rss: 1500 MB
heapTotal: 520 MB
heapUsed: 310 MB
external: 850 MB
arrayBuffers: 780 MB
The important number here is not only heapUsed. external and arrayBuffers often carry the real payload, especially in proxy services, upload handlers, Kafka consumers, and image processing pipelines.
Large object space is also vulnerable to fragmentation because large allocations are handled separately and may not compact like smaller objects.
Mitigations include streaming instead of buffering, bounding concurrent payloads, reusing buffers carefully, and setting memory limits with container headroom:
node --max-old-space-size=1024 server.js
That limit does not cap total RSS. It only constrains V8 old space.
The GC Anomalies That Hurt Production
Garbage collection failures at scale often show up as anomalies rather than steady degradation.
I. Promotion storms
In generational collectors, objects start young. If they survive enough collections, they are promoted to older regions.
A traffic pattern that keeps request objects alive slightly longer than usual can cause mass promotion.
Examples include slow downstream calls, larger queue depth, retry storms, or blocked event loops.
The service may recover functionally after the downstream clears, but the heap is now full of promoted objects. Old-generation collection cost rises. Latency remains poor after the original incident ends.
II. Floating garbage
Concurrent collectors mark live objects while the application continues running. Objects that die after marking starts may not be reclaimed until the next cycle. This is floating garbage.
Under high allocation rates, floating garbage can be large enough to trigger extra cycles or allocation stalls.
Low-pause collectors are not magic; they trade pause time for concurrent work, memory headroom, and CPU.
III. Allocation stalls
A service can pause not because GC wants to run, but because allocation cannot proceed.
This happens when the runtime needs a suitable region, span, or large object area and cannot find one quickly.
Symptoms include:
• Latency spikes with low CPU saturation
• GC logs showing allocation failure
• RSS near container limit
• Heap occupancy below configured maximum
• Increased time in allocator functions in CPU profiles
IV. Finalizer and cleanup delays
Finalizers are unpredictable. So are cleanup hooks tied to GC progress.
If file descriptors, native handles, direct buffers, or GPU memory depend on GC to release them, production will eventually find the worst timing. Explicit close semantics are safer.
In Java, use try-with-resources. In Go, call Close. In Node, release handles deterministically where APIs allow it.
What to Measure
Heap charts alone are not enough.
Track these metrics per process:
• RSS (Really Simple Syndication or Rich Site Summary)
• Virtual memory size
• Runtime heap committed
• Runtime heap used
• Native memory, if available
• Allocation rate
• GC pause time and frequency
• Time spent in GC
• Objects promoted per cycle
• Large object allocation count
• Container memory limit and working set
• Major page faults
• OOM kill count
• Thread count
• Per-endpoint request size distribution
For Linux, these files are useful:
/proc//status
/proc//smaps_rollup
/proc//numa_maps
/sys/fs/cgroup/memory.current
/sys/fs/cgroup/memory.max
For allocator-specific diagnostics:
MALLOC_CONF=stats_print:true # jemalloc
MALLOC_CONF=prof:true # jemalloc profiling
HEAPPROFILE=/tmp/heap # tcmalloc
For JVM:
-Xlog:gc*,safepoint:file=gc.log:time,uptime,level,tags
-XX:NativeMemoryTracking=summary
For Go:
GODEBUG=gctrace=1,madvdontneed=1
For Node:
--trace-gc
--heapsnapshot-near-heap-limit=3
Practical Mitigations
The best fixes change allocation behaviour before adding memory.
1. Bound large allocations
Set hard limits for request bodies, message sizes, batch sizes, and decompressed payloads. A 2 MB Kafka message can become a 40 MB JSON tree after parsing.
2. Prefer streaming
Streaming parsers and chunked processing reduce peak memory. This is especially effective for JSON, CSV, image processing, compression, and upload paths.
3. Separate object lifetimes
Do not store short-lived objects inside long-lived containers unless necessary. Avoid global maps that retain request-derived data. Clear slices and arrays containing pointers if the backing storage will be reused.
In Go:
for i := range items {
items[i] = nil
}
items = items[:0]
4. Tune arenas and pools carefully
Pools can reduce allocation rate, but they can also pin memory permanently.
A pool that grows during a traffic spike may keep peak-sized buffers forever. Prefer capped pools, size buckets, and discard policies for oversized buffers.
5. Leave container headroom
Do not set runtime heap limits equal to container limits.
For a JVM in a 4 GB container, -Xmx4g is unsafe. Native memory, thread stacks, code cache, direct buffers, and GC overhead need room. A safer starting point may be -Xmx2800m, then adjust using real measurements.
6. Test with production allocation patterns
Synthetic benchmarks that use fixed 1 KB objects miss the issue. Replay real payload sizes, Include burst traffic, Include slow downstreams, Include cache warmup, Include high-cardinality labels if production has them.
A good load test tracks RSS after traffic stops. If live heap falls but RSS remains high, investigate allocator retention and fragmentation before shipping.
Treat Memory Shape as a Design Constraint
Memory usage is not a single number. It has shape: object size distribution, lifetime distribution, allocation rate, locality, and release behaviour.
A service with a 500 MB live set can be more dangerous than a service with a 2 GB live set if it allocates at 10 GB/s, retains large backing arrays, or depends on native buffers outside runtime limits.
The fix starts with better questions:
• Which objects survive longer than expected?
• Which allocations cross size-class boundaries?
• Which buffers live outside the managed heap?
• Which pools retain peak traffic memory?
• Which GC cycles reclaim little despite high CPU cost?
• Which metrics ignore RSS?
Memory fragmentation and GC anomalies do not announce themselves as clean defects. They arrive as rising tail latency, expensive nodes, unexplained OOM kills, and services that need twice the memory their live data appears to require.
The next production incident will be easier to diagnose if memory is measured as allocation behaviour, not just heap occupancy.
Top comments (0)