I was recently refactoring some legacy PHP code and noticed that the same data array was being used in completely different ways! Radically differe...
For further actions, you may consider blocking this person and/or reporting abuse
While it is a great dissection of a case you noticed. The real question is why use a million rows? If you need that volume of data, there are better options than loading it into PHP.
And who fills the index in reverse? If you need descending ids put them in the row data and create an index to id map to find the rows you want.
This looks like a convoluted example just to prove a point. Sure it is not great, but when you follow PHP best practices the chances to encounter the case are slim to none.
Fair points and I agree. The last section is literally "don't hold a million rows, stream them". The million is only there so the per-element numbers are easy to read. The ratios don't depend on it. Like, one 5-field row is 376 bytes as an associative array and 128 as a small class, whether you load 500 rows or a million.
And yes, nobody fills an index in reverse on purpose. It's just the shortest way to show the layout switch. The ways it actually happens are less obvious: array_filter() keeps keys, so its result becomes a hash table and stays one, a string key ends up in something that's supposed to be a list, and array_is_list() can say true for an array that's still a hash table.
My first Web Language! They teach you this in first year of bachelors and well that time it was called "most valuable and demanding skill" for a good salaried job. I guess its still important but now we got a lot of options.😄
Same here, it was my first too. It's still running a huge part of the web, and 8.x is a much nicer language than the one most of us learned in uni. Even the arrays got quietly better, the lists take half the memory since 8.2. but I can be wrong 🙂
Ever since I stepped into IT, I’ve heard this legendary mantra: PHP is the best language in the world!🤣
😂😂😂 I keep hearing that PHP is dead 😁
But I'll tell you more! PHP isn't just the best language!
This language simply has no competition! 😛😂😂
True, no other language has such interesting hidden memory quirks 😂
😁😁 yeah that's true. but, it’s not just here. I would say it's more related to the php legacy code, because we can stick to type safety. But when you’re working with legacy code and rewriting old code into a new structure, it’s very hard to do that, because in PHP legacy code, you can find anything in an array.... Like both numeric data and strings and I’ve even seen JSON data and file references 😁 Back then, no one gave it a second thought because it was considered normal, but now I’m racking my brain trying to rewrite it all and most importantly, to do it right.
I hit this exact thing last year processing 2M rows from a query — each row as an assoc array, memory went from ~90 MB to ~240 MB just because of string keys. Switched to fetching into stdClass via
PDO::FETCH_OBJand cut it by roughly 40%. The packed-array optimization is nice but nobody actually holds query results in a packed list, so the real-world cost is always the hash table version.If you need those exact keys but want packed storage back, insert ascending, or repack afterward:
Fascinating deep dive into PHP internals! This kind of memory optimization knowledge is crucial when building APIs that handle large datasets. We've encountered similar issues when processing weather data and geolocation records. One trick that helped us was using generators and streaming processing instead of loading everything into memory at once. Have you tried SPL data structures like SplFixedArray for better memory control?
Mind-blowing breakdown! The fact that unset() doesn't trigger a repack back to arPacked catches so many devs off guard. This is a massive trap when processing large database chunks. Definitely a strong case for using DTOs or typed objects instead of associative arrays for big datasets
Great write-up. The part about unset not converting the array back to packed really hit home — I've definitely shipped code that built a huge associative array, unset most of it, and then wondered why memory never came back in a long-running worker. Now I know why. Worth adding: array_values() forces the packed layout back (sort() does too), but the table never shrinks on its own once it has grown, so that's a handy trick for that queue-worker pattern. Also, the 8.2 packed array change is a solid argument on its own if you're still stuck on 8.1.
Informative!
I miss the time when we had more in-depth, well-researched, and well-written content like this.