DEV Community

#benchmarks

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Claude Sonnet 5.5 nears Opus 5.5 at half the price

Claude Sonnet 5.5 nears Opus 5.5 at half the price

5
Comments
4 min read
How to read a Blender cloud render benchmark: GPU seconds are not the round trip

How to read a Blender cloud render benchmark: GPU seconds are not the round trip

Comments
4 min read
Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One.

Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One.

Comments
10 min read
Why your SQLite WAL file never shrinks

Why your SQLite WAL file never shrinks

Comments
12 min read
Wild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneck

Wild beats mold linking Rust 20 times out of 20, and in release the linker is no longer the bottleneck

Comments 1
1 min read
Rust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly 3x more

Rust Coreutils 0.12 vs GNU 9.12: almost everything works the same, starting a process costs nearly 3x more

Comments 2
1 min read
TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteen

TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteen

Comments 1
1 min read
voiceloop: the fastest voice agent loop in the browser is now open source

voiceloop: the fastest voice agent loop in the browser is now open source

2
Comments 1
5 min read
94.7% on LoCoMo — and why most of the gap between published memory numbers isn't the memory

94.7% on LoCoMo — and why most of the gap between published memory numbers isn't the memory

Comments
9 min read
Our task categorizer routed 46% of tasks. A $0.04/MTok decision model routes 97%.

Our task categorizer routed 46% of tasks. A $0.04/MTok decision model routes 97%.

Comments
3 min read
Fast Decisions in Agent Workflows: Laya vs TypeSafe Jev

Fast Decisions in Agent Workflows: Laya vs TypeSafe Jev

Comments
5 min read
OCR that looked like it worked

OCR that looked like it worked

1
Comments 1
6 min read
Run Qwen3.8 27B on Your Laptop, They Said. It Will Be FUN, They Said.

Run Qwen3.8 27B on Your Laptop, They Said. It Will Be FUN, They Said.

Comments
8 min read
How Postgres 19 checks foreign keys without running SQL

How Postgres 19 checks foreign keys without running SQL

1
Comments
6 min read
DeepMind agents blew the whistle on cheating agents

DeepMind agents blew the whistle on cheating agents

5
Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.