DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Hugging Face now has 48 official benchmarks. Here is what the map looks like

Hugging Face now has 48 official benchmarks. Here is what the map looks like

Comments
3 min read
Hugging Face official benchmarks: the complete list (48) and how their leaderboards work

Hugging Face official benchmarks: the complete list (48) and how their leaderboards work

Comments
4 min read
Gemini 4 Argon ขึ้นที่ 1 Arena AI แล้ว แต่มีตัวเลขสามตัวที่บทความต้นทางไม่ได้บอก

Gemini 4 Argon ขึ้นที่ 1 Arena AI แล้ว แต่มีตัวเลขสามตัวที่บทความต้นทางไม่ได้บอก

Comments
3 min read
โมเดล 27B ย่อ 4 บิต ชนะโมเดลคลาวด์ในโจทย์เดียว แต่ทำไมยังไม่ควรเชื่อ

โมเดล 27B ย่อ 4 บิต ชนะโมเดลคลาวด์ในโจทย์เดียว แต่ทำไมยังไม่ควรเชื่อ

Comments
2 min read
Understanding Tokens per Second: A Practical Benchmark Guide

Understanding Tokens per Second: A Practical Benchmark Guide

Comments
5 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

2
Comments
4 min read
We recompute TypeSafe's 444x claim — here's what we found

We recompute TypeSafe's 444x claim — here's what we found

Comments
1 min read
Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First

Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First

Comments
1 min read
One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.

One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.

7
Comments
3 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

1
Comments
4 min read
In Search of the Best A.I. Literary Translator – Detailed

In Search of the Best A.I. Literary Translator – Detailed

Comments
15 min read
Claude Opus 5.5: 40% cheaper, frontier-grade performance

Claude Opus 5.5: 40% cheaper, frontier-grade performance

Comments
5 min read
Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

7
Comments 1
2 min read
Enterprise Vector Database 2026: Qdrant vs Milvus vs pgvector vs Pinecone

Enterprise Vector Database 2026: Qdrant vs Milvus vs pgvector vs Pinecone

Comments
23 min read
TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper?

TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper?

2
Comments 1
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.