DEV Community

#evaluation

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances

Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances

Comments
12 min read
Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements

Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements

Comments
12 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

2
Comments
4 min read
I rejected a model that passed everyone else's benchmark

I rejected a model that passed everyone else's benchmark

Comments 1
5 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

1
Comments
4 min read
Your Agent Has Observability. It Doesn't Have Evals.

Your Agent Has Observability. It Doesn't Have Evals.

Comments
10 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Comments 2
5 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

Comments
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

Comments
4 min read
Using Execution Traces to Evaluate AI Agent Behavior

Using Execution Traces to Evaluate AI Agent Behavior

2
Comments
5 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

24
Comments 5
5 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

Comments
4 min read
The benchmark that disproved its own result

The benchmark that disproved its own result

Comments 1
6 min read
Green Scores Without a Pinned Harness Are Screenshots

Green Scores Without a Pinned Harness Are Screenshots

2
Comments 1
4 min read
How to Evaluate AI Agents

How to Evaluate AI Agents

2
Comments
7 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.