Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances
orca forge
orca forge
orca forge
Follow
Oct 2
Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances
#
llm
#
ai
#
evaluation
#
japanese
Comments
Add Comment
12 min read
Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements
Artyom Kornilov
Artyom Kornilov
Artyom Kornilov
Follow
Oct 1
Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements
#
encryption
#
security
#
cryptography
#
evaluation
Comments
Add Comment
12 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark
RESK
RESK
RESK
Follow
Oct 2
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark
#
llm
#
evaluation
#
benchmark
#
reverseengineering
2
reactions
Comments
Add Comment
4 min read
I rejected a model that passed everyone else's benchmark
Priyansh Kansara
Priyansh Kansara
Priyansh Kansara
Follow
Oct 3
I rejected a model that passed everyone else's benchmark
#
machinelearning
#
deepfake
#
evaluation
#
python
Comments
1
comment
5 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs
RESK
RESK
RESK
Follow
Sep 30
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs
#
llm
#
benchmark
#
politics
#
evaluation
1
reaction
Comments
Add Comment
4 min read
Your Agent Has Observability. It Doesn't Have Evals.
Jason Lau
Jason Lau
Jason Lau
Follow
Sep 24
Your Agent Has Observability. It Doesn't Have Evals.
#
agents
#
observability
#
evaluation
#
llm
Comments
Add Comment
10 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges
Modusensus
Modusensus
Modusensus
Follow
Sep 25
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges
#
ai
#
oss
#
evaluation
#
llm
Comments
2
comments
5 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
Amirul Cyber
Amirul Cyber
Amirul Cyber
Follow
Sep 17
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
#
ai
#
security
#
evaluation
#
llm
Comments
Add Comment
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
saaro
saaro
saaro
Follow
Sep 17
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
#
rag
#
ai
#
evaluation
#
metrics
Comments
Add Comment
4 min read
Using Execution Traces to Evaluate AI Agent Behavior
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 17
Using Execution Traces to Evaluate AI Agent Behavior
#
ai
#
opensource
#
evaluation
2
reactions
Comments
Add Comment
5 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)
Debashish Ghosal
Debashish Ghosal
Debashish Ghosal
Follow
Sep 24
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)
#
ai
#
evaluation
#
llm
#
programming
24
reactions
Comments
5
comments
5 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta
Cristian Gormaz
Cristian Gormaz
Cristian Gormaz
Follow
Sep 7
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta
#
ai
#
testing
#
llm
#
evaluation
Comments
Add Comment
4 min read
The benchmark that disproved its own result
the kilted dev
the kilted dev
the kilted dev
Follow
Sep 10
The benchmark that disproved its own result
#
localmodels
#
llm
#
evaluation
#
buildinpublic
Comments
1
comment
6 min read
Green Scores Without a Pinned Harness Are Screenshots
Igor Eduardo
Igor Eduardo
Igor Eduardo
Follow
Sep 28
Green Scores Without a Pinned Harness Are Screenshots
#
evaluation
#
rag
#
ai
#
llm
2
reactions
Comments
1
comment
4 min read
How to Evaluate AI Agents
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 4
How to Evaluate AI Agents
#
ai
#
agentskills
#
agents
#
evaluation
2
reactions
Comments
Add Comment
7 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account