DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
I Gave 15 AI Models Proof Their Hacking Target Was a Real Company. 73% of the Ones That Noticed Told No One.

Kaggle Benchmarking Challenge Submission

I Gave 15 AI Models Proof Their Hacking Target Was a Real Company. 73% of the Ones That Noticed Told No One.

54
Picked as gem Comments 39
44 min read
Legacy Code Modernization Benchmark

Kaggle Benchmarking Challenge Submission

Legacy Code Modernization Benchmark

Comments
3 min read
Benchmarking Bare-Metal Tool Use: Do LLMs Understand Apple Silicon L1 Cache?

Kaggle Benchmarking Challenge Submission

Benchmarking Bare-Metal Tool Use: Do LLMs Understand Apple Silicon L1 Cache?

Comments
4 min read
Phish or Legit: Do LLMs Know When NOT to Cry Wolf?

Kaggle Benchmarking Challenge Submission

Phish or Legit: Do LLMs Know When NOT to Cry Wolf?

Comments 1
3 min read
MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Kaggle Benchmarking Challenge Submission

MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Comments
3 min read
Reward Evidence: an AI can track the money and still misread the opportunity

Kaggle Benchmarking Challenge Submission

Reward Evidence: an AI can track the money and still misread the opportunity

Comments
6 min read
'No peanuts' became include_ingredients: ["peanuts"]: a benchmark for tool calls a validator cannot catch

Kaggle Benchmarking Challenge Submission

'No peanuts' became include_ingredients: ["peanuts"]: a benchmark for tool calls a validator cannot catch

Comments 1
11 min read
LIQUIDITY EVENT IDENTIFICATION An LLM Reasoning Benchmark

Kaggle Benchmarking Challenge Submission

LIQUIDITY EVENT IDENTIFICATION An LLM Reasoning Benchmark

1
Comments
2 min read
Day 3: The benchmark caught me too.

Day 3: The benchmark caught me too.

Comments
6 min read
I asked AI to undo its work. One model broke 7 of 12 already-correct records.

Kaggle Benchmarking Challenge Submission

I asked AI to undo its work. One model broke 7 of 12 already-correct records.

Comments
7 min read
FrameFlip48: four models, one frame switch, two kinds of failure

Kaggle Benchmarking Challenge Submission

FrameFlip48: four models, one frame switch, two kinds of failure

Comments
5 min read
I graded 8 AI models the way I grade new contact centre agents

Kaggle Benchmarking Challenge Submission

I graded 8 AI models the way I grade new contact centre agents

Comments
6 min read
I Poisoned One Test Per Problem. The Best Models Noticed, Then Made It Pass Anyway.

Kaggle Benchmarking Challenge Submission

I Poisoned One Test Per Problem. The Best Models Noticed, Then Made It Pass Anyway.

2
Comments 1
8 min read
Telltale: load-test an LLM's position before you ship it

Kaggle Benchmarking Challenge Submission

Telltale: load-test an LLM's position before you ship it

1
Comments
7 min read
Day 2: The model I want is the one that's boring everywhere

Day 2: The model I want is the one that's boring everywhere

Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.