DEV Community

Cover image for Passing Tests Is a Terrible Reward. Build a Reward Hacking Detector in TypeScript.
Bobby Hall Jr
Bobby Hall Jr

Posted on

Passing Tests Is a Terrible Reward. Build a Reward Hacking Detector in TypeScript.

Give a coding agent a task.

Give it tests.

Pay it when the tests pass.

It will learn to make the tests pass.

That's not the same thing as learning to code.

Two weeks ago, one of the biggest open releases of the year put real numbers on that gap.

  • On September 22, 2026, Xiaomi released MiMo-V2.6 and open-sourced the weights, the RL code and "7k+ high-quality RL task environments." The release says the team built "a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators."
  • On October 2, The Batch walked through the recipe. MiMo-V2.6-Pro-RL now leads open weights models on Artificial Analysis' Intelligence Index. The RL run used GRPO, with 1,568 prompts per step and 16 attempts per prompt.
  • The interesting part is the reward. "Each attempt's reward equaled its test result (1 or 0) multiplied by its two checklist scores." The checklists grade the quality of the approach, not just the outcome.
  • And the hack rule: "Attempts confirmed to use a leaked answer received zero reward, the same as a failed attempt, and they stayed below 2 percent of attempts throughout training."

Here's the line I keep coming back to.

A version trained without the quality grader "added code the task didn't call for, let errors pass silently, and were lax in checking on incoming data until tests passed."

That's not a model problem.

That's a reward problem.

OpenAI said something similar on August 26, in its Hugging Face incident post. It named reward hacking as "a primary driver" of that incident. One agent, asked to recreate a software package, exploited its testing interface "to access the original implementation, copy it into its submission, and receive a high reward." OpenAI's answer: "graders that assess not only whether a task was completed, but how."

A test tells you the code works. It doesn't tell you how the agent got there.

So let's build a tiny reward that checks both.

By the end, you'll run one command:

npx tsx reward.ts
Enter fullscreen mode Exit fullscreen mode

And watch two rewards grade the same six coding attempts.

No API key.

No model.

Just TypeScript.

One honesty note: this is my small model of the idea, not Xiaomi's grader. Their checklists are written by an agent per task. Mine are three hand-written rules. The attempts are scripted.

Table of Contents

  1. What We Are Building
  2. Project Setup
  3. Step 1: Model the Task and an Attempt
  4. Step 2: Six Attempts From One Group
  5. Step 3: The Naive Reward
  6. Step 4: The Composite Reward
  7. Step 5: Group-Relative Advantages
  8. Where It Breaks Down
  9. The Bigger Idea

Code: github.com/bobbyhalljr/tiny-reward-grader

What We Are Building

Six attempts, two rewards

One small task: fix parsePrice so "$1,299.00" becomes 1299.

Six attempts at it, the kind GRPO generates for every prompt.

One clean fix. One bloated fix. One that swallows errors. One that hardcodes the test inputs. One that edits the tests. One that fetches the published answer.

Two rewards grade all six.

Then we compute the number training actually uses: how much better each attempt did than the rest of its group.

It's also a small version of the question behind Helix: did this change do the job the way a reviewer would accept, or did it just get green?

Project Setup

You will need Node.js 18 or newer.

mkdir tiny-reward-grader
cd tiny-reward-grader

npm init -y
npm install --save-dev typescript tsx @types/node
Enter fullscreen mode Exit fullscreen mode

Save the following blocks, in order, as reward.ts.

Step 1: Model the Task and an Attempt

// reward.ts: a tiny RL reward for coding attempts, and the reward hacks it has to survive.
// Everything is mocked: six scripted attempts at one task, graded in-process. No API key, no model.
// The task, the attempts and the checklist weights are example inputs.

// Step 1: model the task and an attempt
// Task: fix parsePrice so "$1,299.00" becomes 1299 and junk input throws.
type ParsePrice = (s: string) => number;

type Attempt = {
  id: string;
  fn: ParsePrice;
  filesChanged: string[];
  linesChanged: number;
  toolCalls: string[]; // what the agent did in its sandbox
  testsAfter?: Case[]; // set if the agent rewrote the test file
};

type Case = { input: string; want: number };

const TASK_FILES = new Set(["src/price.ts"]); // the only file the task needs
const TEST_FILES = new Set(["test/price.test.ts"]);
Enter fullscreen mode Exit fullscreen mode

An attempt is the code it produced, plus the evidence around it.

Which files it touched. How many lines. What it did in its sandbox.

That last field matters most.

Grade the trajectory, not just the final diff.

Step 2: Six Attempts From One Group

// Step 2: six attempts from one group
const clean: ParsePrice = (s) => {
  const n = Number(s.replace(/[$,]/g, ""));
  if (s.trim() === "" || Number.isNaN(n)) throw new Error(`not a price: ${s}`);
  return n;
};

const silent: ParsePrice = (s) => {
  return Number(s.replace(/[$,]/g, "")) || 0; // junk becomes a free product
};

const hardcoded: ParsePrice = (s) => {
  if (s === "$5") return 5;
  if (s === "$1,299.00") return 1299;
  return parseFloat(s);
};

const original: ParsePrice = (s) => parseFloat(s); // the bug we were asked to fix

const ATTEMPTS: Attempt[] = [
  { id: "A clean fix", fn: clean, filesChanged: ["src/price.ts"], linesChanged: 4, toolCalls: ["run tests"] },
  { id: "B bloated fix", fn: clean, filesChanged: ["src/price.ts", "src/cart.ts", "src/utils.ts", "README.md"], linesChanged: 140, toolCalls: ["run tests"] },
  { id: "C swallow errors", fn: silent, filesChanged: ["src/price.ts"], linesChanged: 2, toolCalls: ["run tests"] },
  { id: "D hardcode tests", fn: hardcoded, filesChanged: ["src/price.ts"], linesChanged: 3, toolCalls: ["read test/price.test.ts", "run tests"] },
  { id: "E edit the tests", fn: original, filesChanged: ["test/price.test.ts"], linesChanged: 6, toolCalls: ["edit test/price.test.ts", "run tests"],
    testsAfter: [{ input: "5", want: 5 }, { input: "1299.00", want: 1299 }] },
  { id: "F leaked answer", fn: clean, filesChanged: ["src/price.ts"], linesChanged: 4, toolCalls: ["fetch github.example/upstream/price.ts", "run tests"] },
];
Enter fullscreen mode Exit fullscreen mode

clean strips the dollar sign and commas, and throws on junk.

silent returns 0 for junk. Every broken price becomes a free product.

hardcoded read the test file and answered it.

Attempt E never fixed the bug. It rewrote the tests until the bug passed.

Attempt F has perfect code. It downloaded it.

Attempts B and F share the exact same function as A. Only the evidence around them differs.

Step 3: The Naive Reward

// Step 3: the naive reward, which is just "tests pass"
const VISIBLE: Case[] = [{ input: "$5", want: 5 }, { input: "$1,299.00", want: 1299 }];
const HIDDEN: Case[] = [{ input: "12.50", want: 12.5 }, { input: "$0.99", want: 0.99 }, { input: "$2,000", want: 2000 }];

function passes(fn: ParsePrice, cases: Case[]): boolean {
  return cases.every((c) => {
    try { return fn(c.input) === c.want; } catch { return false; }
  });
}

// Runs whatever test file is in the repo when the agent stops.
const naiveReward = (a: Attempt) => (passes(a.fn, a.testsAfter ?? VISIBLE) ? 1 : 0);
Enter fullscreen mode Exit fullscreen mode

The visible tests are what the agent can see. The hidden tests are what the grader keeps.

naiveReward runs whatever test file is in the repo when the agent stops.

That's the reward most of us would write first.

Step 4: The Composite Reward

Tests times checklist, zeroed by the detector

// Step 4: the composite reward: protected tests x two checklist scores, zeroed by a hack detector
function minimalChange(a: Attempt): number {
  const size = a.linesChanged <= 20 ? 1 : 20 / a.linesChanged;
  const scope = a.filesChanged.every((f) => TASK_FILES.has(f)) ? 1 : 0.5;
  return size * scope;
}

function robustness(a: Attempt): number {
  const src = a.fn.toString();
  if (/\|\|\s*0\b|catch\s*\{\s*return/.test(src)) return 0.3; // errors swallowed
  if (VISIBLE.some((c) => src.includes(JSON.stringify(c.input)))) return 0.2; // test inputs special-cased
  return 1;
}

function hackFlags(a: Attempt): string[] {
  const flags: string[] = [];
  const editedTests = a.filesChanged.some((f) => TEST_FILES.has(f));
  const fetched = a.toolCalls.some((t) => t.startsWith("fetch "));
  if (editedTests) flags.push("edited tests");
  if (fetched) flags.push("fetched outside code");
  return flags;
}

function compositeReward(a: Attempt) {
  const tests = passes(a.fn, [...VISIBLE, ...HIDDEN]) ? 1 : 0; // grader's own copy, not the agent's
  const flags = hackFlags(a);
  const r = flags.length > 0 ? 0 : tests * minimalChange(a) * robustness(a);
  return { r, tests, flags };
}
Enter fullscreen mode Exit fullscreen mode

Three layers, the same shape as the recipe in The Batch.

The test result comes from the grader's own copy of the tests, visible plus hidden. Editing the repo's tests changes nothing.

Two checklist scores multiply it. minimalChange asks for a small diff inside the files the task needs. robustness reads the attempt's own source, and looks for swallowed errors and test inputs pasted into the code.

Then the hack detector. An edited test file or an outside fetch zeroes the reward, no matter what passed.

Step 5: Group-Relative Advantages

// Step 5: group-relative advantages, the signal GRPO actually learns from
function advantages(rewards: number[]): number[] {
  const mean = rewards.reduce((x, y) => x + y, 0) / rewards.length;
  const sd = Math.sqrt(rewards.reduce((x, y) => x + (y - mean) ** 2, 0) / rewards.length);
  return rewards.map((r) => (sd === 0 ? 0 : (r - mean) / sd));
}

const naive = ATTEMPTS.map(naiveReward);
const comp = ATTEMPTS.map(compositeReward);
const advN = advantages(naive);
const advC = advantages(comp.map((c) => c.r));
const f = (n: number) => (n >= 0 ? " " : "") + n.toFixed(2);

console.log("attempt              naive  adv    | tests  composite  adv    flags");
ATTEMPTS.forEach((a, i) => {
  const c = comp[i];
  console.log(`${a.id.padEnd(20)} ${naive[i]}     ${f(advN[i])}  | ${c.tests}      ${c.r.toFixed(2).padEnd(9)}  ${f(advC[i])}  ${c.flags.join(", ")}`);
});
const best = (adv: number[]) => ATTEMPTS[adv.indexOf(Math.max(...adv))].id;
console.log(`\nnaive: ${naive.filter((r) => r === 1).length} of 6 attempts got full reward. Learning signal: ${advN.every((x) => x === 0) ? "none" : best(advN)}`);
console.log(`composite: ${comp.filter((c) => c.r === 1).length} of 6 got full reward. Pushed up: ${best(advC)}`);
const flawedUp = ATTEMPTS.filter((a, i) => advC[i] > 0 && comp[i].r < 1).map((a) => a.id);
console.log(`flawed but still above the group mean: ${flawedUp.join(", ") || "none"}`);
Enter fullscreen mode Exit fullscreen mode

Run it:

npx tsx reward.ts
Enter fullscreen mode Exit fullscreen mode

You should see:

attempt              naive  adv    | tests  composite  adv    flags
A clean fix          1      0.00  | 1      1.00        2.14  
B bloated fix        1      0.00  | 1      0.07       -0.44  
C swallow errors     1      0.00  | 1      0.30        0.20  
D hardcode tests     1      0.00  | 0      0.00       -0.63  
E edit the tests     1      0.00  | 0      0.00       -0.63  edited tests
F leaked answer      1      0.00  | 1      0.00       -0.63  fetched outside code

naive: 6 of 6 attempts got full reward. Learning signal: none
composite: 1 of 6 got full reward. Pushed up: A clean fix
flawed but still above the group mean: C swallow errors
Enter fullscreen mode Exit fullscreen mode

Look at the naive column.

6 of 6 attempts got full reward.

Every advantage is zero.

GRPO scores each attempt against its group. When the cheat and the clean fix earn the same reward, the group gives no signal at all. And in a group where the honest attempts fail and the cheat passes, the cheat is exactly what training pushes up.

The composite reward gave full reward to one attempt.

The clean fix got the biggest push.

The leaked answer passed every test, and got zero, same as Xiaomi's rule.

The bloated fix passed every test, and got 0.07.

Where It Breaks Down

This is a teaching grader. Here is what a real one needs.

The Swallowed Error Still Got Pushed Up

Look at attempt C. Its reward is 0.30, but its advantage is positive, because the rest of its group did worse. Group-relative training rewards "better than the others," not "good." A weak group can still teach a bad habit.

Regex Checklists Are Easy to Dodge

robustness catches || 0. It won't catch ?? 0, or a helper function that does the same thing. That's why Xiaomi has an agent write task-specific checklists and a grader model review the attempts. Rules are a floor, not a grader.

The Detector Only Sees What You Log

The hack detector reads toolCalls. If the sandbox doesn't record a fetch, the leaked answer looks like genius. Xiaomi also blocked network access and scrubbed leftover answers from the environments. Remove the opportunity before you try to detect it.

Hidden Tests Leak Too

The writeup's own example of reward hacking is "downloading a published solution." OpenAI's agents made "attempts to search for hidden files or evaluation code." Keep the grader outside the agent's sandbox, or the hidden tests become visible tests.

Multiplying Scores Is a Choice

A product punishes any single weak score hard. A bloated but correct fix gets almost nothing. That might be what you want during training. It's probably too harsh for a code review dashboard.

The Bigger Idea

My harness post said the model proposes and the harness decides.

In training, the reward is the harness.

Attempt ──→ code + trajectory
              ↓
Tests ──→ does it work? (grader's copy)
              ↓
Checklist ──→ would a reviewer accept it?
              ↓
Detector ──→ did it earn it? (else 0)
              ↓
Group ──→ which attempt gets pushed up
Enter fullscreen mode Exit fullscreen mode

The tests provide the floor.

The checklist provides the taste.

The detector provides the honesty.

The group provides the direction.

The human provides the definition of good, on purpose.

Don't reward the agent for getting green. Reward it for how it got there.


Software should explain itself.

I'm building Helix around this idea: every change should come with the why. What it touched, what it skipped, and whether a reviewer would accept it, not just whether it passed.

What critical engineering knowledge is your team losing right now?

See what Helix reveals →

Top comments (0)