Give a coding agent a task.
Give it tests.
Pay it when the tests pass.
It will learn to make the tests pass.
That's not the same thing as learning to code.
Two weeks ago, one of the biggest open releases of the year put real numbers on that gap.
- On September 22, 2026, Xiaomi released MiMo-V2.6 and open-sourced the weights, the RL code and "7k+ high-quality RL task environments." The release says the team built "a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators."
- On October 2, The Batch walked through the recipe. MiMo-V2.6-Pro-RL now leads open weights models on Artificial Analysis' Intelligence Index. The RL run used GRPO, with 1,568 prompts per step and 16 attempts per prompt.
- The interesting part is the reward. "Each attempt's reward equaled its test result (1 or 0) multiplied by its two checklist scores." The checklists grade the quality of the approach, not just the outcome.
- And the hack rule: "Attempts confirmed to use a leaked answer received zero reward, the same as a failed attempt, and they stayed below 2 percent of attempts throughout training."
Here's the line I keep coming back to.
A version trained without the quality grader "added code the task didn't call for, let errors pass silently, and were lax in checking on incoming data until tests passed."
That's not a model problem.
That's a reward problem.
OpenAI said something similar on August 26, in its Hugging Face incident post. It named reward hacking as "a primary driver" of that incident. One agent, asked to recreate a software package, exploited its testing interface "to access the original implementation, copy it into its submission, and receive a high reward." OpenAI's answer: "graders that assess not only whether a task was completed, but how."
A test tells you the code works. It doesn't tell you how the agent got there.
So let's build a tiny reward that checks both.
By the end, you'll run one command:
npx tsx reward.ts
And watch two rewards grade the same six coding attempts.
No API key.
No model.
Just TypeScript.
One honesty note: this is my small model of the idea, not Xiaomi's grader. Their checklists are written by an agent per task. Mine are three hand-written rules. The attempts are scripted.
Table of Contents
- What We Are Building
- Project Setup
- Step 1: Model the Task and an Attempt
- Step 2: Six Attempts From One Group
- Step 3: The Naive Reward
- Step 4: The Composite Reward
- Step 5: Group-Relative Advantages
- Where It Breaks Down
- The Bigger Idea
Code: github.com/bobbyhalljr/tiny-reward-grader
What We Are Building
One small task: fix parsePrice so "$1,299.00" becomes 1299.
Six attempts at it, the kind GRPO generates for every prompt.
One clean fix. One bloated fix. One that swallows errors. One that hardcodes the test inputs. One that edits the tests. One that fetches the published answer.
Two rewards grade all six.
Then we compute the number training actually uses: how much better each attempt did than the rest of its group.
It's also a small version of the question behind Helix: did this change do the job the way a reviewer would accept, or did it just get green?
Project Setup
You will need Node.js 18 or newer.
mkdir tiny-reward-grader
cd tiny-reward-grader
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as reward.ts.
Step 1: Model the Task and an Attempt
// reward.ts: a tiny RL reward for coding attempts, and the reward hacks it has to survive.
// Everything is mocked: six scripted attempts at one task, graded in-process. No API key, no model.
// The task, the attempts and the checklist weights are example inputs.
// Step 1: model the task and an attempt
// Task: fix parsePrice so "$1,299.00" becomes 1299 and junk input throws.
type ParsePrice = (s: string) => number;
type Attempt = {
id: string;
fn: ParsePrice;
filesChanged: string[];
linesChanged: number;
toolCalls: string[]; // what the agent did in its sandbox
testsAfter?: Case[]; // set if the agent rewrote the test file
};
type Case = { input: string; want: number };
const TASK_FILES = new Set(["src/price.ts"]); // the only file the task needs
const TEST_FILES = new Set(["test/price.test.ts"]);
An attempt is the code it produced, plus the evidence around it.
Which files it touched. How many lines. What it did in its sandbox.
That last field matters most.
Grade the trajectory, not just the final diff.
Step 2: Six Attempts From One Group
// Step 2: six attempts from one group
const clean: ParsePrice = (s) => {
const n = Number(s.replace(/[$,]/g, ""));
if (s.trim() === "" || Number.isNaN(n)) throw new Error(`not a price: ${s}`);
return n;
};
const silent: ParsePrice = (s) => {
return Number(s.replace(/[$,]/g, "")) || 0; // junk becomes a free product
};
const hardcoded: ParsePrice = (s) => {
if (s === "$5") return 5;
if (s === "$1,299.00") return 1299;
return parseFloat(s);
};
const original: ParsePrice = (s) => parseFloat(s); // the bug we were asked to fix
const ATTEMPTS: Attempt[] = [
{ id: "A clean fix", fn: clean, filesChanged: ["src/price.ts"], linesChanged: 4, toolCalls: ["run tests"] },
{ id: "B bloated fix", fn: clean, filesChanged: ["src/price.ts", "src/cart.ts", "src/utils.ts", "README.md"], linesChanged: 140, toolCalls: ["run tests"] },
{ id: "C swallow errors", fn: silent, filesChanged: ["src/price.ts"], linesChanged: 2, toolCalls: ["run tests"] },
{ id: "D hardcode tests", fn: hardcoded, filesChanged: ["src/price.ts"], linesChanged: 3, toolCalls: ["read test/price.test.ts", "run tests"] },
{ id: "E edit the tests", fn: original, filesChanged: ["test/price.test.ts"], linesChanged: 6, toolCalls: ["edit test/price.test.ts", "run tests"],
testsAfter: [{ input: "5", want: 5 }, { input: "1299.00", want: 1299 }] },
{ id: "F leaked answer", fn: clean, filesChanged: ["src/price.ts"], linesChanged: 4, toolCalls: ["fetch github.example/upstream/price.ts", "run tests"] },
];
clean strips the dollar sign and commas, and throws on junk.
silent returns 0 for junk. Every broken price becomes a free product.
hardcoded read the test file and answered it.
Attempt E never fixed the bug. It rewrote the tests until the bug passed.
Attempt F has perfect code. It downloaded it.
Attempts B and F share the exact same function as A. Only the evidence around them differs.
Step 3: The Naive Reward
// Step 3: the naive reward, which is just "tests pass"
const VISIBLE: Case[] = [{ input: "$5", want: 5 }, { input: "$1,299.00", want: 1299 }];
const HIDDEN: Case[] = [{ input: "12.50", want: 12.5 }, { input: "$0.99", want: 0.99 }, { input: "$2,000", want: 2000 }];
function passes(fn: ParsePrice, cases: Case[]): boolean {
return cases.every((c) => {
try { return fn(c.input) === c.want; } catch { return false; }
});
}
// Runs whatever test file is in the repo when the agent stops.
const naiveReward = (a: Attempt) => (passes(a.fn, a.testsAfter ?? VISIBLE) ? 1 : 0);
The visible tests are what the agent can see. The hidden tests are what the grader keeps.
naiveReward runs whatever test file is in the repo when the agent stops.
That's the reward most of us would write first.
Step 4: The Composite Reward
// Step 4: the composite reward: protected tests x two checklist scores, zeroed by a hack detector
function minimalChange(a: Attempt): number {
const size = a.linesChanged <= 20 ? 1 : 20 / a.linesChanged;
const scope = a.filesChanged.every((f) => TASK_FILES.has(f)) ? 1 : 0.5;
return size * scope;
}
function robustness(a: Attempt): number {
const src = a.fn.toString();
if (/\|\|\s*0\b|catch\s*\{\s*return/.test(src)) return 0.3; // errors swallowed
if (VISIBLE.some((c) => src.includes(JSON.stringify(c.input)))) return 0.2; // test inputs special-cased
return 1;
}
function hackFlags(a: Attempt): string[] {
const flags: string[] = [];
const editedTests = a.filesChanged.some((f) => TEST_FILES.has(f));
const fetched = a.toolCalls.some((t) => t.startsWith("fetch "));
if (editedTests) flags.push("edited tests");
if (fetched) flags.push("fetched outside code");
return flags;
}
function compositeReward(a: Attempt) {
const tests = passes(a.fn, [...VISIBLE, ...HIDDEN]) ? 1 : 0; // grader's own copy, not the agent's
const flags = hackFlags(a);
const r = flags.length > 0 ? 0 : tests * minimalChange(a) * robustness(a);
return { r, tests, flags };
}
Three layers, the same shape as the recipe in The Batch.
The test result comes from the grader's own copy of the tests, visible plus hidden. Editing the repo's tests changes nothing.
Two checklist scores multiply it. minimalChange asks for a small diff inside the files the task needs. robustness reads the attempt's own source, and looks for swallowed errors and test inputs pasted into the code.
Then the hack detector. An edited test file or an outside fetch zeroes the reward, no matter what passed.
Step 5: Group-Relative Advantages
// Step 5: group-relative advantages, the signal GRPO actually learns from
function advantages(rewards: number[]): number[] {
const mean = rewards.reduce((x, y) => x + y, 0) / rewards.length;
const sd = Math.sqrt(rewards.reduce((x, y) => x + (y - mean) ** 2, 0) / rewards.length);
return rewards.map((r) => (sd === 0 ? 0 : (r - mean) / sd));
}
const naive = ATTEMPTS.map(naiveReward);
const comp = ATTEMPTS.map(compositeReward);
const advN = advantages(naive);
const advC = advantages(comp.map((c) => c.r));
const f = (n: number) => (n >= 0 ? " " : "") + n.toFixed(2);
console.log("attempt naive adv | tests composite adv flags");
ATTEMPTS.forEach((a, i) => {
const c = comp[i];
console.log(`${a.id.padEnd(20)} ${naive[i]} ${f(advN[i])} | ${c.tests} ${c.r.toFixed(2).padEnd(9)} ${f(advC[i])} ${c.flags.join(", ")}`);
});
const best = (adv: number[]) => ATTEMPTS[adv.indexOf(Math.max(...adv))].id;
console.log(`\nnaive: ${naive.filter((r) => r === 1).length} of 6 attempts got full reward. Learning signal: ${advN.every((x) => x === 0) ? "none" : best(advN)}`);
console.log(`composite: ${comp.filter((c) => c.r === 1).length} of 6 got full reward. Pushed up: ${best(advC)}`);
const flawedUp = ATTEMPTS.filter((a, i) => advC[i] > 0 && comp[i].r < 1).map((a) => a.id);
console.log(`flawed but still above the group mean: ${flawedUp.join(", ") || "none"}`);
Run it:
npx tsx reward.ts
You should see:
attempt naive adv | tests composite adv flags
A clean fix 1 0.00 | 1 1.00 2.14
B bloated fix 1 0.00 | 1 0.07 -0.44
C swallow errors 1 0.00 | 1 0.30 0.20
D hardcode tests 1 0.00 | 0 0.00 -0.63
E edit the tests 1 0.00 | 0 0.00 -0.63 edited tests
F leaked answer 1 0.00 | 1 0.00 -0.63 fetched outside code
naive: 6 of 6 attempts got full reward. Learning signal: none
composite: 1 of 6 got full reward. Pushed up: A clean fix
flawed but still above the group mean: C swallow errors
Look at the naive column.
6 of 6 attempts got full reward.
Every advantage is zero.
GRPO scores each attempt against its group. When the cheat and the clean fix earn the same reward, the group gives no signal at all. And in a group where the honest attempts fail and the cheat passes, the cheat is exactly what training pushes up.
The composite reward gave full reward to one attempt.
The clean fix got the biggest push.
The leaked answer passed every test, and got zero, same as Xiaomi's rule.
The bloated fix passed every test, and got 0.07.
Where It Breaks Down
This is a teaching grader. Here is what a real one needs.
The Swallowed Error Still Got Pushed Up
Look at attempt C. Its reward is 0.30, but its advantage is positive, because the rest of its group did worse. Group-relative training rewards "better than the others," not "good." A weak group can still teach a bad habit.
Regex Checklists Are Easy to Dodge
robustness catches || 0. It won't catch ?? 0, or a helper function that does the same thing. That's why Xiaomi has an agent write task-specific checklists and a grader model review the attempts. Rules are a floor, not a grader.
The Detector Only Sees What You Log
The hack detector reads toolCalls. If the sandbox doesn't record a fetch, the leaked answer looks like genius. Xiaomi also blocked network access and scrubbed leftover answers from the environments. Remove the opportunity before you try to detect it.
Hidden Tests Leak Too
The writeup's own example of reward hacking is "downloading a published solution." OpenAI's agents made "attempts to search for hidden files or evaluation code." Keep the grader outside the agent's sandbox, or the hidden tests become visible tests.
Multiplying Scores Is a Choice
A product punishes any single weak score hard. A bloated but correct fix gets almost nothing. That might be what you want during training. It's probably too harsh for a code review dashboard.
The Bigger Idea
My harness post said the model proposes and the harness decides.
In training, the reward is the harness.
Attempt ──→ code + trajectory
↓
Tests ──→ does it work? (grader's copy)
↓
Checklist ──→ would a reviewer accept it?
↓
Detector ──→ did it earn it? (else 0)
↓
Group ──→ which attempt gets pushed up
The tests provide the floor.
The checklist provides the taste.
The detector provides the honesty.
The group provides the direction.
The human provides the definition of good, on purpose.
Don't reward the agent for getting green. Reward it for how it got there.
Software should explain itself.
I'm building Helix around this idea: every change should come with the why. What it touched, what it skipped, and whether a reviewer would accept it, not just whether it passed.
What critical engineering knowledge is your team losing right now?


Top comments (0)