Our judging software ranked a project second because a judge who scored only that project gave it a flat 3.
Two honest judges had it first. The third judge scored one project, once, and the maths treated "I have no opinion about relative quality" as "I think this one is exactly average." That quietly pulled the best project toward the middle and handed first place to the runner-up.
DOGFOOD's brief was one line: build the platform that will judge you. So we built HackPulse in 72 hours, a self-hosted portal that takes a hackathon from registration to certificates, and the hardest part was the one that has to be fair: the maths that turns a pile of judge scores into a ranking you can defend. If that maths is wrong, the platform judges its own builders wrong, and nobody can tell by looking.
This post is about three bugs in that maths. They look unrelated. They're the same mistake.
A missing signal is not a neutral signal. If you fill the hole with a default, the default becomes data.
Why raw averages were never an option
Judge A gives everything 4 or 5. Judge B gives everything 2 or 3. They aren't necessarily disagreeing about quality; they anchor to different points on the scale. If a project happens to draw the generous judge, it wins.
So we normalize per judge. Take each judge's own average (μ) and how spread out their scores are (σ, the standard deviation). A raw score r becomes z = (r − μ) / σ (a "z-score"): how many of this judge's own "spreads" above or below their own average the score sits. A 4 from a judge who averages 4 becomes 0, an ordinary score for them. A 4 from a judge who usually gives 2 becomes a large positive number, a rave. A project's final score is the mean of its z-scores across judges.
We tested it against the shared reference dataset (40 projects, 30 judges, 122 scores). Normalization changed 37 of 40 ranks, by up to 24 places, and only 3 of the top 5 survived. That's what normalization is for, and it's why getting the edge cases wrong matters: the output is no longer a number anyone can eyeball.
Bug 1: the judge with no spread
σ is zero when a judge gave every score the same value, or gave only one score. Dividing by it is undefined, so the first version defined z = 0 for those judges and averaged it in. It felt safe: zero means "average", and average is neutral.
It isn't. Here is a constructed case that reproduces it with the real code (two honest judges, one flat judge who scored only Q1):
OLD (flat judge's 0 averaged in): Q2 0.910 (#1) | Q1 0.690 (#2) | Q3 -0.661 | Q4 -1.284
FIXED (flat judge excluded): Q1 1.035 (#1) | Q2 0.910 (#2) | Q3 -0.661 | Q4 -1.284
Q1's normalized score fell from 1.035 to 0.690, because a z of 0 was averaged against two real signals of about +1. Q2 never changed. In the live test where we first saw this, the numbers were 1.283 → 0.855.
The fix is to stop pretending. A judge with σ < 0.001 gave no relative-quality signal, so they're excluded from the normalized mean entirely, their raw scores still show in rawMean, and a project scored only by such judges gets a neutral 0, because then there really is no evidence either way. The judge is also flagged for the organizer, because "excluded silently" would be the same sin in a different place.
What I need to be honest about. On the real fixture, the fix changed 10 ranks but not the winner. The fixture has three flat judges (scores per judge: 1, 1 and 3) and the old behaviour shuffled projects at ranks 10–13, 16–18 and 30–32, not the podium. The "unanimous winner gets demoted" case is real and reproduces, but I had to construct it. I won't claim the seeded data does it.
Two more things the fix exposed:
-
The seed loader had its own copy of the maths. It still averaged the zero in, so seeded rankings disagreed with the running app after the fix. The loader now imports the same
computeNormalizedRanking. If your seed data goes through a copy of your logic, it will lie about your logic. - Our docs argued with the code. Our own JUDGING.md claimed such a judge "stops moving anyone's rank," while the code averaged the 0 in, so they pulled everything they touched toward the middle. We corrected the doc to describe what the code really did, then fixed the code itself, so now they agree.
Bug 2: a perfect record has no maximum
Scoring isn't the only way to judge. In pairwise mode a judge never gives a number: we show two projects and ask "which is better?", and recover a ranking from the pile of answers. The standard model is Bradley-Terry: every project has a hidden strength π, and the chance that i beats j is πᵢ / (πᵢ + πⱼ). It's the idea behind chess ratings. We wrote the solver by hand (about 100 lines), because no trustworthy package existed and we had to defend the maths in writing anyway.
Test it with the smallest transitive case: A beats B, A beats C, B beats C. Three comparisons. Here is the unregularized fit:
iterations=50 A=2.71e+5 B=3.68e+3 C=1.00e-9
iterations=200 A=5.45e+5 B=1.83e+3 C=1.00e-9
iterations=1000 A=1.22e+6 B=8.17e+2 C=1.00e-9
A never settles on an answer. A has won every match, and nothing in the data says how strong is "strong enough", so each extra iteration just pushes it higher. C, which lost everything, is pinned at our 1e-9 floor. The database column that stores strengths is numeric(10,6), whose largest value is 9999.999999. Even the 50-iteration row above is far past it, and on the live stack the third comparison threw a server error (HTTP 500). A perfect record is the common case on a new track, not an edge case.
The fix is one fictional win and one fictional loss for every project against a phantom opponent pinned at strength 1:
regularized (200 it): A=2.49 B=1.00 C=0.40
Same ordering, finite numbers, and the prior only matters when real data is thin, which is exactly when the unregularized answer is least trustworthy. The regression test reproduces the three-comparison case.
The same bug has a mirror in the iteration itself: projects with zero comparisons must stay pinned at the prior and be excluded from the geometric-mean rescaling. An earlier version rescaled everything, so an unjudged project's strength drifted upward without bound. A unit test asserting "an unjudged item stays exactly at 1" caught it.
Bug 3: the loop that was our own tiebreak
Pairwise mode picks the next pair for a judge by preferring the least-compared projects. Once every pair was compared, an early version fell back to re-serving a pair, on the theory that more comparisons sharpen the ranking.
It didn't settle on one repeat. Because "least compared" is recomputed from fresh counts on every call, each repeat vote changed which pair now looked least compared. The judge rotated through every pair in the track forever. A real user reported it as "the judge is stuck in a loop... these comparisons keep repeating."
The fix isn't smarter selection. It's admitting there's nothing left: the API now answers 409 ALL_PAIRS_COMPARED (HTTP "conflict"), and compare() independently rejects a repeated pair so a stale client can't reintroduce the loop. It landed on the evening of Sep 28, in the same commit as a fix for its sibling: the organizer dashboard only read the rubric-mode assignment table, so a pairwise event showed "No assignments yet" no matter how much comparing had happened. Two bugs, one root cause: pairwise mode wasn't a first-class citizen anywhere except the maths.
The benchmark that disappointed us
Our architecture doc says pairwise recomputation is "cheap enough per event" to run inside the request that triggers it. I put a number on that, with synthetic random comparisons on 40 projects, on my laptop:
40 projects, 100 comparisons: ≈ 60 ms per compare()
40 projects, 780 comparisons: ≈ 140 ms
40 projects, 5000 comparisons: ≈ 323 ms
40 projects, 23400 comparisons: ≈ 583 ms
The surprise is where the time goes. Fitting the model is the cheap part (4–13 ms). The cost is the bootstrap confidence interval, our "how much could this ranking wobble?" estimate: re-fit the model on 80 reshuffled copies of the comparison history, 80 iterations each, every copy as large as the whole history. We added the intervals because ranking uncertainty is the honest thing to show, and it turns out the statistics that make a ranking trustworthy cost far more than the ranking itself, growing with every comparison. Even at 23,400 comparisons, the equivalent of 30 judges each comparing every pair in a 40-project track, a click costs about 0.6 s, which is why running it synchronously held up at hackathon scale.
Smaller things, same family
- Editing one criterion on a submitted score recomputed the raw score from only the criteria in the edit body. With stored values 5/2/3 (true 3.7), an Innovation-only edit produced 2.5, and that wrong number fed into normalization. The raw score is now recomputed from stored rows.
- A weight like 0.00001 was accepted, stored as 0.0000, and the criterion could never affect anything. And the ±0.001 tolerance on weight sums let an all-5 evaluation reach 5.0045, so the raw score is now divided by the actual weight total.
- A second active rubric for the same track returned 201 and was silently never shown to any judge. It now returns 409.
-
A
javascript:URL passedz.string().url(), which accepts any scheme, in every user-supplied link field (repo, demo, banner, profile links) and could become a clickable script in the embeddable widget. Now http/https only, enforced once in a shared helper plus a render-time check in the widget. - Cross-event leak: an organizer of event B could pass event A's rubric id and read A's embargoed rankings. Lookups are now scoped to the event in the URL.
Most of these are the thesis again: a value that was absent (an unlisted criterion, a zero weight, an unseen pair) got silently replaced by something that looked fine.
What I'd tell someone building scoring software
- Ask what your code does with "I don't know." Zero variance, one score, a perfect record, an unseen pair, an omitted criterion. Each is a place where a default value silently turns into evidence.
- Reproduce the live failure in a unit test with the real inputs. Ours are named after what broke: "stays finite and bounded for a perfect record on sparse data", "uninformative judges".
- Don't let seed data fork your logic. If the loader has its own copy of the maths, it will agree with the docs and disagree with the app.
- Measure the sentence "it's cheap enough" before you print it.
- Your docs are code that never ran. Several of our "found live" notes are fixes to the architecture doc, which described a worker process and an audit interceptor that didn't exist.
Repo: github.com/Abhishekjha18/hackpulse (frozen for judging). Maths: apps/api/src/judging/normalization-math.ts and pairwise/bradley-terry.ts. Every number in this post can be reproduced with three commands, no install, from the gist: the two maths files copied unmodified, the shared fixture dataset, and one script per claim. Worked examples and the full reasoning for z-score over min-max, percentile and trimmed means: JUDGING.md.
Built for DOGFOOD by Hackathon Raptors, Sep 26–29, 2026. Submitted to the Write Up Quest. #hackathonraptors
Top comments (0)