GitHub: github.com/MENG-COOLMAN/PitchQuant
For academic research and technical exchange only. Not betting advice.
What's new in this update
In my previous post I wrote about building a complete football data analysis pipeline from 220,000 matches, pushing direction accuracy from 48% to 55%. But one question kept bugging me — how should you actually read odds movement?
Over the last two weeks I figured it out: the odds aren't set by one person. Two groups are buying them.
Retail bettors chase Chinese lottery odds, piling into whichever direction is dropping, convinced "dropping odds = smart money." Institutions move the Asian/European market with real money, pushing lines toward what they believe. These two groups disagree all the time — and when they do, who's right?
The answer: most of the time, retail is wrong.
So in this update (V3.5.76) I did one thing: split these two groups into two independent channels that cross-validate — and expose — each other. This post walks through the "dual-channel" mechanism and the uncomfortable finding it surfaced.
Why two channels are necessary
First, a measured data point. I backtested the tick-by-tick Chinese lottery odds: when a direction's odds keep dropping before kickoff (retail reads this as "the bookmaker likes it"), that direction actually hits only 25% of the time.
Four-match sample — the dropping direction won a quarter of them.
The drop isn't wrong; the cause of the drop is wrong. The Chinese lottery vig is as high as 12.8%, and its water levels mostly follow retail sentiment, not information. Retail piles in, odds drop, and the bookmaker is happy — because with lower odds, retail's expected value evaporates and the bookmaker's exposure shrinks.
The Asian/European market, by contrast, has only ~6.9% vig. Line and water movement there is backed by real hedging pressure. Institutions don't buy retail sentiment; they buy their own positions.
So I split the model's data flow:
Lottery tick-by-tick odds → Retail channel (sentiment·12.8% vig)
Asian/European time series → Institutional channel (intent·6.9% vig)
↓ cross-validate ↓
Quadrant decision (Q1 consensus/Q2 trap/Q3 upset/Q4 block)
The four quadrants: how the two groups fight
Crossing the two channels produces four scenarios, ranked by priority:
Q2 Trap alarm (most dangerous): retail floods direction A with dropping odds, institutions stay still or push the other way.
→ Classic "trap." Retail is moving exactly where the bookmaker wants. Handling: knock the favorite down one confidence notch and force the second direction into a tie.
Q1 Consensus (most trustworthy): retail drops odds on A, institutions follow by deepening the line.
→ Both groups agree — only now is the signal actually worth something. Confidence up half a notch.
Q4 Block signal: retail sees odds rise on A (too scared to buy), institutions quietly move toward A.
→ The bookmaker is "blocking" — deliberately inflating odds to scare off retail while quietly favoring A itself. Follow the institutional direction, +0.5 notch.
Q3 Upset hint (weakest): retail hasn't moved, institutions move alone.
→ Flag for attention only, no adjustment. Wait for confirmation.
One hard rule here: the quadrants may only adjust confidence, by at most ±1 notch, and may never reverse the hardcore direction. I learned this the hard way — letting weak signals overturn strong ones just hands control to noise.
The time dimension: when is a signal real?
The other dimension I added is time.
Lines older than 48 hours before kickoff I call "smoke screens" — they're shaped by capital flows and the bookmaker's exploratory positioning, and carry little signal. Lines within 90 minutes of kickoff carry the most information density, because by then the money that was going to arrive has arrived, and the line basically reflects real views.
>48h early line → smoke screen·institutional channel downgraded to weak reference
≤90min late line → highest information density·full weight
in between → normal window
It sounds simple, but this grading filtered out a lot of "fake signals" in backtesting. Many seemingly bizarre odds movements, zoomed out on the time axis, turn out to be early-line noise.
Engineering debt I paid off along the way
Beyond the algorithm, the engineering got a lot of attention too. The model now runs 413 automated checks — if rules and data disagree, it fails loudly. "Looks about right" doesn't pass.
A few fixes worth mentioning:
- Cross-language team matching: snapshots store English team names, but the lottery txt files use Chinese names, and they used to never match. Now there are 191 bilingual alias mappings, covering all top-5 leagues + UCL.
- The auditor audits itself: my check rules used to have a loophole — "keyword present = pass." Now a check needs the keyword plus a number to count, so the auditor can't be lazy.
- Fixture regression: every parsing bug we've ever hit is now frozen as a real input sample, and the model re-runs all of them on every pass — so the same bug can't come back in a different function.
The other big stuff this fortnight
Besides the dual channel, several new modules landed that weren't in v1:
- UEFA Nations League sub-model: national-team football is a completely different game — no transfers, no form continuity, mystery motivation in friendlies. I built a dedicated tiered rule set for the NL, backtested on 658 national-team matches, and set a "domain discipline": top-5 league thresholds don't apply to national teams, and vice versa.
- European competition three-layer calibration: UCL knockout two-legged ties, group-stage round-robins, and neutral-venue finals all have completely different motivation and formats. Now each is calibrated by tier — the first-leg leader parks the bus in the second leg, group-stage dead rubbers get tanked, all quantified into the model.
- Per-league sub-models: each of the top-5 leagues gets its own distilled rule set (7,000-9,000 backtested matches each). Bundesliga is a goals machine, Serie A is draw-heavy, Ligue 1 is low-scoring — applying one universal threshold to every league is wrong by construction.
- Post-match learning loop: after every match, the model writes the result, prediction, and actual odds back to the case library, then periodically re-runs backtests to recalibrate the lookup tables. It doesn't retrain everything on every update — only the parameters whose backtests show genuine drift.
The honest numbers
As always, the cold-water section:
- Direction accuracy: ~55-58% (clear gains on weak-signal matches, stable on strong-signal matches)
- Dual-channel core scenario: low European/Asian dispersion (spread<2%) hits **52.7%** direction accuracy; high dispersion (spread>10%) drops to 44% — the more aligned the institutions, the more trustworthy the signal
- Deep favorite + institutional alignment (home <1.50 and spread<2%): 80.0% (smaller sample, reference only)
- Top-2 score hit rate: ~30% (strong-signal matches)
- All numbers are time-split backtests on 220k historical matches. They do not predict the future.
And the line from the project's front page: football is chaotic, markets are efficient. This project helps you make fewer mistakes — it doesn't help you beat probability. Negative long-term EV on sports lottery is structural, independent of any model.
Open source
The project is on GitHub under MIT license:
- Core scripts are all Python, stdlib-heavy
- Distilled JSON lookup tables (de-vig calibration, score depth, handicap tiers, etc.) are included
- The raw 220k-match CSV is too large and has source terms, so it's not included
- You'll need your own API keys (odds-api and api-football both have free tiers)
- Calibrated for top-5 European leagues + UCL; other leagues aren't calibrated
One last thing: for academic research and technical exchange only. Not investment advice.
If you're curious about telling retail and institutional money apart in odds, star it and drop an issue. Happy to discuss.
A student studying for grad school in financial math. This project was built bit by bit in my spare time. All the data, rules, and code are on GitHub. Come build with me.
Top comments (4)
The "dropping direction wins only 25%" figure is from four matches, and one win in four is well within noise: if the true rate were 45%, you'd see one or fewer wins out of four about 39% of the time. With 220,000 matches in the pipeline, that one is worth re-running on the full tick history before it drives the Q2 trap rule.
The other number I'd pin down is the baseline for the 55–58% direction accuracy. In three-way football markets, simply picking the bookmaker's favourite is right somewhere around half the time, so the comparison that says whether the model adds anything is its accuracy against the de-vigged market favourite on the same matches, or better, its log loss against the closing-line probabilities. Accuracy against 50% or 33% overstates it.
And with per-league sub-models, Nations League tiers, quadrants and time windows all tuned on backtests, it's worth counting how many rule variants were tried. A full season held out and never touched during tuning is the cleanest check that the lift comes from the rules and not from the search.
Thanks !Genuinely the most useful feedback this project has gotten, and all three points are blind spots on our side. Fair hits, here's the honest version.
The 25% figure.Fair. Four matches is noise, and we shouldn't have led with it. That number was a mechanism demo ("lottery drops ≠ smart money"), not a calibrated stat — Q2 is actually driven by the bigger Asian/European dispersion backtest (spread<2%: 52.7% vs >10%: 44%). But a headline should survive its own sample size: I'll re-run it on full tick history and add the binomial CI, or drop the claim.
The 55-58% baseline.Also fair. The number that actually means something from my backtests: naive lookup tables were ~84.4% vs ~86.1% random — basically nothing — and the full pipeline gained ~+6.4pp on top. I should have led with that. Going forward I'll report against the de-vigged market favorite on the same matches, and switch the primary metric to log loss vs closing lines. Accuracy hides calibration error; you're right.
Overfitting. You nailed the one I worry about most. I never counted rule variants tried — that alone is a red flag. Current defenses (time-split, fixture regression, drift gate) don't substitute for a clean holdout. Plan: one full season per league, never touched during tuning, and "lift survives the untouched season" becomes a release gate. Until then, sub-model numbers stay labeled "in-sample, holdout pending."
If you have advice, I'd genuinely welcome it. Fighting overfitting and the structural risk in the model design is something I've wanted to do properly for a long time, and your comment reads like you've thought about this more than I have. Specifically, I'd love your take on: how you track rule-variant experiments cheaply (without building a whole experiment framework), how many held-out seasons you'd consider enough before trusting a release gate, and whether you'd fix the eval around log loss first or around the de-vigged-favorite baseline first. No pressure — even a sentence on any of these would help.
Glad it helped. Short answers to the three:
Tracking variants cheaply: wrap the one function that scores a configuration so every call appends a line to a file: time, git commit, a hash of the rule/config, the data slice, and the metrics. No framework needed. The number of distinct config hashes ever scored on the tuning data is your trial count, and the file doubles as the record of what you tried and dropped.
How much holdout: size it from the lift you need to detect rather than counting seasons. One season across the top five leagues is about 1,750 matches. For accuracy near 50% the standard error is about 1.2 points, so that season can only confirm a lift of roughly 2.5 points or more. Comparing the model and the market favourite on the same matches helps a lot, because only the matches where they disagree carry information.
Which eval first: log loss against de-vigged closing probabilities. It uses every match, rewards calibration, and the per-match difference between your log loss and the market's is a paired statistic you can put a confidence interval on directly. Accuracy against the favourite then becomes the readable headline, computed on the same matches.
Thanks for the constructive feedback. I'll attempt to modify and backtest the model based on your suggestions.