This project is a submission for the Hacktoberfest: Build for a Friend DEV Challenge, targeting the **Best Use of TabPFN* category.*
1. Context & Motivation
Me and my brother play in the same competitive fantasy football league, which means my default setting on Sundays is actively rooting for his team to implode.
That said, watching him agonize every single week over two flex players projected within 0.3 points of each other got painful to witness.
Default platform projections have known limitations:
- Static estimates: They rarely adjust quickly for game scripts (Vegas totals and spreads).
- Positional matchups: Opponent defensive strength against specific positions is often averaged out or ignored.
- Point estimates only: They show a single expected number rather than a distribution. In reality, whether you need a high floor (P10) to protect a lead or a high ceiling (P90) to chase upside completely changes who you should start.
I built StressFreeFantasy to automate these decisions. It syncs with private ESPN leagues, conditions an in-context tabular foundation model on historical NFL data, and outputs an optimized lineup with floor/ceiling estimates. (Though whether I let him use it against me when we play head-to-head remains an open question.)
2. Why TabPFN?
Standard tabular models (like XGBoost or LightGBM) require training pipelines, feature scaling, and hyperparameter tuning. More importantly, when leagues use non-standard scoring rules (Full PPR, Half PPR, TE premium, big-play bonuses), a traditional model has to be retrained from scratch for those specific point values.
TabPFN handles this differently:
- In-Context Adaptation: As a prior-data fitted foundation model, it runs in-context learning in a single forward pass. By feeding it 2,000 historical NFL player-weeks recalculated to the user's specific scoring settings, TabPFN adapts instantly without gradient steps.
-
Quantile Outputs: TabPFN natively supports querying specific quantiles (
quantiles=[0.1, 0.9]), providing calibrated P10 (floor) and P90 (ceiling) values rather than just a noisy mean.
Why Open Innovation Matters Here
Relying on an open-weight, locally runnable tabular foundation model makes all the difference compared to closed cloud APIs:
- Zero API Cost & Unlimited Inference: Simulating thousands of matchup permutations and testing historical holdouts incurs zero per-token API charges.
-
Data Privacy: League authentication tokens (
espn_s2,SWID) and private league rosters stay completely local on the user's machineβnever routed through a third-party server. - Offline-Ready In-Context Learning: TabPFN runs locally on a consumer laptop CPU/GPU without depending on external proprietary server uptimes right before Sunday kickoffs.
3. Architecture & Data Pipeline
The project is built with Streamlit and ties together three data sources:
ββββββββββββββββββββββββββββββ
β nflverse / nflreadpy β
β 5 Seasons Historical Data β
βββββββββββββββ¬βββββββββββββββ
β
βΌ
ββββββββββββββββββββββββ ββββββββββββββββ ββββββββββββββββββββββββ
β Private ESPN League ββββΆβ TabPFN ββββΆβ Streamlit Dashboard β
β (espn-api / S2&SWID)β β In-Context β β Lineup + Floor/Ceil β
ββββββββββββββββββββββββ β Regressor β ββββββββββββββββββββββββ
ββββββββββββββββ
Feature Engineering
Features are computed without data leakage, using strictly pre-game rolling historical signals:
-
Player baseline: 16-game rolling average (
Avg_Pts), previous week's score (Last_Week_Pts), and short-term form (Form_L3). -
Volume: Rolling 3-game touches and targets (
Volume_Proj), which is historically more stable than point totals. -
Game environment: Opponent defensive ranking against the player's specific position (
Opp_Def_Rank), dome status, and Vegas implied team total derived from over/under and point spread: $$\text{Implied Team Total} = \frac{\text{Vegas Total} + \text{Spread}}{2}$$
4. Benchmark & Validation
Fantasy projections have high variance due to random touchdown variance, making raw MAE an incomplete metric. What matters for a manager is whether the model ranks the right player higher in a head-to-head Start/Sit call.
We set up an out-of-time temporal test, fitting on earlier seasons and testing on an unseen holdout season ($N = 600$ test player-weeks):
| Metric | Baseline (Avg) | Baseline (Form) | TabPFN (Ours) |
|---|---|---|---|
| MAE (lower is better) | 6.84 pts | 7.12 pts | 5.91 pts |
| Correlation ($r$) | 0.44 | 0.38 | 0.58 |
| Pairwise Start/Sit Accuracy | 52.8% | 51.4% | 63.7% |
In pairwise head-to-head matchups between players at the same position, TabPFN selected the higher-scoring option 63.7% of the time.
5. Interface & Usage
The Streamlit UI provides:
-
ESPN Sync: Imports rosters, schedules, and custom scoring tables using league credentials (
espn_s2,SWID). - Manual Adjustments: An interactive data editor to tweak expected volume, health status, or weather/stadium conditions.
- Lineup Optimizer: Greedily fills position slots (QB, RB, WR, TE, FLEX) sorted by projected output and matchup edge.
6. Testing & Feedback
Since we play in the same league, I initially tested the engine on my own roster to see how the recommendations differed from ESPN's defaults. Rather than relying on a single static projection, having direct access to calibrated P10 (floor) and P90 (ceiling) intervals made marginal Start/Sit calls immediately clearer.
After seeing the floor/ceiling spreads and the pairwise accuracy numbers, I showed the tool to my brother to help him with his own weekly lineup headaches. He now uses it as a sanity check before kickoff rather than overthinking marginal callsβwhich unfortunately means his roster is noticeably harder to beat when we play each other.
7. Repository & AI Transparency
- GitHub: lucasantonsson/StressFreeFantasy
-
Stack: Python, Streamlit, TabPFN,
nflreadpy,espn-api - AI Collaboration: Code generation, pipeline scaffolding, and debugging were assisted by Claude and Gemini, while system design, logic requirements, and validation were directed independently.

Top comments (0)