A research platform should not make it easy to stop at a good-looking number. That applies to negative results too. Transaction-cost checks, held-out evaluation, and parameter sweeps are useful only when their results can be traced back to their runs.
| Candidate | Historical report | Current audit status |
|---|---|---|
| Price mean reversion | Gross Sharpe 7–13; rejected after costs | Historical revision range not independently reproduced here |
| Funding carry | +48.6% development; −0.5% holdout | Rerun: +45.62% development; +1.31% holdout; discrepancy unresolved |
| Daily momentum | +97% best case; sensitive to lookback | Matching run/configuration not located; sweep remains unverified |
Edge 1: Mean reversion and the cost problem
The historical summary reports cross-sectional reversal on 15-minute bars with gross Sharpe between 7 and 13, positive in over 80% of out-of-sample windows across 76 test periods. These figures span research revisions; they are not measurements from one controlled run.
That summary also reports friction at 5–10 times the available edge when rebalancing every bar, and neighboring settings ranging from +45% to −60% net. Those revision-specific ratios and parameter results have not been independently reproduced in this audit. They should not be treated as verified evidence that every implementation must fail.
The reported decision was to reject the tested implementations because a robust net result had not been established. The engineering question is whether a pre-cost pattern survives a specified execution model, not whether mean reversion is universally untradeable.
A separate deep-dive describes a controlled evaluation. Its metrics should not be mixed with this historical revision summary as if they came from one run.
Edge 2: Funding carry and a failed reproduction
The tested carry construction combines long spot with short perpetual exposure and funding payments. Lower turnover reduces transaction costs; it does not eliminate them or other risks.
The original report gives +48.6% net development return, Sharpe 1.17, and −0.5% net holdout return, Sharpe −0.02. On September 30, I reran the current implementation against the current local cache. Both the original evaluation script and a separate audit runner produced different results:
| Metric | Historical report | Current rerun |
|---|---|---|
| Development net cumulative return | +48.6% | +45.62% |
| Development Sharpe | 1.17 | 1.079 |
| Holdout net cumulative return | −0.5% | +1.31% |
| Holdout Sharpe | −0.02 | 0.208 |
Historical report: −0.5%. Current rerun: +1.31%. These are different run versions, not a controlled experiment measuring an improvement. The cause of the discrepancy is not established.
The current configuration uses FundingCarry(384), rebalances every 1344 15-minute bars (14 days), normalizes to fixed gross exposure, and applies modelled fees of 4 bps plus slippage of 2 bps. It selects 45 symbols and scores 76 development windows and 17 holdout windows.
The 17 complete 21-day holdout windows cover September 1, 2025 through August 24, 2026 (exclusive): 357 days, not the entire configured interval ending September 1. The original saved window list was not found, so historical coverage cannot be independently certified.
In the current rerun, gross holdout return is +2.14% and net return is +1.31% on the same windows. Costs reduce the result under this model. This comparison does not explain the historical −0.5% return. Development returns are positive in each represented calendar year, including +3.44% in 2022; that does not recover the original annual figures.
The audit runner fingerprints source code and cached input files before and after execution, saves scenario returns, and independently recalculates cumulative return and Sharpe from those returns. Those checks passed, as did 14 relevant engine tests. They check these calculations, not every modelling assumption or the historical result.
The current universe filter uses funding-record counts from the cache before date restriction. Applying the threshold using only pre-holdout records yielded the same 45-symbol set in this cache. That does not establish point-in-time eligibility throughout development. No original run manifest or committed HFM code history was available to recover the historical code, universe and dataset. Nor does this audit prove the reserved interval was originally untouched.
The earlier development null comparison, reported at about +28%, was not reproduced in this audit. It is not used here to infer a market-wide premium or selection skill.
Null diagnostics on the current rerun
Two structure-preserving null tests were run post-hoc on the 17 holdout windows of the current reproduction result (net +1.31%, Sharpe +0.208). These diagnostics do not apply to the historical result and do not resolve the discrepancy. They ask two separable questions: did the temporal alignment between position decisions and outcomes contribute, and did the specific coin-selection rule add anything beyond random assignment?
CircularShiftNull — temporal alignment. For each coin i, a shift τᵢ is drawn uniformly from [0, T). The weight path becomes:
w̃ᵢ(t) = wᵢ((t + τᵢ) mod T)
What this preserves: per-asset weight distribution, autocorrelation structure, holding period, and per-asset turnover Σₜ |wᵢ(t) − wᵢ(t−1)|. What it destroys: the alignment between a position decision and the specific return interval it faced. H₀: the observed Sharpe is no better than an arbitrarily time-shifted version of the same weight paths. N = 500 seeds.
RebalanceSelectionNull — coin selection. Rebalance events are detected as:
T_R = {t : ‖w(t) − w(t−1)‖₁ > ε}
At each t ∈ T_R, the weight vector is permuted across universe symbols. What this preserves: rebalance schedule, gross exposure Σᵢ|wᵢ|, number of active positions, weight magnitudes. What it destroys: which symbol receives which weight at each rebalance. N = 500 seeds.
Both use the Davison-Hinkley finite-sample p-value to avoid the p = 0 artifact from small N:
p = (1 + #{θ_null ≥ θ_obs}) / (N + 1)
Turnover inflation. CircularShiftNull wraps each coin's weight path at the window boundary. When τᵢ > 0, the series at time T − τᵢ jumps discontinuously to the value at time 0, creating large simultaneous |Δw| spikes at the seam. Mean null turnover reached 1.82× frozen. RebalanceSelectionNull avoids the wrap-around artifact but permutes across the full universe rather than the positive-funding eligible set, changing the effective cost profile. Mean null turnover: 1.34× frozen. In both cases the frozen strategy received a structural cost advantage independent of signal quality. The CircularShift p-value (0.010) reflects a mix of genuine temporal misalignment loss and artificial cost inflation that cannot be cleanly separated. These p-values are reported for transparency, not as evidence.
The right next step is a funding-matched null: at each rebalance, permute weight assignments only within the positive-funding eligible set. This would hold funding-exposure loading constant and isolate the coin-ranking question directly. It was not implemented because the execution API does not expose the contemporaneous panel signal needed to define the eligible set.
Sharpe uncertainty via block bootstrap. The 17 holdout windows are serially dependent: consecutive windows share overlapping formation data, and 21-day non-overlapping hold periods are far too short for serial correlation to decay. The naive standard error 1/√17 ≈ 0.24 treats windows as independent and understates uncertainty by a material factor.
A moving-block bootstrap was run on the 15-minute holdout return series. Block length L = 4032 bars corresponds to approximately 3 full rebalance cycles (~63 days), chosen to span the dominant autocorrelation horizon of the strategy. N = 2000 resamples; annualization factor √(365 × 24 × 4).
point estimate: Sharpe +0.208
95% CI (L = 4032): [−0.21, +0.44]
p(Sharpe > 0): 78.2%
The 95% interval spans zero. With 17 non-overlapping windows, wide CIs are the honest result, not a flaw in the method.
The historical REJECT decision is part of the research record, but a negative holdout is no longer a verified rationale in this article. The candidate remains on hold pending reconciliation. The positive rerun alone does not establish a tradeable edge. Analysis of this already-opened interval is exploratory, not a new out-of-sample confirmation.
Edge 3: Momentum and an unverified parameter sweep
The historical summary describes daily cross-sectional momentum with weekly rebalancing and a best net return of approximately +97%, including taker fees. It gives four lookback results: 7 days +33%, 10 days −23%, 14 days +97%, and 21 days +35%.
Values transcribed from the frozen narrative report. A matching saved run/configuration was not located in this check. These are not four independently validated results.
If reproduced, that sweep would raise a parameter-sensitivity concern: the best setting is not surrounded by comparably strong results. It would not, by itself, prove the peak was luck. The historical summary also attributes about half the return to 2021 and reports a negative 2022. Those regime figures remain unverified here; the report did not estimate a formal beta decomposition.
The reported WEAK verdict is a historical assessment, not a newly verified finding from this audit.
What the audit actually established
The current carry calculation is reproducible locally with recorded code and input hashes. Its headline values do not match the historical report. For the mean-reversion revision range and momentum sweep, narrative evidence is available but numerical reproduction remains incomplete.
That is a less tidy conclusion than three precise diagnoses. It is also the result the evidence supports. A negative backtest deserves the same provenance checks as a profitable one.
No candidate is being advanced to live trading on the basis of this article. The next step is to recover the historical run artifacts or explicitly retire unsupported historical numbers, not silently substitute new results into the old story.
The research package preserves the original narrative evidence. This is not a complete independent reproduction of the private data pipeline. SHA-256 checks establish file identity, not economic validity. The frozen source retains stronger claims that this corrected article does not endorse.
Evidence source SHA-256: cf9ce73542af24be3e5c2872864b2b41c6ce864e3b2f07b760e5434a1bdb23ba
Research system: HFM · Historical reports and a post-publication rerun · Bybit spot and perpetuals · Python


Top comments (6)
Three guards that each catch a different failure is the useful structure here; any one of them alone would have passed something.
One thing about the lockbox verdict on carry: a single year is a weak test at that Sharpe. For a strategy whose true annual Sharpe is 1.17, the realized Sharpe over one year has a standard error of roughly 1 to 1.3, so about one year in six to eight comes out at −0.02 or worse with nothing having changed. The held-out year is consistent with a regime shift, but it is also consistent with an ordinary bad year of the development edge, and one year can't tell them apart.
The null you already built can split that further. Run the shuffled-funding strategy through the same lockbox year: if it also drops to about zero, the shared funding premium faded in that year, and your coin selection wasn't what failed. If the null still earns while the frozen strategy doesn't, the selection is what broke, and that is the stronger reason to reject it.
The momentum sweep is the clearest picture of a spike I've seen: four lookbacks, and the +97% sits between −23% and +35%. Counting those four (and any others tried in earlier revisions) as trials when judging the best one would put a number on how much of it is selection.
Thank you! It pushed me further than the original article went.
Running a reproduction audit: the holdout changed sign. Historical report: −0.5%, Sharpe −0.02. Current rerun: +1.31%, Sharpe +0.208. Cause not established. The REJECT rationale is no longer verifiable from the original number — updated the article with a full correction.
On your null question: CircularShiftNull gave p = 0.010 but with 1.82x turnover inflation from wrap-around artifacts, so the p-value is not usable. The clean test (permute only within positive-funding eligible set at each rebalance) is not implemented yet. Block bootstrap CI on the holdout: [−0.21, +0.44], spans zero. Full methodology with formulas in the updated article.
Publishing the correction instead of quietly updating the number is the right call, and it's rarer than it should be.
Two things that might help with the audit:
The fastest way to find the cause is usually to diff the two runs trade by trade rather than compare the totals. The first date where the position or fill differs points straight at it: a data revision (funding or price history re-downloaded), a changed universe, or a config default that moved. For the future, storing the data snapshot hash and the commit next to every reported number makes this a one-line check.
And the decision may survive the flip. Your block bootstrap interval [-0.21, +0.44] contains both the old -0.02 and the new +0.208, so the two runs aren't distinguishable from each other, and neither is distinguishable from zero. The rationale changes from "negative on the holdout" to "no evidence either way on the holdout", but that's still a reject for a strategy that needed the holdout to confirm it.
On the null: your permute-within-eligible-set design is the cleaner one. If you want something quicker in the meantime, circular shifts by whole rebalance periods, with the wrap-around segment dropped from the P&L, usually remove most of the turnover artifact.
SE is what the article understated. With one year and Sharpe 1.17, the realized distribution is wide enough that −0.02 sits inside it. Leaning on the negative sign as evidence was wrong.
The null you described is what was attempted. It broke differently: shuffled-funding inflated turnover 1.82x and costs ate everything. The funding-matched version (permute only within eligible coins at each rebalance) would do the split. Not built yet.
On momentum: four lookbacks. p ≈ 0.02 for 14-day does not survive Bonferroni at four tests.
One thing on the momentum test: the four lookbacks are strongly correlated with each other, so Bonferroni treats them as four independent chances when they're closer to one or two, and it ends up stricter than it needs to be. A permutation test on the maximum over the four gives the exact correction: shuffle or circularly shift the returns, recompute all four lookbacks each time, and record the best p (or t) of the four. The share of shuffles whose best beats your 14-day result is the adjusted p-value. It will land somewhere between 0.02 and 0.08, and with correlated lookbacks usually nearer the low end, so it's worth running before calling it a miss.
Fair point. Bonferroni assumes independence and overcorrects here - daily momentum at 7, 10, 14, 21 days on the same assets will be highly correlated. The max-T permutation procedure you described gives the exact family-wise correction without that assumption. Worth running before closing the verdict on momentum.