732 Trades That Were Really 37 Matches
A case study in auditing my own trading-research platform: three flattering artifacts removed, the evaluation rebuilt, and a promotion gate that went on to govern 53 market-making variants.
Status (July 2026): paper research. This article withholds specific strategy parameters, data sources and current figures. The methods it produced carried into the research desk, which went on to place 1,092 real orders.
The setup: numbers that demanded scrutiny
By mid-2026 the platform had what looked like a mature research result: hundreds of settled trades, a calibration curve that said the model's probabilities could be trusted, and an edge estimate worth funding. Everything a solo quant wants to see.
The correct response to a result you want to be true is to attack it. So I audited my own lab the way I'd audit a stranger's, and that decision is what everything good that followed came from.
Catch #1, the 732 trades that were really 37 matches
The "settled trades" ledger contained 732 rows. Tracing row provenance revealed they collapsed to 37 unique matches, each counted ~19.8 times over, a classic fan-out where per-strategy and per-snapshot records were being treated as independent bets. Every downstream statistic (sample size, significance, edge-per-trade) was inflated by the same factor.
A sample of 732 justifies confidence. A sample of 37 justifies humility. They were the same data, and the audit caught it before a single decision was made on it.
Catch #2, the calibration curve that graded its own homework
The calibration curve ("when the model says 70%, does it win 70% of the time?") had been fit on the platform's own bets. That's self-reference: the curve inherits every selection effect of the betting policy it's supposed to validate.
The tell was mathematical: the curve was non-monotonic, which is impossible for genuine calibration on adequate data. Real calibration can be noisy; it cannot systematically bend backwards.
The rebuild:
- Bet-side-invariant calibration, evaluated over every observed market, not just the ones the model chose to bet, so the betting policy can't flatter the curve.
- Wilson confidence intervals on every calibration bucket, with explicit monotonicity integrity flags in the output.
- An exogenous match-observation ledger, an append-only record of every observed market, bet or not, so future evaluation has a denominator the model can't curate.
- The recomputed output ships with explicit
bias_removed/bias_remainingfields, because honesty is a schema, not a mood.
Catch #3, the "best result" that was survivorship bias
The most exciting-looking finding in the fleet was that removing stop-losses appeared to triple profits. The audit traced it to an exit-recording rule that booked positions at different times depending on how they closed, so variants were being compared on unequal footing. Marking every open position at fair value put all of them on the same scale.
The fix went beyond correcting the number. The strategy leaderboard now prints its own bias warnings in every snapshot: which metrics are flattered by open positions, which variants have too few settlements to mean anything, and how concentrated each result is in its single best market. The report argues against itself before anyone else can, which is exactly why its conclusions can be trusted.
Redirect, don't defend
With clean data and a rigorous evaluation in place, I pointed the platform at a sharper question: market-making, where settlement decomposition could measure spread capture separately from directional exposure. That is the question professional liquidity providers are paid to answer, and it is a better fit for these markets than picking winners.
The infrastructure, the instrumentation and every control the audit produced went with the move, and that is what made the next phase fast.
Where it led: the program that polices itself
The market-making program was evaluated from day one under the controls this audit created: exogenous denominators, bootstrap CIs, multiple-testing correction, promotion gates.
That program grew into a 53-variant fleet whose every result is evaluated by a statistical promotion gate: 20,000-resample confidence intervals over hundreds of settled markets, spread-vs-direction decomposition on every settlement, and a hard rule enforced in code: no strategy touches live capital until its 95% CI clears the bar. The full story is in the mm-trader case study, and the next generation of the work is the research desk.
The researcher and the platform stayed the same. What changed is that every result now has to pass the same checks the audit created.
The discipline that carried over
The standing controls it put in place:
- Multiple-testing correction everywhere: with dozens of strategy variants, a couple will always look brilliant by luck. A result only counts if it survives being one of many tries (Benjamini-Hochberg FDR gating).
- A statistical screen for external signals tested against an efficient-market null with out-of-sample requirements strict enough that a false signal cannot sneak through, so when something passes, it means something.
- Pessimistic sizing via Monte-Carlo Kelly on the p25 tail of log-growth, using clustered and cost-repriced returns.
- Promotion gates as code: nothing moves from paper toward live without surviving out-of-sample screens, and every variant carries a hard
live_promotion_allowed = Falselock.
What the audit changed
Sample size now means unique matches. Calibration uses an exogenous ledger. Open positions are marked instead of left outside the result. Every strategy stays locked until its confidence interval clears zero, so capital only ever follows evidence.
That discipline became the foundation of everything I have built since, including a research desk with 280+ pre-registered experiments and a live execution lane measured on 1,092 real orders.
Related: The research desk · mm-trader: the research program this discipline governed · Catching silent data corruption