mm-trader, Event-Driven Market-Making Research Platform
A high-performance, event-driven paper-trading engine that runs a fleet of 53 quantitative market-making strategies across four prediction-market sports (CS2, Dota 2, Valorant, LoL) on Polymarket, with a rigorous, anti-overfitting evaluation system. Built solo in Python.
Status (July 2026): paper research. Every strategy is hard-locked behind a live_promotion_allowed = False flag, and a statistical promotion gate decides from confidence intervals what may ever trade live. Its methods carried into the research desk, which went on to place 1,092 real orders.
What it is
mm-trader is a self-contained Python package that acts as an automated dealer / market maker on Polymarket binary prediction markets for esports matches. For each market it:
- Discovers active markets from the Polymarket gamma API / on-chain feed.
- Anchors a fair-value probability from de-vigged sharp sportsbook references, blended across books.
- Quotes two-sided bids/asks (or, in taker mode, crosses the spread) around that fair value using a configurable volatility/skew model.
- Manages each position with a per-strategy exit policy (take-profit, stop-loss, trailing stop, fair-reversion scratch, time stop, anchor-move exit, close-before-start, or hold-to-settlement).
- Records every fill, exit, and settlement to append-only JSONL ledgers and periodic atomic snapshots.
- Evaluates all 53 strategies on their real track record, marking open positions and flagging concentration and staleness.
It ran continuously under PM2 as four independent processes (one per sport), each maintaining its own isolated state directory.
The strategy fleet (53 variants, 14 families)
The fleet is a factorial experiment: each variant is a named, versioned combination of an entry (how and where to quote) and an exit policy (how to leave a position). The 14 families span dealer baselines, bracketed and reversion-based exits, information- and flow-driven exits, quote shading, sizing models, and specialist regimes, plus an overlay layer. The specific entry and exit recipes are intentionally omitted here.
Architecture & engineering
The codebase is deliberately split into a pure, synchronous core (testable, no I/O) wrapped by an asynchronous I/O shell, the classic functional-core / imperative-shell pattern. 24 modules, 10,316 LOC total.
engine.py, Pure core:(state + event) -> (state', fills)fleet_runner.py, Orchestrates all 53 variants per marketfeed.py, Async REST poller: trades 3s / market 20s / anchor 60sfair_anchor.py, De-vig + multi-book blended fair valuequoting.py, Volatility/skew quote mathexits.py, Exit-policy state machinerisk.py, Async watcher: staleness, bad-book, drawdown, kill switchevaluate.py, Bias-corrected strategy leaderboardstore.py, Append-only JSONL + debounced atomic snapshots
Performance engineering (measured wins)
| Component | Before | After | Gain |
|---|---|---|---|
| Fill detection cadence | 180 s polling loop | 3 s async trade poll | 60× faster |
| Anchor read | Full 3.3 MB file scan / tick | Tail-read + position cache | O(n) → O(1) |
| HTTP | Blocking, sequential | httpx.AsyncClient, concurrent | parallel I/O |
| JSON parse | stdlib json | orjson | 5-10× faster |
| Event loop | sync loop | uvloop | 2-4× asyncio throughput |
Reliability & risk engineering
- Kill switch,
touch KILLhalts all trading instantly while holding positions. - Staleness guards, positions on a stale snapshot are excluded from headline PnL rather than marked at a fictional price.
- Stale-anchor protection, the engine pauses on a stale or bad-book anchor so it never buys a falling knife against a frozen fair value.
- Drawdown + daily-loss kill switches in the async risk watcher.
- Promotion lock, every variant carries
live_promotion_allowed = False; the catalog validator fails the build if any variant tries to flip it.
Software quality
10,316 lines of typed Python across 24 modules, with 271 test functions across 28 test files, all passing (pytest, async mode) covering the engine, quoting, exits, risk, discovery, market selection, realism, robustness, data integrity, early settlement, fleet orchestration, sports profiles, and taker logic. Clean dependency surface (httpx, orjson, uvloop); ruff-linted; reproducible pyproject.toml.
Scale & metrics (real run data)
| Metric | Value |
|---|---|
| Strategies simulated concurrently | 53 (40 maker + 13 taker), 14 families |
| Sports covered | 4 (CS2, Dota 2, Valorant, LoL) |
| Market state directories | 150 |
| Total simulated fills (lifetime) | 19,309 |
| Settled markets in the verdict corpus | 277 (376 settlement records) |
| Lines of code | 10,316 |
| Automated tests | 271 (all passing) |
| Fill-detection latency | 3 s (60× faster than baseline) |
Artifacts from the running system (real, sanitized)
A settlement record, decomposed at write time. Every settlement is split into spread and direction, per variant, the moment it books; the decomposition lives in the ledger schema, not in a notebook:
{
"type": "settlement",
"pnl": 90.0,
"match_slug": "cs2-***-2026-07-10",
"outcome": 1.0,
"decomp": {
"variant-A": {"spread": 1.8619, "direction": 28.1381, "pnl": 30.0, "n": 3},
"variant-B": {"spread": 1.8619, "direction": 28.1381, "pnl": 30.0, "n": 3},
"variant-C": {"spread": 1.9809, "direction": 28.0191, "pnl": 30.0, "n": 3}
},
"_ts": "2026-07-10T20:00:48.535789+00:00"
}
The orphan-position tripwire, hourly. A "leak" (a resolved market still holding unsettled inventory) would page immediately. Running record: zero.
in-flight (correct): 9 leaks: 0 audit-errors: 0
[2026-07-13T00:47:01Z] exit=0
Quantitative research method
Scoring a strategy fleet without flattering it is harder than running it. The evaluation system catches the standard traps:
- Open-position bias correction, "realized PnL" alone would favour strategies whose positions are still open. The system always marks open positions at the current fair so every strategy is judged on the same footing.
- Concentration tracking, a large ROI on one unsettled market is not edge. Every number is reported alongside market count, settlement count, and top-market PnL share.
- Spread-vs-direction P&L decomposition; every settled fill is split, at write time, into spread captured at entry versus direction (which way the game went). This is the fleet's core diagnostic: it separates the dealer's business (spread) from the implicit bet (direction), and it is how adverse selection gets measured directly in the ledger instead of guessed at.
- Bootstrap confidence intervals, strategy edges are tested with 20,000-resample bootstrap CIs on per-market PnL, so "positive total" is never confused with "statistically real."
- Data-driven pruning, capital attention concentrates on the families that earn it; underperformers are paused reversibly (kept in catalog so open positions still settle, but posting no new orders).
What the evaluation found
- Exit design matters more than entry design. Pairing every entry with several exit policies showed that how a dealer leaves a position shapes its results more than where it quotes, which is why the fleet is built as a factorial experiment rather than a list of strategies.
- Adverse selection is measurable, not hypothetical. The spread-versus-direction decomposition made the informed-flow effect visible directly in the ledger, and showed it growing with the gap between the fair-value anchor and live match state. Diagnosing the mechanism points straight at the next engineering lever: a faster, sub-second live match-state feed.
- The promotion gate works. The standing rule (a variant's 95% bootstrap confidence interval must clear zero before any live capital) is enforced in code and evaluated over the full settled corpus of 277 markets at 20,000 resamples per verdict. Capital only moves on evidence. (Exact strategy parameters and figures are withheld.)
Built to be trusted
Most research platforms cannot tell you whether their own results are real. This one answers that question with arithmetic, in code, before any money moves: a 20,000-resample bootstrap over 277 settled markets, per-variant confidence intervals, and a promotion flag the build refuses to flip by hand.
Holding it up: 13 ledger-invariant tests (append-only, reward income never mixed into PnL, snapshot reconciliation) and an hourly orphan-position audit that has recorded zero leaks. When the exchange announced a new taker fee schedule, the engine reproduced it within hours and re-based every A/B epoch so results from before and after the change were never mixed.
The platform's instrumentation, its evaluation discipline and its decomposition diagnostics carried straight into the next generation of the work, the research desk, which now records the full order book on two venues and has run real orders end to end.
Related: The research desk: the next generation of this work · 732 trades that were really 37 matches · Catching silent data corruption