Back to portfolio
Case Study · Prediction-market research

The Research Desk

I run a research desk on esports prediction markets. It records the Polymarket and Kalshi order books with its own recorders, files one directory per experiment with the hypothesis written before the data, and promotes a result only when it clears a bar declared in advance. In September 2026 it went live on Polymarket with real orders.

Figures on this page were counted on 3 October 2026. Most of them grow every day, because the desk runs every day.

What it sees

On 22 August 2026, Polymarket's book for FURIA vs FUT Esports went from a 92¢ bid to no bids at any price in 25 minutes. The desk holds 548,830 full-depth snapshots of that one book from match day, and at the scheduled start 92% of the money resting within 10¢ of the price was gone in 30 seconds. The full story, with the chart.

The question it asks

Most people building models on a betting market ask who wins. That question is already priced, and the market is very good at it. The desk asks a sharper one:

What does a resting order earn, given the state of the book it is resting in?

Written as a quantity, that is E[maker PnL | market state] rather than P(team A wins). The two questions need completely different evidence. A forecast is scored against an outcome. A resting order is scored against the spread, the fee, the queue it sits in and the information that arrives while it waits. Answering the second question well is what professional market makers are paid for, and it is the question this desk is built around.

What is in it

Order-book records from the desk's own recorders2.63B: 1.12B from Polymarket over 44 days and 1.51B from Kalshi over 34 days
Data lake657 GB across 44 datasets, every file content-addressed and checked by sha256 or ETag
Experiments filed280+, one directory each, hypothesis committed before the data
Replay344 strategy versions over 3,023 matches, byte-identical on two independent runs
Tests4,500+ test functions across nearly 300 files
Commits1,350+ since 16 August 2026
Live execution1,092 real orders across 171 markets, 10 to 17 September 2026
VenuesPolymarket and Kalshi, each with its own cost model

The layout is deliberate. pmx/ holds importable library code with no import-time side effects, scripts/ holds thin entry points that parse arguments and make one call inward, and live/ is the only directory in the repository allowed to hold an API key or place an order. A test asserts that layout, so structural drift fails the build automatically instead of relying on review.

Six rules the desk runs on

Each rule came out of real measurement work. I keep the reason attached to the rule, because a rule with its reason survives much longer than a rule on its own.

Population before percentile. Every number is stated with its population and its n in the same sentence. This one rule made the spread measurement exact: by pinning down precisely which books were being measured, and checking what each instrument was quantised by, the desk got a spread figure it could trust rather than one that only looked tidy.

Condition on liquidity, never pool. Split by how much liquidity a book actually carries, the same population reads 34.5c in thin books, 6.0c in the middle band and 2.0c in thick ones, on 26, 28 and 15 resolved fixtures. A single pooled figure would have described none of them. Every cost model on the desk is conditioned on the band it applies to.

Ask what a stage did not read. A pipeline stage that exits cleanly can still skip rows. The desk checks inputs as well as outputs, and that habit recovered 80,525 backfilled trades into the canonical build and moved one key join rate from 0.0117 to 0.2321 over the same 74 markets.

Report the statistical edge and the net edge together. Predictability and economics are measured side by side, with the fee and the spread charged in every result. Polymarket and Kalshi get separate fee schedules, queue assumptions and book bindings, and every fee is read from one module rather than typed into a strategy.

Report both weightings. Event-weighted and notional-weighted results are always published together, so no single view can dominate a conclusion.

Correct in place. When a figure is refined, the earlier figure stays visible next to the new one with the reason. The repository reads as a complete, dated record of the research, which is what makes it auditable by someone else.

The promotion gate

A result is promoted only when it clears a fixed bar, declared before the run:

  • The hypothesis, the primary metric and the stopping rules are committed to git before the data is touched.
  • Confidence intervals are clustered by market, and folds are purged and embargoed in time.
  • Multiplicity is corrected across the whole family of tests, not per test.
  • Placebos and negative controls run on the same opportunities as the strategy.
  • A leakage alarm checks every feature for point-in-time validity.

Every experiment on the desk runs under this bar. A 14,222-market, 31-day Kalshi holdout is kept sealed for a single final confirmation, so whatever the desk promotes will have been tested on data nobody has touched.

Live execution

In September 2026 the desk moved from simulation to real execution on Polymarket. From 10 to 17 September 2026, a dedicated live lane placed 1,092 entry orders across 171 markets, every one journalled with its full lifecycle.

That gave the research something no simulation can: observed queue behaviour, real time-to-fill, and real guard behaviour on a production venue. The engineering held up well. In the first readout, measured with 95% market-clustered intervals at 2,000 resamples, infrastructure cancels came in at 0.42% of placed orders, far inside the 10% threshold registered in advance, and the safety guards pulled orders from stale or off-grid books in well under a second, exactly as designed.

The lane is built around operator control: a supervised scheduler, a stop file that only an operator can clear, a flat-position certifier, and a reconciliation tool that checks the desk's own records against the venue. Those observed fills now feed directly into the next generation of the quoting policy.

The plumbing underneath

  • First-party WebSocket recorders capture the order book as it happens. Polymarket recording runs on a dedicated Linux server, and the Kalshi capture holds 1.51B records over 34 days. A stream that was captured can be studied forever, so the recorder is treated as the most important process on the machine.
  • Raw capture is immutable. Canonical facts are built from it with sources kept separate and admitted explicitly, derived datasets are hash-pinned, and every result keeps pointing at the exact version it was computed on. Every file in the lake is content-addressed and checked by sha256 or ETag, and 144.8 GB of event logs are stored as 7.1 GB, each file kept only after it decompressed back to its original sha256.
  • A clock measured from the market. A trade has to hit the price that was resting just before it, so the order book dates the trade tape. Measured that way across 14.5M prints in 25,765 markets, the desk found the one day, 7 May 2026, when Polymarket's tape-to-book offset shifted.
  • Deterministic replay. The replay engine runs 344 strategy versions over 3,023 matches, and two independent runs produce the same hash on every one of them.
  • Admission gates. New sources land in quarantine and are admitted only after passing identity, accounting, decoding and clock-order checks.
  • Population gates. A rebuilt fixture classifier recovered 969 real two-competitor fixtures carrying $25.7M of volume, and now proves 1,260 of 1,260 events across 3,187 raw snapshots.
  • Performance budgets are asserted on the machine that owns them, with a benchmark command that reports every budget as met or unverified.
  • Rules for the AI agents. A pre-commit hook keeps each agent's commits to its own work and refuses kills aimed at the recorder before they run. How that works.
  • Continuous integration runs the repository law, a correctness lint and the full offline suite on Windows and Linux on every push.

What I take from it

Doing research on markets well is mostly about building machinery you can trust: a layout test, an immutable data directory, a pre-registered hypothesis, a bootstrap clustered the right way and a clear record of every measurement.

Those same instincts are what pricing, trading, risk and fraud teams need every day. Know the population behind every number, check what a pipeline did not read, and make a decision only on evidence that has been attacked properly.

I built the desk as one engineer directing AI coding agents, under repo rules that run as tests. Strategy specifications, signal definitions and quoting recipes are deliberately not published here.


Related: 92¢ to zero in 25 minutes · How I run AI coding agents on one repo · The gate that keeps capital safe · 732 trades that were really 37 matches · Build evidence, reproducible from git