Catching Silent Data Corruption
A case study in data-integrity engineering: how an unattended research platform caught a silent data fault that every health check missed, and the structural guards that now prevent that whole class of problem, on a system where numbers drive money-shaped decisions.
Status: All figures below are from the platform's own telemetry and audit artifacts. Data-source names are withheld.
The setting
A solo-built esports research platform that runs itself: 15 supervised services and 49 distinct scheduled jobs moving data through ingestion, prediction, settlement and publishing, around the clock, with no daily manual operations. Over a recent 30-day window the pipeline logged 68,874 job executions, each writing one JSONL telemetry row, and ran a 22-day stretch with zero restarts.
At that scale of automation, crashes are easy, because crashes are loud. The hard problems are the quiet ones: data that keeps arriving and keeps parsing after it has stopped being current. Catching those is what this case study is about.
The catch
An upstream source for one odds-archiving job went quiet, and the archiver kept writing the same snapshot under fresh, dated filenames for nineteen consecutive days. Every file parsed, every job reported success, and the pipeline stayed green.
I caught it before any study used it, by checksumming archives against each other instead of trusting their timestamps. Nineteen files, one hash. That check protected a planned cross-source timing study that would otherwise have measured the gap as a real market effect.
The structural fix
Two rules, now enforced in code rather than vigilance:
- Stale-source refusal. If a source's newest row is older than its freshness budget, the archiver refuses to archive (hard exit, recorded in a health file) rather than laundering old data under a new date. A missing file is an honest signal; a fresh-looking stale file is a lie.
- Provenance manifest sidecars. Every archived file now ships with a machine-readable manifest recording the true data-timestamp range of its contents, so no consumer ever trusts a filename. What a file is called and what a file contains are now independently verifiable claims.
The guard is live. This is a real refusal record from the archive's health manifest, verbatim except the source name:
{
"source": "[withheld]",
"checked_at": "2026-07-12T23:50:01Z",
"rows": 22359,
"data_min_ts": "2026-05-12T03:09:36Z",
"data_max_ts": "2026-06-20T01:20:02Z",
"age_hours": 550.5,
"status": "stale",
"reason": "newest row is 550.5h old (> 24.0h); scraper is dead, refusing to archive a frozen snapshot under today's date"
}
The rest of the integrity stack
That catch hardened one pipeline; the same philosophy runs through the whole platform:
- Append-only ledgers with lockfile discipline, all money-state lives in append-only JSONL with sidecar locks; nothing edits history.
- A typed write gate on the money path, direct writes to the prediction log are hard-blocked; every record passes schema validation to get in.
- Nightly reconciliation of a ~2.2M-row paper-trade ledger and a 134k-row prediction log. When rows went missing, they were rebuilt from git history with full provenance, and every reconstructed record is labeled as reconstructed, never silently backfilled.
- Scheduler-truth auditing, an automated job reconciles the live crontab against the job registry and flags drift, so a job can't silently stop existing.
- Five independent watchdog layers on different cadences (health, orphaned settlements, pipeline liveness, circuit-breaker state, feed health), plus an integrity-event ledger so every catch becomes institutional memory.
- Duplicate-counting defenses, born from the audit that found 732 "settled trades" were really 37 matches counted ~19.8× (that story here); uniqueness is now asserted, not assumed.
The operating principle
Every pipeline assumes two things: the source is messy, and the consumer is downstream of money.
From those assumptions the rules follow: validate at the boundary, quarantine what fails, reconcile what matters nightly, record provenance for everything, and surface problems visibly, because a system that says "this data is not fresh" is one people can trust.
Boring, disciplined data engineering is what makes the interesting work (models, pricing, research) mean anything at all.
Related: The research desk · mm-trader: the market-making research platform · 732 trades that were really 37 matches · Scoring 75,000 profiles a session: the same discipline against adversarial input