---
title: "Catching Silent Data Corruption, Data-Integrity Engineering Case Study"
description: "How an unattended research platform (15 services, 49 scheduled jobs, 68,874 logged runs in 30 days) caught nineteen days of silently frozen data being re-archived under fresh dates, and the structural fixes: stale-source refusal, provenance manifest sidecars, typed write gates, nightly reconciliation of a 2.2M-row ledger, and visible data-freshness signals."
url: https://www.afonsomartins.com/data-integrity
author: Afonso Martins
published: 2026-07-10
updated: 2026-10-03
---

# Catching Silent Data Corruption

A case study in data-integrity engineering: how an unattended research platform caught a silent data fault that every health check missed, and the structural guards that now prevent that whole class of problem, on a system where numbers drive money-shaped decisions.

**Status:** All figures below are from the platform's own telemetry and audit artifacts. Data-source names are withheld.

## The setting

A solo-built esports research platform that runs itself: **15 supervised services and 49 distinct scheduled jobs** moving data through ingestion, prediction, settlement and publishing, around the clock, with no daily manual operations. Over a recent 30-day window the pipeline logged **68,874 job executions**, each writing one JSONL telemetry row, and ran a 22-day stretch with zero restarts.

At that scale of automation, crashes are easy, because crashes are loud. The hard problems are the quiet ones: data that keeps arriving and keeps parsing after it has stopped being current. Catching those is what this case study is about.

## The catch

An upstream source for one odds-archiving job went quiet, and the archiver kept writing the **same snapshot under fresh, dated filenames for nineteen consecutive days**. Every file parsed, every job reported success, and the pipeline stayed green.

I caught it before any study used it, by checksumming archives against each other instead of trusting their timestamps. Nineteen files, one hash. That check protected a planned cross-source timing study that would otherwise have measured the gap as a real market effect.

## The structural fix

Two rules, now enforced in code rather than vigilance:

1. **Stale-source refusal.** If a source's newest row is older than its freshness budget, the archiver refuses to archive (hard exit, recorded in a health file) rather than laundering old data under a new date. A missing file is an honest signal; a fresh-looking stale file is a lie.
2. **Provenance manifest sidecars.** Every archived file now ships with a machine-readable manifest recording the true data-timestamp range of its *contents*, so **no consumer ever trusts a filename**. What a file is called and what a file contains are now independently verifiable claims.

The guard is live. This is a real refusal record from the archive's health manifest, verbatim except the source name:

```json
{
  "source": "[withheld]",
  "checked_at": "2026-07-12T23:50:01Z",
  "rows": 22359,
  "data_min_ts": "2026-05-12T03:09:36Z",
  "data_max_ts": "2026-06-20T01:20:02Z",
  "age_hours": 550.5,
  "status": "stale",
  "reason": "newest row is 550.5h old (> 24.0h); scraper is dead, refusing to archive a frozen snapshot under today's date"
}
```

## The rest of the integrity stack

That catch hardened one pipeline; the same philosophy runs through the whole platform:

- **Append-only ledgers with lockfile discipline**, all money-state lives in append-only JSONL with sidecar locks; nothing edits history.
- **A typed write gate on the money path**, direct writes to the prediction log are hard-blocked; every record passes schema validation to get in.
- **Nightly reconciliation** of a **~2.2M-row paper-trade ledger** and a 134k-row prediction log. When rows went missing, they were rebuilt from git history with full provenance, and every reconstructed record is *labeled* as reconstructed, never silently backfilled.
- **Scheduler-truth auditing**, an automated job reconciles the live crontab against the job registry and flags drift, so a job can't silently stop existing.
- **Five independent watchdog layers** on different cadences (health, orphaned settlements, pipeline liveness, circuit-breaker state, feed health), plus an integrity-event ledger so every catch becomes institutional memory.
- **Duplicate-counting defenses**, born from the audit that found 732 "settled trades" were really 37 matches counted ~19.8× ([that story here](https://www.afonsomartins.com/the-audit)); uniqueness is now asserted, not assumed.

## The operating principle

Every pipeline assumes two things: the source is messy, and the consumer is downstream of money.

From those assumptions the rules follow: validate at the boundary, quarantine what fails, reconcile what matters nightly, record provenance for everything, and **surface problems visibly**, because a system that says "this data is not fresh" is one people can trust.

Boring, disciplined data engineering is what makes the interesting work (models, pricing, research) mean anything at all.

---

*Related: [The research desk](https://www.afonsomartins.com/research-desk) · [mm-trader: the market-making research platform](https://www.afonsomartins.com/mm-trader) · [732 trades that were really 37 matches](https://www.afonsomartins.com/the-audit) · [Scoring 75,000 profiles a session: the same discipline against adversarial input](https://www.afonsomartins.com/fraud-scoring)*
