How I Run AI Coding Agents on One Repo
The rules that let several agents share a research codebase and a live recorder
I started the research desk's codebase on 16 August 2026, and AI coding agents wrote most of it. As of 3 October 2026 it has 1,350+ commits and about 620,000 lines of Python, and most of those commits carry an AI co-author line. I decide what gets built, set the rules and own every number the desk publishes.
The engineering I am proudest of is the set of rules that lets several agents work in one repository, on one machine, next to a market-data recorder that must keep running, without stepping on each other.
The repo has a constitution
A hook checks every command an agent is about to run, and the test suite checks the code it leaves behind.
Commits stay in their lane. The hook refuses git add -A, git add . and git commit -a, the commands that take the whole working tree instead of the files named. Another agent's half-finished edits never ride along in someone else's commit.
The recorder is protected. The hook refuses any kill command aimed at the market-data recorder. A live stream that is not captured is gone for good, so the recorder is treated as the most important process on the machine.
Raw data is read-only. Redirects, moves, deletes and truncation aimed at the raw capture or the quarantine area are refused before they run. Everything downstream is rebuilt from the raw layer, so it never changes.
The machine is shared fairly. Jobs that would take every core are refused, and the tests fail the build on n_jobs=-1, so two agents' parallel jobs cannot squeeze the recorder between them.
Corrections stay in place. When a figure is refined, the earlier figure stays visible next to the new one with the reason. The docs carry 341 dated corrections.
The layout is law. pmx/ holds library code, scripts/ holds thin entry points, and live/ is the only directory allowed to hold an API key or place an order. A test asserts it, and continuous integration runs that test, a correctness lint and the full offline suite on Windows and Linux on every push.
What the rules buy
| Tests | 4,500+ test functions across nearly 300 files |
| Replay | 344 strategy versions over 3,023 matches, byte-identical on two independent runs |
| Data lake | 657 GB, every file content-addressed and checked by sha256 or ETag |
| Archive | 144.8 GB of event logs stored as 7.1 GB, each file kept only after it decompressed back to its original sha256 |
Agents write code quickly. The rules are what make the output something I can stand behind: every change runs against the same tests, the data underneath cannot be edited, and a result has to replay to the same bytes before it counts.
What I take from it
Working this way is closer to running a small team than to typing code. Most of the job is writing down what good looks like in a form a machine can check, and keeping the record honest when the work moves fast. That habit carries over to any team that ships with AI agents.
Related: The research desk · 92¢ to zero in 25 minutes · Build evidence