RAG That Cannot Fake a Citation
A technical case study on a retrieval engine I built and ran: hybrid retrieval over PostgreSQL/pgvector, span-exact and version-aware citation validation, an answer router where every lane declares itself, red-team gates in CI, and latency/cost budgets written as code.
Status: this describes the retrieval engine from an earlier generation of 21Core, a multi-tenant decision workspace. The same discipline, making a system that cannot claim more than its evidence supports, now powers the provenance tiers behind 21Core's prediction-market data. Snippets are simplified and sanitized. Figures are from that platform's own eval artifacts.
The premise
The engine answered questions over a customer's own documents and data. A wrong answer could reach a reader inside a document that looked finished, so it followed a narrow rule: an answer may be uncited, but it may never look cited when it isn't. A fabricated citation is especially dangerous because it can survive a quick review.
So the engine was designed around a contract, not a vibe:
Every claim is either backed by a citation that provably exists in tenant evidence, computed deterministically from approved data, or explicitly labeled as ungrounded.
The contract also changed by surface. In chat, a general lane could exist if it labelled itself clearly. In an exported document or an unattended review job, that lane was removed. Unsupported claims stayed out of the output.
Retrieval: hybrid by default
Single-leg vector search fails in predictable ways: exact references and document identifiers embed poorly, and long documents drown their own sections. The pipeline runs multiple recall legs and fuses them:
Question
-> embed query (local ONNX, 384-dim)
-> vector leg (pgvector HNSW)
-> keyword leg (GIN full-text)
-> exact-token leg (identifiers, codes)
-> trigram leg (fuzzy recall)
-> reciprocal rank fusion
-> rerank
-> parent-child section expansion
-> citation validation
-> sanitized retrieval trace
Two decisions matter more than the diagram:
- Embeddings run locally (ONNX). Tenant text never leaves the box to be embedded. That removes an external processor from the document path.
- Parent-child section routing. Large documents are chunked with parent sections; retrieval can land on a child and expand to its parent, so answers get coherent context instead of orphaned fragments.
The grounding contract
Retrieval quality is necessary but not sufficient; the generation side has to be held to the evidence:
- Citation validation. Finding the quote somewhere is not enough. A citation is rejected before it is stored or displayed unless the chunk was retrieved for this question, the tenant matches, the document and version match, the version is active, the excerpt appears in the chunk, and the stored character offsets resolve to that excerpt. The version check rejects a citation from a source that was valid last quarter but has since been replaced. The offset check catches a model quoting real words from a real document while pointing at the wrong place.
- Insufficient-evidence refusal. When a lane has no evidence to work with, the system says so, as a designed outcome with its own UI state, not as an apology. In a money domain, "I don't have evidence for that" is a feature.
- Sanitized traces. Every retrieval stores a trace (candidate chunks, latency, cache state, routing decisions) without storing raw prompt text, observability and privacy in the same schema.
The answer router
Not every question should touch the RAG path at all. A router classifies each question into a lane, and the lane is part of the answer's contract:
| Lane | Use case | Behaviour |
|---|---|---|
| Cited evidence | Grounded in tenant/approved sources | Citations required, refusal on weak evidence |
| Deterministic calculation | Price per m², time on market, document totals | Computed from approved rows, never generated |
| Search-only | User wants sources, not prose | No generation at all |
| Deep review | Bounded scoring/review jobs | Async worker, cited output |
| General | Non-evidence context | Explicitly labeled unverified |
| Web research | Current public info | Explicit public-web route |
Numbers come from a calculation engine over approved data, so the LLM can explain a figure but cannot invent it. Zero LLM-generated document figures is a checkable property.
Evals before routing decisions
Model choices were made by measurement, not fashion:
- Embedding bakeoff (4 models): recall@1 ranged 0.00 to 0.89 across candidates on the platform's own retrieval cases, with an explicit small-sample caveat recorded next to the result, because an honest eval names its own weakness.
- LLM route bakeoff (16 live cases, real providers): the adopted route cut mean latency 61% and estimated cost 84% versus the incumbent. The same artifact records every citation check alongside it.
Red-team gates in CI
The failure modes of LLM apps are known, so they're tested like any other regression. A safe-slice suite runs gates for:
- Prompt injection (instructions embedded in documents)
- Fabricated citations
- Wrong-tenant retrieval
- Unsupported questions (must refuse, not improvise)
- Metric hallucination (numbers outside the deterministic lane)
- Public-claim safety (no overclaiming in outward-facing text)
- Private-data blocking (the fail-closed gate stays closed)
A change that breaks a gate doesn't ship. That's the whole point of writing safety as tests instead of policy documents.
Latency and cost as code
Two production disciplines round out the engine:
- Per-path p75 / p95 response budgets are written as code-level contracts and checked, so a retrieval regression is a failing check, not a slow week.
- A spend firewall wraps every model call: tenant credit reservation before the call, provider-cent ceilings, route envelopes, local fallbacks, settlement after, refund on failure. The model can be down or expensive; the accounting stays correct.
What this demonstrates
- RAG engineering past the demo: hybrid retrieval, reranking, section routing, caching, on boring, operable infrastructure (PostgreSQL).
- Safety as engineering: grounding, refusal, and red-team gates as testable properties in CI.
- Product judgment: deterministic calculations, local embeddings for privacy, evals with recorded caveats and outputs an agency can inspect.
Related: How I built 21Core around prediction markets · 732 trades that were really 37 matches: the same checks applied to trading research