---
title: "What I Chose Not to Build, Method Selection"
description: "A method-selection case study for a real-estate intelligence product: gradient boosting over neural networks on tabular data, spatially weighted conformal prediction over block bootstrap, a discrete-time hazard model over a hand-weighted score, uplift over propensity, and learned visual embeddings over perceptual hashing for entity resolution. Roughly fifty sources, three of my own decisions overturned, and four things the evidence said not to build at all."
url: https://www.afonsomartins.com/method
author: Afonso Martins
published: 2026-08-23
updated: 2026-10-03
---

# What I Chose Not to Build

A method-selection case study: how the estimator, the interval, the ranking model and the entity-resolution layer inside 21Core were chosen from the literature rather than from habit, and the four things the evidence said not to build at all.

**Status:** This is a decision record from an earlier generation of 21Core, a real-estate intelligence product. The method discipline in it carried straight into [21Core's prediction-market work](https://www.afonsomartins.com/how-i-built-21core). Every item below is labelled *built*, *designed* or *not started*, and the status table at the end is the authoritative version. The market figures are from INE, idealista and ECO for 2026; the method findings are from the papers named beside them.

## The premise

The expensive decisions in a machine-learning product are rarely about what to build. They are about what not to build, and they get made early, quietly, by whoever is most enthusiastic in the room.

I spent August 2026 doing the opposite: reading roughly fifty papers and benchmarks across valuation, liquidity, seller behaviour, agency incentives, targeting and deep learning, and writing down which of my own design decisions the evidence contradicted. Three of them it did. This is that record.

Reading a NeurIPS benchmark costs an afternoon. Discovering the same result after shipping a graph neural network can cost a quarter.

## The market that set the constraints

Portugal in 2026 is an unusual market, and the shape of it decides what the product has to be.

| Indicator | 2026 | Source |
|---|---|---|
| Median price per m², Q1 | 2,337 € · **+17.8%** | INE |
| Transactions, Q1 | 37,745 · **−8.7%** | INE |
| Stock for sale, Q1 then Q2 | −14%, then −6% | idealista |
| Agents in the five largest networks | 24,825 · **+8.5%** | ECO |
| Transactions per agent per year | **5.2** · −1.6% | ECO |
| Listings sitting unsold past 3 months | **41%** | idealista |
| Listings that have ever cut price | **8%** | idealista |

Read those last two together. Forty-one per cent of Portuguese listings are stuck and eight per cent have adjusted. Owners are not repricing. They are waiting.

And the agent cannot tell them to reprice, because of a result that is thirty years old and still the most useful thing in the literature on this trade. Levitt and Syverson compared the houses agents sell for clients with the houses they sell for themselves: **their own sell for 3.7% more and stay on the market 9.5 days longer**, and the distortion is largest where the agent's informational advantage is largest. The owner does not have to have read the paper to have internalised it. They know the agent earns more from a fast cheap sale than a slow expensive one, so the agent's opinion about price is the one opinion in the room that cannot be trusted.

The problem is credibility. The product makes the price conversation possible without asking the owner to take the agent's opinion on trust, and every method decision below follows from that.

## Decision 1: Trees, not neural networks

The tempting build is a neural network over property attributes. The benchmarks say no, and they say it repeatedly.

| Work | Finding |
|---|---|
| Grinsztajn et al., NeurIPS 2022 | Tree ensembles still outperform deep learning on tabular data, with the structural reasons identified |
| Shwartz-Ziv & Armon, *Information Fusion* | Deep learning is not all you need |
| Graph neural networks for house prices, IJDSA | Benchmarked on six housing datasets: GNNs do not beat the baselines. Explicit recommendation for LightGBM and CatBoost |

What the evidence *does* support is abandoning the pure hedonic model, and the margins there are uncomfortable:

| Comparison | Hedonic | Machine learning |
|---|---|---|
| Zaki et al., *Concurrency and Computation* | 42% accuracy | 84.1% with XGBoost |
| Neural vs hedonic, PRRES | MAPE 38.23% | MAPE 15.94% |

**The call:** LightGBM with SHAP. Tree ensembles keep the higher accuracy while SHAP provides enough explanation for the document. The hedonic model remains as an explanation layer that produces the implicit price per attribute. The gradient booster produces the estimate.

## Decision 2: Where deep learning does earn its place

Deep learning loses on the tabular columns and wins on the modalities the table cannot see. A multimodal study (arXiv 2409.05335) finds that text embeddings of the description and image embeddings of the photographs improve accuracy significantly beyond raw attributes and geospatial embedding. Satellite imagery alone distinguishes the top 15% of properties from the bottom 15% at 91% accuracy. And there is a lovely small finding in the *Journal of Real Estate Finance and Economics*: price rises at a decreasing rate with the number of photographs, while time on market rises with interior photos and is unrelated to exterior ones.

So: SBERT for the text, CLIP for the images, concatenated with the raw attributes, a tree on top.

The highest-value application of deep learning in this system is entity resolution: recognising that the property listed by agency A in March and the property listed by agency B in September are the same property. My architecture notes estimate that this is eighty per cent of the engineering work. The method I had chosen was perceptual hashing.

Perceptual hashing is fragile in exactly the places this product needs robustness. It breaks under cropping, under a new agency watermark, under colour correction, and under re-photography. Those are not edge cases. **They are the precise signature of a property changing agency, which is the single highest-value event the system can detect.** I had chosen a method that fails hardest on the signal worth the most.

**The call:** a learned visual embedding, CLIP or DINOv2, indexed for approximate nearest-neighbour search, with the perceptual hash retained as a cheap first-pass filter. An embedding recognises the same living room shot from another angle, in different light, under a different watermark. A hash does not.

## Decision 3: Conformal prediction, not bootstrap

The system already does something most valuation products do not: it refuses to output a point estimate without an interval, computed by block bootstrap. The intent was right. The method was not quite.

A block bootstrap gives you an interval. It does not give you a **coverage guarantee**. The literature shows that non-conformal quantile-regression intervals undercover badly, while conformalised quantile regression achieves exact coverage in finite samples. There is a housing-specific wrinkle too: because prices are spatially dependent, applying conformal prediction naively produces sets that are not calibrated everywhere. The fix is to calibrate against a locally weighted version of the non-conformity scores (arXiv 2312.06531).

**What that buys is one sentence.** Instead of *"here is an interval"*, the document can say *"here is a 90% interval, and the measured empirical coverage of these intervals in this region last quarter was 89.4%"*. The first sentence is a claim. The second is a result, in finite samples, that anyone can check next quarter.

No competitor in this market publishes the empirical coverage of its own interval. It costs one table.

## Decision 4: A hazard model, not a weighted score

My own scoring engine ranks properties on five weighted signals that sum to 100 by construction: withdrawn without a sale, change of agency, repeated price revisions, age on market, relaunch at an unchanged price. It is a reasonable design and the weights have never been calibrated against a real outcome.

Which means the number it emits is an **ordering, not a probability**. "82 points" is not comparable between one parish and another, cannot be validated against anything, and cannot be wrong in any way you could detect. My own design system already carries the rule this violates: do not present a hand-weighted priority score as if it were a calibrated probability.

**The call:** a discrete-time hazard model over the archive, unit of observation the property-month, response *listed within the next 90 days*. The five signals stop being weights and become features. The agent stops reading "82 points" and starts reading "38% probability this owner lists in the next 90 days". The second sentence can be checked against what happens.

The migration is deliberately slow: run the deterministic score in parallel for a full season and only switch when the calibration holds.

## Decision 5: Uplift, not propensity

This one contradicts the design most directly. The ranking engine orders by the probability that an owner wants to sell. That is propensity, and it is the wrong question.

The benchmark work on uplift modelling and heterogeneous treatment effects (Rößler & Schoder, *Journal of Interactive Marketing*, 2022) finds consistently that targeting by uplift beats targeting by ordinary supervised learning. Uplift estimates whose behaviour the contact changes. An owner who was going to list anyway should never receive the letter because that spends budget on an outcome that was already going to happen.

Uplift cannot be estimated without randomisation. There is no shortcut. The minimum viable design is to hold back ten per cent of each batch at random and not contact them, and after two or three quarters there is enough data for a causal forest.

That ten per cent is a real cost and has to be sold to the client honestly, which turns out to be easier than it sounds: *we hold back one opportunity in ten so we can prove to you what the other nine were worth.* It also upgrades the monthly report from "three listings appeared on your site after we flagged them", which is observational, to "three listings, against 0.4 expected in the control group", which is an effect.

## Decision 6: Publish the calibration

Every document the system produces is a dated prediction. When the property leaves the market, the archive observes the outcome and the prediction gets scored. After a quarter that produces a reliability chart: *we said 60% to 340 properties, 203 of them sold, so 59.7%.*

Alongside it, the IAAO standard on automated valuation models already defines the metrics this class of product should be judged by: COD for variability, PRD and PRB for regressivity, and FSD for projected precision. The standard publishes tolerances. Reporting them by quarter and municipality costs a table.

**This is the strongest competitive move available and it is irreversible.** No competitor in this market publishes a calibration chart. The moment one does, the others inherit the burden of explaining why they do not. But once you start you publish every quarter, including the bad ones, and a broken series is worth less than never having started. So: two quarters measured internally first, and every published figure carries a confidence interval on the calibration itself.

## Decision 7: An explicit "I don't know"

The Harvard Business School field experiment with Boston Consulting Group, published in *Organization Science*, ran 758 consultants through tasks with and without AI assistance.

| Condition | Result |
|---|---|
| Tasks inside the model's capability frontier | +12.2% tasks completed · +30% quality · −25% time |
| One complex task outside it | **−19 percentage points** on correctness |
| Where the frontier sits | Not visible to the user |

The paper calls that invisible boundary the **jagged frontier**. Capability does not fall off smoothly at the edge of what the
model can do; it is ragged, and two tasks that look equally hard to a person can sit
on opposite sides of it. The tool is enormously helpful right up until it is silently harmful, and the person using it cannot see the line.

The system already has explicit *insufficient* bands and a coverage report that refuses to score when a parish has fewer than eight comparable histories. I had been treating those as engineering hygiene. They are not. **They are the only mitigation with controlled-experiment evidence behind them for the one failure mode that makes this category of product actively worse than no product at all**, and they belong on the front of the document, visible to the customer, not buried in a config file.

## The four I refused

| Refused | Why |
|---|---|
| Neural network for price from tabular features | NeurIPS 2022 and *Information Fusion*: trees win, and the reasons are structural |
| Graph neural network for price | Benchmarked on six housing datasets; does not beat LightGBM or CatBoost |
| Our own foundation model, trained from scratch | Neither the data nor the reason exists. It is where startups in this category die |
| iBuying, in any form | Two papers on adverse selection: owners accept the overpriced offers. And iBuyers deliberately operate where markets are liquid and valuation error is low, which is the opposite of this region |

The last one has a corollary worth keeping: **knowing where your valuation error is large is worth as much as the valuation itself.**

## What is built, what is designed, what is not started

| Component | Status |
|---|---|
| Multi-tenant platform, 133 API operations, 607 tests, row-level security | **Built** |
| Deterministic scoring engine, five signals, every point traceable to its event | **Built**, and uncalibrated by admission |
| Minimum-sample floor and explicit *insufficient* bands | **Built** |
| Deterministic tax and duty engine, 51 tests | **Built** |
| Mandatory interval on every estimate | **Built**, by block bootstrap |
| Learned visual embeddings for entity resolution | **Designed**, hash still in place |
| Spatially weighted conformal prediction | **Designed** |
| Discrete-time hazard model and probability-by-price curve | **Designed** |
| Uplift targeting and the 10% holdout | **Designed** |
| Published calibration chart and IAAO metrics | **Not started**, two internal quarters required first |
| Licensed longitudinal price archive | **Not secured** |

## What this cost, and what it bought

Three weeks of reading, and three of my own decisions overturned before any of them reached a customer: a score presented as if it were a probability, an interval without a coverage guarantee, and an entity-resolution method that failed hardest on the most valuable event in the system.

None of the three would have announced itself. The score would have kept emitting confident numbers. The interval would have kept looking like an interval. The hash would have kept matching the easy cases and quietly missing the agency changes, which is to say quietly missing the product.

That is the argument for doing this at all. The failures worth engineering against are never the loud ones.

---

*Related: [How I built 21Core around prediction markets](https://www.afonsomartins.com/how-i-built-21core) · [RAG that cannot fake a citation](https://www.afonsomartins.com/citation-first-rag) · [732 trades that were really 37 matches](https://www.afonsomartins.com/the-audit)*
