About the role
We are a modelling company with an application in front of it. Almost everything a customer reads on a screen is a prediction: a value, a rent, a rate, a probability, a grade. Someone has to be accountable for whether those predictions survive contact with reality, and that is this role.
You will not be handed a clean dataset and a metric to move. You will decide what the target should be, whether the data can answer it honestly, and what we should refuse to predict.
What you will do
- Build, train and retrain supervised models on large tabular data with a heavy geospatial component. Gradient boosted trees are the workhorse (XGBoost), with regularised linear and GLM baselines kept alive because they win more often than people expect.
- Own feature engineering end to end: location, structure, condition, time, and the joins across public registers that make those features possible at all.
- Design the validation, not just run it. Hold out by geography and by period, never at random, because a random split on this data quietly leaks the answer and flatters the model.
- Calibrate. A number we put in front of a paying customer needs an interval around it, and the interval has to mean what it says.
- Watch for drift and decide when a model is retired. Retiring your own model is part of the job.
- Write the model card: what it was trained on, where it is weak, who should not use it and for what.
- Turn a modelling result into a decision a non-modeller can act on without being misled.
What we look for
- Three or more years building supervised models on tabular data that other people then depended on in production.
- Fluent Python: pandas, numpy, scikit-learn, and a gradient boosting library you know the failure modes of.
- SQL you write yourself against a large database, including the query plan when it is slow.
- You can explain why your model is wrong before somebody else finds out, and you volunteer it.
- Comfort with genuinely messy input: missing fields, duplicate records, noisy labels, and sources that disagree.
Nice to have
- Geospatial modelling: PostGIS, H3 or similar indexing, raster data.
- Hedonic pricing, index construction, or mix adjustment.
- Actuarial, credit risk, or another domain where a wrong number costs money directly.
- Having shipped a model behind a live API rather than leaving it in a notebook.






