All roles
Careers

Senior Data Scientist

Own the models that turn a property into a number, and the evidence that the number holds.

Team
Modelling
Type
Full time
Location
Remote, CET hours
Level
Senior
Stack
PythonXGBoostscikit-learnpandasPostgreSQLPostGIS
Apply

About the role

We are a modelling company with an application in front of it. Almost everything a customer reads on a screen is a prediction: a value, a rent, a rate, a probability, a grade. Someone has to be accountable for whether those predictions survive contact with reality, and that is this role.

You will not be handed a clean dataset and a metric to move. You will decide what the target should be, whether the data can answer it honestly, and what we should refuse to predict.

What you will do

  • Build, train and retrain supervised models on large tabular data with a heavy geospatial component. Gradient boosted trees are the workhorse (XGBoost), with regularised linear and GLM baselines kept alive because they win more often than people expect.
  • Own feature engineering end to end: location, structure, condition, time, and the joins across public registers that make those features possible at all.
  • Design the validation, not just run it. Hold out by geography and by period, never at random, because a random split on this data quietly leaks the answer and flatters the model.
  • Calibrate. A number we put in front of a paying customer needs an interval around it, and the interval has to mean what it says.
  • Watch for drift and decide when a model is retired. Retiring your own model is part of the job.
  • Write the model card: what it was trained on, where it is weak, who should not use it and for what.
  • Turn a modelling result into a decision a non-modeller can act on without being misled.

What we look for

  • Three or more years building supervised models on tabular data that other people then depended on in production.
  • Fluent Python: pandas, numpy, scikit-learn, and a gradient boosting library you know the failure modes of.
  • SQL you write yourself against a large database, including the query plan when it is slow.
  • You can explain why your model is wrong before somebody else finds out, and you volunteer it.
  • Comfort with genuinely messy input: missing fields, duplicate records, noisy labels, and sources that disagree.

Nice to have

  • Geospatial modelling: PostGIS, H3 or similar indexing, raster data.
  • Hedonic pricing, index construction, or mix adjustment.
  • Actuarial, credit risk, or another domain where a wrong number costs money directly.
  • Having shipped a model behind a live API rather than leaving it in a notebook.

Apply

Read by a person, not by a filter.

RoleSenior Data Scientist
GitHub, LinkedIn, a portfolio, anything you would rather we read than a CV.
CV optional
PDF, DOC, DOCX, ODT, RTF or TXT. Up to 5 MB.