Practice
Reading a calibration you did not run
A short field guide for decision-makers: five questions that separate a model fitted to evidence from one fitted to the past.
AlphaIQ · 30 June 2026 · 5 min read
Senior decision-makers are routinely asked to act on models they did not build and cannot inspect line by line. The realistic goal is not to audit the code. It is to ask a small number of questions whose answers reveal whether the work underneath is sound.
One: what was held fixed, and why?
Every calibration fixes some parameters and fits others. The fixed ones carry assumptions that never appear in a goodness-of-fit statistic. Ask which were fixed and on what basis. If the answer is a literature value, ask whether the population in that literature resembles the one being modelled here.
Two: how wide is the acceptable region?
A single best-fitting parameter set tells you almost nothing on its own. What matters is the shape of the region that fits acceptably well. A narrow region means the data is informative. A wide one means many different worlds are consistent with what has been observed — which is worth knowing before you treat any single one as the answer.
Three: was it tested out of sample?
Fitting a flexible model to a historical period is not difficult, and a good in-sample fit is weak evidence. The question is whether the model reproduces behaviour in a period, a regime or a population it was not fitted on. If the honest answer is that no such test was possible, that is acceptable — but it should change how the result is used.
Four: what would falsify this?
Ask what observation would have made the modelling team abandon the conclusion. A specific answer indicates the conclusion has content. A vague one suggests the model has been built in a way that accommodates most outcomes, which makes it a description rather than a prediction.
Five: where does it break?
Ask for the conditions under which the conclusion reverses, and judge for yourself whether those conditions are plausible. This is the single most useful question on the list, because it converts a number into a boundary — and a boundary is the thing you can actually manage against.
You are not trying to verify the model. You are trying to establish whether the people who built it know where it stops working.

