Method
What a simulation owes a reviewer
A model that cannot show its calibration, its sensitivities and its failure modes is a story with numbers in it. Three artifacts that should exist before any result is presented.
AlphaIQ · 4 August 2026 · 6 min read
There is a moment in every modelling engagement where the model starts producing plausible output, and the temptation is to present it. The output is legible. It moves in the right direction. It agrees, roughly, with intuition. This is the most dangerous point in the entire exercise, because a plausible result is indistinguishable from a correct one until someone tries to break it.
The discipline that separates the two is not more sophistication in the model. It is the existence of three artifacts, produced before the result is shown to anyone.
One: the specification, written down
Every model embeds interpretation. A paper says agents are risk-averse; the code says something specific about a utility function and a parameter range. That translation is a decision, and if it is not written down separately from the code, it cannot be reviewed — a reader is left inferring intent from implementation, which is the least reliable form of review available.
A specification is not documentation written afterwards. It is the object the code is generated from: agent types, state variables, behavioural rules, parameters, and the intended meaning of each. When a reviewer disagrees, they should be able to disagree with a line in the specification rather than a line in a loop.
Two: the calibration record
Calibration is where the model meets evidence, and it is routinely reported as a single sentence — that parameters were fitted to data. That sentence conceals every question worth asking. Which data, over which period. What the objective was. Which parameters were free and which were fixed by assumption. How wide the region of parameter space that fits acceptably well turned out to be.
- The target series, its source, and the period used
- The distance measure, and why it was chosen over the alternatives
- The accepted parameter region, not just the best-fitting point
- The parameters that were fixed by assumption rather than fitted, and on what grounds
The last of these is the one most often omitted, and the one a good reviewer asks about first. A parameter fixed by assumption is a judgement wearing a number's clothes.
Three: the failure map
Sensitivity analysis is usually presented as reassurance — the result is robust to perturbation. That is the less interesting half. The more useful output is the inverse: the specific combinations of inputs under which the conclusion reverses, and whether those combinations are plausible.
A model that has never been shown to fail has not been tested. It has been demonstrated.
This reframes what a modelling team delivers. Not a number with error bars, but a map: here is the conclusion, here is the region of the world in which it holds, and here is the boundary. Decision-makers are considerably better at using that than at using a point estimate, because it matches how they already think about risk.
The practical objection
All three artifacts take time, and time is the binding constraint. This is true, and it is the actual argument for automating the modelling lifecycle rather than the model itself. If specification, calibration records and sensitivity sweeps are produced as a by-product of how the model is built, the trade-off between rigour and deadline mostly stops being a trade-off.
That is the bet: not that automation makes models smarter, but that it makes the evidence around them affordable.

