Model Validation and Calibration Status
Review what Aiovel has tested, what the current evidence does not establish and the standard required for a calibration claim.
Current evidence status
Aiovel has not yet published a prospective, held-out calibration study for the Probability Map. It therefore does not claim that displayed percentages are calibrated forecasts or that the model has demonstrated forecasting skill. The evidence currently supports calculation integrity and transparent descriptive comparison—not predictive validation.
What the model estimates
The Probability Map converts historical price behaviour into scenario estimates for whether selected levels may be touched or exceeded over stated horizons. It does not incorporate every source of market information, and real markets can depart materially from its statistical assumptions through jumps, regime changes, liquidity effects and other discontinuities. See Data Provenance for source categories and scope.
Calculation-integrity evidence
Automated tests cover core probability identities, boundary conditions, incomplete-data handling, time-window alignment and rejection of invalid feeds. Release checks cover probability bounds, internal consistency, expected directional behaviour and public/private content boundaries. These controls are intended to catch calculation and data-contract regressions. They do not measure forecast accuracy.
Historical cross-check—and why it is not calibration
The product compares model output with descriptive historical outcomes when enough usable history is available. Those observations can overlap, share inputs with the model and span different regimes, so the comparison is contextual evidence rather than an independent test. It must not be described as prospective or out-of-sample calibration.
Required prospective protocol
A defensible calibration study would freeze the evaluated model version, record forecasts before outcomes are known, retain unsuccessful and missing cases, predefine labels and evaluation groups, prevent future information from entering the inputs, compare against suitable baselines and report proper scoring rules such as Brier score and log loss with uncertainty. Model development and final evaluation would use separate periods.
Publication threshold
Aiovel will only describe the model as calibrated when a dated scorecard with sample definitions, frozen version, evaluation window, missing-data policy, baselines and uncertainty is public. Until then, percentages should be treated as assumption-dependent scenarios and used alongside source data, options markets where relevant, and independent judgement.