Forecast Evaluation Scorecard
A versioned framework for measuring financial forecasts against realized outcomes without hindsight or selective reporting.
What this resource does
A forecast is useful only when its target, horizon, information cut-off, version, and evaluation rule are fixed before the outcome. Retrospective narrative is not a substitute for measurement.
The scorecard supports point forecasts, ranges, probabilities, and directional classifications. It records every eligible forecast, including misses and unavailable outcomes.
Methodology
- Freeze target definition, horizon, data cut-off, model version, and publication timestamp.
- Prevent revised data or later information from entering the original forecast record.
- Choose error, coverage, calibration, and baseline comparisons appropriate to the output type.
- Report results by period, regime, universe, and confidence bucket with sample sizes.
How to interpret it
A lower average error may coexist with poor tail behavior. Range coverage can be high because ranges are too wide. Directional accuracy can look strong in a one-direction market.
Compare against simple baselines such as no change, historical average, or last observation. Complexity is justified only when it improves a pre-declared objective.
Limitations and failure modes
- Small samples produce unstable conclusions.
- Survivorship and look-ahead bias can invalidate results.
- Regime changes reduce the relevance of older observations.
- Past model performance does not guarantee future accuracy.
Research workflow
- Register the forecast before outcome time.
- Lock inputs and version.
- Score on schedule.
- Publish complete results and model changes.
Questions and answers
Why compare with a simple baseline?
A sophisticated model can appear accurate while adding no value beyond persistence or the historical average. The baseline exposes that.
Should failed forecasts be removed?
No. Removing misses creates selection bias. Corrections should remain linked to the original version and reason.
