Skip to content

Models & validation · 7 min read

Weighted models need chronological evidence, not a winning chart

Build a transparent weighted forecast, then evaluate it with chronological holdouts, leakage controls, calibration and realistic trading costs.

By Stakesift editorial

Published · Updated

A weighted model makes assumptions explicit. It does not make them correct. Its usefulness depends on whether the inputs were available when the decision was made, whether the output has a defensible probability meaning, and whether the evaluation survives unseen data and realistic costs.

Know whether you are combining probabilities or scores

Suppose three forecasts for the same binary outcome are 58%, 54% and 50%. With nonnegative weights 0.50, 0.30 and 0.20 that sum to one, the weighted probability is 0.50 × 0.58 + 0.30 × 0.54 + 0.20 × 0.50 = 0.552, or 55.2%. This is an explicit blending rule, not proof that the event will happen 55.2% of the time.

The forecasts must refer to the same outcome and horizon. Their errors may be highly correlated: three systems using the same market price are not three independent pieces of evidence. Adding more similar forecasts can create an impression of agreement without adding information.

Raw factors are different. Combining a ranking, a percentage and a count directly makes the result depend on their units. If standardized factors are 0.8, −0.4 and 0.5, the same weights produce a score of 0.50 × 0.8 + 0.30 × (−0.4) + 0.20 × 0.5 = 0.38. That is not a 38% probability. Even mapping it through a logistic function gives about 59.4% only under chosen coefficients; the mapping still needs to be fitted and evaluated.

Freeze the information set at the decision time

Each historical row needs a decision timestamp, inputs known by that time, the quote available then, and an eventual outcome. A result can be present in today's database without having been available to yesterday's model. Store observation time as well as the time the underlying event occurred.

  • Do not use closing odds to simulate a decision made hours before the close.
  • Do not use an injury report, lineup or corrected statistic published after the decision timestamp.
  • Compute rolling features using only earlier observations. Exclude the target event's own result.
  • Fit scaling, imputation, feature selection and probability calibration on the permitted training or calibration data, not the whole dataset.
  • Keep repeated rows for the same event together. Do not let one market snapshot train the model while another snapshot of the same eventual result becomes an allegedly unseen test case.

These are leakage controls. A model that accidentally sees its answer can look excellent while being unusable in real time. The scikit-learn guide to data leakage explains why preprocessing must respect train/test separation.

Validate in the direction time actually moves

For an illustrative historical dataset, use January through April to fit the model, May to choose weights and decision thresholds, and June as an untouched final holdout. These are example partitions, not reported Stakesift results. Only include training labels resolved before the next evaluation period; a January contract that settles in June cannot supply a known outcome to a May model.

For a walk-forward evaluation, fit on earlier resolved data, predict the next time block, then advance and refit according to a rule specified in advance. A gap may be needed where outcomes or information windows overlap. For irregular sporting events, choose calendar or event-aware boundaries rather than assuming each row represents the same duration.

The TimeSeriesSplit documentation describes chronological folds and gaps, including its equal-spacing condition for comparable fold durations. It is a methodological reference, not evidence that a generic splitter automatically handles every event dataset.

Once you use June results to change the model, June is no longer untouched. Record each attempted specification and preserve another later holdout. Trying dozens of weight combinations and showing only the winner measures your selection process as well as your forecasting skill.

Check probability quality, not only hit rate

For binary outcomes, one common Brier score convention is the average of (p − y)², with y equal to 1 for a win and 0 for a loss. For forecasts 0.60 and 0.30 with outcomes 1 and 0, the score is [(0.60 − 1)² + (0.30 − 0)²] / 2 = (0.16 + 0.09) / 2 = 0.125. Lower is better under this convention, but two observations prove almost nothing.

Compare against simple baselines on the exact same holdout: a constant base rate and, where available, the contemporaneous no-vig market benchmark. Also examine reliability bins. Among many forecasts near 60%, roughly 60% should occur if the probabilities are well calibrated. Report the count and uncertainty in each bin; sparse bins can look misleadingly good or bad.

Brier score reflects more than calibration alone. A lower score can result from improved discrimination even if calibration worsens. The official probability calibration guide explains this distinction and why a calibrator needs data independent of model fitting. Hit rate alone ignores how confident predictions were and what price was paid.

Evaluate a decision rule separately from the forecast

A forecasting score is not a trading return. To evaluate a decision rule, specify the minimum estimated edge, stake policy, concurrent exposure limits, executable prices, fees, void handling and settlement conventions before viewing the holdout. Include skipped and unfilled decisions rather than selecting only convenient fills.

Report sample size, time period, turnover, net result, drawdown and performance across time blocks, alongside probability metrics. Correlated bets and regime changes make uncertainty larger than a simple independent-trial calculation suggests. Past favorable results do not guarantee future profit.

Open the models workspace (signup required) to work with your own assumptions and data. Read no-vig probability and EV to connect a forecast to a price, then fees and execution risk before interpreting a backtest as an achievable return. These examples teach evaluation principles; they are not a claim of a profitable or independently certified Stakesift model.