Define the Half Time/Full Time target before reading momentum
The first task is to specify when the forecast is made. A prediction before kick-off, at half time and during the second half uses a different information set and therefore answers a different probability question.
Before kick-off, there are nine mutually exclusive Half Time/Full Time paths: H/H, H/D, H/A, D/H, D/D, D/A, A/H, A/D and A/A. The first symbol describes the half-time state and the second the full-time result, where H, D and A mean home win, draw and away win. A pre-match joint probability can be decomposed as P(HT = i, FT = j) = P(HT = i) × P(FT = j | HT = i).
For another angle on the same subject, see football betting analysis.
That decomposition identifies two distinct sources of error. A method may estimate the half-time state poorly, the transition from half time to full time poorly, or both. A strong model of second-half reversals does not compensate for an unrealistic estimate of how often the required half-time state occurs.
At the interval, the first state is known. If the home team leads, only H/H, H/D and H/A remain possible, and their conditional probabilities must sum to 100%. The question is no longer whether the home side will lead at half time, but how the observed match information changes the distribution of the three possible full-time results.
After the second half has started, remaining time and the current score become explicit conditioning variables. The target is then the full-time result conditional on the known half-time state, current score, current minute and information available at that moment. Describing every live forecast as an HT/FT prediction can conceal this change in the target and information set.
State the momentum claim as a falsifiable hypothesis
The useful hypothesis is not merely that momentum exists. It is that a pre-defined set of recent match events contains incremental predictive information about the full-time result.
This framework also connects with How to Analyse BTTS With a Balanced Probability Framework.
The null hypothesis is stricter: after team strength, venue, score state, remaining time, dismissals and other structural information are included, the momentum variables do not improve out-of-sample probability estimates. This is the appropriate test because much apparent momentum is already explained by baseline conditions.
For example, a trailing side may dominate territory because it must attack while the leading side protects space. The territorial observation may be accurate, but the inference that an equaliser is now highly likely does not follow automatically. The pressure must predict outcomes better than a baseline model that already recognises the score-driven tactical pattern.
Evidence, interpretation and inference should remain separate. An increase in dangerous entries is evidence recorded by the method. The claim that the opponent is losing control is an interpretation. A higher probability of a comeback is an inference only if the evidence improves a conditional forecast relative to the baseline.
| Observed signal | Proposed mechanism | Competing explanation | Required test |
|---|---|---|---|
| Repeated dangerous entries | The defence is failing to prevent meaningful access | The trailing side is taking score-driven risks | Test incremental value after controlling for score, time and strength |
| High recent possession | One team is sustaining territorial pressure | The opponent is deliberately conceding harmless possession | Separate possession location and progression from raw share |
| Several recent shots | Attacking threat is recurring | Attempts are low quality or blocked from poor locations | Compare shot count with chance-quality variables |
| Pressure after a dismissal | Match control has shifted | Numerical advantage explains the change | Model the dismissal directly and test residual momentum |
Build a baseline, then update it rather than replacing it
A defensible analysis starts with baseline probabilities for the feasible remaining paths. These may come from a statistical model using pre-match team strength, venue, score state, time and known player availability. The momentum read is then introduced as an update, not as a substitute for the baseline.
For a related perspective, see How to Analyse Both Teams to Score with a Probability Framework.
For mutually exclusive outcomes, one update form is p′j = pj × Lj / Σ(pk × Lk). Here, pj is the baseline probability of outcome j and Lj is the evidence weight associated with the observed match information. The denominator renormalises the updated probabilities so that the complete outcome distribution still sums to one.
The evidence weights should be estimated and calibrated on historical observations that were not used to evaluate the model. They should not be created from verbal confidence or selected after an outcome is known. Where likelihood ratios have not been empirically estimated, the values must be labelled as scenario assumptions rather than findings.
An illustrative half-time update
Suppose, purely for illustration, that a home team leading at the interval has baseline probabilities of 55% for H/H, 27% for H/D and 18% for H/A. Assume a defined group of first-half signals receives scenario weights of 1.25, 0.95 and 0.75 respectively. The unnormalised values are 0.55 × 1.25 = 0.6875, 0.27 × 0.95 = 0.2565 and 0.18 × 0.75 = 0.1350. Their total is 1.0790. Dividing each value by that total gives updated probabilities of approximately 63.7% for H/H, 23.8% for H/D and 12.5% for H/A.
This calculation does not establish that positive home momentum deserves those weights. It demonstrates two constraints: increasing H/H requires competing outcomes to fall, and every update remains anchored to the original conditional distribution.
A simpler binary update can be used for sensitivity analysis. If H/H has a baseline probability p and the evidence multiplier is L, then p′ = Lp / (1 − p + Lp). This is useful for showing how dependent a conclusion is on an assumed evidence strength. A production HT/FT model, however, should retain all mutually exclusive outcomes rather than carelessly reducing the task to one-versus-rest.
Starting from a fictional 55% H/H baseline, the update uses p′ = Lp / (1 − p + Lp). The chart shows how strongly the forecast depends on the assumed evidence multiplier; it does not claim that any multiplier is empirically estimated.
Illustrative scenario only. Values are rounded to one decimal place: 0.7 × 0.55 / (0.45 + 0.7 × 0.55) = 46.1%; 1.0 gives 55.0%; 1.3 gives 61.4%; and 1.6 gives 66.2%.
Operationalise momentum as variables, windows and context
Momentum cannot be validated while it remains a visual impression. The method needs observable variables, a defined measurement window and rules set before the result is known.
- Chance quality: distinguish low-value attempts from opportunities created in dangerous locations. Raw shot totals can exaggerate pressure when attempts are speculative, blocked or heavily contested.
- Territorial access: measure entries into threatening areas, progression and the ability to keep the opponent from exiting. Possession without penetration should not automatically receive the same weight.
- Defensive disruption: identify whether pressure follows a dismissal, injury, forced substitution or tactical reorganisation. These are structural changes and should usually be modelled directly rather than hidden inside a momentum score.
- Tempo and recurrence: separate one isolated chance from a repeated pattern. A short burst may matter, but its effect should be allowed to decay if the pattern does not continue.
- Game-state context: compare behaviour with what is expected from teams in the same score and time situation. A trailing team taking more risks is not, by itself, unusual evidence.
The measurement window is a modelling choice. A very short window reacts quickly but is noisy. A longer window is more stable but can blend together different tactical phases. Rather than declaring one window universally best, define several plausible windows before evaluation and test whether the conclusion survives that choice.
Variable overlap also matters. Shots, expected chance value, penalty-area entries and final-third possession may all describe part of the same attacking sequence. Adding each as though it were independent can count one pattern multiple times. Regularisation, dimension reduction or a deliberately small variable set can reduce this problem.
Direction must also be specified. Positive attacking momentum for one team is not necessarily equivalent to negative momentum for the other. A team may concede territory deliberately while retaining a meaningful counterattacking threat. A balanced model represents attacking threat and defensive vulnerability separately instead of collapsing them into a single emotional label.
Challenge the read with confounders and counter-cases
The main confounder is score effect. Teams that lead often change risk, width, pressing intensity and possession priorities, while teams that trail are encouraged to advance more players. A model trained without adequate score-state controls may learn urgency rather than predictive momentum.
Team quality is another confounder. The same sequence can have different implications depending on which team produces it and whether the opponent is comfortable defending deeper. Pre-match strength should not disappear merely because live observations are available. Live evidence modifies the prior; it does not erase established differences.
Other threats include red cards, injuries, fatigue, weather, stoppage time, tactical substitutions and formation changes. Some are difficult to measure consistently, but omitting them can make a momentum variable appear more informative than it is. Territorial dominance after a dismissal, for example, may principally be evidence of numerical advantage.
Measurement error enters through event definitions and data timing. Different providers may classify chances, pressures or possession sequences differently. A live feed may also revise an event after the initial observation. If a method is intended for real-time use, it must be tested using only information genuinely available at the stated forecast time. Using later corrections, final event classifications or post-match line-up information would introduce look-ahead bias.
Narrative selection creates a further danger. Analysts tend to remember pressure that preceded a goal and forget equally intense periods that produced nothing. Recording every qualifying momentum episode under a fixed rule is necessary. Otherwise, the sample is selected partly by the outcome the method is supposed to predict.
Test incremental value, calibration and robustness
The cleanest test compares nested models. A baseline model contains pre-match strength, venue, score, time and major structural events. A second model adds the proposed momentum variables. The relevant question is whether the expanded model improves forecasts on chronologically later matches that were not used to choose variables, tune windows or estimate calibration.
Accuracy alone is inadequate because HT/FT outcomes are probabilistic and unevenly distributed. Proper scoring rules such as multiclass log loss or the Brier score assess the full probability vector. Calibration should also be inspected: among forecasts assigned similar probabilities, does the outcome occur at a comparable rate? A model can rank possibilities reasonably well while remaining systematically overconfident.
Evaluation should be conditional as well as aggregate. Performance may differ when the home side leads, the match is level or the away side leads at half time. It may also vary by remaining time, pre-match strength difference and dismissal status. Pooling all states can conceal a method that appears acceptable overall but fails in the reversal situations where momentum is most often invoked.
Useful robustness checks include:
- changing the recent-event window without redesigning the hypothesis after seeing results;
- removing one correlated variable group at a time;
- testing alternative definitions of chance quality or territorial pressure;
- comparing compatible event definitions across data providers where possible;
- excluding dismissals and other extreme states, then evaluating them separately;
- recalibrating on one period and evaluating on a later period to detect drift.
Rare HT/FT transitions create uncertainty even when the full archive is large. Reversal paths may contain far fewer observations than continuation paths. Confidence intervals, bootstrap variation or Bayesian posterior ranges can communicate this instability. Hierarchical shrinkage can also prevent small subgroups from generating implausibly extreme probabilities.
If market probabilities are used as an external benchmark, the comparison should remove the bookmaker margin and align timestamps. A model observed after new match information arrived cannot fairly be compared with an earlier price. Such a benchmark can be informative, but it does not replace outcome-based calibration and held-out evaluation.
Interpret momentum as a probability modifier, not a verdict
A methodologically coherent workflow fixes the forecast time and feasible HT/FT paths, generates a baseline distribution, records only pre-defined momentum variables and applies an empirically calibrated update. It should report the revised probability distribution rather than a single narrative conclusion.
The source of any probability change should also be explicit. A shift caused by a dismissal is substantively different from one caused by repeated high-quality chances, even if both increase the same outcome probability. Transparent attribution makes double counting easier to detect and allows the model's assumptions to be challenged.
The central inference is deliberately limited. Momentum can support an HT/FT forecast when it is operationally defined, tested out of sample and shown to improve conditional probabilities beyond a credible baseline. It cannot make a path certain. Low-frequency events, finishing variance and tactical changes remain only partly observable.

