Define the 1X2 value hypothesis before testing it
A useful working hypothesis is: when a calibrated model assigns an outcome a probability above the break-even probability of an executable price, by more than a justified uncertainty allowance, that price has positive expected value under the model. Each part of this statement is necessary.
First, the model should be calibrated. If outcomes assigned probabilities near 50% do not occur at roughly that rate over a suitable out-of-sample sample, its probability scale cannot be interpreted literally. A model can rank matches well while still producing probabilities that are systematically too extreme or too conservative.
For another angle on the same subject, see football betting analysis.
Second, the price must have been genuinely available when the decision could have been made. Comparing a current forecast with an expired opening quote, or evaluating an earlier forecast against a closing quote that was unavailable at forecast time, does not test an executable decision. Model forecasts and market observations need aligned timestamps.
Third, the difference must be large enough to matter. A model probability of 48.0% against a break-even probability of 47.8% is positive under a point estimate, but a 0.2-percentage-point difference is unlikely to be robust to calibration error, uncertain team information, data revisions or minor price movement.
Finally, the hypothesis concerns expected return across repeated, comparable decisions. It does not claim that the selected outcome will win the next match. A losing decision can have been made at a value price, while a winning decision can have been made at an unfavourable price. Outcome and decision quality can only be assessed meaningfully over repeated observations.
This definition also states what would weaken the claim. Poor calibration, unstable estimates, unavailable prices, sensitivity to one margin-removal method, selection-rule changes made after results are known or weak out-of-sample performance should all reduce confidence in an apparent edge.
A complementary analysis is available in Half Time Full Time Analysis: A Probability Framework.
Convert 1X2 odds into a defensible market baseline
For decimal odds O, the raw implied probability is 1 ÷ O. Odds of 2.10 therefore imply a break-even probability of approximately 47.62%. At that probability, the expected return at that individual price is zero before any applicable commission or other transaction cost.
In a three-way football market, the reciprocal probabilities for home, draw and away will usually sum to more than 100%. The excess is commonly described as the overround. If the illustrative prices are 2.10, 3.40 and 3.80, their raw implied probabilities are 47.62%, 29.41% and 26.32%. Their total is 103.35%, implying an overround of about 3.35 percentage points.
Those raw figures should not automatically be interpreted as the market's probability forecast because they include margin. One basic adjustment is proportional normalisation: divide each raw implied probability by their total. In this example, that produces approximately 46.08% for home, 28.46% for draw and 25.46% for away. The adjusted probabilities sum to 100%.
Proportional normalisation is transparent, but it assumes that margin is distributed proportionally across outcomes. That need not hold. Margin allocation can vary between favourites, draws and outsiders, while pricing practices can differ across operators and competitions. Power and odds-ratio approaches provide alternative margin-removal rules, but they also rely on assumptions rather than observing a uniquely true market probability.
For a related perspective, see How to Analyse BTTS With a Balanced Probability Framework.
The methodological response is not to label one de-vig method universally correct. It is to test whether the conclusion changes across defensible methods. If a supposed edge appears only under one transformation and disappears under reasonable alternatives, it is method-dependent evidence rather than robust evidence of market disagreement.
| Method | Core assumption | Analytical strength | Main limitation |
|---|---|---|---|
| Raw reciprocal probabilities | Each price is assessed separately as 1 ÷ odds | Directly identifies the break-even threshold for an offered price | The three probabilities normally sum above 100%, so they are not a margin-free market forecast |
| Proportional normalisation | Margin is distributed proportionally across all three outcomes | Transparent, reproducible and easy to compare across markets | Can miss asymmetric margin allocation and favourite–outsider effects |
| Power transformation | A common exponent redistributes implied probabilities until they sum to 100% | Allows a nonlinear adjustment rather than equal proportional scaling | The inferred probabilities depend on the transformation assumption and its interpretation |
| Timestamp-aligned market consensus | Several consistently de-vigged sources provide a more stable reference than one quote | Reduces dependence on a single source or temporary outlier | Can blend incompatible margins, availability conditions or information states if timestamps are not aligned |
Build an independent and internally coherent probability estimate
The model side of the comparison must produce mutually exclusive probabilities for home win, draw and away win that sum to 100%. The choice of modelling route matters less than whether the process is defined, reproducible, trained without information leakage and validated on observations not used to develop it.
A goal-based method may estimate scoring rates and derive 1X2 probabilities from a score distribution. A direct classification model may estimate the three outcomes without first modelling goals. A structured analyst forecast may combine team strength, venue, expected line-ups and tactical information. Each route embeds different assumptions and can fail in different ways.
Goal models can be sensitive to assumptions about scoring independence, low-score dependence and the stability of attacking and defensive strength. Direct models can conceal feature relationships and may become overconfident after optimisation. Human forecasts can incorporate late information, but they are harder to reproduce and more exposed to selective reasoning. A model should therefore document both its inputs and the point in time at which each input was available.
Variables require particular scrutiny. A confirmed absence may be relevant if known before the forecast, but adding it retrospectively after the result contaminates evaluation. Recent-form measures can partly capture opponent quality, fixture difficulty or random finishing variation rather than persistent ability. League position may duplicate information already contained in a team-strength estimate. Highly correlated variables can also make coefficient-based explanations unstable even where predictive performance appears acceptable.
Market odds create a further design question. If the model uses bookmaker prices as an input, it is not fully independent of the market it is intended to challenge. Such a model may still be useful, but the claim becomes narrower: it is testing whether other variables improve or recalibrate the market probability, rather than whether an independent football model has produced a separate estimate.
Probability quality should be assessed before realised return. Calibration, multiclass Brier score and log loss measure different properties. Calibration tests whether stated probabilities correspond with observed frequencies. Brier score penalises squared probability errors across all three outcomes. Log loss penalises overconfident errors particularly heavily. None is sufficient alone, but together they provide evidence on whether the estimates are suitable inputs to expected-value calculations.
Separate market disagreement from expected return
Two comparisons are useful, and they answer different questions. The first compares the model with a margin-adjusted market estimate. The second compares the model probability with the break-even probability of the actual offered price.
Continue with the explicitly illustrative prices of 2.10, 3.40 and 3.80. After proportional margin removal, the market baseline is approximately 46.08%, 28.46% and 25.46%. Suppose a model produces 49%, 28% and 23%. The model-market difference for the home outcome is about 2.92 percentage points. This is a measure of disagreement with one declared market baseline, not direct proof of an expected return.
The offered home price of 2.10 has a raw break-even probability of 47.62%. The model's 49% estimate exceeds that threshold by about 1.38 percentage points. Expected return per unit staked is probability × decimal odds − 1. The illustrative calculation is 0.49 × 2.10 − 1 = 0.029, or 2.9%.
The model's fair decimal price for a 49% event is approximately 2.04, calculated as 1 ÷ 0.49. An offered price of 2.10 is higher, which is favourable under the model. However, this inference is conditional on the 49% estimate being sufficiently accurate. It does not establish that 2.10 is objectively generous or that the estimate is independent of the market.
This distinction prevents two common errors. First, a model can disagree with a de-vigged market while still failing to beat the available break-even threshold. Second, a price can appear positive under a point estimate while becoming negative after a small probability revision. An edge should therefore be recorded with its assumptions: model probability, available odds, timestamp, market source, de-vig method, market reference and uncertainty treatment.
Reviewing only the highlighted outcome also discards useful information. In the illustration, the complete model vector is 49%, 28% and 23%. Its expected returns at the quoted prices are approximately 2.9% for home, −4.8% for draw and −12.6% for away. Inspecting the full probability vector can expose incoherent probability transfer between outcomes that a single selected price would conceal.
Treat market movement as context, not automatic confirmation
Market analysis extends beyond one set of odds. Prices can differ by source, observation time, margin, information availability, liquidity and operating constraints. A useful comparison records where and when each price was observed rather than combining unmatched snapshots.
A shortening price may reflect new information, order flow, risk management, imitation of another source or correction of an earlier quote. A drifting price can have similarly varied causes. Movement alone does not identify which explanation applies. It becomes analytically useful only when linked to a documented information timeline.
For example, if a model forecast is frozen before team news, appropriate reference points include the contemporaneous price and a later closing price. A later move in the model's direction is consistent with information becoming more fully reflected in the market. It is not proof that the model identified the cause of the move: the change may be unrelated, and an individual observation carries little evidential weight.
Closing-price comparison can be a diagnostic because the closing market has incorporated more time and, potentially, more information. Consistently obtaining prices above a later, consistently defined reference may support the narrower claim that a process identifies favourable numbers. It does not replace outcome-based probability validation, and it can be distorted if the reference source, closing timestamp or margin-removal method changes across observations.
Outlier prices require particular care. A high quote may represent genuine competition, but it may also be stale, available only briefly, subject to restrictive limits or attached to different settlement terms. The research record should preserve the executable price and relevant constraints rather than substituting a cleaner retrospective number.
Market consensus is also a methodological choice rather than an observed fact. Averaging several de-vigged sources may reduce dependence on one operator or temporary outlier, but only when observations refer to the same time and market definition. Otherwise, the average combines different information states and creates a baseline that no participant could have observed.
Test whether the apparent value survives plausible error
A point estimate conceals uncertainty. The practical question is not merely whether expected return is positive at 49%, but how easily that conclusion changes when the probability, available price or market-adjustment method changes.
For the illustrative home price of 2.10, expected return is negative at 46% and 47%, becomes positive at 48%, and rises as the assumed probability increases. The break-even point is approximately 47.62%. The accompanying sensitivity chart applies the same formula—probability × 2.10 − 1—to probabilities from 46% to 52%. These are arithmetic scenario values, not measured results.
A robustness process should challenge several dimensions:
- Probability sensitivity: vary the model estimate within a range justified by calibration error, parameter uncertainty and plausible team-information changes.
- Price sensitivity: recalculate using the price that was realistically available, including commission or effective transaction costs where relevant.
- Margin sensitivity: compare proportional, power-based and other defensible de-vig approaches.
- Specification sensitivity: rerun the model without fragile variables, with alternative recency weights or without subjective adjustments.
- Timing sensitivity: repeat comparisons using consistently defined opening, forecast-time and closing snapshots.
- Segment sensitivity: check whether calibration deteriorates for draws, strong favourites, outsiders, particular competitions or broad price ranges.
A probability interval is not automatically valid because software produced it. Statistical intervals may omit model misspecification, data errors, omitted variables and structural changes in teams or competitions. A model can report a narrow interval while being confidently wrong about the process that generates football outcomes.
One conservative decision rule is to require the lower end of a justified probability range to exceed the break-even probability. This rejects many marginal cases, but reduces dependence on small and unstable differences. Another approach is to apply a predeclared probability haircut or minimum expected-return threshold. Whatever rule is used should be specified before results are reviewed; otherwise, thresholds can be adjusted retrospectively to preserve preferred conclusions.
The line shows expected return calculated as probability × 2.10 − 1. It demonstrates how a small probability revision can reverse the conclusion near the 47.62% break-even point.
Illustrative scenario only. Values are calculated directly from probability × 2.10 − 1.
Validate the process out of sample, not just the selections
A valid test freezes the model, selection rule, feature definitions and data-availability rules before evaluating unseen matches. Repeatedly modifying the process in response to the same results turns a test set into training data and makes reported performance increasingly optimistic.
The validation record should include every forecast, not only outcomes that later met a selection rule. For each match, retain the three model probabilities, forecast timestamp, quoted prices, source, de-vigged market probabilities, decision threshold and any exclusion reason. This makes it possible to separate model quality from price availability and selective reporting.
Evaluation should proceed at several levels. First, test overall calibration and proper scoring rules. Second, examine calibration by outcome and broad probability region, while recognising that smaller segments create wider uncertainty. Third, test whether the predefined value rule identifies prices that outperform the chosen market reference. Fourth, assess realised returns cautiously because football outcomes generate substantial variance and observations may not be independent across leagues, teams or shared model regimes.
Return alone is a weak short-run diagnostic. A small number of high-odds outcomes can make a poorly calibrated method look successful, while a sound process can experience an extended losing sequence. Conversely, good Brier-score or log-loss performance does not guarantee favourable prices because a model can forecast accurately without disagreeing with the market in a useful direction.
Benchmark comparisons help identify where any improvement originates. Relevant baselines include raw market probabilities, a consistently de-vigged market, a simple team-strength model and the full proposed method. If the more complex method does not improve probability scores or value discrimination out of sample, its additional variables may be adding noise rather than information.
Uncertainty should remain visible in the final interpretation. Strong evidence requires calibrated probabilities, stable findings across reasonable specifications, executable prices, consistent timestamp rules and performance that persists on genuinely unseen observations. A positive point estimate without those conditions is a candidate hypothesis, not a validated 1X2 edge.
Interpret value as a conditional research finding
A probability framework makes 1X2 analysis testable, but it does not remove uncertainty. The market probability is inferred through a margin-removal model, the football probability is estimated through an imperfect forecasting process, and the expected return depends on a price that may move, be restricted or disappear.
The strongest conclusion is therefore conditional: a price may represent value if the forecast is calibrated, the odds were executable, the comparison uses aligned information and the edge survives reasonable sensitivity tests. Recording those conditions is not administrative detail. It is what separates a reproducible market-analysis method from a retrospective explanation of outcomes.

