Why Shapley values matter in Indian football
Football performance models can rank players, estimate win probabilities, or predict expected goals, but a prediction alone rarely answers the question coaches and sporting directors care about: why did the model reach this conclusion? Shapley values provide a structured way to attribute a model output to its inputs.
The method comes from cooperative game theory. In machine learning, the “players” are features such as progressive passes, pressures, carries, shot locations, minutes, opponent strength, or rest days. A Shapley value estimates how much each feature changes a prediction, averaged across the feature combinations in which it could appear.
For Indian clubs, academies, analysts, and sports-tech startups, this is useful when data is uneven across competitions and decision-makers need evidence they can inspect. It can complement high-performance AI pipelines by making the final outputs easier to audit and communicate.
Start with the decision, not the explanation method
Before calculating any values, define the football decision the model supports. Examples include:
- Estimating a team’s probability of winning or drawing.
- Predicting expected goals or expected threat generated by a player.
- Identifying players likely to succeed in a particular tactical role.
- Comparing recruitment targets after adjusting for league, minutes, and opponent quality.
- Flagging fatigue or injury-risk indicators for workload management.
The target determines what a positive contribution means. A feature that raises predicted pressing success may not raise a team’s win probability in every match. Avoid presenting one universal “player impact” score when the model actually answers a narrower question.
Also separate player identity from player actions. If a model contains both a player ID and event statistics, the ID may absorb team, coach, position, or competition effects. That can make explanations look precise while hiding confounding. For recruitment, prefer role-relevant, context-adjusted variables and evaluate whether the model generalises to unseen matches or seasons.
Build a reliable football dataset
A Shapley explanation cannot repair poor event data. Assemble a table with one clearly defined observation per player-match, player-possession, or team-match, depending on the task. Useful fields may include:
- Minutes played and starting status.
- Goals, assists, shots, and shot quality.
- Progressive passes, carries, final-third entries, and key passes.
- Tackles, interceptions, pressures, blocks, and recoveries.
- Pass completion split by zone and pressure state.
- Position, formation, role, and substitution timing.
- Opponent strength, venue, score state, travel, rest, and weather where available.
- Competition and season, including the Indian Super League, I-League, domestic cups, or academy competitions.
Normalise volume statistics by minutes or possessions where appropriate, but retain raw counts when workload itself matters. Treat missingness explicitly: a missing tracking metric is not necessarily a zero. Check whether coverage differs by competition or venue, because that can create a model that rewards better-instrumented teams.
Use time-based validation. Training on later matches and testing on earlier matches can produce misleading confidence, particularly when squads, coaches, and tactical systems change. A stronger design trains on earlier periods and tests on future matches, with a separate evaluation across competitions when possible.
Choose the right SHAP approach
The Python shap library supports several explainers, but the choice should follow the model and data structure:
- TreeExplainer: efficient for decision trees, Random Forests, XGBoost, and LightGBM.
- LinearExplainer: suitable for linear and logistic regression with a carefully chosen background dataset.
- KernelExplainer: model-agnostic but slower; useful when no specialised explainer fits.
- Deep or gradient explainers: relevant to neural models, but require extra care with baselines and correlated inputs.
SHAP values are calculated relative to a baseline, usually the average model output over a background sample. A player’s explanation therefore means “higher or lower than this reference prediction,” not “this player caused this result.” Select a background dataset that reflects the use case—such as recent league matches or comparable player roles—and document it.
For a binary outcome, decide whether explanations are reported in probability space or log-odds. Probability-space explanations are easier for coaches to read, while log-odds may add more cleanly for some models. State the choice in reports.
Interpret a player explanation correctly
Suppose a model predicts a team’s chance of creating a high-quality chance. A positive SHAP value for progressive carries means that, given the other inputs and the selected baseline, the observed carry value pushed the prediction upward. It does not prove that the player independently created that outcome.
Use three views together:
- Local explanation: why one player-match or team-match prediction moved above or below baseline.
- Global importance: which features matter most across the evaluation set.
- Dependence and interaction views: whether a feature’s effect changes with position, score state, opponent, or another feature.
Do not equate a large absolute value with overall player quality. A centre-back may receive a large positive contribution in a defensive-transition model but little value in a chance-creation model. Compare players only within a defined role, competition, sample size, and tactical context.
Correlated variables require particular caution. Tackles, pressures, defensive actions, and possession may encode overlapping information. SHAP can distribute credit among correlated features in ways that are mathematically consistent but difficult to interpret as independent football contributions. Group related features—such as defensive actions or progression—and report grouped importance alongside individual values.
A practical workflow for an Indian club or startup
1. Define the outcome. Write down the decision, prediction horizon, unit of analysis, and acceptable error.
2. Create a data dictionary. Record source, coverage, units, missing-value rules, and whether each field is available before the prediction moment.
3. Establish a baseline. Compare the model with simple benchmarks such as league-average rates, position averages, or a regularised regression.
4. Train with leakage controls. Exclude post-event variables and use chronological validation.
5. Select a representative background set. Stratify by position, competition, or match state if the dataset is heterogeneous.
6. Calculate SHAP values on held-out data. Explanations on training data can look persuasive while reflecting memorisation.
7. Audit stability. Refit the model across seasons, competitions, and random seeds. Track whether rankings and explanations change materially.
8. Translate findings into actions. Convert a pattern into a scouting question, training intervention, or tactical hypothesis that staff can test.
Teams building this capability should also invest in system design for high-performance AI startups, especially when analysts, coaching staff, and data providers need shared access to versioned models and reports.
Common mistakes to avoid
- Calling attribution causation: SHAP explains model behaviour; it does not establish that a player caused a win.
- Ignoring exposure: A substitute with strong per-minute numbers may have a small and selective sample.
- Mixing leagues without adjustment: Event definitions and playing styles differ across competitions.
- Using only aggregate rankings: A global feature chart cannot explain a specific match or role.
- Reporting unstable values: Wide variation across folds should be shown rather than hidden.
- Overlooking fairness: Age, nationality, salary, or club reputation can become proxies that reproduce selection bias.
- Skipping human review: Coaches should challenge explanations that conflict with video, role instructions, or known tactical constraints.
For operational deployments, monitor explanation drift alongside predictive performance. If a tracking provider changes its event definitions or a club changes its playing style, the model may remain accurate while the meaning of its features changes. Practices from LLM application performance monitoring in India are not specific to football, but the same principles—versioning, alerting, and monitoring inputs and outputs—apply.
How to present results to coaches and scouts
A useful report should lead with the football question, not the algorithm. Show the prediction, baseline, top positive and negative factors, confidence or stability indicators, and a short caveat. Use plain language such as “progression actions increased the model’s estimate relative to comparable matches,” rather than “the player generated X points of causal impact.”
Pair each explanation with video clips and match context. If the model flags a full-back’s defensive positioning, analysts can inspect whether the value came from genuine anticipation, a deep block, or an opponent error. This makes SHAP a tool for structured review rather than an automated replacement for expertise.
Conclusion
Shapley values can make Indian football performance models more transparent, but their value depends on disciplined modelling. Define the outcome, control for context, validate chronologically, choose a meaningful baseline, and report uncertainty. Then use explanations to generate testable coaching and recruitment hypotheses—not to produce simplistic player verdicts.
As of 2026, the strongest use case is a repeatable loop: reliable match data, a validated model, transparent explanations, video review, and feedback from football staff. That combination can help Indian clubs make better decisions while keeping analytical claims proportionate to the evidence.
FAQ
Are Shapley values the same as player ratings?
No. They explain how input features move a particular model prediction. A player rating is a separate scoring framework and may combine many matches or objectives.
Can SHAP prove that a player caused a team to win?
No. SHAP describes the model’s attribution, not causal impact. Causal claims require a different research design and stronger assumptions.
Which model works best with SHAP?
There is no universal best model. Tree-based models are often practical for structured event data, while linear models offer easier baseline interpretation. Compare accuracy, calibration, stability, and usability.
How can small Indian clubs begin?
Start with a narrow target, a clean match-level dataset, a transparent baseline, and held-out evaluation. Expand to richer tracking data only after the workflow is reliable.
Apply for AI Grants India
Are you building explainable sports analytics, scouting systems, or other AI applications in India? Explore support through AI Grants India, and connect the project to measurable outcomes such as model reliability, analyst productivity, or player-development impact.