What this model can—and cannot—predict
A decision tree can estimate whether an Indian football player is likely to change clubs during a defined future window, such as the next transfer window or the next 12 months. It cannot know confidential negotiations, agent relationships, injuries not yet reported, or a club’s last-minute budget decision. Treat the output as a recruitment signal, not a verdict.
That distinction matters in India’s football ecosystem, where the Indian Super League, I-League, domestic competitions, loans, player releases, registration rules, and short-term contracts create different meanings of “transfer”. Define the outcome before collecting data:
- Binary target:
1if the player joins a different club during the next window;0otherwise. - Prediction date: the date on which the model is allowed to use information.
- Forecast horizon: for example, 90 days, one transfer window, or one season.
- Transfer definition: permanent move, loan, free-agent signing, or any registration change.
If you are building this as a scouting product, document these definitions in the interface and model card. Clear labels are more valuable than a sophisticated algorithm trained on ambiguous records.
Build an Indian football dataset
Start with one row per player per prediction date, rather than one row per player for their entire career. Useful columns may include:
- Age, position, preferred foot, nationality, and eligibility status.
- Minutes played, starts, goals, assists, progressive actions, cards, and injury availability.
- Club level, league, squad role, contract end date, and prior loan history.
- Recent performance trend, such as minutes and goal contributions over the last 5, 10, and 20 matches.
- Club results, coaching changes, financial signals where publicly available, and squad congestion.
- Historical transfer outcome within the chosen forecast window.
Use official competition records, club announcements, player registries, and carefully documented public sources. Store the source and retrieval date for every important field. Do not scrape or republish personal data without checking the source’s terms and applicable privacy obligations.
The target should be created after the prediction date. For example, a row dated 1 May 2025 can use information available on 1 May, while its label records whether a qualifying move occurred by 31 August 2025. Never use the announced transfer itself, a later match statistic, or a post-window salary estimate as an input for that row.
Teams building a broader sports intelligence stack can borrow practices from Indian open-source AI developer projects, especially around reproducible datasets, documentation, and community review.
Prepare features without leaking future information
Clean duplicate player records and standardise club names before training. Resolve players who share names by using a stable identifier, but avoid exposing sensitive identifiers in dashboards. For missing values, distinguish between “zero”, “not recorded”, and “not applicable”. A missing contract end date is not the same as an expired contract.
Categorical variables such as position, club, and competition can be one-hot encoded. A decision tree does not require feature scaling, but numerical features should still be checked for impossible values and inconsistent units. Useful derived features include:
days_to_contract_endat the prediction date.minutes_sharewithin the player’s club.starts_last_10andavailability_rate_last_season.recent_form_changecompared with the previous period.club_change_rateacross the player’s prior seasons.
The largest risk is target leakage. If a feature is only known after the transfer window closes, it cannot be used in a legitimate forecast. Randomly splitting rows can also leak information when the same player appears in both training and test sets across nearby dates. Prefer a chronological split: train on earlier windows, validate on a later window, and reserve the latest window as an untouched test set.
Train a constrained decision tree in Python
A simple baseline is valuable because it gives analysts something interpretable to compare against. The following example assumes that categorical columns have already been encoded and that transfer_next_window is a binary target:
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import (
classification_report, confusion_matrix,
roc_auc_score, average_precision_score
)
features = [
"age", "minutes_last_10", "starts_last_10",
"availability_rate", "days_to_contract_end",
"club_change_rate", "competition_level"
]
df = pd.read_csv("indian_football_transfers.csv")
df["prediction_date"] = pd.to_datetime(df["prediction_date"])
df = df.sort_values("prediction_date")
cutoff = pd.Timestamp("2025-01-01")
train = df[df["prediction_date"] < cutoff]
test = df[df["prediction_date"] >= cutoff]
X_train, y_train = train[features], train["transfer_next_window"]
X_test, y_test = test[features], test["transfer_next_window"]
model = DecisionTreeClassifier(
max_depth=4,
min_samples_leaf=25,
class_weight="balanced",
random_state=42
)
model.fit(X_train, y_train)
probabilities = model.predict_proba(X_test)[:, 1]
predictions = (probabilities >= 0.5).astype(int)
print(classification_report(y_test, predictions))
print("ROC-AUC:", roc_auc_score(y_test, probabilities))
print("PR-AUC:", average_precision_score(y_test, probabilities))
print(confusion_matrix(y_test, predictions))The depth and leaf-size limits reduce overfitting. class_weight="balanced" is useful when transfers are less common than non-transfers, but it does not solve every imbalance problem. Select a decision threshold based on the operational goal: a scouting team may prefer higher recall, while a club with limited analyst time may prioritise precision.
Evaluate what matters to a recruitment team
Accuracy alone can be misleading. A model that predicts “no transfer” for everyone may appear accurate when actual moves are rare. Report:
- Precision: among players flagged, how many transferred?
- Recall: of all players who transferred, how many were flagged?
- F1 score: a balance between precision and recall.
- ROC-AUC: ranking quality across thresholds.
- PR-AUC: often more informative for rare transfer events.
- Calibration: whether a 70% prediction happens roughly 70% of the time.
- Top-k lift: how many transfers appear among the top 10, 25, or 50 ranked players?
Break results down by position, competition, age group, club size, and transfer type. Performance that looks strong overall may fail for goalkeepers, younger players, or leagues with sparse reporting. Compare the tree with a simple baseline such as “contract expires soon” and with a regularised logistic regression. A complex model should earn its place by improving decisions, not just benchmark scores.
Interpret and operationalise predictions
Export the tree and inspect its paths with sklearn.tree.plot_tree, but do not present a branch as causal evidence. If the first split is contract proximity, the model has identified an association in your dataset—not proof that contract timing caused the move. Use feature importance cautiously; correlated variables can divide importance between one another.
For each flagged player, show the probability, forecast window, top contributing conditions, data freshness, and confidence limitations. Keep a human review step involving scouting, football operations, and legal or registration specialists. A transfer recommendation should also consider role fit, wage budget, language and relocation support, injury risk, and squad rules—factors that a public dataset may not capture.
If you are turning this into a service, design the workflow like a practical analytics product rather than a static notebook. Lessons from cost-effective recruitment platforms for Indian founders can help with permissions, recruiter workflows, audit trails, and candidate-data governance.
Common failure modes
- Small samples: combine several seasons carefully, while accounting for rule and competition changes.
- Duplicate players: use stable IDs and verify transfers manually where possible.
- Random validation: replace it with time-based backtesting.
- Overgrown trees: tune depth, leaf size, and pruning through time-aware validation.
- Unclear labels: separate loans, releases, retirements, and club closures.
- False certainty: display probabilities and intervals, not definitive claims.
- Selection bias: public data may overrepresent prominent clubs and announced transfers.
A random forest or gradient-boosting model may improve ranking, but it will usually be less transparent. Start with the decision tree baseline, record every experiment, and only add complexity after proving that it improves out-of-sample performance.
Responsible use in Indian football
Transfer predictions can affect a player’s reputation, bargaining position, and employment prospects. Do not publish speculative individual rankings as fact. Limit access to sensitive contract or injury information, explain how scores are produced, and provide a correction process for inaccurate records. Keep predictions advisory and require qualified human review before a club acts on them.
The strongest project is not the one with the most features. It is the one with defensible labels, clean time boundaries, transparent evaluation, and a workflow that helps football professionals ask better questions.