Why t-SNE belongs in a football analytics workflow
Indian clubs, academies, and scouting teams increasingly work with event data, GPS outputs, video tags, and competition reports. The challenge is not merely collecting metrics; it is finding comparable player profiles across positions, match minutes, and playing styles. t-SNE (t-distributed Stochastic Neighbor Embedding) can help create a two-dimensional map of those profiles.
The map is useful for exploration: nearby points may represent players with similar statistical signatures, while visible groups can suggest archetypes such as ball-winning midfielders, progressive full-backs, or high-volume forwards. It is not, however, a ranking model or proof that a player is better. Treat it as an exploratory layer in a broader workflow that includes domain expertise, video review, and out-of-sample validation.
Teams building this workflow should also plan for reproducibility. A modular, testable pipeline—similar to the practices described in how to build high-performance AI pipelines—makes it easier to rerun analyses when new matches or seasons arrive.
Define the football question first
Do not begin with a colourful scatter plot. Start with a decision the analysis must support:
- Which academy players resemble a senior-team role profile?
- Which domestic players could replace a departing midfielder?
- How do players change after moving between the Indian Super League, I-League, state leagues, or youth competitions?
- Which opponents use similar role distributions or attacking patterns?
The question determines the unit of analysis. A single row might represent a player-season, player-competition-season, or player-match. For scouting, player-season records are usually more stable than individual matches. For opposition preparation, match-level or team-phase records may be appropriate.
Build features that account for Indian football contexts
Collect metrics from event data, tracking systems, manually coded video, training records, and club databases. Useful categories include:
- Possession: progressive passes, carries, final-third entries, pass completion under pressure, and turnovers.
- Chance creation: expected assists, key passes, shot assists, touches in the penalty area, and crossing outcomes.
- Defending: pressures, interceptions, tackles, recoveries, aerial-duel success, and defensive actions after losing possession.
- Physical output: high-speed running, accelerations, repeated sprints, and distance by match phase where tracking data is available.
- Context: position, minutes, team possession, competition level, match state, and opponent strength.
Raw totals favour players who play more minutes. Convert count metrics to per-90 values, but apply a minimum-minute threshold and show the sample size in scouting reports. Adjusting for team possession and role can prevent a defensive midfielder in a low-possession side from being compared unfairly with a possession-dominant midfielder.
For Indian competitions, metadata matters. Record competition, season, venue conditions where relevant, and whether a player operated in a different role. Avoid mixing youth and senior data without an explicit modelling decision. Standardise player and club names before joining files, and remove duplicated fixtures.
Prepare the dataset before t-SNE
A defensible preprocessing sequence is:
1. Remove identifiers and leakage-prone fields such as future transfer fees or labels derived from the target decision.
2. Impute missing values using football-aware rules, then add missingness indicators where absence itself is informative.
3. Winsorise extreme values only when they reflect recording errors or unstable tiny samples.
4. Standardise numerical features with StandardScaler or a robust alternative.
5. Reduce very wide feature sets with PCA before t-SNE. Keeping 20–50 principal components often reduces noise and computation time.
6. Run separate analyses by broad role when comparing roles directly would be misleading.
Feature correlation is a common problem. Goals, shots, expected goals, and touches in the box may all describe a similar attacking dimension. Use domain knowledge or correlation checks to avoid allowing one concept to dominate the embedding. A strong data pipeline should also log feature definitions and transformations; system design for high-performance AI startups offers useful principles for keeping analytical systems observable and maintainable.
A practical Python implementation
The following example uses scikit-learn. It includes scaling, PCA, a fixed seed, and a perplexity that must be tested rather than accepted blindly.
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
metrics = [
"progressive_passes_p90", "shot_assists_p90",
"pressures_p90", "interceptions_p90",
"aerial_duel_rate", "high_speed_distance_p90"
]
df = pd.read_csv("indian_football_players.csv")
work = df.dropna(subset=["player", "minutes"]).copy()
work = work[work["minutes"] >= 900]
X = work[metrics]
X = SimpleImputer(strategy="median").fit_transform(X)
X = StandardScaler().fit_transform(X)
X = PCA(n_components=min(30, len(metrics))).fit_transform(X)
embedding = TSNE(
n_components=2,
perplexity=20,
init="pca",
learning_rate="auto",
max_iter=1500,
random_state=42
).fit_transform(X)
work[["tsne_1", "tsne_2"]] = embedding
plt.figure(figsize=(11, 7))
for label, group in work.groupby("broad_position"):
plt.scatter(group.tsne_1, group.tsne_2, label=label, alpha=0.75)
plt.legend(title="Position")
plt.xlabel("t-SNE dimension 1")
plt.ylabel("t-SNE dimension 2")
plt.title("Indian football player performance profiles")
plt.show()For larger datasets, consider a high-performance implementation and track runtime, memory, and data version. Teams can borrow deployment discipline from building high-performance AI applications with open-source tools, especially when analysts need repeatable runs on shared infrastructure.
Tune and validate the embedding
The visual appearance of t-SNE changes with perplexity, learning rate, initialisation, distance metric, feature set, and random seed. Run several configurations—for example, perplexities of 5, 15, 30, and 50—and compare whether meaningful neighbourhoods persist. Do not interpret the distance between two far-apart clusters as a precise football distance: t-SNE preserves local neighbourhoods better than global geometry.
Use complementary checks:
- Compare embeddings across multiple random seeds.
- Colour the same points by position, competition, age band, and team possession.
- Inspect nearest neighbours in the original standardised feature space.
- Test cluster candidates with silhouette scores or density-based methods, while recognising that t-SNE itself does not produce definitive clusters.
- Hold out a season and check whether player neighbourhoods remain stable.
- Ask coaches and scouts whether the nearest-neighbour examples make football sense.
If clusters disappear when one metric is removed or the seed changes, report that uncertainty. Stability is more valuable than a visually dramatic plot.
Turn clusters into scouting decisions
A cluster should be described by its feature profile, not just a label. Calculate the median and interquartile range for each metric within a candidate group, then name it provisionally: “high-progressive-pass midfielders” is more useful than “Cluster 3.” Review outliers separately; they may be versatile players, data errors, or genuinely unusual profiles.
Practical applications include:
- Shortlisting: Find domestic players near a club’s tactical role profile, then review minutes, injury history, contract status, and video.
- Academy development: Identify which technical and physical dimensions separate a youth player from the intended senior role.
- Recruitment gaps: Compare the club’s current squad with the distribution of available players.
- Opponent preparation: Map opposition players by role and identify substitutions that preserve a team’s playing pattern.
Never use t-SNE alone to make selection or contract decisions. Combine it with age, availability, language and relocation considerations, tactical fit, and qualitative reports. For model and dashboard reliability, maintain monitoring for missing feeds, delayed match data, and schema changes; principles from LLM application performance monitoring in India also apply to broader AI analytics operations.
Common mistakes to avoid
- Treating axes as meaningful football dimensions.
- Assuming every visible island is a real player archetype.
- Comparing raw totals across unequal minutes.
- Mixing positions without role controls.
- Using too many correlated metrics.
- Presenting a single seed as definitive evidence.
- Publishing player names without consent, data rights, or appropriate access controls.
A responsible operating model for Indian clubs
Document data provenance, metric definitions, minimum sample rules, and the exact software environment. Restrict personally sensitive data, especially medical and biometric information, to authorised staff. Give players and coaches an understandable explanation of how the visualisation informs—not determines—decisions.
Start with a small pilot: one competition, one season, a clearly defined role group, and a review panel of an analyst plus coaching staff. Measure whether the map improves shortlist quality or reduces review time. Expand only after the workflow produces stable, actionable results.
FAQs
Is t-SNE a clustering algorithm?
No. t-SNE creates a low-dimensional visual representation. If you need formal groups, apply clustering to the original or PCA-reduced features and use t-SNE only for visual inspection.
What perplexity should I use?
There is no universal value. Test several values based on dataset size and check whether neighbourhoods remain stable across seeds and preprocessing choices.
Can I compare players from different positions?
You can, but broad role differences may dominate the result. Separate role-specific analyses or include role-aware features and interpret cross-position comparisons cautiously.
What should a club do after finding a promising cluster?
Profile the group statistically, verify nearest neighbours in the source data, review video, confirm data quality, and evaluate the player against the club’s tactical and operational requirements.
How often should the map be refreshed?
Refresh it when the data distribution changes—typically each competition phase or season—and preserve previous versions so staff can distinguish real player development from a changed dataset.