What you are actually predicting
A useful set-piece model should answer a precise question, not simply label a corner or free kick “successful”. Define the outcome before collecting data. For example:
- Immediate shot: Did the set piece produce a shot within 15 seconds?
- Shot quality: What was the expected-goals value of the first shot?
- Goal: Did the attacking team score within 30 seconds?
- Tactical outcome: Did the team retain possession, win a second ball, or force a defensive error?
For local matches, begin with one target—usually shot creation or goal probability. Goals are relatively rare, so a model trained only on goals may learn unstable patterns. Recording intermediate outcomes gives coaches more feedback and provides enough positive examples for a small dataset.
Keep the unit of analysis consistent. One row should represent one corner, direct free kick, indirect free kick, or long throw. Record the match, minute, scoreline, venue, attacking team, defending team, and outcome. Do not combine corners and free kicks until you have tested whether they behave similarly.
Build a realistic local-football dataset
You do not need an elite tracking system to start. A structured video-tagging workflow can produce a valuable first dataset. Use a shared spreadsheet or a small annotation application, and create a fixed codebook so different volunteers record events in the same way.
Capture features available before the delivery, including:
- Set-piece type, side, distance, and approximate angle
- Scoreline, match minute, home/away status, and weather
- Number of attackers and defenders in the penalty area
- Delivery zone, intended target, and whether the kick is inswinging or outswinging
- Server identity, preferred foot, and recent availability
- Heights or aerial ratings of likely targets, if reliably recorded
- Defensive structure: zonal, man-marking, hybrid, or unclear
- Short-corner option, screen, decoy run, or rehearsed movement
- Outcome labels such as shot, goal, clearance, second-ball recovery, and turnover
Avoid using information that becomes available after the event. A feature such as “number of defenders who missed the header” may explain the outcome, but it cannot be used to predict it before the kick. This is data leakage, and it can make a model look excellent in testing while failing in matches.
For Indian local competitions, conditions can vary sharply between grounds. Store pitch dimensions where possible, surface type, lighting, rain, wind, and whether the match was played on grass or artificial turf. These fields may appear minor, but they help separate tactical decisions from environmental effects.
Prepare the data before choosing a neural network
Start with a baseline rather than assuming a neural network is automatically superior. Compare it with a goal-rate estimate, logistic regression, or a gradient-boosted tree. If the neural network does not improve calibrated predictions or coaching usefulness, use the simpler model.
A practical preparation pipeline includes:
1. Remove duplicates and impossible records. Check that event times, teams, and set-piece types are valid.
2. Handle missing values explicitly. “Unknown defensive structure” is different from “zonal defence”. Add a missingness indicator where appropriate.
3. Encode categories. One-hot encoding works for small datasets; embeddings can help only when you have enough examples per category.
4. Scale numeric variables. Standardise distance, angle, player counts, and recent rates using training-set statistics only.
5. Group rare players and tactics. A local club may have too few observations for individual player embeddings to generalise.
6. Split by match or date. Never place events from the same match in both training and test sets.
If you are new to model design, review customizable neural network architectures for beginners before selecting layers. For most local datasets, a small multilayer perceptron with two or three hidden layers is a sensible starting point. Use ReLU activations, dropout cautiously, early stopping, and a sigmoid output for binary outcomes. A compact model is easier to train, explain, and run on an ordinary laptop.
Train and validate without fooling yourself
Use a time-based evaluation design. Train on earlier matches, validate on the next period, and test on the most recent matches. This reflects deployment, where the model predicts future fixtures rather than randomly selected historical events.
Track more than accuracy. If only 5% of set pieces lead to goals, a model that predicts “no goal” every time can appear accurate while being useless. Report:
- Log loss or Brier score for probability quality
- Precision-recall AUC for rare positive outcomes
- Calibration to test whether predicted 20% events happen roughly 20% of the time
- Recall at a practical threshold if analysts need to flag high-value routines
- Lift over baseline for specific delivery zones or defensive systems
Use bootstrapped confidence intervals because local datasets are often small. Evaluate separately by set-piece type, home and away status, weather, and opponent defensive structure. A model that performs well overall but fails on corners may still be misleading for tactical planning.
Calibration is particularly important. Coaches need to know whether a routine has a 12% or 35% chance of producing a shot—not merely which routine receives the highest score. A calibration plot, reliability table, and a plain-language explanation are more useful than a single accuracy figure.
Turn predictions into coaching decisions
The model should support a set-piece process, not replace the coach. Before a match, generate a short report showing:
- The predicted shot and goal probability for each planned routine
- The conditions under which the routine works best
- The likely defensive response and confidence level
- The recommended fallback if the first delivery is blocked
- The number of historical examples behind each estimate
For example, the system might find that a near-post corner creates more shots against zonal defences but loses value in heavy wind. That is a coaching hypothesis to test, not an automatic instruction. Run controlled trials in training, tag the results, and update the model after competitive matches.
Use explainability tools carefully. Feature importance, partial-dependence plots, and counterfactuals can reveal that delivery location or attacker numbers drive predictions. They do not prove causation. Pair every model finding with video review and feedback from the set-piece coach.
A practical technical stack in India
A small team can build the first version with Python, pandas, scikit-learn, and PyTorch or TensorFlow. Store raw annotations separately from cleaned features, keep model versions in a repository, and log every training run. Use a simple dashboard—such as Streamlit—for match reports rather than building a full platform immediately.
If video files remain on a club computer, a local-first workflow can reduce bandwidth and privacy risks. Guidance on how to deploy large language models locally is not specific to football, but its principles around local inference, access control, and hardware planning are relevant to an on-premise analytics setup. For larger academies, lightweight GPU machines may speed up video processing, but the prediction model itself should remain small enough to run offline.
Protect player data. Limit access to identifiable performance records, obtain consent where required, retain only necessary information, and avoid publishing individual “failure” scores. Aggregate reporting is often sufficient for opponents and public-facing analysis.
Common failure modes
- Too little data: Treat early predictions as exploratory and show uncertainty.
- Inconsistent tagging: Re-train annotators and audit a sample of events every month.
- Target leakage: Freeze all input fields at the moment before delivery.
- Random train-test splits: Split by match and time to measure real-world performance.
- Overfitting player identities: Test whether the model works for new players and opponents.
- Optimising for goals alone: Include shots, second balls, and possession outcomes.
- Ignoring selection bias: Planned routines are not random; record when a routine was available but not chosen.
- Presenting probabilities as certainty: Include sample size, confidence, and a human review step.
A 30-day implementation plan
Week 1: Define the outcome, create the annotation manual, and tag 100–200 historical events. Review disagreements with coaches.
Week 2: Expand the dataset, clean fields, and build logistic-regression and baseline-rate benchmarks.
Week 3: Train a compact neural network with time-based validation. Compare calibration and lift, not just accuracy.
Week 4: Pilot the dashboard in training, collect coach feedback, and document where predictions were useful or wrong. Do not automate tactical selection until the system has survived several match cycles.
The strongest local solution is usually not the most complex one. Reliable tagging, honest evaluation, calibrated probabilities, and a coach-friendly workflow will create more value than a large neural network trained on inconsistent video labels. Indian academies and community clubs can start with modest hardware, improve their data discipline, and scale only when the evidence justifies it.