Video summary

How to Build An Expected Goals Model 1: Data and Model

Main summary

Key takeaways

Educational

Summary

The lecture introduces the data and basic ideas behind building an expected goals (xG) model. The presenter explains why xG is useful, examines shot and goal locations, and develops two important features for predicting scoring probability: shot angle and distance from goal. The statistical model that combines these features is left for the next lecture.

What Expected Goals Measure

Expected goals estimate the probability that a shot of a particular type, taken from a particular situation, will result in a goal. The estimate is based on many previous shots and their outcomes. For example, a model can compare a shot’s location and other characteristics with similar shots in the data.

In this lecture, the focus is limited to open-play shots, excluding headers, free kicks, and penalties.

Why xG Is Useful

  • It adds context to a scoreline. A team can create more and better chances yet still lose, perhaps because of a goalkeeper’s performance or the randomness of finishing. An xG map can show the quality and locations of each team’s chances.
  • It can be more informative about future performance than goals alone. Goals are relatively rare, so short-term goal totals can be noisy. xG captures the quality of a larger set of chances. The lecture cites analysis showing that a team’s recent xG performance predicts its future performance better than recent goals do.
  • It can inform coaching and player decisions. Chance probabilities by location can help explain why some shooting positions are more promising than others.
  • It is a foundation for other models. The lecture presents xG as an early example of a model-based approach to football analysis, which can later support models of action impact and possession value.

Data and Visualizing Shots

The demonstration uses Wyscout event data from top divisions in England, Spain, Germany, France, and Italy. The presenter selects the relevant shots and uses two-dimensional histograms to map:

  • Shot counts: Most shots come from relatively close to goal, particularly around the penalty area. Very few come from far away or directly beside the goal.
  • Goal counts: Goals are more concentrated in areas close to goal than shots are.
  • Goal frequency by location: Dividing the number of goals in each location bin by the number of shots there gives the observed scoring proportion for that area.

These raw proportions are uneven and “pixelated.” A location where one of very few attempts resulted in a goal may appear unusually promising, even though that result could be due to chance. A useful model should smooth the noisy observations and identify broader, more reliable patterns.

Building an Explanatory Model

The presenter proposes that two geometric features help explain scoring probability:

  • Shot angle: The angle between lines from the shooting position to the two goalposts. A larger angle means the shooter can see more of the goal. It can be calculated from the shot’s coordinates and the goal’s width.
  • Distance to goal: In general, shots become less likely to score as the shooter moves farther away.

The plotted data show that scoring probability tends to rise with shot angle and fall with distance. Both features matter, and they are related: moving closer to goal often changes the angle as well.

Limitations to Keep in Mind

  • Sparse data can produce misleadingly high or low scoring rates at particular locations.
  • Event-data coordinates may reflect coding conventions or errors. The presenter notes that shots appear unusually rare directly on the edge of the penalty area, possibly because coders tend to place events just inside or outside the line.
  • Pitches can differ in size, so detailed models may need to account for pitch dimensions.
  • The lecture simplifies the problem to angle and distance as a starting point. The next step is to fit a statistical model that estimates whether a shot will become a goal.

Speakers and Sources Featured

  • David Sumpter — lecturer and presenter.
  • Michael Caley — cited for expected-goals visualizations and analysis comparing xG with future goals.
  • César Morales — cited for work applying shot-angle geometry to the design of the penalty area.
  • Wyscout — provider of the event data used in the demonstration.
  • Friends of Tracking — the educational project within which the lecture is presented.

Tobias, the lecturer’s dog, is mentioned in an aside but does not speak.

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video