scieee AI-readable full text Open interactive document viewer

Adjusting Expected Goals (xG) and Shots on Target for Game Context in Soccer

Skripnikov, Andrey; Cemek, Ahmet; Gillman, David

Abstract

With advancements in soccer analytics, considerable attention has been devoted to developing sophisticated measures for quality scoring opportunities —such as expected goals and expected threat. Far less effort, however, has gone toward adjusting these and other performance metrics for game context, including factors like score differential, red cards, and time remaining. It is well known that certain match scenarios prompt systematic tactical adjustments—for example, teams often defend more conservatively when leading or playing shorthanded. In this study, we employ generalized additive mixed-effects models (GAMMs) on minute-by-minute match event data from Europe’s five major leagues to quantify how such contextual variables influence offensive production metrics like shots on target and expected goals. This approach allows us to estimate global effects of contextual variables on offensive production, while also allowing for league-specific distinctions to be captured via random effects. We then use these estimates to adjust observed team statistics, projecting each team’s performance onto a standardized baseline scenario—a tied home game played at even manpower. We believe the resulting adjusted measures to better isolate underlying team quality by removing distortions arising from transient tactical responses to game state.

Full text

Supplementary Materials for ”Adjusting Expected Goals (xG) and Shots on Target for Game Context in Soccer” Andrey Skripnikov1∗ , Ahmet Cemek1, and David Gillman1 1New College of Florida, Sarasota, FL, USA October 11, 2025 1 Goodness-of-Fit Tests 1.1 Expected Goals (xG) as response The DHARMa goodness-of-fit tests for the models we used for expected goals (xG) clearly showcase the inadequacy of Gaussian approaches (top row of Figure 1). The quantile–quantile plots show large deviations from assumptions, and the residuals-vs-fitted plots reveal systematic skew. For Tweedie models, the quantile–quantile plots demonstrate much better alignment with the theoretical quantiles, while the residuals-vs-fitted plots show a more adequate fit, especially for the log-link Tweedie model. ∗Corresponding author: askripniko[email protected] 1 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.984 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.88 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Log−Link Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Log−Link Tweedie 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Inverse−Link Tweedie Figure 1: Visual representation of DHARMa model diagnostics with expected goals (xG) as the response variable. Diagnostics include quantile–quantile plots (left) and residuals-vs-fitted plots (right) for four models: Gaussian (top left), log-link Gaussian (top right), log-link Tweedie (bottom left), and inverse-link Tweedie (bottom right). 2 1.2 Shots on Target as Response After running models for each league–season combination with shots on target as the response variable, Figures 2 and 3 display the DHARMa goodness-of-fit results. Figure 2 shows Holm-adjusted p-values for various tests (uniformity, outlier, overdispersion, and zero-inflation), while Figure 3 shows the quantile–quantile and residuals-vs-fitted plots. Both figures point to Poisson-family models (Poisson and Negative Binomial) as superior to Gaussian models, with a slight edge to Negative Binomial, particularly for outlier tests. Figure 2: P-values from DHARMa model diagnostics for each model type predicting shots on target. Tests include uniformity (Kolmogorov–Smirnov), outlier, overdispersion, and zero-inflation tests. The null hypothesis in each test is that model assumptions are satisfied. 3 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.984 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.984 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Log−Link Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0.65993 Deviation n.s. Outlier test: p= 0.22118 Deviation n.s. Dispersion test: p= 0.032 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Poisson 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0.49806 Deviation n.s. Outlier test: p= 0.57962 Deviation n.s. Dispersion test: p= 0.04 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Negative Binomial Figure 3: Visual representation of DHARMa diagnostics with shots on target as the response variable. Diagnostics include quantile–quantile plots (left) and residuals-vs-fitted plots (right) for four models: Gaussian (top left), log-link Gaussian (top right), Poisson (bottom left), and Negative Binomial (bottom right). 4