Showing posts with label logistic regression. Show all posts
Showing posts with label logistic regression. Show all posts

Saturday, August 24, 2019

Returning Punts and Losing Field Position

Fans of collegiate and professional football teams have seen it. The opposing team punts. Arrival of the coverage unit is imminent as your return man situates to catch the ball. He shows nary a handwave, telling everyone there will be no fair catch on this punt. No, yours is an enterprising returner. Upon catching the punt, he will begin to explore the prospect of negative return yardage whilst attempting to evade the coverage unit. Perhaps, he will pick up some punctual blocks from his teammates or move quickly enough to elude would-be tacklers before reaching open grass and improving field position for your offense. Sometimes this risk produces minimal gains and on other occasions, the returns are huge. Yet, to the displeasure of fans and the hypertension of coaches, sometimes many yards are lost, and offenses start drives closer to their own endzone. 

There are other ways that field position is lost. I’m less interested in these, but we can examine them too. Punt returners can muff the catch or fumble the ball during the return. Although neither muffs nor fumbles guarantee lost field position, both create a risk for lost field position. Moreover, both risk turnovers--let alone the detriment to field position. Penalties. Specifically, the holding, block in the back, and clipping varieties, which can negate returns and start the offense closer to their own endzone. 

Who is to blame for lost return yardage? The ability of the return team to pressure the punter and the extent to which the coverage unit protects the punter. The skill of the punter to both focus and execute as well as the distance (and hangtime) of the punt matter, too. It is the punt returner who chooses to run toward his own endzone. It’s also on him if he muffs the catch, and he needs to protect the pigskin to prevent fumbles. Penalties just suck, I'm sorry. Nonetheless, regardless of how it occurs, lost field position is created by an interaction between individual players and their emergent units. One simple way we can look at who is responsible for lost field position on punt returns is intraclass correlations (ICCs; though my methods differ).

Our data are (primarily) 19,363 punts that were returned during 2002-18 NFL seasons (regular and some playoffs; holding-type penalties included). We include in the model return teams and coverage units both by season and across seasons to account for seasonal personnel changes and season-to-season consistency, respectively. Season itself was included to account for League-wide fluctuations in gameplay. In each model, we shall also account for the line of scrimmage and the punt yards. 

Table 1. ICCs of Team Units for ways Field Position is Lost on Punt Returns in NFL, 2002-18
On Punt Returns All Punts
Unit Negative Yards Muffs Fumbles Penalties Penalties
Returner 0.059 0.042 0.028 0 NA
Return Team by Season 0.001 0.014 0 0 0
Punt Team by Season 0.021 0.025 0 0.001 0.004
Punter 0.012 0.008 0 0 0
Return Team in all Seasons 0 0.001 0.006 0.001 0.001
Punt Team in all Seasons 0 0.005 0.008 0 0
Season 0.005 0.019 0 0.006 0.007
Unit R² (sum of ICCs) 0.098 0.114 0.042 0.009 0.012
Line of Scrimmage & Punt Yards R² 0.043 0.092 0.004 0.019 0.073
Total R² 0.141 0.206 0.046 0.028 0.085
Across all seasons, 721 non-muffed punts were returned for negative yardage, or 3.72% of returns, with an average of -3.53 yards (SD = 2.25). Table 1 contains ICCs for each unit. The ICC value means that 5.94% of negative yardage is due to some qualities of punt returners, 2.11% is due to some qualities of the punting teams, and 1.23% is due to the punter. In other words, the ICCs can be summed to obtain an approximate R². The effect of punting team is not statistically significant (p = 0.32) but the effect of punter tends to be (p = 0.07), and the effect of returner is (p < 0.001) (compared to models with each excluded). Together, the remaining factors account for 0.54 %. That only 9.82% of the responsibility for negative returns is meaningfully explained speaks to the stochastic nature of punt returns and special teams in general. 

Unsurprisingly, returners bare the most responsibility for muffs. However, the punting team and the return team appear to contribute to this meaningfully as well. Returners appear to be mostly responsible for fumbles. Penalties appear to be mostly random based on the ICCs all being < 1%. 

Summarily, the present report showed that punt returners carry the most responsibility for negative return yardage, but qualities of the punting team and punter are likely involved. Conceivably then, some punt returners should be more likely than others to have returns for negative yardage. In other words, a subset of returners may attempt to evade tacklers despite the risk of compromising field position for their offensive units. How such a tendency relates to punt return outcomes (e.g., yards gained or touchdowns) is a matter for future study. 



Methods
For analysis we’ll use generalized linear mixed models and specify binomial distribution. Essentially, we are estimating the likelihood that there is a return of negative yards, a muff, a fumble, or a penalty on a given punt and how much of that can be attributed to returners, punter, return teams, coverage units, and the season. Return teams and coverage units were examined by season and overall to account for seasonal personnel changes and season-to-season consistency, respectively. Season was included to account for League-wide trends in gameplay. We also include the punt spot and punt yards. For each GLMM, we'll use the icc() function of sjstats package in R to compute ICCs.

Bulleted below are definitions for each of the ways field position is lost by punt returners and units. 

  • I define negative returns as returns of ≤ -2 yards on an attempted return without a muff. Muffing should be should be considered separately from a decision to run into negative yardage. I set the threshold at -2 because I felt that returns of -1 yard could occur inadvertently, whereas ≤-2 yards are more likely the result of volitionally moving into the negative.
  • Muffs occur when the returner botches the catch. Muffs do not necessarily result in lost yardage, but they risk lost yardage and turnovers.
  • Fumbles occur when the returner loses possession of the ball during the return. Same caveats as muffs.
  • Penalties are holding, block in the back, and clipping penalties committed by the return team. We’ll look at penalties with and without considering returners, that is, on punt returns only (i.e., 19363 punts) and then on all punts (i.e., 41912 punts). This is because the play-by-play data only tell me when a returner was on the field for punt returns and fair catches and so we exclude returner from the model with all punts. 




Saturday, August 17, 2019

How Meteorological Conditions Affect Punting and Punt Outcomes

How does weather affect punting and punt outcomes? We know from prior studies that decreasing temperature is associated with reduced accuracy for field goals from the 25-yard line and farther.  Likewise, longer field goals tend be more accurate in the high altitude of Denver.  Regarding punts, there is evidence suggesting wind reduces punt yards. 

In short, we’re using 37,253 or so NFL punts from 2002-16. A weather data set culled from NFL Savant covers only 28,000 or so of those punts, through 2013, or about 75% of the data set. 
Figure 1. Average Punt Yards by Altitude

We can first see in Figure 1 that altitude has a limited effects on punt yards (PY) with the exception of the highest altitudes. The second highest altitude group includes Atlanta and Arizona, which average nearly 1 yard more on punts (p = 0.001; Atlanta is a dome) and Denver averages nearly 3 yards more per punt (p <0.001). This is consistent with findings on field goals. 

I used a generalized additive regression with smoothing splines to examine weather effects on punting. The punt spot (PS), wind (in MPH), temperature (Fahrenheit), and precipitation (%, 0-1) as well as all interactions between the meteorological variables were all fit with splines. I included the categorical variable for altitude instead of a smooth line for altitude because Denver distorts the altitude spline. I suppose I could have transformed the variable, but the laziness vice is king for the day. I also included a variable indicating if the punt was in a dome or open stadium. 

Figure 2. Modeling punt yards as a function of temperature, precipitation, and wind

As shown in Figure 2, weather appears to influence punt distance. Lower temperatures result in shorter punts. Wind appears to be most influential when precipitation is greatest. Maximum precipitation appears to reduce punts by about 3 yards, on average, compared to no precipitation. The influence of temperature is diminished when wind and precipitation increase. That punt distances are reduced in increasingly inclement meteorological conditions is consistent with the existing literature on field goals and punts in the NFL. The effect of the Denver altitude is consistent in this model, but the effect of Atlanta and Arizona is diminished likely because the model accounts for dome conditions. The upper rightmost panel is weird, though, perhaps because having only a few cases with higher wind speed influences this finding?

Figure 3. Punt return yards by altitude

There appears to be negligible effects of altitude on average punt return (PR) yards on punts that were actually returned (R2 < 0.001, that is R-squared not p!); see Figure 3. Not shown is an ecologically meaningless but statistically significant effect of temperature increasing PR yards on returned punts by about 0.13-yard for every 30° increase in temperature. Ah!, the frivolity that emerges from large data sets.

Figure 4. Punt outcomes by altitude. dd = defense downed/declared dead. fc = fair catch. oob = out of bounds. pr = punt return. tb = touchback.

It appears that there are more touchbacks in Denver, χ² = 92.15, df = 28, p < 0.001. Not much else to say here.

Figure 5. Secondary punt events by altitude. blk = blocked/tipped punt. fum = fumble. muff = returner muffed catch. pen = holding, blocking in back, or clipping penalty on return team. td = touchdown.

There appears to be more penalties in Atlanta and Arizona, but I am unsure why this is. Arizona had six seasons with 5 or fewer wins from 2002-13. ATL had three such seasons. All-around poor team play could have evidenced in more block in the back type penalties on punt returns. 


Figure 6. Punt outcomes as a function of temperature.

I used binary logistic regressions to assess the probability of several punt outcomes associated with several meteorological variables. The meaningful differences (to me) for outcomes due to temperature are between 25° and 75°. Specifically, there is a 5% greater probability of punts being declared dead or downed by the defense (DD) as it gets colder and 5% greater probability of punts being returned when it is warmer. 


Figure 7. Punt outcomes as a function of wind

For wind, I’m looking at the probability difference between no wind and 20mph. The probabilities of fair catches (FCs) decrease and DDs increase as it gets windier. This suggests to me that returners are less likely to even attempt to field the punt when it’s windier. OOBs also increase when it is windier. 

Figure 8. Punt outcomes as a function of precipitation

For precipitation, I’m looking at the change from none to maximum where there is a 5% less probability of a FC when it’s wetter, a 5% greater probability of TBs when it’s wetter, and a 5% greater probability of DD when it’s wetter. Together, these amount to there being fewer punt returns in wetter weather.

In short, the probabilities shown in Figures 6-8 demonstrate to me that punt returners are less inclined to even attempt catching a punt in colder and wetter conditions, and rightfully so. I’m unwilling, however, to conclude exactly the same for windier weather because [a] there are interactions between the meteorological variables not accounted for in these analyses; [b] the analysis accounts for the direction of neither the wind nor the punt; [c] steady winds and, more so, powerful wind gusts could dramatically alter the trajectory of a punt, and leave a return man far out of position. However, as shown above, windier, colder, and wetter conditions reduce punt distance meaning that the coverage unit is approaching the returner much quicker. 

Then I identified 7, 6, and 9 classes, respectively, for temperature, wind, and precipitation using an estimation-maximization procedure. I used these classes to examine the probabilities between meteorological variables and several secondary events: blocks, muffed catches, fumbles, penalties, turnovers, and TDs.  There was no difference in the distribution of PR TDs, fumbles, or turnovers between the classes of any meteorological variable (not shown). 


Figure 9. Muffs as a function of temperature and wind

For muffs, see Figure 9. It appears there is no difference in the distribution across precipitation (χ² = 12.1, p = 0.15; not shown) but the distribution does differ across wind (χ² = 15.9, p = 0.007) and temperature (χ² = 30.92, p < 0.001). Specifically, muffs increased in windier and colder conditions.


Figure 10. Blocked/tipped punts as a function of precipitation and wind

Shown in Figure 10 are blocked punts, which I’m wary of even broaching since it is such a rare event. There is no difference for temperature but there is a difference in the distribution across wind and precipitation. Blocks appear to be slightly less random when it is windier and wetter, but this could be due to adverse conditions affecting punt trajectory or increased pressure due to the expectations that punting is complicated by such weather conditions. However, we must be mindful that there are fewer samples at the meteorological extremes and the results very well could be spurious.


Figure 11. Block in the back, holding, or clipping penalties on the return team as a function of temperature

Distributions of block in the back, holding, or clipping penalties on the return team are no different for wind and precipitation. However, the distributions do differ across temperature such that penalties become more likely in warmer temperatures (χ² = 23.4 , p < 0.001). Penalties likely increase as temperature increases not because of some pressure exerted by warmer conditions per se but, rather, because punt returns are more likely as temperature increases. The odds of a penalty occurring on a punt that is returned are 4.6 times greater than on a punt with no return (z = 26.7, p < 0.001) whereas the odds of a penalty increase by about 0.004 for 1° increase in temperature (z = 3.03, p = 0.002), or by about 0.12 for an increase of 30°.

Summarily, very high altitudes increase punt yards. Colder, wetter, and windier weather reduce punt yards. There is a negligible influence of meteorological variables on punt return yards of returned punts. Punt returners, I subsume, are less likely to attempt to catch a punt during inclement weather. Fumbles, turnovers, and TDs appear to be stochastic and independent of the influence of meteorological conditions. Muffs, however, do appear to increase when it is colder and windier but not in greater precipitation. It seems blocked punts are slightly less random as precipitation and wind increase but these are the rarest of rare events. Penalties are slightly more likely to occur as temperature increases but this is likely due to there being more punt returns in warmer weather. So, that covers meteorology and punting with a healthy dose of chart gluttony. 

Wednesday, February 13, 2019

How do College Quarterbacks Perform After Throwing an Interception?

So, you’re watching the game on any day that is or should be Saturday and your QB throws an interception. Waves of fury and grief crash through your viscera. Moments earlier, the stodgy commentator was waxing anachronistically about the merits of establishing the running game but has since transitioned to divining the psychological state of your beloved QB. You will be haunted by the potential of him throwing an interception when your offense retakes the field. You are a fan, after all, and the potential for your QB to throw interceptions looms over your every daily activity and dreamscape from August to January, anyways. But what about your QB? The objective of this report is to examine how (collegiate) QBs perform after they throw an interception.

Did I conduct a literature review? Yes. We know that, in the NFL at least, team interception rates are weakly (r = 0.08)  to modestly correlated (0.27)  within seasons. Among individual NFL QBs, there is a modest relationship (0.27)  in year-to-year INT rates with the same team but it approaches negligibility after changing teams (0.10); but the best NFL QBs do consistently throw interceptions at a rate slightly less than the League average in a given year.  Other work shows there is a moderate relationship (0.43) between a QB’s INT rate and his true quarterbacking ability.  We also know that a QB’s year-to-year completion percentage is pretty stable with the same team (0.58)  but less so when changing teams (0.25).  Together, the prior research suggests that, [a] although INTs are wildly random, some QBs are inevitably more prone to throwing INTs and [b] that some QBs are more accurate than others. Given that interceptions are random, the majority of collegiate QBs are (relatively) well-practiced, and that completion percentage is moderately stable, we might expect that a QBs performance is minimally affected after throwing an INT. 

For the present study, we shall use a massive data set with plays from NCAA games in years 2001 and 2003 through mid-November 2018. I removed all offensive plays of FCS teams but retained plays with FBS offenses against FCS defenses. Because I do not own a supercomputer or even a particularly powerful machine, I removed all plays without a pass attempt or sack in the description. This means we cannot account for QB scrambling ability/threat in the analysis. 

Table 1. Descriptive Data for Cumulative Interceptions
Cumulative INTs Completion % Interception % Sack % Completions Interceptions Sacks Pass Attempts Drop Backs
0 0.609 0.029 0.052 286334 13575 25643 470461 496104
1 0.589 0.032 0.057 101146 5511 10372 171704 182076
2 0.577 0.037 0.060 28199 1785 3130 48866 51996
3 0.563 0.040 0.061 6828 490 785 12121 12906
4 0.560 0.034 0.064 1435 87 174 2564 2738
5 0.559 0.043 0.051 209 16 20 374 394
6 0.455 0.036 0.068 25 2 4 55 59
7 0.333 0.000 0.000 2 0 0 6 6

Nevertheless, let our dependent variable—how we’ll measure performance—be whether a given play was a completion. Table 1 shows that completion % decreases after throwing 1, 2, and 3 interceptions but it remains constant at about 56% after the second INT (ignoring the 6th and 7th INTs which are rare scenarios). Now, these differences would be significant statistically because of our sample size; however, is the difference in completion rate between 1 and 2 INTs meaningful? I would wager that a difference of 5 percentage-points, such as between no interceptions (~61%) and 3 interceptions (~56%), might be meaningful. Likewise, interceptions become slightly more likely once a QB has thrown an interception. Although a multitude of factors contribute to QB performance after an interception, some of those factors probably led to cumulative INTs, anyhow. Many of the factors are simply unknowable including individual tendencies of each QB such as perseverance or resilience, self-confidence, and intelligence. There are also in-game factors such as down, distance, field position, time remaining, score differential, team and opponent quality, and others. Let us account for these on-field factors (as best we can), which are listed and defined at the close of the post.

We’ll use a logistic regression to predict the probability of a completion before and after throwing an interception, while controlling for all of the factors we can. Indeed, we should actually use a (generalized) mixed model because there are many plays from each season, game, team unit, and QB—as well as plays for each team unit and QB nested within seasons and games, and QBs nested within teams—but my computer would explode and burn down this apartment building, destroying all of my stuff. 

Our model indicated no main effect of cumulative interceptions on completion %. However, there is a significant effect of cumulative interceptions when approaching the defense’s endzone (i.e., when the offense is getting closer to scoring). Interestingly, QBs were slightly but significantly more likely to complete passes when close to the defense endzone after having thrown an interception. The estimated (marginal) probabilities of completion can be seen in Figure 1. The difference in expected completion % is, at most, 1-2 percentage points from one INT to the next. To me, this is virtually meaningless. Additionally, both poorer offensive lines and better defenses reduced the likelihood of completing a pass. Having thrown more interceptions in the season to date also reduced the likelihood of a completion, albeit meagerly. Completions were less likely on 3rd and 4th downs and as games progressed. QBs at home were slightly more likely to complete passes. The model has a mess of interactions between down, distance, etc., but a summary can be seen here.
Figure 1. Expected Marginal Completion percentage as a function of field position, Time Remaining in the Game, and Cumulative INTs, with other covariates held constant at mean.

Summarily, this report provides evidence that quarterback performance is essentially unchanged after throwing an interception. Critical in-game factors, home-status of QB, and team quality were controlled for in the analysis. The sample size was >700K and the slight effects that were observed are probably ecologically meaningless. That said, I would like to note that if performance is reduced after throwing an INT, say, in a subset of QBs or for some QBs in some games, it is at least somewhat related to psychological processes. For example, some neuroscience research suggests that performance would be impaired immediately following an error, such as on the QB’s next few passes after an INT.  This research also suggests that for well-learned tasks, such as quarterbacking in our case, errors in performance may be more related to activity in the prefrontal cortex (located in your forehead, above your eyes, until about your temple) than in the circuitry of the brain largely responsible for volitional motor movements (but, in the study, electrodes were not placed on the circuits more responsible for learned motor movements). The prefrontal cortex is related to, among myriad other activity, planning and decision making. Nonetheless, you are probably best-served by disregarding that atavistic commentator. Your precious quarterback will probably be fine, but it will never feel that way for you and that is part of what makes college football so delightful.

Acknowledgements
On behalf of all of us here at POTH, I would like to thank my colleague CK for her insight and suggestions.
Covariates
Below are variables included in the analysis. OS% and ONR% are thought to reflect offensive line quality. Passes dropped by receivers and yards after the catch are not described in the play-by-play and thus there was no way to measure the ability receiving units. DS%, DC%, and DPD% are thought to reflect defensive unit quality. Specifically, %DS might reflect D-Line pass rush quality; DPD% might reflect ability of LBs and secondary to defend passes; and DC% might reflect the overall ability of a defensive unit to prevent successful passing. 

  • Offensive Sack % (OS%): the proportion of the offensive team’s QB drop-backs that resulted in sacks, for the season on which a play occurred. Adjusted for sack % and FBS/FCS-status of opposing defenses.
  • Offensive Negative Rush % (ONR%): the proportion of the offensive team’s rushing plays that resulted in negative yards, for the season. Adjusted for negative rush % and FBS/FCS-status of opposing defenses.
  • Defensive Sack % (DS%): the proportion of the defensive team’s opponents’ QB drop-backs that results in sacks, for the season on which a play occurred. Adjusted for sack % of all offenses a defense faced.
  • Defensive Completion % (DC%): the season average completion percentage against the defense on a given play. Adjusted for completion % of opposing defense.
  • Defensive Passes Defensed % (DPD%): the proportion of passes broken up or intercepted by the defense, in the season on which the play occurred. Bayesian-average adjustment for quantity of passes faced by the defense in that season based on FBS averages.
  • Cumulative Interceptions, Game: the quantity of interceptions thrown by the offensive team in the current game up to but not including a given play. That is, if a team had thrown 2 interceptions before a given play, the cumulative value would be 2 for this play. This way we can compute the probability of throwing an interception when none have been thrown. I had to settle for cumulation at the team level instead of the QB level because, again, I do not have a supercomputer.
  • Cumulative Interceptions, Season: the quantity of interceptions thrown by the offensive team in the season up to but not including a given play. Used team to be consistent with above.
  • Whether QB is playing at home or away.
  • A variety of on-field factors noted in the main body of the text. 




Wednesday, January 16, 2019

Icing the Kicker in NCAA Football 2005-18

In gridiron football, the icing the kicker phenomenon is thought to occur when the defending team calls a TO just before the ball is snapped on a FG attempt (FGA) that could tie, win, or otherwise sway the outcome of the game in favor of the kicking team. The motivation for calling the TO is that it could somehow disturb, or ‘ice’ the kicker in a way that he will be more likely to miss the FGA. 

Other authors have endeavored to examine icing the kicker. Some have reported that, in the NFL, calling a TO before a FG does not reduce the likelihood of making a FGA, whether controlling for FGA length or not. Other studies suggest suggests there is indeed an effect of reducing likelihood on longer NFL FGAs that is absent on shorter FGAs, when controlling for FGA length and other factors. At the collegiate level, it appears that icing the kicker may be effective on longer FGs; specifically, greater than 45 yards.  However, this study had a small sample of iced kicks.

Table 1. Descriptive Statistics for NCAA FGAs 2005-18
This Many Attempted Fields Goals were
Quarter FG% uFG% Attempted Made Blocked Home Attempts Last 2min Attempts ≤15s after TO Attempts
1st 0.727 0.753 7051 5123 248 3559 1129 468
2nd 0.702 0.732 11473 8049 475 5957 4413 3264
3rd 0.740 0.766 6629 4906 224 3417 1033 425
4th & OT 0.715 0.747 7176 5129 308 3764 1566 1754
TOTAL 0.718 0.747 32329 23207 1255 16697 8141 5911
We here at POTH sought to reexamine icing the kicker at the collegiate level using a much larger data set. This includes 32,329 FGAs from NCAA Division I FBS vs FBS and FBS vs FCS games from 2005 through mid-November 2018. Table 1 has the breakdowns of some data we’ll refer to throughout. The last 2 minutes refers to FGAs during the last two minutes of quarters 1 through 4 and any FGA occurring in OT. 

Let us start with blocked FGAs, though. Notably, as seen in Table 1, blocked FGAs were more likely to occur in the 2nd quarter and 4th quarter and OT (χ² = 12.3, p = 0.007)—the situations in games most relevant to icing the kicker. Longer FGAs were more likely to be blocked regardless of the quarter (p < 0.001). FGAs were also more likely to be blocked in the last 2 minutes of quarters and OT, but especially in the last 2 minutes of the 4th quarter and OT (p = 0.06). For these reasons, we shall include in our analyses only unblocked FGAs. This leaves 31,074 FGAs for analysis.

Table 2. Proportional Statistics for NCAA FGAs 2005-18
Proportion of Field Goals Made
Quarter % Blocked Home Team Away Tem Last 2min Before Last 2m ≤15s after TO No TO Before
1st 0.035 0.743 0.710 0.731 0.751 0.726 0.752
2nd 0.041 0.714 0.688 0.677 0.746 0.680 0.742
3rd 0.034 0.755 0.724 0.743 0.765 0.701 0.769
4th & OT 0.043 0.728 0.700 0.676 0.754 0.694 0.749
OVERALL 0.039 0.757 0.736 0.725 0.754 0.721 0.752

About 74.6% of (unblocked) FGAs are made. Figure 1 shows that FG% declines as the length of the FGA increases. There is some variation in FG% between quarters, with 3rd-quarter FGAs being most successful. Only differences between the 3rd and 2nd (p = 0.03) and the 4th and 3rd (p = 0.03) are significant when we account for length of the FGA, which is, by far, the most significant predictor of FG success. Longer FGAs are less successful at all points in the game. 
Figure 1. Likelihood of Making a FGA, by Length (using binomial smooth)
FGAs by the home team (75.6%) are about 2.8% more likely to be made than FGAs by the road team (73.5%) (χ² = 17.9, p < 0.001). When controlling for FGA length and quarter, home FGAs are 6.6% more likely to be made (p < 0.001). However, this advantage of home FG% is relatively constant at all FGA lengths. That is, home-team FG kickers tend to be slightly more successful than road-team kickers on FGAs of any length, and at any point in the game. 

What about the FG% in the last 2 minutes of quarters, when icing the kicker usually occurs? Table 1 shows that it clearly drops in the 4th quarter and OT (in the 2nd too). This drop in FG% in the last two minutes is, however, diminished when controlling for FGA length, quarter, and home/away (p = 0.52). It should be noted that FGAs in the last 2 minutes of the 2nd and 4th quarters are 1-2 yards longer than FGAs at other times in the game (ps < 0.002). 

How do the stakes of the game effect FG%? The opportunity to tie the game seems to have a general effect of increasing the likelihood of making a FG (p = 0.05). Otherwise, though, there is no effect of stakes on FG% when controlling for length, quarter, home/away, and being in the last 2 minutes or not. 

FGAs 15 seconds or less after a TO are made 72.1% of the time whereas other FGAs are made 75.2% of the time (χ² = 24, p < 0.001). Now, this is just if any TO is called; that is, by the offense, the defense, or some other TO that was not attributed to either team in the data. Really, we have a variable that indicates whether the TO was called by the offense, the defense, was unattributed, or if no TO was called. If we were to continue the analysis as we have been doing it, we would examine a four-way interaction between quarter, last 2 minutes or not, stakes, and who called the TO before or not. Four-way interactions are messy. And three of those variables have four levels. We should do something else.
Figure 2a. FG% by TO TypeFigure 2b. FGA Length by TO Type
Let us narrow our focus to FGAs in the last 2 minutes of the 4th quarter and OT where the offense can either tie the game or take the lead with a FG, which leaves 2,173 FGAs for analysis. In the two figures we see that iced FGAs (i.e., those after a defense TO) [a] are the least successful, at ~70%, but [b] are also, on average, the longest FGAs in this game situation. Thus, when we model the likelihood of making a FG, while accounting for length and home/away, there is no effect of icing the kicker (p = 0.24). Like, icing the kicker has no statistically differentiable effect of decreasing the likelihood of making longer FGAs (p = 0.25). However, the estimated marginal probabilities in the figure below suggest that the likelihood of iced FGAs declines slightly more at longer distances, although, again, this is not statistically significant. 
Figure 3. Estimated probabilities of FG% by length and who called TO in last 2 minutes of 4th & OT

Whereas we have used the raw yardage value for FGA length, the one previous study of icing the kicker in NCAA football split length into ‘bins’: distances of 18-25 yards, 26-35 yards, 36-45 yards, and >45 yards. The author of the previous study used only data from 2017-18 and found that of 38 iced FGAs in the last 2 minutes of the 4th quarter and OT, only 26% were made. If I examine only data from 2017-18, I find these same numbers (38 iced FGAs, 10 made, 28 missed). Below, using all data, I went ahead and show the FG% for each of these length-bins by who called the TO, for the sake of comparison across studies. The quantities of FGAs are shown parenthetically. Longer iced FGAs appear to be made lower rates.
Figure 4. Proportion of FGAs Made by TO Type by Yardage Bins used in Dalen (2018)
Summarily, the present report examined icing the kicker in NCAA football. This study used a sizable data set which would enhance the generalizability of the findings. However, the primary analysis indicated there was no effect of icing the kicker. Additional examination suggested that there might be an effect of icing the kicker at FGAs longer than 45 but such a conclusion is limited by there being fewer FGAs attempted from these lengths (i.e., smaller sample) and the variability of success at increasing lengths. Likewise, other potentially influential factors such as meteorological conditions, team FG kicking/defensing quality, and on-field activity were not accounted for in the analysis. NCAA football coaches should continue utilizing icing the kicker so they may endure the rancor of punditry, boosters, delusional fans, etc., when their teams lose games on last-second field goals.  

Sunday, January 21, 2018

Examination of Success Thresholds in College Football

Anyone familiarizing themselves with gridiron football analytics will quickly acquaint with success rate. Success is widely defined by an offense gaining 40-50% of yards to go on 1st down, 60-70% on 2nd down, and 100% on 3rd and 4th downs; preventing gains of said percentages defines success for defenses. Counting all the successful plays in a drive, game, or season and dividing by the total quantity of plays for that period yields the success rate. Personally, I am more interested in whether an activity was productive, unproductive, or counterproductive but I’ll save that for another post. However, curious as to how the thresholds for success may have been established and if there are nuances to current definitions, I examine it here. 

Myself and others before me, suspect that success rate is derived from traditional football notions of ‘staying ahead of the chains’ or ‘setting up for third and short’. Popularized by Football Outsiders, success, by our definition—like many gridiron analytic concepts—can be traced at least to the mid-1980s when it was outlined on p. 69 of the Hidden Game of Football. The authors used 40%, 60%, and 100% to benchmark ‘wins’ and ‘failures’, as well as a derived qualitative measure of success that awards more credit (i.e., >1-point) for big plays and penalizes turnovers and lost yardage (i.e., negative points). 


How do we determine what is a successful play? That is not, how success is defined according to X-amount of yardage gained or lost on a given play, per se, but how X-amount of yardage on a given play portends future success when aggregated from many, many plays in a similar context. Given this notion, let us define success as a play occurring on a drive that ends in scoring either a TD or FG.


Let us define success another way, too: a first down occurring on or after a given play on a drive (or series, in this case, really). For example, take a 2nd and 8; if there is a first down obtained on that play or subsequent play in the drive, that 2nd and 8 would be considered as having occurred on a successful drive (or, series). Alternatively, imagine a 1st and 10 which is, say, the fifth play of a drive and occurs after obtaining having at least one first down on the drive; if there is not a first down obtained on that 1st and 10 or any play later in the drive, that 1st and 10 would be considered as having occurred on an unsuccessful drive (or series).


For data, I have all pass and rush plays from games played by Division 1 college football teams from 2005-13—995,895 plays. For each play, I included an indicator of whether the play occurred on a scoring drive and whether there was another first down on that drive. As these are a binary variables, indicating yes or no, logistic regression is suitable. As the predictor variable we will use yards gained on a play divided by the yards to go on the play. This way we can say gaining X% of yards on a given down down is the threshold of success. Oh, so since we’re using college data, we’ll use the thresholds utilized by Football Study Hall of 50%, 70%, and 100% on 1st, 2nd, and 3rd and 4th downs, respectively.


Using logistic regression and ROC curves, we identify thresholds for the proportion of yards gained on each down that correctly predicts both the maximum quantity of plays on successful drives while minimizing the quantity of plays on unsuccessful drive wrongly predicted as successful (in our data set). This becomes our threshold of success. Figure 1 shows the success thresholds from these analyses for scoring drives in purple, drives with another first down in green, and the commonly applied success thresholds in orange.

Figure 1. Thresholds for Success

That the thresholds for 3rd and 4th down are essentially identical for scoring and first downs is unsurprising because scoring requires gaining at least the yards to go. The disparity in thresholds on second downs is also intuitive. It suggests that gaining a greater portion of the yards to go on 2nd down portends a more successful drive. The lower threshold for scoring drives on 1st down is interesting, however. It may be that obtaining 40% of the yards to on first downs typically setups a 2nd and 6 with offensive being in neither a definitive rush nor definitive pass situation. This, in turn, could conceivably lead to future success and the disparity here compared to the commonly used threshold. 


I was curious also how field position affects success. Let us focus only first downs, for convenience. I computed whether each play was a success based on the threshold for scoring drives described above; we’ll call this the fixed threshold. A mixture model was used to segment the field into 8 segments. Several logistic regression models were blended to generate thresholds for each segment, which we’ll call blended thresholds.i This is shown in Figure 2. Yard line 1-9 is closest to the defense’s end zone. The bottom row of panels are successful plays based on the fixed threshold and the blended threshold on the top row.
On the X-axis are Yes or No to indicate whether a play actually occurred on a scoring drive or not. Green indicates a play was predicted to occur on a non-scoring drive and orange indicates a play was predicted to occur on a scoring drive. We can see the fixed threshold emerged because it accurately predicts so many plays on unsuccessful drives in opponent’s territory.

Figure 2. Comparing Fixed and Blended Success Thresholds by Field Position on First Down

Summarily, this report showed that, at least in college football, success thresholds are relatively constant whether success is defined as a drive ending in a score or whether there is a first down after a given play. Secondarily, the report provides evidence that statistically-derived success thresholds vary by field position, at least on first down. Thus, future work should examine how adjusting thresholds by field position affects the valuation of player and team performance when using success rates.






iTo do this, I averaged the threshold from three logistic regressions. For each group I obtained thresholds from three logistic regressions with the following subset of the data: [a] plays in each field position segment, [b] all plays in each field position segment and all plays from field positions closer to the defense's end zone, and [c] all plays in each field position segment and all plays from field positions farther from the defense's end zone.