Showing posts with label R. Show all posts
Showing posts with label R. Show all posts

Sunday, September 8, 2019

How do NFL Kickers Age?

I was watching games on Saturday while chatting with another college football diehard. We were both enamored by the ongoing failure (relatively speaking) that is field goal kicking at perennial powerhouse Alabama. Juxtaposed against their otherwise prolific success, conjecture proceeded about the underlying cause(s) of ‘Bama’s FG kicking woes over the years. 

FG kicking troubles pervade the college game, frustrating fans, and our conjecturous chat led me to wonder if and how NFL kicking is better than college kicking. This led me to wonder if kickers just get better (or more consistent and reliable) as they get older. We did some Google searching but couldn’t find any NFL kicker aging curves for accuracy. So, we made our own.
Figure 1. Histogram of career lengths for NFL kickers 1960-2018


First, we obtained a bunch of NFL kicker data from the wonderful resource known as PFR. This includes season-by-season data for 369 NFL kickers from 1960 through present. These kickers made 33558 of 45777 FGs (73.3%) and 54658 of 56369 (97%) extra points. Based on the distribution of career-lengths shown in Figure 1, there were concerns that the large amount of kickers with 3 or fewer NFL seasons would skew the analysis. Our concerns were reinforced when we looked at Figure 2. 
Figure 2. Mean Field Goal % by length of career in seasons for NFL Kickers 1960-2018


Kickers with 3 or fewer NFL seasons have notably lower career FG% than kickers with lengthier careers. This in itself is not surprising, but it would confound the interpretation of the data. The lower accuracy of kickers with 3 or fewer seasons might lead to exaggerated year-to-year increases in accuracy in the early stage of the kicker career. To better convey this, displayed in Figure 3 is the average FG% in each season of the careers of kickers with ≤3 NFL seasons and >3 seasons. 
Figure 3. Mean FG% in each season for NFL kickers 1960-2018 with career lengths of <4 or >3 seaons


Figure 3 also suggests that FG% increases linearly as kickers age; as if kickers just keep getting more accurate. However, recall the smaller quantities of kickers with lengthier careers seen in Figure 1. The continued increases in accuracy by kickers with lengthier careers may be obscuring the declining accuracy in later seasons of kickers with shorter careers. This is exactly what is shown in Figure 4.
Figure 4. Thick lines are LOESS curves of the average FG% in each season of careers of NFL kickers 1960-2018 with various career lengths; fainter, thinner lines are raw mean FG% in each season


There is a group of ‘super agers’, kickers with careers longer than 16 years, whose annual FG% seems to level out and remain constant around their 12th season—which is about when kickers with careers of 11-16 seasons begin to experience slight declines in accuracy. Likewise, kickers with 11-16 seasons appear to peak around their 7th season—which is about when kickers with careers of 4-10 seasons start to decline.

Let us look at the data another way. Figure 5 contains average FG% through the course of the career normalized such that 0.50 (on the X-axis) represents a season halfway through the course of the kicker career. Figure 5 shows that, aside from kickers with ≤3 NFL seasons, NFL kickers start to experience a downward trend in accuracy about 75% of the way through their career. 
Figure 5. Career length is normalized such that 0.00 = rookie season, 1.00 = final season, and 0.50 = halfway through career. Thick lines are LOESS curves of the average FG% in each season of careers of NFL kickers 1960-2018 with various career lengths; fainter, thinner lines are raw mean FG% in each season


Summarily, the (slightly manipulated) raw data indicate that NFL kickers experience declines in accuracy late in their career (Figure 5). However, using the percent of the way through the career (as in Figure 5) does not conduce toward a prospective aging curve for NFL kickers. That is, a predictive model could not know beforehand how long a kicker’s career will be. In other words, future analyses will need to model an NFL kicker aging curve based on seasons in the League (or perhaps age). Future analyses should also account for era—FG kicking has improved dramatically over the years—and FG accuracy by distance. PAT% might also be informative (more so since 2015). Likewise, some measure of consistency (e.g., coefficient of variation) may provide a more alternative measure (than accuracy) of kicker performance.  

Monday, August 26, 2019

Punt Returner Personalities: The Enterprising Risk-Taker, the Dependable Risk-Averter, and the Consummate-Moderate

Let us examine how the frequency with which punt returners produce negative yardage can be viewed as a sort of personality trait. Moreover, we’ll examine how such a trait can provide insight into on-field performance. The data set (initially) includes 19,363 punt returns from 2002-18 NFL seasons, both regular season and playoffs. We’ll examine the career data of punt returners who spent at least one season as the primary returner for a team. Without snap count data for the whole data set, I defined primary returner as anyone who returned the most punts (plus fair catches) for a team in at least one season; within-season ties for a team were permitted (i.e., could have more than one primary returner from one team in a season). That totaled 227 punt returners, with a range of 8 to 331 career punt returns (only, not fair catches). I then excluded returners with 30 or fewer career returns to have a decent sample size of returns for each returner. This leaves 170 primary returners who, together, returned 16,234 punts and have a median of 79 career returns (25th percentile = 47; 75th = 119). 

The first step was classifying returners based on tendency for negative yards. I started with the proportion of career returns for negative yards based on the findings of the previous post. My criterion for negative return yardage is ≤ -2, excluding returns with muffed catches. I set the threshold at -2 because I felt that returns of -1 yard could occur inadvertently, whereas ≤-2 yards are more likely the result of volitionally moving into the negative. Then I binned returners into three groups using cutoffs at the 33rd and 66th percentiles, or 2.1% and 4.43% of career returns being negative, respectively. I thought this segmentation would provide three groups of risk-preference: risk-averse, moderate, and risk-takers.
Several variables were selected to examine how this conceptualization of risk-preference might relate to on-field performance. For each returner, I computed the following variables to explore relationships between risk-preference and on-field outcomes.

  • % of career returns >6 yards 
  • % of career returns with a TD
  • % of career returns + fair catches that were fair catches
  • % of career returns where the returner muffed the catch
  • % of career returns where the returner fumbled the ball
  • % of career returns where there was an illegal blocking penalty called against a member of the return team


Figure 1. Median career punt returns by risk-preference group

Figure 1 shows that moderate returners have the highest median number of career returns, followed by risk-takers and then risk-averters. One possible explanation is that guys with fewer opportunities to return punts may be more averse to risk, perhaps, in hopes of securing roster spots. There is some potential evidence for this assertion in the data. Returners who were ever a primary returner were less likely to call for a fair catch (32.7%; 7900 of 24129 returns and fair catches) than those who were never a primary (35.3%; 1710 on 4844), χ² = 11.9, p < 0.001. That is, I’m saying that guys who are less experienced returning punts may be more cautious.


Figure 2. For visualization, I split returners into groups above and at or below the median of 53.5% of career returns being >6 yards (2 groups). I split returners into groups above or at and below the median of 1.18% of career returns with a TD (2 groups).

Figure 2 indicates how likely a returner in each risk-preference group is to return for more yards than would be expected by chance alone and return for a TD. Indeed, compared to the risk-averse, moderates (p = 0.02) and risk-takers (p < 0.001) returned a higher proportion of their career punt returns for TDs. Likewise, compared to risk-takers, the risk-averse (p < 0.001) and moderates (p = 0.04) returned a higher proportion of their career returns for >6 yards. If we exclude negative returns and returns for TDs and look at the % of returns >6 yards, the difference between risk-averters (57.9%) and risk-takers (54.7%) is significant (p = 0.01); but moderates (54.8%) are no different than risk takers (p = 0.43).

Figure 3 shows probabilities and standard errors of other variables by risk-preference. Moderates (p < 0.001) and the risk-averse (p < 0.001) had a higher proportion of fair catches than the risk-takers. This suggests that risk-takers were less likely to call for a fair catch, but this is largely my own conjecture as we cannot account for whether returners had more punts out of bounds, downed, declared dead, or touchbacks. Also, we cannot account for how often returners returned a punt when they should have called for a fair catch. 


Figure 3. Proportion of career punt returns are fair catches, muffed, fumbled, or had a holding-type penalty, by group.

Compared to the risk-takers, the risk-averse had significantly fewer returns with penalties (p = 0.005), and there was a similar trend for the moderates (p = 0.12). This finding is potentially due to some quality of risk-takers because the results are essentially unchanged if we control for number of career returns, career average return yards, and career touchdown return %. Likewise, using all of the data, penalties are called less often on negative returns (9.9%; 7 of 720) than positive returns (12.3%; 2292 of 18643), χ² = 3.83, p = 0.05 (penalties enforced and declined are included).

There were no significant group differences in the proportion of fumbles (ps > 0.24) and muffs (ps > 0.32). If we control for the number of career returns, average yards, and TD%, the risk-averse tend to have fewer fumbles than the moderates (p = 0.13) but otherwise, the proportions of fumbles and muffs are unchanged. 

There is a shortcoming of my thesis to consider. I am assuming that returners who are more often tackled for a loss of ≤-2 yards (i.e., negative returns) on returns are also more likely to run into the negative area overall. Based on the available data we cannot determine if this is the case. It may be that the risk-averse and moderates run into the negative just as often, but the risk-takers just are more likely to be tackled after running into negative return yardage space. A caveat to this is that risk-taking returners tended to be less likely to call for fair catches. However, only if we have data indicating that the risk-takers are more likely to forgo fair catches when the coverage unit is closing in on them can it be demonstrated that they are more likely to take risks.

Importantly, these findings show that there appears to be a balance to productive punt returning: Risk-takers may produce more TDs, but they also produce return yardage less consistently, whereas risk-averters may produce return yardage more dependably, they also produce fewer TDs. Ultimately, punt returners who take risks in moderation are probably the most productive in that they consistently produce decent return yardage while still producing TDs at a relatively high rate.


Methods 
We used generalized linear models (GLMs), specifying Poisson distributions, to compare on-field outcomes between the risk-preference groups. There were six GLMs. The dependent variable was the quantity of career returns with a given outcome, for each returner. The independent variable was risk-preference. The DV was offset by the total career punt returns (or punt returns + fair catches for the model of fair catches, this yields a proportional value. The variables are described below.

  • The proportion of career returns >6 yards. I used >6 yards because 7 is the median of 90% of the punt returns in the data (range of -1 to 32) and it is a decent guess at the return yards we would expect to occur randomly. Then I split returners into groups above and at or below the median of 53.5% of career returns were >6 yards (2 groups). In other words, returners with a lower proportion of returns >6 yards are more often returning punts below what we would expect based on chance alone.
  • The proportion of career returns with a TD. I split returners into groups above or at and below the median of 1.18% of career returns with a TD (2 groups). My thought was that risk takers should return TDs at a comparable rate as the other groups, despite having more negative returns.
  • The proportion of career fair catches, which is the number of fair catches divided by the sum of fair catches and returns. Ideally, the number of fair catches would be divided by the number of punts on which the returner was on the field to return the punt. Nevertheless, the thought here is, risk-takers should be less likely to call for a fair catch overall. 
  • The proportion of career punt returns where the returner muffed the catch. I included this as a measure of conscientiousness. That is, can the returner do the most critical and fundamental part of successful punt returning: catch the ball?
  • The proportion of career punt returns where the returner fumbled the ball on the return. I included this as another measure of conscientiousness, perhaps, although fumbles tend to be random events. 
  • The proportion of career punt returns where there was a block in the back or illegal block called against a member of the return team. 




Saturday, August 24, 2019

Returning Punts and Losing Field Position

Fans of collegiate and professional football teams have seen it. The opposing team punts. Arrival of the coverage unit is imminent as your return man situates to catch the ball. He shows nary a handwave, telling everyone there will be no fair catch on this punt. No, yours is an enterprising returner. Upon catching the punt, he will begin to explore the prospect of negative return yardage whilst attempting to evade the coverage unit. Perhaps, he will pick up some punctual blocks from his teammates or move quickly enough to elude would-be tacklers before reaching open grass and improving field position for your offense. Sometimes this risk produces minimal gains and on other occasions, the returns are huge. Yet, to the displeasure of fans and the hypertension of coaches, sometimes many yards are lost, and offenses start drives closer to their own endzone. 

There are other ways that field position is lost. I’m less interested in these, but we can examine them too. Punt returners can muff the catch or fumble the ball during the return. Although neither muffs nor fumbles guarantee lost field position, both create a risk for lost field position. Moreover, both risk turnovers--let alone the detriment to field position. Penalties. Specifically, the holding, block in the back, and clipping varieties, which can negate returns and start the offense closer to their own endzone. 

Who is to blame for lost return yardage? The ability of the return team to pressure the punter and the extent to which the coverage unit protects the punter. The skill of the punter to both focus and execute as well as the distance (and hangtime) of the punt matter, too. It is the punt returner who chooses to run toward his own endzone. It’s also on him if he muffs the catch, and he needs to protect the pigskin to prevent fumbles. Penalties just suck, I'm sorry. Nonetheless, regardless of how it occurs, lost field position is created by an interaction between individual players and their emergent units. One simple way we can look at who is responsible for lost field position on punt returns is intraclass correlations (ICCs; though my methods differ).

Our data are (primarily) 19,363 punts that were returned during 2002-18 NFL seasons (regular and some playoffs; holding-type penalties included). We include in the model return teams and coverage units both by season and across seasons to account for seasonal personnel changes and season-to-season consistency, respectively. Season itself was included to account for League-wide fluctuations in gameplay. In each model, we shall also account for the line of scrimmage and the punt yards. 

Table 1. ICCs of Team Units for ways Field Position is Lost on Punt Returns in NFL, 2002-18
On Punt Returns All Punts
Unit Negative Yards Muffs Fumbles Penalties Penalties
Returner 0.059 0.042 0.028 0 NA
Return Team by Season 0.001 0.014 0 0 0
Punt Team by Season 0.021 0.025 0 0.001 0.004
Punter 0.012 0.008 0 0 0
Return Team in all Seasons 0 0.001 0.006 0.001 0.001
Punt Team in all Seasons 0 0.005 0.008 0 0
Season 0.005 0.019 0 0.006 0.007
Unit R² (sum of ICCs) 0.098 0.114 0.042 0.009 0.012
Line of Scrimmage & Punt Yards R² 0.043 0.092 0.004 0.019 0.073
Total R² 0.141 0.206 0.046 0.028 0.085
Across all seasons, 721 non-muffed punts were returned for negative yardage, or 3.72% of returns, with an average of -3.53 yards (SD = 2.25). Table 1 contains ICCs for each unit. The ICC value means that 5.94% of negative yardage is due to some qualities of punt returners, 2.11% is due to some qualities of the punting teams, and 1.23% is due to the punter. In other words, the ICCs can be summed to obtain an approximate R². The effect of punting team is not statistically significant (p = 0.32) but the effect of punter tends to be (p = 0.07), and the effect of returner is (p < 0.001) (compared to models with each excluded). Together, the remaining factors account for 0.54 %. That only 9.82% of the responsibility for negative returns is meaningfully explained speaks to the stochastic nature of punt returns and special teams in general. 

Unsurprisingly, returners bare the most responsibility for muffs. However, the punting team and the return team appear to contribute to this meaningfully as well. Returners appear to be mostly responsible for fumbles. Penalties appear to be mostly random based on the ICCs all being < 1%. 

Summarily, the present report showed that punt returners carry the most responsibility for negative return yardage, but qualities of the punting team and punter are likely involved. Conceivably then, some punt returners should be more likely than others to have returns for negative yardage. In other words, a subset of returners may attempt to evade tacklers despite the risk of compromising field position for their offensive units. How such a tendency relates to punt return outcomes (e.g., yards gained or touchdowns) is a matter for future study. 



Methods
For analysis we’ll use generalized linear mixed models and specify binomial distribution. Essentially, we are estimating the likelihood that there is a return of negative yards, a muff, a fumble, or a penalty on a given punt and how much of that can be attributed to returners, punter, return teams, coverage units, and the season. Return teams and coverage units were examined by season and overall to account for seasonal personnel changes and season-to-season consistency, respectively. Season was included to account for League-wide trends in gameplay. We also include the punt spot and punt yards. For each GLMM, we'll use the icc() function of sjstats package in R to compute ICCs.

Bulleted below are definitions for each of the ways field position is lost by punt returners and units. 

  • I define negative returns as returns of ≤ -2 yards on an attempted return without a muff. Muffing should be should be considered separately from a decision to run into negative yardage. I set the threshold at -2 because I felt that returns of -1 yard could occur inadvertently, whereas ≤-2 yards are more likely the result of volitionally moving into the negative.
  • Muffs occur when the returner botches the catch. Muffs do not necessarily result in lost yardage, but they risk lost yardage and turnovers.
  • Fumbles occur when the returner loses possession of the ball during the return. Same caveats as muffs.
  • Penalties are holding, block in the back, and clipping penalties committed by the return team. We’ll look at penalties with and without considering returners, that is, on punt returns only (i.e., 19363 punts) and then on all punts (i.e., 41912 punts). This is because the play-by-play data only tell me when a returner was on the field for punt returns and fair catches and so we exclude returner from the model with all punts. 




Saturday, August 17, 2019

How Meteorological Conditions Affect Punting and Punt Outcomes

How does weather affect punting and punt outcomes? We know from prior studies that decreasing temperature is associated with reduced accuracy for field goals from the 25-yard line and farther.  Likewise, longer field goals tend be more accurate in the high altitude of Denver.  Regarding punts, there is evidence suggesting wind reduces punt yards. 

In short, we’re using 37,253 or so NFL punts from 2002-16. A weather data set culled from NFL Savant covers only 28,000 or so of those punts, through 2013, or about 75% of the data set. 
Figure 1. Average Punt Yards by Altitude

We can first see in Figure 1 that altitude has a limited effects on punt yards (PY) with the exception of the highest altitudes. The second highest altitude group includes Atlanta and Arizona, which average nearly 1 yard more on punts (p = 0.001; Atlanta is a dome) and Denver averages nearly 3 yards more per punt (p <0.001). This is consistent with findings on field goals. 

I used a generalized additive regression with smoothing splines to examine weather effects on punting. The punt spot (PS), wind (in MPH), temperature (Fahrenheit), and precipitation (%, 0-1) as well as all interactions between the meteorological variables were all fit with splines. I included the categorical variable for altitude instead of a smooth line for altitude because Denver distorts the altitude spline. I suppose I could have transformed the variable, but the laziness vice is king for the day. I also included a variable indicating if the punt was in a dome or open stadium. 

Figure 2. Modeling punt yards as a function of temperature, precipitation, and wind

As shown in Figure 2, weather appears to influence punt distance. Lower temperatures result in shorter punts. Wind appears to be most influential when precipitation is greatest. Maximum precipitation appears to reduce punts by about 3 yards, on average, compared to no precipitation. The influence of temperature is diminished when wind and precipitation increase. That punt distances are reduced in increasingly inclement meteorological conditions is consistent with the existing literature on field goals and punts in the NFL. The effect of the Denver altitude is consistent in this model, but the effect of Atlanta and Arizona is diminished likely because the model accounts for dome conditions. The upper rightmost panel is weird, though, perhaps because having only a few cases with higher wind speed influences this finding?

Figure 3. Punt return yards by altitude

There appears to be negligible effects of altitude on average punt return (PR) yards on punts that were actually returned (R2 < 0.001, that is R-squared not p!); see Figure 3. Not shown is an ecologically meaningless but statistically significant effect of temperature increasing PR yards on returned punts by about 0.13-yard for every 30° increase in temperature. Ah!, the frivolity that emerges from large data sets.

Figure 4. Punt outcomes by altitude. dd = defense downed/declared dead. fc = fair catch. oob = out of bounds. pr = punt return. tb = touchback.

It appears that there are more touchbacks in Denver, χ² = 92.15, df = 28, p < 0.001. Not much else to say here.

Figure 5. Secondary punt events by altitude. blk = blocked/tipped punt. fum = fumble. muff = returner muffed catch. pen = holding, blocking in back, or clipping penalty on return team. td = touchdown.

There appears to be more penalties in Atlanta and Arizona, but I am unsure why this is. Arizona had six seasons with 5 or fewer wins from 2002-13. ATL had three such seasons. All-around poor team play could have evidenced in more block in the back type penalties on punt returns. 


Figure 6. Punt outcomes as a function of temperature.

I used binary logistic regressions to assess the probability of several punt outcomes associated with several meteorological variables. The meaningful differences (to me) for outcomes due to temperature are between 25° and 75°. Specifically, there is a 5% greater probability of punts being declared dead or downed by the defense (DD) as it gets colder and 5% greater probability of punts being returned when it is warmer. 


Figure 7. Punt outcomes as a function of wind

For wind, I’m looking at the probability difference between no wind and 20mph. The probabilities of fair catches (FCs) decrease and DDs increase as it gets windier. This suggests to me that returners are less likely to even attempt to field the punt when it’s windier. OOBs also increase when it is windier. 

Figure 8. Punt outcomes as a function of precipitation

For precipitation, I’m looking at the change from none to maximum where there is a 5% less probability of a FC when it’s wetter, a 5% greater probability of TBs when it’s wetter, and a 5% greater probability of DD when it’s wetter. Together, these amount to there being fewer punt returns in wetter weather.

In short, the probabilities shown in Figures 6-8 demonstrate to me that punt returners are less inclined to even attempt catching a punt in colder and wetter conditions, and rightfully so. I’m unwilling, however, to conclude exactly the same for windier weather because [a] there are interactions between the meteorological variables not accounted for in these analyses; [b] the analysis accounts for the direction of neither the wind nor the punt; [c] steady winds and, more so, powerful wind gusts could dramatically alter the trajectory of a punt, and leave a return man far out of position. However, as shown above, windier, colder, and wetter conditions reduce punt distance meaning that the coverage unit is approaching the returner much quicker. 

Then I identified 7, 6, and 9 classes, respectively, for temperature, wind, and precipitation using an estimation-maximization procedure. I used these classes to examine the probabilities between meteorological variables and several secondary events: blocks, muffed catches, fumbles, penalties, turnovers, and TDs.  There was no difference in the distribution of PR TDs, fumbles, or turnovers between the classes of any meteorological variable (not shown). 


Figure 9. Muffs as a function of temperature and wind

For muffs, see Figure 9. It appears there is no difference in the distribution across precipitation (χ² = 12.1, p = 0.15; not shown) but the distribution does differ across wind (χ² = 15.9, p = 0.007) and temperature (χ² = 30.92, p < 0.001). Specifically, muffs increased in windier and colder conditions.


Figure 10. Blocked/tipped punts as a function of precipitation and wind

Shown in Figure 10 are blocked punts, which I’m wary of even broaching since it is such a rare event. There is no difference for temperature but there is a difference in the distribution across wind and precipitation. Blocks appear to be slightly less random when it is windier and wetter, but this could be due to adverse conditions affecting punt trajectory or increased pressure due to the expectations that punting is complicated by such weather conditions. However, we must be mindful that there are fewer samples at the meteorological extremes and the results very well could be spurious.


Figure 11. Block in the back, holding, or clipping penalties on the return team as a function of temperature

Distributions of block in the back, holding, or clipping penalties on the return team are no different for wind and precipitation. However, the distributions do differ across temperature such that penalties become more likely in warmer temperatures (χ² = 23.4 , p < 0.001). Penalties likely increase as temperature increases not because of some pressure exerted by warmer conditions per se but, rather, because punt returns are more likely as temperature increases. The odds of a penalty occurring on a punt that is returned are 4.6 times greater than on a punt with no return (z = 26.7, p < 0.001) whereas the odds of a penalty increase by about 0.004 for 1° increase in temperature (z = 3.03, p = 0.002), or by about 0.12 for an increase of 30°.

Summarily, very high altitudes increase punt yards. Colder, wetter, and windier weather reduce punt yards. There is a negligible influence of meteorological variables on punt return yards of returned punts. Punt returners, I subsume, are less likely to attempt to catch a punt during inclement weather. Fumbles, turnovers, and TDs appear to be stochastic and independent of the influence of meteorological conditions. Muffs, however, do appear to increase when it is colder and windier but not in greater precipitation. It seems blocked punts are slightly less random as precipitation and wind increase but these are the rarest of rare events. Penalties are slightly more likely to occur as temperature increases but this is likely due to there being more punt returns in warmer weather. So, that covers meteorology and punting with a healthy dose of chart gluttony. 

Wednesday, January 16, 2019

Icing the Kicker in NCAA Football 2005-18

In gridiron football, the icing the kicker phenomenon is thought to occur when the defending team calls a TO just before the ball is snapped on a FG attempt (FGA) that could tie, win, or otherwise sway the outcome of the game in favor of the kicking team. The motivation for calling the TO is that it could somehow disturb, or ‘ice’ the kicker in a way that he will be more likely to miss the FGA. 

Other authors have endeavored to examine icing the kicker. Some have reported that, in the NFL, calling a TO before a FG does not reduce the likelihood of making a FGA, whether controlling for FGA length or not. Other studies suggest suggests there is indeed an effect of reducing likelihood on longer NFL FGAs that is absent on shorter FGAs, when controlling for FGA length and other factors. At the collegiate level, it appears that icing the kicker may be effective on longer FGs; specifically, greater than 45 yards.  However, this study had a small sample of iced kicks.

Table 1. Descriptive Statistics for NCAA FGAs 2005-18
This Many Attempted Fields Goals were
Quarter FG% uFG% Attempted Made Blocked Home Attempts Last 2min Attempts ≤15s after TO Attempts
1st 0.727 0.753 7051 5123 248 3559 1129 468
2nd 0.702 0.732 11473 8049 475 5957 4413 3264
3rd 0.740 0.766 6629 4906 224 3417 1033 425
4th & OT 0.715 0.747 7176 5129 308 3764 1566 1754
TOTAL 0.718 0.747 32329 23207 1255 16697 8141 5911
We here at POTH sought to reexamine icing the kicker at the collegiate level using a much larger data set. This includes 32,329 FGAs from NCAA Division I FBS vs FBS and FBS vs FCS games from 2005 through mid-November 2018. Table 1 has the breakdowns of some data we’ll refer to throughout. The last 2 minutes refers to FGAs during the last two minutes of quarters 1 through 4 and any FGA occurring in OT. 

Let us start with blocked FGAs, though. Notably, as seen in Table 1, blocked FGAs were more likely to occur in the 2nd quarter and 4th quarter and OT (χ² = 12.3, p = 0.007)—the situations in games most relevant to icing the kicker. Longer FGAs were more likely to be blocked regardless of the quarter (p < 0.001). FGAs were also more likely to be blocked in the last 2 minutes of quarters and OT, but especially in the last 2 minutes of the 4th quarter and OT (p = 0.06). For these reasons, we shall include in our analyses only unblocked FGAs. This leaves 31,074 FGAs for analysis.

Table 2. Proportional Statistics for NCAA FGAs 2005-18
Proportion of Field Goals Made
Quarter % Blocked Home Team Away Tem Last 2min Before Last 2m ≤15s after TO No TO Before
1st 0.035 0.743 0.710 0.731 0.751 0.726 0.752
2nd 0.041 0.714 0.688 0.677 0.746 0.680 0.742
3rd 0.034 0.755 0.724 0.743 0.765 0.701 0.769
4th & OT 0.043 0.728 0.700 0.676 0.754 0.694 0.749
OVERALL 0.039 0.757 0.736 0.725 0.754 0.721 0.752

About 74.6% of (unblocked) FGAs are made. Figure 1 shows that FG% declines as the length of the FGA increases. There is some variation in FG% between quarters, with 3rd-quarter FGAs being most successful. Only differences between the 3rd and 2nd (p = 0.03) and the 4th and 3rd (p = 0.03) are significant when we account for length of the FGA, which is, by far, the most significant predictor of FG success. Longer FGAs are less successful at all points in the game. 
Figure 1. Likelihood of Making a FGA, by Length (using binomial smooth)
FGAs by the home team (75.6%) are about 2.8% more likely to be made than FGAs by the road team (73.5%) (χ² = 17.9, p < 0.001). When controlling for FGA length and quarter, home FGAs are 6.6% more likely to be made (p < 0.001). However, this advantage of home FG% is relatively constant at all FGA lengths. That is, home-team FG kickers tend to be slightly more successful than road-team kickers on FGAs of any length, and at any point in the game. 

What about the FG% in the last 2 minutes of quarters, when icing the kicker usually occurs? Table 1 shows that it clearly drops in the 4th quarter and OT (in the 2nd too). This drop in FG% in the last two minutes is, however, diminished when controlling for FGA length, quarter, and home/away (p = 0.52). It should be noted that FGAs in the last 2 minutes of the 2nd and 4th quarters are 1-2 yards longer than FGAs at other times in the game (ps < 0.002). 

How do the stakes of the game effect FG%? The opportunity to tie the game seems to have a general effect of increasing the likelihood of making a FG (p = 0.05). Otherwise, though, there is no effect of stakes on FG% when controlling for length, quarter, home/away, and being in the last 2 minutes or not. 

FGAs 15 seconds or less after a TO are made 72.1% of the time whereas other FGAs are made 75.2% of the time (χ² = 24, p < 0.001). Now, this is just if any TO is called; that is, by the offense, the defense, or some other TO that was not attributed to either team in the data. Really, we have a variable that indicates whether the TO was called by the offense, the defense, was unattributed, or if no TO was called. If we were to continue the analysis as we have been doing it, we would examine a four-way interaction between quarter, last 2 minutes or not, stakes, and who called the TO before or not. Four-way interactions are messy. And three of those variables have four levels. We should do something else.
Figure 2a. FG% by TO TypeFigure 2b. FGA Length by TO Type
Let us narrow our focus to FGAs in the last 2 minutes of the 4th quarter and OT where the offense can either tie the game or take the lead with a FG, which leaves 2,173 FGAs for analysis. In the two figures we see that iced FGAs (i.e., those after a defense TO) [a] are the least successful, at ~70%, but [b] are also, on average, the longest FGAs in this game situation. Thus, when we model the likelihood of making a FG, while accounting for length and home/away, there is no effect of icing the kicker (p = 0.24). Like, icing the kicker has no statistically differentiable effect of decreasing the likelihood of making longer FGAs (p = 0.25). However, the estimated marginal probabilities in the figure below suggest that the likelihood of iced FGAs declines slightly more at longer distances, although, again, this is not statistically significant. 
Figure 3. Estimated probabilities of FG% by length and who called TO in last 2 minutes of 4th & OT

Whereas we have used the raw yardage value for FGA length, the one previous study of icing the kicker in NCAA football split length into ‘bins’: distances of 18-25 yards, 26-35 yards, 36-45 yards, and >45 yards. The author of the previous study used only data from 2017-18 and found that of 38 iced FGAs in the last 2 minutes of the 4th quarter and OT, only 26% were made. If I examine only data from 2017-18, I find these same numbers (38 iced FGAs, 10 made, 28 missed). Below, using all data, I went ahead and show the FG% for each of these length-bins by who called the TO, for the sake of comparison across studies. The quantities of FGAs are shown parenthetically. Longer iced FGAs appear to be made lower rates.
Figure 4. Proportion of FGAs Made by TO Type by Yardage Bins used in Dalen (2018)
Summarily, the present report examined icing the kicker in NCAA football. This study used a sizable data set which would enhance the generalizability of the findings. However, the primary analysis indicated there was no effect of icing the kicker. Additional examination suggested that there might be an effect of icing the kicker at FGAs longer than 45 but such a conclusion is limited by there being fewer FGAs attempted from these lengths (i.e., smaller sample) and the variability of success at increasing lengths. Likewise, other potentially influential factors such as meteorological conditions, team FG kicking/defensing quality, and on-field activity were not accounted for in the analysis. NCAA football coaches should continue utilizing icing the kicker so they may endure the rancor of punditry, boosters, delusional fans, etc., when their teams lose games on last-second field goals.  

Sunday, January 21, 2018

Examination of Success Thresholds in College Football

Anyone familiarizing themselves with gridiron football analytics will quickly acquaint with success rate. Success is widely defined by an offense gaining 40-50% of yards to go on 1st down, 60-70% on 2nd down, and 100% on 3rd and 4th downs; preventing gains of said percentages defines success for defenses. Counting all the successful plays in a drive, game, or season and dividing by the total quantity of plays for that period yields the success rate. Personally, I am more interested in whether an activity was productive, unproductive, or counterproductive but I’ll save that for another post. However, curious as to how the thresholds for success may have been established and if there are nuances to current definitions, I examine it here. 

Myself and others before me, suspect that success rate is derived from traditional football notions of ‘staying ahead of the chains’ or ‘setting up for third and short’. Popularized by Football Outsiders, success, by our definition—like many gridiron analytic concepts—can be traced at least to the mid-1980s when it was outlined on p. 69 of the Hidden Game of Football. The authors used 40%, 60%, and 100% to benchmark ‘wins’ and ‘failures’, as well as a derived qualitative measure of success that awards more credit (i.e., >1-point) for big plays and penalizes turnovers and lost yardage (i.e., negative points). 


How do we determine what is a successful play? That is not, how success is defined according to X-amount of yardage gained or lost on a given play, per se, but how X-amount of yardage on a given play portends future success when aggregated from many, many plays in a similar context. Given this notion, let us define success as a play occurring on a drive that ends in scoring either a TD or FG.


Let us define success another way, too: a first down occurring on or after a given play on a drive (or series, in this case, really). For example, take a 2nd and 8; if there is a first down obtained on that play or subsequent play in the drive, that 2nd and 8 would be considered as having occurred on a successful drive (or, series). Alternatively, imagine a 1st and 10 which is, say, the fifth play of a drive and occurs after obtaining having at least one first down on the drive; if there is not a first down obtained on that 1st and 10 or any play later in the drive, that 1st and 10 would be considered as having occurred on an unsuccessful drive (or series).


For data, I have all pass and rush plays from games played by Division 1 college football teams from 2005-13—995,895 plays. For each play, I included an indicator of whether the play occurred on a scoring drive and whether there was another first down on that drive. As these are a binary variables, indicating yes or no, logistic regression is suitable. As the predictor variable we will use yards gained on a play divided by the yards to go on the play. This way we can say gaining X% of yards on a given down down is the threshold of success. Oh, so since we’re using college data, we’ll use the thresholds utilized by Football Study Hall of 50%, 70%, and 100% on 1st, 2nd, and 3rd and 4th downs, respectively.


Using logistic regression and ROC curves, we identify thresholds for the proportion of yards gained on each down that correctly predicts both the maximum quantity of plays on successful drives while minimizing the quantity of plays on unsuccessful drive wrongly predicted as successful (in our data set). This becomes our threshold of success. Figure 1 shows the success thresholds from these analyses for scoring drives in purple, drives with another first down in green, and the commonly applied success thresholds in orange.

Figure 1. Thresholds for Success

That the thresholds for 3rd and 4th down are essentially identical for scoring and first downs is unsurprising because scoring requires gaining at least the yards to go. The disparity in thresholds on second downs is also intuitive. It suggests that gaining a greater portion of the yards to go on 2nd down portends a more successful drive. The lower threshold for scoring drives on 1st down is interesting, however. It may be that obtaining 40% of the yards to on first downs typically setups a 2nd and 6 with offensive being in neither a definitive rush nor definitive pass situation. This, in turn, could conceivably lead to future success and the disparity here compared to the commonly used threshold. 


I was curious also how field position affects success. Let us focus only first downs, for convenience. I computed whether each play was a success based on the threshold for scoring drives described above; we’ll call this the fixed threshold. A mixture model was used to segment the field into 8 segments. Several logistic regression models were blended to generate thresholds for each segment, which we’ll call blended thresholds.i This is shown in Figure 2. Yard line 1-9 is closest to the defense’s end zone. The bottom row of panels are successful plays based on the fixed threshold and the blended threshold on the top row.
On the X-axis are Yes or No to indicate whether a play actually occurred on a scoring drive or not. Green indicates a play was predicted to occur on a non-scoring drive and orange indicates a play was predicted to occur on a scoring drive. We can see the fixed threshold emerged because it accurately predicts so many plays on unsuccessful drives in opponent’s territory.

Figure 2. Comparing Fixed and Blended Success Thresholds by Field Position on First Down

Summarily, this report showed that, at least in college football, success thresholds are relatively constant whether success is defined as a drive ending in a score or whether there is a first down after a given play. Secondarily, the report provides evidence that statistically-derived success thresholds vary by field position, at least on first down. Thus, future work should examine how adjusting thresholds by field position affects the valuation of player and team performance when using success rates.






iTo do this, I averaged the threshold from three logistic regressions. For each group I obtained thresholds from three logistic regressions with the following subset of the data: [a] plays in each field position segment, [b] all plays in each field position segment and all plays from field positions closer to the defense's end zone, and [c] all plays in each field position segment and all plays from field positions farther from the defense's end zone.

Sunday, February 19, 2017

2016 WNBA FG% Distribution by Shot Location

So this is a first for POTH: I am posting twice in one day, or twice in one uninterrupted span of wakefulness. However, it is a diminutive post. Below is a shot chart for the 2016 WNBA regular season with field goal percentage. Greener means higher shooting percentage, navy-er means lower percentage. 

Rather unexpectedly, I was able to amass some data rather quickly and it happened that the shot locations were included. I got excited as I often do when there are discoveries at 00:18. Of course, the chart would would be more informative if the hexagons were sized according to the quantity of shots therein (e.g., here). But it's late and my belly aches.

Chart 1: Distribution of FG% by Shot Location in the WNBA, 2016

A previous post revealed some differences in WNBA compared to NBA league-wide aggregate statistics. I am now able to address these and other topics.

FBS vs FCS Score Differentials Equated to FBS vs FBS Score Differentials

Every matchup is unique. Either team could win. Although related to the outcomes of the other matchups of either team, the outcome of any one matchup is somewhat independent of the others. This notion underlies the nature of competition, the allure of sports betting, and the precedence for retold stories of unlikely winners. For football, because of its small sample size relative to other games, this notion underlies the complexity of numerating many activities on the gridiron and is, to some extent, the topic of this post. 

A recent undertaking at work portends a new analytic technique: observed-score linking and equating. I will undoubtedly seek guidance from our expert colleagues, but, of course, I prefer to be informed before that day is upon us. Linking and equating have distinct definitions, applications, and procedures but I will refer to these casually as equating. Equating allows us to generate uniform score-ranges between sections or items belonging to different versions of a single assessment, two unique assessments, or an old and a new version. 


More practically, consider the ACT, for example. Let us imagine that ACT Inc (the ACT developer) develops 20 versions of the ACT Reading Section. ACT Inc needs the scores for each version to be equitable so that a 36 is always a 36. Of the imaginary 20 versions, let us focus on Versions 6 and 12, or V6 and V12, for short. So, to test these versions, ACT Inc has 200 freshmen in college complete both versions. Say, 100 freshmen completed V12 in the first test session and V6 in the second whereas the other 100 freshmen completed V6 in the first and V12 in the second session. Afterwards, ACT Inc realizes that the average score for V6 is 18.5 and the average for V12 is 20.5, whoops. However, the average for all tests completed in the first session is 20.4 and all tests completed in the second session is 20.3, so ACT Inc knows that the disparity in Version-scores is not due to sequence of test administration. Likewise, because the same freshmen completed both versions, the 2-point disparity in Version-scores is not due to differences in the test-takers. ACT Inc must conclude that the disparity is due to differences in V6 and V12. Then, ACT Inc could use equating procedures to develop uniform scores to ensure little Johnny sets realistic standards for his future based on an ACT Reading Version 12 score of 30 instead of the inflated 36 it would have been without equating.


Here, I use equating to generate equivalency score-differentials for interdivisional college football games. That is, a 35-point win (or, +35 score differential) by a FBS team over a FCS team, for instance, is equivalent to what differential in an FBS versus FBS matchup. Let us relate this to the above example. This analysis would get restrictively complex if we sought to equate scores between all FBS and FCS teams—ACT Reading V6 and V12 would be tantamount to FBS Teams 1, 2, 3, …, 128! However, for FBS and FCS programs alike, most matchups each season are versus FBS and FCS foes, respectively. Like many FBS teams face a smattering of inferior opponents with FCS status, many FCS teams face a few inferior opponents with DII or NAIA memberships. So, we can consider two types (or versions) of games: [i] intradivisional games and [ii] interdivisional games. Intradivisional games are FBS vs FBS or FCS vs FCS whereas interdivisional games are FBS vs FCS or FCS vs non-DI. Thus, if the distributions for score differentials of FBS-FBS and FCS-FCS games are similar, and the same is true for FBS-FCS and FCS-non-DI, we can generate FBS-FCS scores that equate to FBS-FBS scores.

Chart 1: Distributions of Score Differentials

First, I obtained all Division I NCAA football game scores for 2012-2016 from this vast resource hosted by Kenneth Massey, that includes 8,349 games in which either an FBS or FCS team played. Second, I specified whether the home team won each game because home advantages are well-documented (here, here, here, but cf. here). Third, I specified one of four classifications for each game, the first two of which are intradivisional and the second two, interdivisional:

•    FBS vs FBS,
•    FCS vs FCS,
•    FBS vs FCS, or

•    FCS vs non-DI teams.

Fourth, I removed all games in which both teams did not play in an interdivisional game in that season, leaving 6,397 games for the analysis. For example, in 2012, neither UCLA nor USC played an FCS team so, the UCLA vs USC game was excluded from the analysis. However, the 2012 USC versus Washington game was included because Washington played a FCS team (Portland St.). The data was prepared in this manner because I only want to analyze score differentials of teams that played both types of games. That is, although it is only one FCS game, we know about FBS-FBS and FBS-FCS games that involve ’12 Washington whereas we only know about FBS-FBS games that involve ’12 USC.

Chart 2: D1 Teams Ranked by Win% and Mean Score Diff.
Fifth, I prepared Chart 1. It shows the distributions of score differentials for the four categories of games. Chart 1 demonstrates that FBS vs FBS scores (green) differentials are distributed almost identically to FCS vs FCS (brown) score differentials. Likewise, the score differentials are similarly distributed for the interdivisional games, but with some distinct dissimilarities. I attribute the dissimilarity in interdivisional distributions to the similar talent levels of lesser-FBS/better-FCS teams and lesser-FCS/non-DI teams while, concurrently, more better-FBS teams play FCS opponents (green) than better-FCS teams play non-DI opponents (orange). Hence, there are more 35-point blowouts in FBS-FCS games. This is evident in the ad hoc chart below, which was the sixth thing I did. 

Anyhow, because the distributions for FBS vs FBS and FCS vs FCS are nonetheless similar, we will consider in the analysis only home field advantage and whether a game was intra- or inter-divisional (i.e., we will ignore whether a team was FCS or FBS). I do this for simplicity—mostly for me, but maybe also for you. 

Seventh, the equating procedure was performed using a nonequivalent-groups design with one anchor, a home team win. Here, the anchor informs the equating procedure that differences in these games might be due to home-field advantage. The influence of including home team victory is evident in Chart 3. The black line represents the intradivisional score differential and the other lines are the corresponding interdivisional scores with or without home advantage. Some descriptive statistics appear in the table below. A table with unadjusted and adjusted score differentials and SEs appears at the close of the post.

Table 1. Descriptive Statistics for NCAA 1 D1 Intra- & Inter-Division Games, 2012-16
mean sd skew kurt min max n
Intradivisional 17.49 13.5 0.96 3.53 1 78 5493
Interdivisional 30.52 19.69 0.36 2.29 1 86 904
Intra- Home Wins 0.54 0.5 -0.15 1.02 0 1 5493
Inter- Home Wins 0.87 0.34 -2.16 5.67 0 1 904
Chart 3: Equated Interdivisional Score Differentials
Controlling for home-advantage—the green line—produces equated scores which are more sound, in my estimation. Notice how the green line equals the black line in the bottom left corner. The green line diverges at the 7-point differential. So, with this equating procedure, if an FBS team wins by 7 or fewer points over an FCS, it is the same differential as an FBS-FBS victory. To this author, this validly reflects in the score differential the competitiveness of an FBS-FCS game decided by one touchdown or less. Without adjusting for home winning, there are inflated point differentials in this range. Also, compared to the orange and the black lines, there is less of a difference between the green and black lines as the score differential increases (if such a feat were meaningful, Baylor). Likewise, Iowa St. is not additionally penalized for succumbing to a last-second field-goal whereas the orange line equates a 7-point FBS-FCS victory to 16 FBS-FBS points and a field-goal lead at 00:00 in the 4th quarter to 5 points. 

Now, there are of course shortcomings to this study, primarily one. Recall in the verbose example I provided earlier that the same 200 college freshmen completed both V6 and V12 of the ACT Reading sections. By doing so, we could be relatively certain that any disparity in V6 and V12 averages was not due to the test takers.  In the analysis, however, I included only games involving at least one team that played in intra- and inter-division games in the season. Thus, this analysis rests on the potentially fallible assumption that all intra- or inter-divisional opponents to these teams are identical—which is patently untrue. Hence, the reason we considered the distribution of different classifications of games in Chart 1.


Summarily, an equating procedure was used to generate score-differential equivalencies for FBS-FCS games to FBS-FBS games. This author concluded that adjusting for well-documented home field advantages provided more valid equivalencies. Secondarily, an ad hoc analysis demonstrated that upper echelon FBS teams more frequently play FCS opponents than upper echelon FCS teams play non-DI teams.



Adjusted Unadjusted
FBS Scr Diff Est. SE Est. SE
1 0.974 0.2 1.358 0.175
2 1.743 0.285 2.672 0.21
3 2.81 0.208 5.106 0.779
4 3.762 0.756 7.546 0.79
5 5.054 0.896 9.956 1.156
6 6.127 0.8 12.3 1.166
7 7.359 0.891 15.918 1.072
8 10.068 1.461 19.907 1.127
9 10.968 1.578 20.872 0.771
10 13.186 1.415 22.593 1.161
11 14.47 1.015 24.394 0.867
12 15.391 1.017 25.328 0.954
13 16.529 1.173 26.464 1.014
14 18.299 1.438 28.122 1.024
15 20.713 1.223 30.582 0.942
16 21.174 1.141 31.068 0.82
17 23.077 1.324 32.161 0.962
18 24.61 1.284 34.025 0.9
19 26.413 1.334 35.059 1.109
20 27.782 1.262 36.855 1.117
21 30.308 1.197 38.215 0.688
22 31.551 0.981 39.434 0.943
23 32.615 1.131 40.542 1.061
24 34.24 1.375 41.97 1.075
25 37.279 1.458 44.231 1.228
26 38.21 1.093 45.077 1.165
27 39.011 1.182 45.894 1.214
28 41.563 1.188 47.865 1.147
29 42.478 1.111 49.033 1.238
30 44.187 1.182 49.589 1.415
31 45.563 1.132 51.843 1.423
32 47.876 1.165 53.61 1.266
33 48.85 1.181 54.631 1.102
34 50.254 1.454 55.331 0.775
35 52.735 1.534 56.108 0.747
36 54.72 1.304 56.876 1.063
37 55.346 1.134 58.072 1.234
38 56.116 1.085 59.217 1.359
39 57.228 1.251 61.627 1.452
40 58.761 1.543 62.474 1.195
41 59.231 1.681 62.889 1.011
42 61.718 1.732 63.532 1.122
43 62.719 1.574 64.919 1.179
44 62.99 1.426 65.649 1.247
45 63.557 1.384 66.142 1.177
46 65.6 1.338 66.812 1.522
47 65.788 1.319 67.265 1.673
48 65.991 1.261 68.845 1.777
49 66.467 1.093 69.941 1.934
50 67.188 1.348 71.796 2.045
51 68.502 1.659 72.619 2.149
52 69.592 1.949 73.615 1.891
53 70.118 2.188 74.109 1.831
54 70.359 2.166 74.335 1.861
55 72.129 2.096 75.157 1.77
56 73.833 2.102 76.604 1.884
57 74.595 1.891 77.468 1.617
58 75.772 1.74 77.78 1.724
59 77.016 1.535 78.302 2.118
60 77.7 1.329 78.892 2.151
61 77.876 1.252 79.221 2.14
62 78.095 1.368 79.633 2.24
63 78.405 1.623 80.209 2.572
64 78.877 1.726 81.62 2.747
65 79.142 1.818 81.785 2.716
66 79.672 1.996 82.114 2.65
67 80.336 2.151 83.525 2.608
68 81.733 2.146 83.772 2.547
69 82.131 2.244 84.019 2.447
70 83.794 2.309 84.43 2.217
71 84.192 2.288 85.677 1.879
72 84.325 2.255 85.759 1.872
73 85.59 2.218 85.924 1.778
74 85.855 2.275 86.089 1.748
75 85.988 2.341 86.171 1.739
76 86.116 2.343 86.253 1.714
77 86.244 2.251 86.335 1.633
78 86.372 2.17 86.418 1.613