Saturday, April 6, 2019

Summarizing factors that influence scoring in NBA Slam Dunk Contests


In late May 2017, a colleague and I finished writing a manuscript about a study of factors influencing scores in NBA Slam Dunk Contests (SDC). We submitted the manuscript for peer review, which means that several sports analytics experts read our manuscript, pointed out its weaknesses and provided insightful critiques. Peer reviewed enabled us to produce a better manuscript. A revised version of the manuscript was recently accepted for publication in the Journal of Sports Analytics. I will summarize the findings of the study in this post. 

Let us first review what the SDC is for unfamiliar readers. I’m assuming you know what a slam dunk is. Well, the SDC is a competition of who can do the ‘best’ dunks. The definition of ‘best’ is subject to interpretation. It may be more apt to say that the SDC is a competition of whose dunks get the highest scores. Scores are awarded by a panel of 5 judges who almost exclusively give scores on a scale of 0 to 10. All the judges’ scores are added together to get a total score, usually, 0 to 50, with higher scores meaning it is a ‘better’ dunk. One contestant will do a dunk and then that dunk is scored. Then the next contestant dunks and it is scored, and so on. The highest scoring contestants in a round move on to the next round. The highest scoring contestant in the final round is declared the winner.

We focused our study on three broad factors that should influence SDC scoring:
  • Dunk elements are things like where the contestant jumped from or what they did with the ball and their body while in the air.
  • Contestant formatting and rules. This is things like the order that contestants dunk in, replacing missed dunks, how many rounds there are, and experience of the judging panel.
  • Superlatives. Superlative factors include how popular a contestant is, having ‘home field advantage’, how unique a dunk is, and contestant height. Another is things contestants do to create excitement, usually before dunks. For example, Blake Griffin staging a singing choir before his dunk (in 2011) or JaVelle McGee having a second basket set up (also in 2011).
Compared to other judged competitions like gymnastics, there is neither an ‘official’ list of the different dunk elements that can be done nor gradings of difficulty for those elements. Indeed, there are common names for dunks like a free throw line dunk or a tomahawk. More complex dunks might have names like 360° windmill or 180° double-pump reverse. However, these names cannot easily be analyzed. This posed a problem.

So, we segmented the dunk into its elements. We split the dunk elements into [a] things that can only be done while possessing (controlling) the ball and [b] things that can be done with or without the ball. We called these primaries and modifiers, respectively. Primaries require possessing the ball and include well-known maneuvers like windmills, double-pumps, between-the-legs, and others. Two primaries can’t be done at the same time, but one can be done after another while still airborne. A neat loophole is that you could do two primaries at the same time if you’re using two balls. Modifier elements that can be done with or without the ball include things like spinning in the air, jumping from the free throw line, catching a pass, jumping over something, covering the eyes, and more. You can do multiple modifiers at the same time. Like, you can cover your eyes while doing a 360° or you can catch a pass while jumping over something. You can also do (multiple) modifiers while doing a primary, like spinning 360° and covering your eyes while doing a windmill.

Another way to think of the elements is to imagine the most basic dunk. You jump straight up and dunk the ball. Nothing else. This system of elements would call that a dunk with no primary and no modifiers. If, instead, you jumped straight up and did a 360°, the system would call that a dunk with no primary, with a 360° modifier.

With that said, we reviewed all the dunks in NBA SDCs from 1984-2016. Separately, we determined what elements each dunk contained. Twice. Our determinations were acceptably consistent. When they were inconsistent, we reviewed those dunks and came to an agreement on what elements were done. Of 682 dunks, 215 had no sourceable footage, no scores, or were missed dunks that judges scored. These couldn’t be analyzed. That left 467 that could be analyzed.


Figure 1a. 
Figure 1b.

First, we did some fancy math (i.e., logistic regression) to determine which elements are the hardest to do. (Actually, we used all 682 dunks for this part.) We used the likelihood that there would be an execution error when performing each element. Execution errors are things like misses, botched attempts, and replacement dunks. As can be seen in Figure 1, the dunk elements we would expect to be more difficult were also more likely to have errors (1a) and be classified as harder (1b).

Figure 2.

Using some even fancier math (i.e., nonparametric regression) we examined how dunk elements factored into scores. Figure 2 shows how dunk elements affect scores. The primaries are shown on the vertical axis and the modifiers shown on the horizontal axis, at the bottom. Next to each primary is a number. This is the (expected) average score for that primary when there are no modifiers. On the line next to each primary, each bar up or down represents how much the average score is changed by a modifier. Bars above the line means the scores goes up and bars below the line means a score goes done. A white line in a bar means +/- 2 points. We can see that the Basic primary, which means there was no primary, will get a low score when there are no modifiers. However, we then see that doing modifiers will increase the score for a Basic dunk with no primary. This makes sense. No primary and no modifier is the most inanimate dunk. Likewise, more difficult primaries like the between-the-legs are less affected by modifiers. This shows that both primaries and modifiers factor into SDC scores.

The dunk elements explain about 44% of why different scores are awarded for different dunks in NBA SDCs. But what about the other 56% of why dunks are scored they way they are? To figure that out, we again used some fancy math (i.e., linear mixed-modeling) to look at how contest formatting and superlatives affected scores. (Rather than use the actual dunk scores, though, we used the parts of scores that were not explained by dunk elements.)


Figure 3.
In Figure 3 we can see how contest formatting and superlative factors would affect an average dunk with the score of 45. That dunk in the initial round would be a 45, but in the middle and final rounds, it is expected to be closer to a 47. (We later show that this is due to lower scores for dunks by contestants eliminated in the initial round, so the 45 is really only for weaker dunkers.) Likewise, 3 botched attempts or 3 replacements is expected to lower the score of the dunk by about 1 point. Histrionics are things contestants do to create excitement, usually before the dunk. Although many uses of histrionics do not affect the execution of a dunk, histrionics increase scores by nearly 3 points! While we can’t tease apart if the most popular contestants are also the most athletic, popular guys like MJ and Vince are expected to get about a 2-point increase in the score—so the 45 becomes 47. If a 6’6” tall contestant and 5’7” contestant do the same dunk, the shorter contestant would be expected to get the 45 and the taller guy would get a 44. Figure 4 shows how dunking later in the order in the initial round greatly increases scores over dunking first or second in the order.
Figure 4.
Overall, the contest factors, superlatives, and dunk elements together explain about 72% of why different scores are awarded in NBA SDCs. This is a pretty good amount to explain considering that we are not saving lives or anything.

A main finding of our study was that scores go up when there is excitement surrounding a dunk. Popularity and ‘home court advantage’ yield higher scores. Scores are slightly reduced the more judges have judged and the more times a dunk has been done, which is perhaps a case of ‘show me something new’. Scores also go down when there are botched attempts or replacements—this might be a spoiler effect. Singing choirs and bouncing cheerleaders boost scores. We speculated that this was due to excitation transfer. For example, imagine you rated some potential dates on how sexy they are. Excitation transfer happens as you will likely rate them to be sexier after you get off a crazy roller ride than after you walk down a small hill. The roller coaster ride gets you excited, and your excitement is transferred to how you rate your potential dates. Judges get excited by choirs and cheerleaders and their excitement is transferred to scores.

But what does this study mean for dunkers? First—well, try not to go first in the initial round! You should try to do creditable dunks you can make on your first try. So, practice, practice, practice a set of respectable dunks so you can make them. You want these dunks to be effortless for you, like singing Twinkle, Twinkle Little Star. Anyone could sing it at a different tempo or in a loud bus station and never miss a word. You want to be able to do your go-to dunks on wood or blacktop, indoors or out, on a slightly higher rim, or in a crowded half-court space.

Likewise, the scores for the most common primary elements—windmill/cradle, double-pump, and between-the-legs—differ only by 1 or 2 points, between 44 and 47. So, although you can probably do that between-the-legs (see Aaron Gordon in 2017), doing a windmill or double-pump may get you a high enough score to progress to the next round or win. Doing things like catching a pass, spinning in the air, jumping over things, jumping from farther away, covering your eyes, and so on, will likely make up the 1- or 2-point difference between, say, a windmill and a between-the-legs. Lastly, get the crowd involved before you dunk. Do something to create some excitement. Dance. Show us how short you are compared to the height of the rim. Bring that kid with his dad in the third row out to throw you a pass.

Summarily, this is for the dunkers. The NBA SDC competitors of course, but more so for those elite athletes who excel at and train to dunk. They aren't in the NBA. Me? I’m 5’11 and I’m getting old. I can’t dunk anymore. My coauthor, who is the same height, but is several years younger, contends he can still dunk. Okay, maybe I could still dunk on a warm day when I’m well-rested and there's curvy Honduran women courtside wearing sundresses pretending not to watch. Even when I could dunk with ease though, I could never do what the dunkers do. But dunking saved my life. So, I've got to help move the sport forward in the ways that I can. One of my life goals is to be able to (almost) dunk at 40. I’ve got a few years before I get there. To 40, that is. A more important goal to me is seeing a Slam Dunk Contest in the Olympics. For the World to appreciate it. Unfortunately, that cannot happen with how winners are currently decided in dunk contests. In closing, this research was the first step of a much larger effort to legitimize dunk contests as a competition.  

Wednesday, February 13, 2019

How do College Quarterbacks Perform After Throwing an Interception?

So, you’re watching the game on any day that is or should be Saturday and your QB throws an interception. Waves of fury and grief crash through your viscera. Moments earlier, the stodgy commentator was waxing anachronistically about the merits of establishing the running game but has since transitioned to divining the psychological state of your beloved QB. You will be haunted by the potential of him throwing an interception when your offense retakes the field. You are a fan, after all, and the potential for your QB to throw interceptions looms over your every daily activity and dreamscape from August to January, anyways. But what about your QB? The objective of this report is to examine how (collegiate) QBs perform after they throw an interception.

Did I conduct a literature review? Yes. We know that, in the NFL at least, team interception rates are weakly (r = 0.08)  to modestly correlated (0.27)  within seasons. Among individual NFL QBs, there is a modest relationship (0.27)  in year-to-year INT rates with the same team but it approaches negligibility after changing teams (0.10); but the best NFL QBs do consistently throw interceptions at a rate slightly less than the League average in a given year.  Other work shows there is a moderate relationship (0.43) between a QB’s INT rate and his true quarterbacking ability.  We also know that a QB’s year-to-year completion percentage is pretty stable with the same team (0.58)  but less so when changing teams (0.25).  Together, the prior research suggests that, [a] although INTs are wildly random, some QBs are inevitably more prone to throwing INTs and [b] that some QBs are more accurate than others. Given that interceptions are random, the majority of collegiate QBs are (relatively) well-practiced, and that completion percentage is moderately stable, we might expect that a QBs performance is minimally affected after throwing an INT. 

For the present study, we shall use a massive data set with plays from NCAA games in years 2001 and 2003 through mid-November 2018. I removed all offensive plays of FCS teams but retained plays with FBS offenses against FCS defenses. Because I do not own a supercomputer or even a particularly powerful machine, I removed all plays without a pass attempt or sack in the description. This means we cannot account for QB scrambling ability/threat in the analysis. 

Table 1. Descriptive Data for Cumulative Interceptions
Cumulative INTs Completion % Interception % Sack % Completions Interceptions Sacks Pass Attempts Drop Backs
0 0.609 0.029 0.052 286334 13575 25643 470461 496104
1 0.589 0.032 0.057 101146 5511 10372 171704 182076
2 0.577 0.037 0.060 28199 1785 3130 48866 51996
3 0.563 0.040 0.061 6828 490 785 12121 12906
4 0.560 0.034 0.064 1435 87 174 2564 2738
5 0.559 0.043 0.051 209 16 20 374 394
6 0.455 0.036 0.068 25 2 4 55 59
7 0.333 0.000 0.000 2 0 0 6 6

Nevertheless, let our dependent variable—how we’ll measure performance—be whether a given play was a completion. Table 1 shows that completion % decreases after throwing 1, 2, and 3 interceptions but it remains constant at about 56% after the second INT (ignoring the 6th and 7th INTs which are rare scenarios). Now, these differences would be significant statistically because of our sample size; however, is the difference in completion rate between 1 and 2 INTs meaningful? I would wager that a difference of 5 percentage-points, such as between no interceptions (~61%) and 3 interceptions (~56%), might be meaningful. Likewise, interceptions become slightly more likely once a QB has thrown an interception. Although a multitude of factors contribute to QB performance after an interception, some of those factors probably led to cumulative INTs, anyhow. Many of the factors are simply unknowable including individual tendencies of each QB such as perseverance or resilience, self-confidence, and intelligence. There are also in-game factors such as down, distance, field position, time remaining, score differential, team and opponent quality, and others. Let us account for these on-field factors (as best we can), which are listed and defined at the close of the post.

We’ll use a logistic regression to predict the probability of a completion before and after throwing an interception, while controlling for all of the factors we can. Indeed, we should actually use a (generalized) mixed model because there are many plays from each season, game, team unit, and QB—as well as plays for each team unit and QB nested within seasons and games, and QBs nested within teams—but my computer would explode and burn down this apartment building, destroying all of my stuff. 

Our model indicated no main effect of cumulative interceptions on completion %. However, there is a significant effect of cumulative interceptions when approaching the defense’s endzone (i.e., when the offense is getting closer to scoring). Interestingly, QBs were slightly but significantly more likely to complete passes when close to the defense endzone after having thrown an interception. The estimated (marginal) probabilities of completion can be seen in Figure 1. The difference in expected completion % is, at most, 1-2 percentage points from one INT to the next. To me, this is virtually meaningless. Additionally, both poorer offensive lines and better defenses reduced the likelihood of completing a pass. Having thrown more interceptions in the season to date also reduced the likelihood of a completion, albeit meagerly. Completions were less likely on 3rd and 4th downs and as games progressed. QBs at home were slightly more likely to complete passes. The model has a mess of interactions between down, distance, etc., but a summary can be seen here.
Figure 1. Expected Marginal Completion percentage as a function of field position, Time Remaining in the Game, and Cumulative INTs, with other covariates held constant at mean.

Summarily, this report provides evidence that quarterback performance is essentially unchanged after throwing an interception. Critical in-game factors, home-status of QB, and team quality were controlled for in the analysis. The sample size was >700K and the slight effects that were observed are probably ecologically meaningless. That said, I would like to note that if performance is reduced after throwing an INT, say, in a subset of QBs or for some QBs in some games, it is at least somewhat related to psychological processes. For example, some neuroscience research suggests that performance would be impaired immediately following an error, such as on the QB’s next few passes after an INT.  This research also suggests that for well-learned tasks, such as quarterbacking in our case, errors in performance may be more related to activity in the prefrontal cortex (located in your forehead, above your eyes, until about your temple) than in the circuitry of the brain largely responsible for volitional motor movements (but, in the study, electrodes were not placed on the circuits more responsible for learned motor movements). The prefrontal cortex is related to, among myriad other activity, planning and decision making. Nonetheless, you are probably best-served by disregarding that atavistic commentator. Your precious quarterback will probably be fine, but it will never feel that way for you and that is part of what makes college football so delightful.

Acknowledgements
On behalf of all of us here at POTH, I would like to thank my colleague CK for her insight and suggestions.
Covariates
Below are variables included in the analysis. OS% and ONR% are thought to reflect offensive line quality. Passes dropped by receivers and yards after the catch are not described in the play-by-play and thus there was no way to measure the ability receiving units. DS%, DC%, and DPD% are thought to reflect defensive unit quality. Specifically, %DS might reflect D-Line pass rush quality; DPD% might reflect ability of LBs and secondary to defend passes; and DC% might reflect the overall ability of a defensive unit to prevent successful passing. 

  • Offensive Sack % (OS%): the proportion of the offensive team’s QB drop-backs that resulted in sacks, for the season on which a play occurred. Adjusted for sack % and FBS/FCS-status of opposing defenses.
  • Offensive Negative Rush % (ONR%): the proportion of the offensive team’s rushing plays that resulted in negative yards, for the season. Adjusted for negative rush % and FBS/FCS-status of opposing defenses.
  • Defensive Sack % (DS%): the proportion of the defensive team’s opponents’ QB drop-backs that results in sacks, for the season on which a play occurred. Adjusted for sack % of all offenses a defense faced.
  • Defensive Completion % (DC%): the season average completion percentage against the defense on a given play. Adjusted for completion % of opposing defense.
  • Defensive Passes Defensed % (DPD%): the proportion of passes broken up or intercepted by the defense, in the season on which the play occurred. Bayesian-average adjustment for quantity of passes faced by the defense in that season based on FBS averages.
  • Cumulative Interceptions, Game: the quantity of interceptions thrown by the offensive team in the current game up to but not including a given play. That is, if a team had thrown 2 interceptions before a given play, the cumulative value would be 2 for this play. This way we can compute the probability of throwing an interception when none have been thrown. I had to settle for cumulation at the team level instead of the QB level because, again, I do not have a supercomputer.
  • Cumulative Interceptions, Season: the quantity of interceptions thrown by the offensive team in the season up to but not including a given play. Used team to be consistent with above.
  • Whether QB is playing at home or away.
  • A variety of on-field factors noted in the main body of the text. 




Wednesday, January 16, 2019

Icing the Kicker in NCAA Football 2005-18

In gridiron football, the icing the kicker phenomenon is thought to occur when the defending team calls a TO just before the ball is snapped on a FG attempt (FGA) that could tie, win, or otherwise sway the outcome of the game in favor of the kicking team. The motivation for calling the TO is that it could somehow disturb, or ‘ice’ the kicker in a way that he will be more likely to miss the FGA. 

Other authors have endeavored to examine icing the kicker. Some have reported that, in the NFL, calling a TO before a FG does not reduce the likelihood of making a FGA, whether controlling for FGA length or not. Other studies suggest suggests there is indeed an effect of reducing likelihood on longer NFL FGAs that is absent on shorter FGAs, when controlling for FGA length and other factors. At the collegiate level, it appears that icing the kicker may be effective on longer FGs; specifically, greater than 45 yards.  However, this study had a small sample of iced kicks.

Table 1. Descriptive Statistics for NCAA FGAs 2005-18
This Many Attempted Fields Goals were
Quarter FG% uFG% Attempted Made Blocked Home Attempts Last 2min Attempts ≤15s after TO Attempts
1st 0.727 0.753 7051 5123 248 3559 1129 468
2nd 0.702 0.732 11473 8049 475 5957 4413 3264
3rd 0.740 0.766 6629 4906 224 3417 1033 425
4th & OT 0.715 0.747 7176 5129 308 3764 1566 1754
TOTAL 0.718 0.747 32329 23207 1255 16697 8141 5911
We here at POTH sought to reexamine icing the kicker at the collegiate level using a much larger data set. This includes 32,329 FGAs from NCAA Division I FBS vs FBS and FBS vs FCS games from 2005 through mid-November 2018. Table 1 has the breakdowns of some data we’ll refer to throughout. The last 2 minutes refers to FGAs during the last two minutes of quarters 1 through 4 and any FGA occurring in OT. 

Let us start with blocked FGAs, though. Notably, as seen in Table 1, blocked FGAs were more likely to occur in the 2nd quarter and 4th quarter and OT (χ² = 12.3, p = 0.007)—the situations in games most relevant to icing the kicker. Longer FGAs were more likely to be blocked regardless of the quarter (p < 0.001). FGAs were also more likely to be blocked in the last 2 minutes of quarters and OT, but especially in the last 2 minutes of the 4th quarter and OT (p = 0.06). For these reasons, we shall include in our analyses only unblocked FGAs. This leaves 31,074 FGAs for analysis.

Table 2. Proportional Statistics for NCAA FGAs 2005-18
Proportion of Field Goals Made
Quarter % Blocked Home Team Away Tem Last 2min Before Last 2m ≤15s after TO No TO Before
1st 0.035 0.743 0.710 0.731 0.751 0.726 0.752
2nd 0.041 0.714 0.688 0.677 0.746 0.680 0.742
3rd 0.034 0.755 0.724 0.743 0.765 0.701 0.769
4th & OT 0.043 0.728 0.700 0.676 0.754 0.694 0.749
OVERALL 0.039 0.757 0.736 0.725 0.754 0.721 0.752

About 74.6% of (unblocked) FGAs are made. Figure 1 shows that FG% declines as the length of the FGA increases. There is some variation in FG% between quarters, with 3rd-quarter FGAs being most successful. Only differences between the 3rd and 2nd (p = 0.03) and the 4th and 3rd (p = 0.03) are significant when we account for length of the FGA, which is, by far, the most significant predictor of FG success. Longer FGAs are less successful at all points in the game. 
Figure 1. Likelihood of Making a FGA, by Length (using binomial smooth)
FGAs by the home team (75.6%) are about 2.8% more likely to be made than FGAs by the road team (73.5%) (χ² = 17.9, p < 0.001). When controlling for FGA length and quarter, home FGAs are 6.6% more likely to be made (p < 0.001). However, this advantage of home FG% is relatively constant at all FGA lengths. That is, home-team FG kickers tend to be slightly more successful than road-team kickers on FGAs of any length, and at any point in the game. 

What about the FG% in the last 2 minutes of quarters, when icing the kicker usually occurs? Table 1 shows that it clearly drops in the 4th quarter and OT (in the 2nd too). This drop in FG% in the last two minutes is, however, diminished when controlling for FGA length, quarter, and home/away (p = 0.52). It should be noted that FGAs in the last 2 minutes of the 2nd and 4th quarters are 1-2 yards longer than FGAs at other times in the game (ps < 0.002). 

How do the stakes of the game effect FG%? The opportunity to tie the game seems to have a general effect of increasing the likelihood of making a FG (p = 0.05). Otherwise, though, there is no effect of stakes on FG% when controlling for length, quarter, home/away, and being in the last 2 minutes or not

FGAs 15 seconds or less after a TO are made 72.1% of the time whereas other FGAs are made 75.2% of the time (χ² = 24, p < 0.001). Now, this is just if any TO is called; that is, by the offense, the defense, or some other TO that was not attributed to either team in the data. Really, we have a variable that indicates whether the TO was called by the offense, the defense, was unattributed, or if no TO was called. If we were to continue the analysis as we have been doing it, we would examine a four-way interaction between quarter, last 2 minutes or not, stakes, and who called the TO before or not. Four-way interactions are messy. And three of those variables have four levels. We should do something else.
Figure 2a. FG% by TO TypeFigure 2b. FGA Length by TO Type
Let us narrow our focus to FGAs in the last 2 minutes of the 4th quarter and OT where the offense can either tie the game or take the lead with a FG, which leaves 2,173 FGAs for analysis. In the two figures we see that iced FGAs (i.e., those after a defense TO) [a] are the least successful, at ~70%, but [b] are also, on average, the longest FGAs in this game situation. Thus, when we model the likelihood of making a FG, while accounting for length and home/away, there is no effect of icing the kicker (p = 0.24). Like, icing the kicker has no statistically differentiable effect of decreasing the likelihood of making longer FGAs (p = 0.25). However, the estimated marginal probabilities in the figure below suggest that the likelihood of iced FGAs declines slightly more at longer distances, although, again, this is not statistically significant. 
Figure 3. Estimated probabilities of FG% by length and who called TO in last 2 minutes of 4th & OT

Whereas we have used the raw yardage value for FGA length, the one previous study of icing the kicker in NCAA football split length into ‘bins’: distances of 18-25 yards, 26-35 yards, 36-45 yards, and >45 yards. The author of the previous study used only data from 2017-18 and found that of 38 iced FGAs in the last 2 minutes of the 4th quarter and OT, only 26% were made. If I examine only data from 2017-18, I find these same numbers (38 iced FGAs, 10 made, 28 missed). Below, using all data, I went ahead and show the FG% for each of these length-bins by who called the TO, for the sake of comparison across studies. The quantities of FGAs are shown parenthetically. Longer iced FGAs appear to be made lower rates.
Figure 4. Proportion of FGAs Made by TO Type by Yardage Bins used in Dalen (2018)
Summarily, the present report examined icing the kicker in NCAA football. This study used a sizable data set which would enhance the generalizability of the findings. However, the primary analysis indicated there was no effect of icing the kicker. Additional examination suggested that there might be an effect of icing the kicker at FGAs longer than 45 but such a conclusion is limited by there being fewer FGAs attempted from these lengths (i.e., smaller sample) and the variability of success at increasing lengths. Likewise, other potentially influential factors such as meteorological conditions, team FG kicking/defensing quality, and on-field activity were not accounted for in the analysis. NCAA football coaches should continue utilizing icing the kicker so they may endure the rancor of punditry, boosters, delusional fans, etc., when their teams lose games on last-second field goals.  

Sunday, January 21, 2018

Examination of Success Thresholds in College Football

Anyone familiarizing themselves with gridiron football analytics will quickly acquaint with success rate. Success is widely defined by an offense gaining 40-50% of yards to go on 1st down, 60-70% on 2nd down, and 100% on 3rd and 4th downs; preventing gains of said percentages defines success for defenses. Counting all the successful plays in a drive, game, or season and dividing by the total quantity of plays for that period yields the success rate. Personally, I am more interested in whether an activity was productive, unproductive, or counterproductive but I’ll save that for another post. However, curious as to how the thresholds for success may have been established and if there are nuances to current definitions, I examine it here. 

Myself and others before me, suspect that success rate is derived from traditional football notions of ‘staying ahead of the chains’ or ‘setting up for third and short’. Popularized by Football Outsiders, success, by our definition—like many gridiron analytic concepts—can be traced at least to the mid-1980s when it was outlined on p. 69 of the Hidden Game of Football. The authors used 40%, 60%, and 100% to benchmark ‘wins’ and ‘failures’, as well as a derived qualitative measure of success that awards more credit (i.e., >1-point) for big plays and penalizes turnovers and lost yardage (i.e., negative points). 


How do we determine what is a successful play? That is not, how success is defined according to X-amount of yardage gained or lost on a given play, per se, but how X-amount of yardage on a given play portends future success when aggregated from many, many plays in a similar context. Given this notion, let us define success as a play occurring on a drive that ends in scoring either a TD or FG.


Let us define success another way, too: a first down occurring on or after a given play on a drive (or series, in this case, really). For example, take a 2nd and 8; if there is a first down obtained on that play or subsequent play in the drive, that 2nd and 8 would be considered as having occurred on a successful drive (or, series). Alternatively, imagine a 1st and 10 which is, say, the fifth play of a drive and occurs after obtaining having at least one first down on the drive; if there is not a first down obtained on that 1st and 10 or any play later in the drive, that 1st and 10 would be considered as having occurred on an unsuccessful drive (or series).


For data, I have all pass and rush plays from games played by Division 1 college football teams from 2005-13—995,895 plays. For each play, I included an indicator of whether the play occurred on a scoring drive and whether there was another first down on that drive. As these are a binary variables, indicating yes or no, logistic regression is suitable. As the predictor variable we will use yards gained on a play divided by the yards to go on the play. This way we can say gaining X% of yards on a given down down is the threshold of success. Oh, so since we’re using college data, we’ll use the thresholds utilized by Football Study Hall of 50%, 70%, and 100% on 1st, 2nd, and 3rd and 4th downs, respectively.


Using logistic regression and ROC curves, we identify thresholds for the proportion of yards gained on each down that correctly predicts both the maximum quantity of plays on successful drives while minimizing the quantity of plays on unsuccessful drive wrongly predicted as successful (in our data set). This becomes our threshold of success. Figure 1 shows the success thresholds from these analyses for scoring drives in purple, drives with another first down in green, and the commonly applied success thresholds in orange.

Figure 1. Thresholds for Success

That the thresholds for 3rd and 4th down are essentially identical for scoring and first downs is unsurprising because scoring requires gaining at least the yards to go. The disparity in thresholds on second downs is also intuitive. It suggests that gaining a greater portion of the yards to go on 2nd down portends a more successful drive. The lower threshold for scoring drives on 1st down is interesting, however. It may be that obtaining 40% of the yards to on first downs typically setups a 2nd and 6 with offensive being in neither a definitive rush nor definitive pass situation. This, in turn, could conceivably lead to future success and the disparity here compared to the commonly used threshold. 


I was curious also how field position affects success. Let us focus only first downs, for convenience. I computed whether each play was a success based on the threshold for scoring drives described above; we’ll call this the fixed threshold. A mixture model was used to segment the field into 8 segments. Several logistic regression models were blended to generate thresholds for each segment, which we’ll call blended thresholds.i This is shown in Figure 2. Yard line 1-9 is closest to the defense’s end zone. The bottom row of panels are successful plays based on the fixed threshold and the blended threshold on the top row.
On the X-axis are Yes or No to indicate whether a play actually occurred on a scoring drive or not. Green indicates a play was predicted to occur on a non-scoring drive and orange indicates a play was predicted to occur on a scoring drive. We can see the fixed threshold emerged because it accurately predicts so many plays on unsuccessful drives in opponent’s territory.

Figure 2. Comparing Fixed and Blended Success Thresholds by Field Position on First Down

Summarily, this report showed that, at least in college football, success thresholds are relatively constant whether success is defined as a drive ending in a score or whether there is a first down after a given play. Secondarily, the report provides evidence that statistically-derived success thresholds vary by field position, at least on first down. Thus, future work should examine how adjusting thresholds by field position affects the valuation of player and team performance when using success rates.






iTo do this, I averaged the threshold from three logistic regressions. For each group I obtained thresholds from three logistic regressions with the following subset of the data: [a] plays in each field position segment, [b] all plays in each field position segment and all plays from field positions closer to the defense's end zone, and [c] all plays in each field position segment and all plays from field positions farther from the defense's end zone.

Wednesday, April 19, 2017

The Point-Value of a WNBA Live-Ball Turnover

Update: 23 April 2017 10:45pm. I detected an inaccuracy in the table in the subsetting of live ball TOs leading to scoring opportunities. Originally, I accidentally computed the ns and values for the subsets using all live ball TOs. Also, although never explicitly mentioned in the post, the value (or cost) of a live ball turnover is equal to the value of a steal.
 
So, I’ve got this trove of WNBA PBP data and a range of exhausting ideas. However, I’ve been inundated with projects and limited time to pursue said ideas. One of those ideas involves a metric for expressing WNBA player productivity as points. In this post, I’ll discuss the value of a turnover for the purposes of using it in that metric.

My idea of expressing player productivity in points is not a unique notion. It is similar to the concept of Marginal Productivity, described by its author David Sparks here and here and applied to the WNBA, here. And as I contended in a previous post, Sparks notes that because the “…regression coefficients … were fitted for the NBA, it is unclear whether or not their values translate identically to WNBA play…” Sparks does suspect, however, there will be little difference between the leagues. The WNBA MP spreadsheet appears to have been designed to update automatically and since the post is dated 2008, there are no values in the sheet. 


Nonetheless, in the Marginal Productivity model, a steal is valued at ~1.60 points and a non-steal turnover is valued at ~1.45 points; we’ll call these live- and dead-ball turnovers. Here, the live-ball turnover is worth about 10% more points than a dead ball turnover such as a travel or a pass hurled out of bounds. This makes sense intuitively and thus, we expect that live-ball turnovers will be more valuable than dead-ball turnovers.


Probably while in the shower or wading through traffic I proposed to myself borrowing a concept from football analytics, expected points. Scoring in the game of basketball is much more fluid than in football, so I rebutted to myself that instead of average next points we should use average points per next scoring opportunity, which include a field goal attempt (FGA) or freethrow attempt (FTA) but also clear path fouls and flagrant fouls. This contrasts with the caveat in the following paragraph because the live-ball turnover leads directly to a scoring opportunity that is criminally prevented by the foul and thus, the FTAs are the scoring opportunity.The points at the next scoring opportunity for live-ball turnovers that directly resulted in a foul with FTAs equal the points accrued for that trip to the line. 

This approach excludes 2-points scored, say, when a live-ball turnover leads to a missed FG, followed by an offensive rebound put-back layup. Inarguably, the turnover in this example leads to the positioning of the player completing the put-back, the defense being out of position, and the points. However, there are numerous outcomes other than an offensive rebound that could have occurred. Furthermore, we’ll have isolated the value of live-ball turnovers, or keep that value separated when we engage in a similar analysis of rebounds.
 

The data were all plays with FGs, missed FGs, FTs, FTAs, and turnover from every WNBA regular season game 2014-16. Something like 121,834 plays; 35238 FGs, 45864 missed FGs, 23735 FTAs, and16,997 TOs. There were 9012 live-ball TOs and 7985 dead-ball TOs. 7683 of the live-ball TOs directly resulted in a scoring opportunity. Table 1 contains the average points per scoring opportunity. 

Table 1. Average Points Per Next Scoring Opportunity from Turnover WNBA, 2014-16
Turnover n pts
Live-Ball TO 9012 1.160
Live-Ball TO to Scoring Opportunity 7683 1.173
   bad pass 6014 5142 1.167 1.182
   lost ball 2945 2499 1.149 1.161
   possession lost 53 42 0.887 0.833
Dead Ball TO 7985 0.979

As we expected and as is consistent with prior research, live-ball rebounds have a higher average point per scoring opportunity than dead-ball rebounds. For the present findings, the point-value of live- and dead-ball TOs are less than that of prior research. This could be due to the different computations in the analytics employed. It could also be due to differences in NBA and WNBA gameplay styles. That is, the value of a TO is less in the WNBA because a greater proportion of WNBA possessions end in TOs by way of stealing


Summarily, an expected points approach was used to compute the average points per next scoring opportunity directly resulting from a TO in the WNBA. This author proposed that the WNBA live-ball TO may be worth less than the NBA TO because there is a higher rate of TOs in the WNBA (different analytic approaches from prior research notwithstanding).

Sunday, March 19, 2017

Comparison of WNBA and NBA 2-Point Field Goals

Comparisons of WNBA and NBA Association-wide, season-level data were compared in a previous post. For the season of each that was analyzed, disparities in FG% were evident. Initially, the NBA appeared to be more successful shooting, 44.9% to 42.5%. However, when excluding dunks, the WNBA was slightly but significantly more successful, 42.5% to ~40%. Having recently acquired WNBA play-by-play (PBP) data for the 2016 season, we can more granularly analyze FG%.

All WNBA data were extracted from the regular season 2016 PBP which is available upon request. I will dump this data along with the regular season PBP data from the 2014 and 2015 seasons in a future post. The NBA shot type and assisted shot type data for the 2015-16 season was culled from Basketball Reference.


Table 1a. WNBA & NBA 2016 counts
Type Total NBA WNBA
All 2FG Made 75549 9815
Attempted 153768 20676
Shots Made 35680 4769
Attempted 89487 12318
Layups Made 30439 5045
Attempted 53920 8356
Dunks Made 9430 1
Attempted 10361 2
Table 1a contains counts and 1b proportions of 2FGAs segmented by shots, dunks, and layups between the Associations for the regular seasons ending in 2016. The class of ‘shots’ includes not only jump shots but also others identified in the PBP as floaters, hooks, and runners. Table 1b also contains Chi-square test-statistics demonstrating that the is NBA is slightly more successful shooting 2-point shots. The WNBA is significantly more successful with layups. The Chi-square was not performed for dunks because only 2 were attempted by WNBA players.

Table 1b. 2pt-FG% by Shot Types
Type NBA WNBA χ² p
All 0.491 0.475 6.961 0.008
Shots 0.399 0.387 2.623 0.105
Layups 0.565 0.604 12.230 0.000
Dunks 0.910 0.500
So, these findings provide a more nuanced perspective on the findings from the previous post that WNBA is more successful shooting non-dunk shots. The two posts did employ data from different seasons. Nonetheless, given the present data, the two Associations ostensibly shoot 2-point shots (i.e., not dunks and layups) with similar successfulness—the p¬¬-value approaching conventional significance levels is likely a result of large quantities. That is, a 1 percentage-point advantage to the NBA may approach significance statistically, but I suspect it is ecologically meaningless.

Alternatively, the WNBA was significantly more successful than the NBA shooting layups. My initial notional hypothesis was that because males are predisposed to heightened athleticism, there are more contested or blocked layup attempts in the NBA. Ironically (to the impetus for pursuing this line of research), although blocked layup attempts can be extracted from the WNBA PBP, I am unable to locate blocked layup attempt data for the NBA (short of scraping). Likewise, contested shots attempted from with ≤5ft from the basket are available for the NBA but shot contesting is not recorded in the WNBA PBP.

Table 2. Assists on Dunks + Layups
Dunk + Layups NBA WNBA χ² p
Made & Assisted 15536 2749 6.596 0.010
Made 64281 8358
Proportion Assisted 0.242 0.329
Table 2 contains the proportion of dunks and layups, combined, that were assisted. The WNBA assists on significantly more of their successful layups (and dunks) than does the NBA. So, this might also explain why WNBA players are more successful executing layups. Because the WNBA assists on a greater proportion of layups, there may be more floor spacing or player movement such that defenders are less frequently positioned to defend layup attempts. Indeed, this notion is interrelated to there being greater athleticism in the NBA, as well as greater size, and thus less floor space in the NBA. Lastly, WNBA players may execute successful layups in certain scenarios whereas NBA players would likely execute successful dunks such as on uncontested fast breaks.

If athleticism were the sole determinant in explaining WNBA layup FG% superiority, there would be little recourse for defensive strategists other than playing larger or quicker players. However, if it is the result of floor spacing or player movement, I suspect WNBA coaches transiently employ zone defenses to narrow passing and driving lanes to reduce opponents’ layup success.

Summarily, this report indicates that WNBA and NBA players shoot with similar accuracy on 2-point FGs that are not dunks or layups. Also, the WNBA is more successful on layups than the NBA. This author posited three reasons why this may be: (a) greater athleticism and size on NBA players results in more contested layup attempts and passes; (b) relatedly, the WNBA assists on a higher proportion of its layups which may be the result of more floor space or player movement, but is also related to point (a); and, (c) NBA players may execute dunks in many scenarios where most WNBA players would have to execute layups, also related to point (a).

Sunday, February 19, 2017

2016 WNBA FG% Distribution by Shot Location

So this is a first for POTH: I am posting twice in one day, or twice in one uninterrupted span of wakefulness. However, it is a diminutive post. Below is a shot chart for the 2016 WNBA regular season with field goal percentage. Greener means higher shooting percentage, navy-er means lower percentage. 

Rather unexpectedly, I was able to amass some data rather quickly and it happened that the shot locations were included. I got excited as I often do when there are discoveries at 00:18. Of course, the chart would would be more informative if the hexagons were sized according to the quantity of shots therein (e.g., here). But it's late and my belly aches.

Chart 1: Distribution of FG% by Shot Location in the WNBA, 2016

A previous post revealed some differences in WNBA compared to NBA league-wide aggregate statistics. I am now able to address these and other topics.