Friday, 13 May 2016

How much more valuable are first division runs?

England announced their squad to play Sri Lanka this week, with Hampshire's James Vince getting the nod to take up the middle order slot unfortunately vacated by James Taylor. Nick Compton, meanwhile, keeps his place at number 3, at least for the time being. Essex's Tom Westley, who has had a productive start to the season and has been much talked up, was left out (I was hoping he would be picked, but not for any cricketing reason- I just wanted the opportunity to make some Princess Bride jokes).

As England squad selections draw near, with places up for grabs, attention often turns to the county championship averages. One of the few things everyone seems to agree on at this point is that runs made in the first division of the championship should be valued more highly, being made against higher quality attacks. This seems eminently reasonable, but raises a question: how much more valuable are they? Can we make the comparison quantitative?

I'm going to have a go.

What we want is to take a sample of batsmen who played in both divisions in successive seasons and ask, on average, how much did their run output drop/rise on switching divisions. Such a sample is provided to us by the championship's promotion and relegation system.

What I've done is go through the county averages for all the completed seasons since 2010, looking at the performance of players in teams that were relegated or promoted and then comparing their season's batting average before and after the change of divisions. (So, for example, I took the batsmen who played for Kent in division 1 in 2010 and compared each batsman's average to what they managed in division 2 in 2011).

I only included batsmen who played at least 10 matches in both seasons. The results are depicted in the graph below. The batting average in division 2 for each batsman in the sample is one the x-axis, with division 1 on the y-axis. Players in relegated teams are in red, promoted teams in blue. Points below the black line averaged higher in division 2 than division 1, and above vice versa. The green line is the best linear fit to the data.

Of the 81 players in the sample, 52 averaged higher in division 1 and 29 averaged lower. So, the intuition that runs are harder to get in division 1 seems solid, as expected. But how big is the difference?

Well, on average the relegated players in the sample increased their averages by 4.98 runs on going from division 1 to division 2. The promoted players saw their averages drop by an average of 7.12 runs on going from division 2 to division 1. So based on those numbers the difference is moderate but noticeable- able to turn a "very good" set of numbers into merely "good" ones and "good" into merely "acceptable".

The linear fit which I attempted (which should be taken with absolute ladelfulls of salt) gives:

average in div 1=28.2 + 0.12 * (average in div 2)

so it would predict a player who averages 50 in division 2 to average only 34.2 in division 1. (As I say, don't take this equation too seriously, and possibly not seriously at all, not least since it predicts that players averaging less than 32 in div 2 should be expected to do better in div 1).

There is a chance that the difference between divisions is exaggerated in this data by a selection bias. Specifically, looking at players who were promoted from div 2 or relegated from div 1 may bias the sample towards players who under-performed their "true" ability when in div 1 or over-performed in div 2. In this case the shift in batting averages may in part be a case of regression to the mean, on top of the real change in the difficulty of run-getting.

This caveat notwithstanding, the difference in divisions seems quite considerable, and division 1 runs are worthy of their additional praise.

Thursday, 5 May 2016

The candidates

Despite its title, this is not a surprise post about the extraordinary political wranglings currently in full swing in the land of baseball and chilli-dogs. No, this will be about the far weightier matter of whether certain batsmen are especially susceptible to being pinned LBW, and who those current players are.

In cricket commentary, it's common for players whose technique looks somehow prone to leave them trapped in front of their stumps to be described as "lbw candidates". This terminology seems to be applied specially to that particular means of dismissal- batsmen are rarely described as "caught behind candidates".

The questions I want to investigate in today's post stem from this.

Firstly, is "lbw candidate" a worthwhile category- is there a substantial subgroup of modern test batsmen who are especially more lbw prone than their peers?

Secondly, who are these prime candidates in the post-Shane Watson era? I've often heard Alastair Cook described as a "candidate". Does he deserve the title?

We'll also be touching on where in the world lbws are most prevalent.

To tackle this, I took a sample of 45 current test match players, representing all the test nations apart from Zimbabwe, who haven't had much opportunity to play recently. The sample was obtained by taking the most recent test for each nation and including all the batsmen in he top 7 who had played at least 15 tests and who weren't obvious night-watchmen. For each player I looked up the total number of LBW dismissals in their test career and divided it by the number of dismissals overall. This is what is on the x-axis of the graph below, with the batting average of each player on the y-axis. The colour/shape of each point indicates the country for which the batsman plays.

The black dashed line is the sample median (0.155) and the red dashed lines either side are the upper (0.187) and lower (0.125) quartiles. As you can see, the data is quite clustered horizontally suggesting only a fairly small degree of variation in vulnerability to LBW amongst current test batsmen. There's also no significant correlation between the LBWs/dismissal and the batting average, suggesting that having a high proportion of dismissals be LBW doesn't indicate much either way for a batsman's run scoring ability.

There are, however, a few noticeable outliers, far removed from the central cluster to whom we now come:


  • The Shane Watson memorial award for excellence in attracting LBW decisions (I like the idea of this award- we could call it the "iron pad" and award it annually) goes to South Africa's JP Duminy, who is way off to the right of the graph with 39% of his dismissals being LBW. (A lot of these were against spin bowlers).
  • There's a select trio of players to the left of the graph who hardly ever get pinned LBW. Namely Pakistan's Sarfraz Ahmed (0 lbws/28 dismissals), England's Ben Stokes (1/41) and Bangladesh's Tamim Iqbal (2/79). It may not be significant but these are all quite aggressive batsmen, so perhaps more than being good at avoiding LBWs, they're finding other, more exciting, ways to get out first.
  • There's a foursome of Pakistan players separated from the main cluster, at around 0.25 LBWs/dismissal. These are: Younis Khan, Misbah ul Haq, Asad Shafiq and Mohammed Hafeez. It's tempting to wonder whether this might be because they play a lot of tests in the UAE, where the low, slow pitches are thought to be favourable for LBWs. Indeed, in the graph below you can see that the UAE does have the highest rate of LBWs per dismissal of top 6 batsmen amongst test match hosts since 2010. However, this probably doesn't fully account for it- if we exclude tests in the UAE for these four players only Hafeez sees his percentage of LBWs drop significantly.


Overall, modern test batsmen don't vary too much in how frequently their pinned leg before, with a small number of exceptions. For what it's worth, Alastair Cook falls close to the central cluster of data points in our first graph, albeit slightly on the high side, with a rate of 0.19 LBWs/dismissal. And with Pakistan's apparently quite LBW prone top order coming to England this summer, it could be quite a good season for the thump of ball on pad, and the slowly raised finger. Maybe.

Saturday, 16 April 2016

Throwing out the form book?

So it's been quite a while since I posted anything here, but with the thrill of a new English cricket season upon me, I'm strapping on my pads of data, taking up my bat of analysis and striding out to the wicket of the internet.

As I scratch around, hoping to hit a bit of early season form, I'm going to attempt some rudimentary analysis of exactly that concept- "form". The point of this blog is meant to be try and hold up some of cricket's hoariest old cliches and nuggets of received wisdom to the light of some data. The idea of being "in form" is surely one of the foremost such cliches in cricket- perhaps in all of sport.

The eseential claim is this: a player is more likely to perform well at times when they have performed well in the recent past. A player who has performed well recently is usually said to be "in form".

The explanations for this tend to hinge on a player's confidence being high when their recent performances have been good. Or people may speak about players "being in a good rhythm", or "in a good place".

But to what extent is "a run of good form" distinguishable from a run of good luck? You sometimes hear commentators say something along the lines of:

"when you're in good form, it's amazing how the little bits of luck start going your way as well- playing and missing rather than nicking it,  balls in the air going between fielders rather than to them..."

At which I might want to say to them: "Is it amazing? Is it though? Or is it just that you only assign players the property of "good form" when they happen to be on a good run of scores- which requires a certain amount of luck?".

I'm not going to attempt a full analysis of whether form is a "real" phenomenon- in the sense of being meaningfully predictive of future performance- in one blog post. Although I may come back to different aspects of the question later.

I do, however, have some data to show which impacts on this question and I think it's interesting.

To make the question narrower, and therefore more tractable, I asked: "are test match batsmen more likely to score a century when they have already scored a test century in the last month?"

To answer this, I looked at the careers of the 23 most prolific test match century scorers in history. I did this because I needed a sample of players who had scored enough centuries that one could meaningfully compare the games when they hadn't scored one recently, with games where they had. Obviously, this does introduce quite a big selection bias- it's possible that the results I obtain may only be applicable to those players at the very top of cricketing history's tree. So be aware of that when you decide what to think of the results.

The graph below shows the rate of century scoring per match in games within a month of having previously scored a century against the total number of centuries scored per match for each player.
Points above the blue line represent players who had a higher rate of century scoring when they had recently scored a century and those below the blue line represent players who had a lower rate of century scoring when they had recently scored a century. The lone point way off to the top right of the graph is, of course, Sir Donald Bradman.

As a group these batsmen scored an overall total of 723 centuries in 2945 games- a rate of 0.246 centuries per match. In games within a month of having scored a test century my research puts them at a total of 182 centuries in 704 games- for a nearly identical (but slightly higher) rate of 0.258 centuries per match. On an individual level 11 of the players were more prolific when they'd recently hit a hundred and 12 were less so. For most players the difference was minor, as indicated by the fact that most points in the graph fall fairly close to the blue line.

There isn't enough evidence here for me to boldly claim that form makes no difference to batsmen. But it does suggest that form doesn't matter as much as you might imagine, at least for this sample of batsmen who belong among history's greatest.

For those out of form I would say this: take heart- form is an ephemeral thing which can return as suddenly as it departs. And maybe it doesn't matter so much whether you have it or not.

Friday, 25 December 2015

Festive Tidings and the Exceptional AB de Villiers

Today, I bring you a festive look at the data! I mean, not that there's anything particularly Christmassy about the content of this blog post- but hey, it's Christmas, there are mince pies in the oven and I'm writing about the batting statistics of wicket keepers. To me, that's festive.

This time around, the piece of cricketing received wisdom coming under the microscope is the belief that when a batsman who's able to keep wicket has to do so, it impacts negatively on their run scoring ability. The most famous (and, not coincidentally, also the most extreme) example of this is Kumar Sangakkara who averaged an acceptable 40.48 when playing as a wicket keeper and a stellar 66.78 when playing as a specialist batsman. It seems reasonable to believe that the physical and mental strain of long periods of wicket keeping would make run scoring harder, but the same could be said of the pressure of the captaincy- and we saw in the last post that captaincy actually seems not to generally make so much difference to run scoring output.

I actually prepared the research for this post a while ago, but didn't write a post on it because- as you'll see below- there isn't so much to work with in this case, and I worried that there wasn't enough numerical meat to make a satisfying analysis. However, the issue came up on the superb Switch Hit podcast this week- in the context of AB de Villiers' stewardship of the keeper's gloves for South Africa- and I thought that since it's an interesting question, I might as well write it up. Decide for yourselves whether the data justifies the conclusions.

So, what we want to do is take some test match players who've played a decent number of tests both as wicket keeper, and as a specialist batsman and compare their batting averages in those two sets of games. The problem is that there are very few players who fit that description. Specifically, I could find only seven players who played both at least 10 tests as the designated wicket keeper and at least 10 not as the wicket keeper. That rather select club is listed in the table below

In the graph below, I've plotted the batting average when playing as keeper against the average when not playing as keeper for each player. Players falling below the blue line have worse averages when playing as wicket keeper and those above have better batting averages when granted the gloves.

Seven players isn't much to draw a conclusion from but nevertheless, the evidence in this case weighs in favour of the received wisdom- it does seem that having to keep wicket depresses a batsman's average. Of our seven players 2 have better averages when playing as keeper and 5 do worse. That in itself could easily just be chance, but what's more notable is the players who are doing worse as keeper tend to be doing rather a lot worse, suggesting that there is a potentially rather a strong effect at play. The average difference between averages when keeping and not our sample was -10.19 runs- less extreme than Sanga's -26.3 but a pretty big difference all the same.

Which makes AB de Villiers' bucking of the trend all the more special. He averages fully 8.83 runs higher when keeping. Of course, this won't necessarily last. It's quite possible - maybe even likely - that if he stays as South Africa's first choice gloveman for a couple more years his average as keeper will regress back in line with his average when not keeping - or even below. Or perhaps - as he has in many other ways - de Villiers will prove to be exceptional in the truest sense of the word.

I want to finish this post by thanking you all for reading and to particularly thank Chris of the excellent blog "Declaration Game" for kindly promoting my blogging over the last 6 months. I was honoured to be included in his "Select XI" blog posts of the year, which if you haven't seen it yet is well worth a look- providing a very broad cross section of some extremely interesting cricket writing.

Merry Christmas!

Sunday, 15 November 2015

Batsmen and the burden of the captaincy

It must be tough being a test match captain. The potential for days in the field, a mind full of bowling changes and fielding positions. Commentators and fans analysing your every move. Are you being too funky or not funky enough? Then, after all that, you have to go out and bat. As well as being held responsible for the collective success or failure of your team, you have your personal performance to take care of. The burden is heavy. Surely you're exhausted. Something must give, mustn't it?

It seems to be a fairly commonly held belief that the cost of doing what test teams usually do, in making one of their best batsmen the captain, will often come in the form of reduced run output from the player in question. In discussions of Joe Root, England's presumptive captain-in-waiting, I have certainly heard it raised that making him captain will dent his prolific run scoring.

It seems a reasonable enough worry to have. The captaincy certainly carries a lot of pressure with it, and a lot of extra responsibility which one would have thought would make it harder to focus on one's batting. But what does the evidence say? How does the captaincy affect a batsman's performance?

The graph below plots the batting average when playing as captain against the batting average when not playing as captain for all the test captains who have led their side at least 30 times.

Points below the blue line represent players who's batting average was lower when captaining, and points above the line represent players who were more prolific when skippering. There are two things to notice here:
1) Most points fall fairly close to the blue line- i.e. for most of the players in our sample their batting averages with or without the captaincy only differ a little bit.
2) There are more points above the line than below it (26 vs 17 to be precise)- i.e. it's more common for a player's average to improve with the captaincy than to decrease.

On average, the players in this sample increased their batting average by 3.76 runs when carrying the captaincy burden. I wouldn't read too much into that positive shift as it is much smaller than the sample standard deviation. The main take home message is that for most players the captaincy doesn't seem to make much difference to their average, and only for very few does their average significantly decrease.

I have heard it said that the England captaincy may carry peculiar pressures- perhaps due to the often slightly tempestuous relationship between players, media and fans in English sport. So one may wonder about the England captains of recent vintage in our sample. Of those only, Michael Vaughan (36.02 with captaincy vs 50.98 without) shows a big negative shift. Alastair Cook (49.94 vs 46.36), Andrew Strauss (40.76 vs 41.04) and Nasser Hussain (36.04 vs 38.10) all have pretty similar numbers for the two cases. Mike Atherton shows a slightly bigger shift but in the positive direction (40.58 vs 35.25).

So there's really not much compelling evidence to make us think that the captaincy depresses the run scoring of batsmen. But why, then, is this believed? I don't know, but my personal theory is this: when a player is given the captaincy they're usually coming off the back of a pretty good run- since you generally don't want to give the captaincy to a player unsure of their place. But all good runs must end eventually, for all players, captaincy or no. Whenever that does happen, this will be widely attributed to the pressures of captaincy catching up with them and the belief is perpetuated.

Test match captains are made of stern stuff- despite the pressure, they'll just carry on batting.

Tuesday, 10 November 2015

Pakistan's spinners have mastered the UAE where others have failed

For this post, I was reflecting on England's recent performance against Pakistan in the UAE. The consensus, after England's 2-0 defeat, seems to be that they performed fairly well but just came up against a side better suited to the conditions.

There's certainly a lot of truth in that. Despite the fact that Pakistan haven't been able to play tests in their home country in recent years their success in the UAE- where they've played in lieu of home matches- rivals some of the strongest home teams in test cricket. The graph below illustrates the 'home' record of each test match side since November 2010 (the period over which Pakistan have been laying regularly in the UAE). The 'x' axis shows the percentage of wins achieved by the 'home' side and the 'y' axis shows their net batting average - bowling average at home in that period.

By these measures Pakistan's record in the UAE is very close to England's in England- not bad considering they don't actually get to play at home.

The most obvious difference between the two sides was the performance of their respective spin bowlers. While England's batsmen floundered against the legspin of Yasir Shah; Adil Rashid, Moeen Ali and Samit Patel neither took regular wickets nor kept the runs down. I think it's fair to say that, in the main, they rightly haven't been over-harshly criticised, but I also think there's an air of disappointment surrounding the fact that the best spinners England could muster simply didn't seem to cut the mustard.

I would like to offer one point in mitigation of this- the UAE is actually quite a difficult place to be a non-Pakistani spinner. The graph below plots the bowling averages of 'home' and 'away' spin bowlers in each test match hosting nation since November 2010.




There's a (fairly weak) trend in the direction you'd expect- that in places where "home" spinners perform well, so do "away" spinners, at least relatively. But the performance of Pakistan's spinners in the UAE is far, far better than the overall performance of spinners for the touring sides they've been playing against. Pakistan's spinners average 29.65 in the UAE since November 2010, as compared with 44.69 for spinners from other test nations in the UAE. The difference between those two figures is the second highest for any of the test hosting nations. The largest difference between home and away spinners is in Australia, where baggy green spinners have been taking wickets at 41.63, as against 57.49 for touring sides.

So it seems that Pakistan's spinners have been finding a way to succeed in the UAE, where the best spin bowlers of touring sides have generally been struggling. Whether this is because of the pitches, because Pakistan's batsmen are really good at playing spin, something else, or a combination- I don't know. I do think that this to some extent sets the performance of England's spinners in some context- their collective average of 59.85, certainly remains a disappointment but there were always unlikely to be England's match winners in that series (although Rashid nearly was in the first test, but for the bad light). I also think it illustrates that there was never likely to be much tactical value in picking a third spinner for the third test, but perhaps a discussion of how to balance a bowling attack is one for another day.

Saturday, 31 October 2015

How long before a batting average means something?

Since I last posted, England have battled through two thirds of test series against Pakistan, acquitting themselves much better than at least I imagined they would, but still coming out behind. The struggles of England's middle order look set to lead to a test comeback for James Taylor.

In recent years, England's selectors have been praised for giving players a decent run in the side when called up- giving them more than or two chances to show what they can do. I assume, and hope, that the same treatment will be extended to Taylor and that, barring injury, he'll also play in the South Africa tour.

These ruminations lead me on to today's question: if we judge a batsman by their batting average, how many matches will it actually take before that average fairly reflects their ability?

I think most of us understand that quoting someone's batting average after two games isn't going to provide terribly strong evidence either way about how good they'll be in the long term. But how long should we wait before we can suppose that their average gives a strong clue as to their underlying run scoring prowess? In my experience, the conventional wisdom might place this number somewhere around 10 matches or a little more, depending on who you talk to.

To try and answer this question, I've attempted something a little different to my previous posts. Instead of using data from past test matches, I wrote a computer simulation to simulate the run scoring output of two (fictional) batsmen of known ability and looked at the distribution of their averages as a function of the number of innings played. The reason for doing this is that allows me to make a controlled 'experiment' in which I know how good the players in my simulation 'should' be and can see the degree to which statistical fluctuations obscure that in a finite sample of innings.

In my previous post, I argued that a player's vulnerability to getting out is only weakly dependent on how many runs they already have- being slightly elevated right at the very beginning of their innings (and maybe also a little elevated immediately after reaching 100).

I simulated the output of two players:

Player A had a 12% chance of getting out before reaching 5 and an 8% chance of getting out before scoring the next five runs thereafter. To put these numbers in context, this is very good- in the long run Player A could expect to average around 55.

Player B had a 16% chance of getting out before reaching 5 and an 12% of getting out before scoring the next five runs thereafter. This is rather more mediocre- in the long run Player B could expect to average around 35.

The two graphs below illustrate the probability distribution of batting averages for the each player as a function of the number of innings they were given in the simulation. The green points represent their median average after that number of innings and the red and blue points are the 10th and 90th percentile respectively. The region between the blue and red points reflects their likely range of batting averages after a given number of innings.

What's striking is that even after 50 innings the distributions are still quite broad - particularly for the better player (Player A). After 50 innings Player A has a 10% chance of averaging more than 66- making him look like a potential legend and also a 10% of averaging lower than 45 making look much more run of the mill.

Player B meanwhile has a 10% chance of averaging higher than 42 or lower than 28- the difference  between fairly good and pretty poor.

These averages are converging to a fair reflection of the players' abilities but they are doing so rather slowly- a hint that even after a fairly decent number of tests we need to base our judgements of players on more than their bare batting average.

Imagine if you were a selector, who brought these two imaginary players into your imaginary team and after a fixed number of tests had to choose between these two (perhaps you have a star player about to come back from injury and have to drop someone to fit him in). Would their averages be likely to guide you to the right decision?

The graph below shows the probability that the very good player A has a better average than the pretty mediocre player B after a given number of tests.


After 10 innings there's around an 80% chance that the averages will correctly reflect that player A is better than player B. Which sounds kind of okay, until one reflects that selection decisions are often- necessarily- based on fewer innings than that and that these two players are really not evenly matched at all- in the long run one would average a full 20 runs higher than the other.

Of course, in reality selectors have a lot more information available to them than just batting averages. Anyone can look up a players' average but selectors must exercise their judgement on a player's technique, temperament and suchlike using what they've seen in both matches and training. They have to do so because they don't have the luxury of letting a player play 20 test matches before making a decision about whether they're good enough- which is probably the minimum they would need to justify a decision based on batting average alone. To look at Gary Ballance's batting average of 47.76 after 27 innings, it's hard to avoid the conclusion he's been hard done by to not be in the team right now. And maybe he is- but one can't be sure of that from just his average.

It may well be the case that one could find a better way of estimating a batsman's ability from their stats after a small number of tests, which would converge on something fair a bit faster than simple batting average. On the other hand, fans like me should perhaps give selectors a break sometimes- they have rather complicated decisions to make, with rather limited and noisy information.