Monday, March 13, 2017

Cheltenham 2017: Supreme Novices Hurdle Handicapping

Once more unto the breach, dear friends, in our annual attempt to find live longshots to finish in the money in the Supreme Novices Hurdle (G1) at Cheltenham 2017. As ever, our approach is based on the following premises: 
  • Supreme Novices Hurdle is similar to Kentucky Derby - young horses, many attempting graded stakes, championship race for first-time with little form in book.
  • Eliminate non-contenders and whatever remains, no matter how improbable...
    • Avoid horses with pedigree mismatch to former winners [MelonCrack Mome].
    • Avoid horses from small fields [Labaik], [Elgin].
    • Late speed important [Magna CartorGlaringCapitol Force].
    • Poor Cheltenham Form [River Wylde].
    • Poor FPR [Pingshou, High Bridge].
  • Minimum price 10/1 [BallyandyBunk Off Early].
This leaves Beyond Conceit 20/1 and Cilaos Emery 16/1 as our selections!
Note, given the limited exposure of all the runners, we are not saying that those we have eliminated are not going to win - simply that they did not meet our criteria for live longshots to run in the money. The key takeaway, as always, is using a process of elimination not selection for identifying contenders.
Footnote: In their respective next outings, Cilaos Emery won G1 (Punchestown) and Beyond Conceit finished second in G1 (Aintree).

Friday, February 03, 2017

Adaptive Boosting

Machine learning studies the design of automatic methods for making predictions about the future based on past experiences. In the context of classification problems, machine-learning methods attempt to learn to predict the correct grouping of unseen examples. In the mid-1990s, Freund and Schapire introduced the meta-heuristic, Adaptive Boosting (AdaBoost), “…an approach to machine learning based on the idea of creating a highly accurate prediction rule by combining many relatively weak and inaccurate rules… Indeed, at its core, boosting solves hard machine-learning problems by forming a very smart committee of grossly incompetent but carefully selected members…” (Boosting: Foundations And Algorithms). As context for how boosting might work, the authors introduce the following toy problem in the opening paragraph to A Short Introduction To Boosting: “A horse-racing gambler, hoping to maximize his winnings, decides to create a computer program that will accurately predict the winner of a horse race based on the usual information...” As discovered by the early pioneers of expert systems in the 70s and 80s and as acknowledged by the authors, the biggest stumbling block to using experts is that many of them are unable to detail their decision process or even to rank order the importance of key variables. In light of this issue, Freund and Schapire point out that the beauty of boosting is that it builds on what experts can do not on what they cannot do, namely, given a specific scenario they are usually able to make a judgment in favor or against a particular outcome. For instance, we can ask an expert handicapper if a specific scenario - course and distance winner in last outing a week ago - would lead him to believe that it was more likely to win again or to finish out of the money. Note the phrase “more likely to” – this is a key strength of boosting - it asks for the balance of probabilities and not for the highly probable. The boosting phase combines many such simple, scenario-based rules into an overall weighted decision for an upcoming event. In its original specification, the defining quality of boosting was that it aggregated an incomplete set of “if-then” rules (decision stumps) that recursively address unexplored regions (areas for which previously chosen rules would give incorrect predictions) of existing data sets. The inherent strength of this approach is that it automatically leverages the key dimensions of Wisdom Of Crowds, namely, diversity, independence, decentralization, and aggregation For a worked example applied to NFL prediction, see James McCaffrey’s Classification And Prediction Using Adaptive Boosting. For the underlying theory of why wisdom of crowds works, see Scott Page’s excellent The Difference. And finally, Malacaria And Smeraldi explore the relationship between the AdaBoost weight update procedure and Kelly’s theory of betting and also establish a connection between AdaBoost and Information theory in On AdaBoost And Optimal Betting Strategies.

Friday, December 23, 2016

Biased Coin (Haghani & Dewey, 2016)

A recent paper by Haghani & Dewey (2016) sheds an unflattering light on subjects formally trained in finance as to their lack of basic knowledge with respect to probability and uncertainty – “If a high fraction of quantitatively sophisticated, financially trained individuals have so much difficulty in playing a simple game with a biased coin, what should we expect when it comes to the more complex and long-term task of investing one’s savings?” Though an otherwise interesting study, there are a couple of key points which do not receive adequate attention in the paper:
  • Financial: Though the median final bankroll of $10,504 is derived in the footnotes, there is not sufficient attention drawn to it in the paper itself. Time-Value automatically generates this value whereas Expected-Value generates the wholly unrealistic $3,220,637.
  • Psychological: The fallacy of “Playing With House Money” – “…you are offered a stake of $25 to take out your laptop to bet on the flip of a coin for thirty minutes.” What would have happened if the subjects had to pay $25 to play instead of being given it for free? 


No less a luminary in both the financial and gambling worlds than Ed Thorp says: “This is a great experiment for many reasons. It ought to become part of the basic education of anyone interested in finance or gambling.

Monday, November 07, 2016

Evens-Equivalent Trades

Risk of Ruin (RoR) (Epstein 2009) provides an easily understood metric (probability of bankroll depletion before doubling it) with which to compare strategies. By way of illustration, let us assume that both you and your brother are recreational handicappers. You trade baseball, home-underdogs and he trades horse-racing, second-favorites, as follows:

Sport
Bank
Trades
Avg.
Odds
Avg.
Win
Rate
Avg.
Stake
Evens
Odds
Evens
Win
Rate
Evens
Stake
Expect.
Value
Std.
Dev.
EV.SD
Edge
Rsk
Of
Ruin
MLB
5,000
500
2.10
51.00%
250.00
2.00
53.37%
263.05
17.75
262.45
219
7.10%
7.11%
H-R  
2,500
1,000
3.50
31.00%
100.00
2.00
52.62%
162.10
8.50
161.87
363
8.50%
16.53%

In the classic treatment of ruin, there is a working assumption of even-money trades to make the calculations tractable. To that end, we must first transform our real-world trades into their even-money equivalents with the same edge and volatility, see Krigman (1999). Despite having a smaller edge and a larger stake, you have a lower probability of depletion than your brother principally because you are risking a lower percentage of your bankroll per trade. Ideally, your RoR should be below 5% and to achieve this you both would have to either increase your bankroll or decrease your stake, as follows. [(MLB: 5%, 218.20 or 5730); (H-R: 5%, 54.97 or 4,546)].



Note that Edge's impact only equates with that of Volatility after 219 trades for you and 363 trades for your brother. And it takes a minimum of 806 trades for you and 1336 trades for your brother before you can be at least 95% confident that the combined effects of positive edge and mixed-bag volatility work in your favor to guarantee positive bankroll growth. In other words, despite having potentially successful trading strategies, you both will be well into your second season of handicapping before you can be sure of beginning to reap the benefits!


Saturday, October 01, 2016

Juvenile Finish Position Ratings

In horse-racing, 2yos are by definition the least exposed runners at the racetrack. In the context of handicapping 2yo races, we can create a simple model to generate race-specific ratings based solely on finishing positions and number of runners per race in each horse's past performances record. These ratings will guide our elimination process. Remember that, in handicapping, elimination of probable non-contenders is always preferred to the selection of possible contenders.
  • First, calculate "horses beaten" (n-f) and "horses beaten by" (f-1) from finishing positions (f) and number of runners (n) for each race in a horse's past performances. 
  • Then, sum across all races for wins (w=Σ(n-f)) and losses (l=Σ(f-1)) respectively.
  • Next, calculate a horse's posterior probability m=(w+α)/((w+α)+(l+β)). Prior wins (α) and losses (β) are derived from two full seasons of 2yo races and are equivalent to a horse finishing fourth of seven runners in a virtual race. Note that a first-time starter would automatically have a posterior probability of 0.50=(0+3)/((0+3)+(0+3))
  • Finally, convert probabilities to performance ratings (min:112=8-00, max:126=9-00) using the following formula: r=((a+(b-a))*((m-x)/(y-x))), where a=out.min, b=out.max, x=in.min, and y=in.max of all runners in the current race.
The final rank ordering of horses is important, not the absolute performance ratings.
In summary
, this finishing position rating system (fpr) does not take into account the strength of opposition, beaten lengths, weight carried, or finishing times; however, when it is based on a whole season of results, the fpr ranks correlate approximately 0.87 with the equivalent ranks from an Elo rating system.

As luck would have it, Sunday's Prix Marcel Boussac (Fillies' Group 1) - France's top 2yo fillies race - had a 0.92 correlation between fpr ranks and finishing positions, with the 10/1 winner (Wuheida) top-rated! Not scientific, nevertheless I get to keep the winnings.

Thursday, September 01, 2016

Speed-Stamina Fingerprints

In 1981, Peter Riegel formulated an equation for the relationship between distance and time of athletics world-records. In 1982, Steve Roman adapted Riegel's equation to try and resolve the Secretariat Preakness controversy.
In a similar vein, we can generate unique speed-stamina, power-law fingerprints for racecourses (and horses) based on the best times for various distances. The simplest interpretation of these racecourse fingerprints is to confirm our expectations of the demands imposed by similar distances for different course configurations (Epsom is faster than either Ascot or Newmarket (lower y-intercept); Ascot and Newmarket have similar stamina profiles (same slopes)). Another possible insight might be how these course fingerprints reflect the potential impact on horses with different pace profiles (early speed at Epsom). A further analysis might be on how to better baseline and equate speed figures at different racecourses (use standard course with own speed-stamina equation). Finally, more controversially, using speed-stamina fingerprints for classic-generation (3yos) horses to match with course fingerprints in the lead-up to Group 1 contests (Epsom Derby) or for comparing performances from different classic generations (Frankel vs Sea The Stars).

Note the graph only shows best times and power-law equations for five, six, and seven-furlong races at Ascot, Epsom, and Newmarket and are for illustrative purposes only.

Wednesday, August 03, 2016

Wisdom Of Crowd Market Index (WCMI)

In sports markets, the probabilities implied by the prices on offer are a proxy for the wisdom of the crowd for that particular event. By adapting Shannon's Entropy formula, we can generate our own "Wisdom Of Crowd Market Index" (WCMI) to represent this information on a scale from 0 - 1.

First, calculate implied probabilities of prices: x = 1.00 / d (where d = decimal odds). Next, calculate log probabilities: y = log(x, n) (where n = number of entrants). Then, multiply probabilities by log probabilities: z = x * y. Finally, sum products and subtract from one for final index: wcmi = 1 - (-sum(z)).
WCMI Entropy Chart 0

WCMI Entropy Chart 1
WCMI Entropy Chart X

Note that the index is at a minimum (0.00) when the market is completely uninformed about the outcome (all prices are the same) and at a maximum (1.00) when the market has closed (no price for winner and maximum price for all others). In the realistic exchange market above (snapshot of prices taken one minute before going in-play), the crowd is relatively uninformed (0.03) about the likely outcome and presents an excellent opportunity for the informed sports trader. Personally, I do not trade in any market with an index above 0.13 (approx).

It is very gratifying to note that FlatStats are now (29-Dec-2017) using our WCMI as a guide to those markets in which the crowd is less well-informed!

The 'wisdom' in the title refers to the crowd's aggregate opinion, not to its accuracy; the index scores how decisive that opinion is.