Fig 1.
Central phenomenon: non-linear deviation from the power law in individuals.
Top left: performances of world record holders and a selection of random runners. Curves labelled by runners are their known best performances (y-axis) at that event (x-axis). Black crosses are world record performances. Individual performances deviate non-linearly from the world record power law. Top right: a good model should take into account specialization, illustration by example. Hypothetical performance curves of three runners, green, red and blue are shown, the task is to predict green on 1500m from all other performances. Dotted green lines are predictions. State-of-art methods such as Riegel or Purdy predict green performance on 1500m close to blue and red; a realistic predictor for 1500m performance of green—such as LMC—will predict that green is outperformed by red and blue on 1500m; since blue and red being worse on 400m indicates that out of the three runners, green specializes most on shorter distances. Bottom: using local matrix completion as a mathematical prediction principle by filling in an entry in a (3 × 3) sub-pattern. Schematic illustration of the algorithm.
Table 1.
The three components of the low-rank model of Eq (1).
Fig 2.
The three components of the low-rank model, and explanation of the world record data.
Left: the components displayed (unit norm, log-time vs log-distance). Tubes around the components are one standard deviation, estimated by the bootstrap. The first component is an exact power law (straight line in log-log coordinates); the last two components are non-linear, describing transitions at around 800m and 10km. Middle: Comparison of first component and world record to the exact power law (log-speed vs log-distance). Right: Least-squares fit of rank 1-3 models to the world record data (log-speed vs log-distance).
Table 2.
Out-of-sample RMSE for prediction methods on different data setups.
Fig 3.
Matrix scatter plot of the three-number-summary vs performance.
For each of the scores in the three-number-summary (rows) and each event distance (columns), the plot matrix shows: a scatter plot of performances (time) vs the coefficient score of the top 25% (on the best event) runners who have attempted at least 4 events. Each scatter plot in the matrix is colored on a continuous color scale according to the absolute value of the scatter sample’s Spearman rank correlation (red = 0, green = 1).
Fig 4.
Scatter plots exploring the three number summary.
Top left and right: 3D scatter plot of three-number-summaries of runners in the data set, colored by preferred distance and shown from two angles. A negative value for the second score is a indicates that the runner is a sprinter, a positive value an endurance runner. In the top right panel, the summaries of the elite runners Usain Bolt (world record holder, 100m, 200m), Mo Farah (world beater over distances between 1500m and 10km), Haile Gabrselassie (former world record holder from 5km to Marathon) and Takahiro Sunada (100km world record holder) are shown; summaries are estimated from their personal bests. For comparison we also display the hypothetical data of a runner who holds all world records. Bottom left: preferred distance vs individual exponents, color is percentile on preferred distance. Bottom right: age vs. exponent, colored by preferred distance.
Table 3.
Estimated three-number-summary (λi) for a selection of elite runners.