Rating systems
This page describes the package's rating line: systems that maintain a belief about each contestant in a multi-entrant contest (a race, a tournament, a leaderboard) and update it from observed finishing orders. The design differs from the TrueSkill family in two respects. Beliefs are whole densities on a lattice rather than parametric (μ, σ) pairs, and each event is absorbed through the exact likelihood of the full finish order rather than a pairwise or stagewise decomposition. Every claim below is benchmarked against TrueSkill, OpenSkill, Glicko-2 and Elo on twelve datasets, with bookmaker markets as ceilings where odds exist.
Beliefs and updates
A contestant's ability is represented as a density on a lattice,
with no parametric family imposed: skewed and multimodal beliefs
are representable, and the race noise is a kernel of the modeller's
choosing: Gaussian, Student-t, or a mixture with a uniform
"retirement" component for contests whose entrants can fail
outright. The update applies the exact likelihood of the observed
finish order, computed by an O(N) forward chain, and a single
update reproduces exact Bayes to lattice precision. Predictions are
exact winner-of-many probabilities from the same machinery as the
race engine, including tie handling.
The construction is described in the
repository
documentation. The package separately ships
winning.ratings, a Gaussian moment-update line: beliefs
there are (μ, σ) pairs advanced by exact N-way moment
updates, the correlated-race generalization of the TrueSkill-style
factor. This page describes the density line.
Every ranked update takes order, the entrant indices
from first to last finisher. A rank array, the position of each
entrant, is the inverse permutation; both are valid permutations,
they coincide at two entrants, and at three or more the rank array
is silently wrong and yields a plausible walk that barely updates.
order_from_positions, order_from_performance
and order_from_times name the conversion at the call
site and raise on ties, and positions_from_order
inverts. The self-test is to repeat one order thirty times from a
flat prior and check that the means fall along it.
Factor ratings and block correlation in winning.ratings
The shipped winning.ratings line carries two entry
points beyond one number per entity. fit_factor_ratings
estimates an entity's ability as a level plus offsets over
conditions observed on the contest (time control, task category,
surface) by MAP through the exact ranking likelihood, with one
ridge penalty per covariate. The penalty vector is the mechanism,
not a convenience: the level is estimated from every contest the
entity enters, and the offsets are shrunk toward it, so an entity
with few contests under some condition falls back on its level
rather than on zero.
Measured on 65,178 Chatbot Arena battles and 59,399 Lichess games. On chess, with both sides tuned on the same validation split, a partially pooled factor rating beats Lichess's deployed architecture (Glicko-2 maintained separately per time control) by 0.008–0.012 nats of held-out log loss across three independent splits, every interval excluding zero. The win decomposes, and the headline should not be quoted without the decomposition: the estimator, a batch Thurstonian fit against Glicko-2's online update, is worth −0.0080 and the factor structure −0.0033, and the structure is the part this API implements.
Both sides have to be tuned on parameters that bite, and
Glicko-2's tau does not: it enters only the volatility
iteration, sweeping it returns bit-identical loss, and selecting it
is a tie-break rather than a tuning. Tuning on its rating period
and initial deviation instead moves the headline by a third and
the estimator column by three fifths, which is the sensitivity to
carry when reading any such comparison. The period selects at the
top of its grid here, so the margin is an upper bound, though it
saturates rather than running away and the residual is bounded
near 0.0001.
Two thirds of the estimator column is the privilege of a batch fit. Run prequentially over the month with both sides tuned, the online tracker still beats Glicko-2 pooled by −0.0029 [−0.0043, −0.0014] and the deployed per-time-control architecture by −0.0023 [−0.0046, −0.0001], so −0.0051 of the batch column is revisiting the training set and −0.0029 is a better estimator. That leaves the estimator and the factor structure about the same size.
Stratification helps neither side once both are tuned: it is worth −0.0011 to Glicko-2 and costs this estimator's scalar 0.0006, and online it is worth −0.0005 and +0.0015, neither significant. Left at its default rating period Glicko-2 appears to gain +0.0074 from stratification, which is a property of the default rather than of the category information. Partial pooling beats both: within one estimator, pooling the level across conditions and shrinking the offsets beats both the scalar and a per-condition rating. On Arena the advantage grows with cell sparsity. On chess it is roughly constant across player-density thresholds from 100 games down to 10 (2,373 players, 109k games, every interval excluding zero), so it is not concentrated among data-poor players, with the caveat that a 10-game minimum excludes the genuinely sparse regime.
The time-control signal is real and stratification is the lossy
way to extract it, paying sparse-cell variance for a benefit that
partial pooling already captures. That is why the shrunk-offset
form is the recommendation whatever the base estimator. With one
shared penalty the same factor arm loses (+0.0037), so
sweep_offset_ridge tunes the offset penalty on a
validation slice and reports when the optimum is at the scalar
limit, which means the condition axis is unidentifiable in that
data. The gain is real, bounded and channel-limited: about three
percent of what a rating extracts over uniform on binary
comparisons, sixteen percent on cardinal scores.
Two conditions decide whether the factor form transfers, and
both are checkable before fitting. Abilities must be static
relative to the covariate: on Formula 1, constructor identity
looked worth −0.25 nats against a driver-only baseline until
the baseline was given a tuned recency half-life, after which it
was worth +0.04, because an omitted time dimension flatters any
richer parameterisation. The covariate must be exogenously varied:
chess opening family gives a null result because players are their
openings (per-player tactical share sd 0.415 against a no-choice
binomial null of 0.058), not because the axis is absent.
covariate_contrast_report computes that ratio; a null
on a self-selected covariate is a statement about identification,
not about the axis.
from winning.ratings import fit_factor_ratings, predict_factor, covariate_contrast_report
# events: (subset, order, x) -- entity indices, finishing order best first,
# x = [1, bullet, blitz] read off the contest
covariate_contrast_report(events, n_entities, n_cov=3)["ratio"] # near 1: assigned; far above: chosen
B = fit_factor_ratings(events, n_entities, n_cov=3, ridge=[1.0, 30.0, 30.0])
predict_factor(B, subset=[3, 7], x=[1, 0, 1]) # win probabilities under blitz
The second entry point is block correlation: same-group entrants
share a component of performance that is redrawn at every contest,
and at twenty entrants that restructures the whole ordering
distribution. AbilityTracker accepts groups
with a contest and a correlation share rho,
variance-preserving so every marginal keeps its scale; observation
and prediction use the same loadings, and tune_block_rho
selects rho by the filter's own evidence without
consulting any target event. The shared component is a mean-zero
draw; a persistent group advantage is a different object and
belongs in the means (a group column in the design of
fit_design_ratings).
On the Formula 1 data that motivated the feature the correlation is real and is not a noise-scale artefact: same-team 1–2 finishes occur at 0.274 against 0.123 under independence, and at matched idiosyncratic noise the team structure is worth +29.2 nats of training evidence where pure noise reduction loses 5.2. It does not transfer. Fitted and priced under one model on the current engine, the blocked arm's held-out winner log-loss is worse by +0.0427 [+0.0089, +0.0769] over 106 races, its joint advantage −0.0224 [−0.0558, +0.0081] spans zero, and on the held-out order functional no contrast reaches significance.
The cause is a rho that moves between seasons: the same-team 1–2 rate is 0.632 in 2015 and 0.045 in 2021 (chi-square 28.3 on 12 degrees of freedom, p = 0.005), so a rho fitted on one window is wrong on the next, and the damage is on the fit channel, where fitting abilities under an over-strong rho costs +2.86 [+0.16, +5.58] nats with the price held fixed. Correctly specified on synthetic data (six entrants in three teams, rho = 0.4) the blocked arm beats independence on ability recovery, winner and joint log-loss. The term is sound; the Formula 1 result is about an unstable rho, and persistent team quality still belongs in the means.
One caveat is about the API rather than the data. An order
likelihood sees difference variances, and a shared component
changes the ratio of within-group to between-group difference
scale whichever way it is normalised: at fixed marginal variance
the within-group scale shrinks, at fixed idiosyncratic variance the
between-group scale grows, and an equal-loading global factor
cancels entirely (independence at idiosyncratic variance 0.6 and a
global factor at 0.6 return the same evidence to 4e-15 at seventeen
entrants). So rho and beta2 trade off, a rho selected at fixed
beta2 is a blended estimate, and on Formula 1 that coupling was 28
percent of the damage. tune_block_rho takes a
beta2_grid to sweep both.
from winning.ratings import AbilityTracker, tune_block_rho
sel = tune_block_rho(contests, rho_grid=(0.0, 0.2, 0.4, 0.6), drift=0.03)
trk = AbilityTracker(rho=sel["rho"], drift=0.03)
trk.observe(ids, t, order=order, groups=teams)
trk.predict(ids, t, groups=teams)
Benchmark protocol
Twelve datasets: Formula 1, ATP and WTA tennis, chess, sumo, football, Halo 2, horse racing, and synthetic worlds with known generating truth. The protocol is prequential, each event predicted before it is observed, and each table carries a uniform floor, an oracle where the truth is known, and the bookmaker market as a ceiling where odds exist. The tables live in BENCHMARKS.md and regenerate from committed scripts.
Non-Gaussian noise, and a market-efficiency result
Formula 1 retirements make finishing order a mixture: pace when the car survives, a lottery when it does not. A noise density with a uniform disaster component beats every Gaussian system tested on 1,158 grands prix, and the same estimator, handed qualifying sessions in which no retirement process exists, drives its own disaster mass to near zero: the component is measuring the physics, not absorbing slack. Against the 88 races with genuinely pre-qualifying bookmaker odds, however, no rating system built on public results and grid data adds statistical power to the market: the optimal forecast pool puts zero weight on every one of the six systems tested, and the market's edge is plausibly knowledge of current car form. Both results, with their designs and caveats, are in Rating Formula 1.
Quick start
This API is the research line, developed in the repository's src/ directory and shipping with a future release; the results above are from BENCHMARKS.md. The shipped race engine is documented on the home page.
from winning import ThurstoneRating
tr = ThurstoneRating()
tr.observe(names=["ada", "ben", "cid", "dot"], ranks=[1, 2, 3, 4])
tr.observe(names=["ada", "cid", "eve"], ranks=[1, 2, 3])
tr.win_probabilities(["ada", "cid", "eve"]) # exact field win probabilities
tr.rating("ada") # Rating(mu=..., sigma=...)
tr.leaderboard() # best-first, conservative
Every system in the package speaks the same three verbs:
observe(names, ranks),
win_probabilities(names) and rating(name),
including the shims around the third-party comparators, so a
benchmark row is a one-line swap. Install with pip install
winning (the core depends only on numpy) or pip
install winning[benchmarks] for the comparators.