Historical scope reviewed September 8, 2026. The original score names and horizons are preserved. Raw metric magnitudes and correlations do not measure payout importance or prediction diversification. Only the effective round configuration identifies payable metrics; weighted contributions require matched observations. See the current metric definitions.

Numerai’s metrics describe different aspects of prediction scores. MMC measures residual contribution; CORJ60, BMC and FNCv3 provide other diagnostics. Their presence in a chart does not make them payable. Correlation among score series does not identify the diversity of uploaded predictions.

This post measures convergence from four angles: cross-metric correlation, cross-sectional dispersion, metric dominance, and sign agreement.

Cross-Metric Correlation: Are the Metrics Collapsing?

For each round, compute the Spearman rank correlation between every pair of primary metrics across all staked models. A rolling 50-round average smooths the noise and reveals the trend.

Rolling 50-round Spearman correlation between primary Numerai metric pairs (MMC-CORJ60, MMC-BMC, MMC-FNCv3, CORJ60-BMC) from round 300 to 1240, showing MMC-BMC running near 0.95-0.99 while MMC-CORJ60 dips to 0.22 around round 950 then recovers to 0.43

The MMC-CORJ60 correlation is not a one-way climb. It started near 0.45 around round 300, drifted down to roughly 0.22 by rounds 900–1000, then recovered to about 0.43 in the most recent rounds: all of the convergence in this pair is concentrated in the last 250 rounds. The starker signal is MMC-BMC: outside a dip to 0.67 around round 340, it has run between 0.95 and 0.99 since round 550, meaning a model's "originality" score and its "benchmark contribution" score are now nearly the same number, even though they were designed to capture different things. MMC-FNCv3 climbed from roughly 0.45 to 0.68 over the same stretch.

When metrics designed to measure different qualities agree this closely, and the recently-decorrelated pairs turn back upward, the scoring system is rewarding a narrower band of approaches than the metric names suggest.

Dispersion: Is the Field Narrowing?

Cross-sectional standard deviation measures how spread out model scores are within each round (smoothed here with a 20-round average). High dispersion means varied strategies; low dispersion means bunching.

Cross-sectional standard deviation of MMC, BMC, CORJ60, and FNCv3 by round (20-round average), showing MMC and BMC declining from ~0.019 to ~0.0105 while CORJ60 cycles between 0.014 and 0.023 with no downtrend

MMC dispersion dropped from roughly 0.019 near round 200 to 0.0105 in recent rounds (a 45% compression), and BMC tracks it almost point for point. FNCv3 fell from a plateau near 0.018 around round 450 to about 0.013. CORJ60 is the exception: its dispersion cycles between 0.014 and 0.023 with market regime and sits near 0.020 today, showing no compression at all. The field is measurably tighter on three of the four primary metrics.

Narrower score dispersion describes the observed score distribution. Changing targets, market conditions and sample composition can all change that distribution. It does not establish that submitted prediction vectors or underlying strategies have become the same.

Why the raw-magnitude comparison was removed

The original personal-best metric chart compared raw magnitudes across different score definitions. Those scores have different scales and horizons, so the largest raw number does not identify a model's strongest economic contribution or a dominant strategy. That figure and its dominance conclusion have been withdrawn. A payout-contribution comparison must instead use the effective round's configured multipliers on matched observations.

Metric Agreement: The Monoculture Test

A separate descriptive comparison is sign agreement: what percentage of models have all four primary metrics pointing the same direction (all positive or all negative) in a given round? Score sign agreement is not a direct measure of prediction diversity.

Percentage of models per round where MMC, BMC, CORJ60, and FNCv3 all share the same sign, with a 20-round average that collapses from noisy early highs to ~31% near round 300 and then holds between 37% and 48% through round 1240

Sign agreement tells a quieter story. After the noisy early rounds (a tiny field where the 20-round average swung between 80% and 100%), agreement collapsed to about 31% near round 300 and has ranged between 37% and 48% ever since. The average sits near 40% today, almost exactly where it was 900 rounds ago.

This historical stability concerns score signs in the selected sample. Testing prediction diversity would require the prediction vectors themselves and their relationship to the meta-model; this chart cannot establish a monoculture or rule one out.

Takeaways

MMC and BMC have effectively merged. Their rolling correlation has held between 0.95 and 0.99 since round 550. Two of the four primary metrics now rank the field almost identically, and MMC-CORJ60 has turned upward again over the last 250 rounds after a long decline.

Dispersion is compressing on three of four metrics. MMC spread fell roughly 45%, with BMC and FNCv3 close behind. CORJ60 is the holdout: its dispersion cycles with market regime rather than trending down.

Raw score magnitude does not establish economic dominance. The corresponding figure was withdrawn; use the configured payout weights and matched scores instead.

Sign agreement is flat near 40%. The sharpest monoculture test comes back negative: the share of models with all four metrics pointing the same way is where it was 900 rounds ago. The evidence for convergence is real but partial: the metrics are collapsing into each other faster than the models are collapsing into each other. Track these trends on the Trends page and per-round detail in Rounds.