MODEL
Every number below is measured, not claimed
NO MODEL VERSION RECORDED · SCORE v0.1
SEASONS
2021–2025
GAMES SCORED
100.0% of 1,424 completed games
SIGNALS
across 8 feature families
ACCURACY VS BASELINES
- ELO 67
- ARCADIA 62.8 One model, shown once: logreg-ma and logreg-mf are the same probability function for every possible input row (registry probability_fingerprint 5c984a4821f9…), so the two rows this used to draw were one measurement counted twice. Variants: market_aware, market_free. Production artifact: v1-logreg-mf-2026.08c.
Model accuracy measured against the benchmarks below. The de-vigged market is the bar; it is not present in the run these figures come from, so no gap can be reported from them.
CALIBRATION
EACH POINT IS A PROBABILITY BIN · DOT AREA GROWS WITH BIN POPULATION · THE DIAGONAL IS PERFECT CALIBRATION
EXPECTED CALIBRATION ERROR 0.0844
| PREDICTOR | ECE |
|---|---|
| ARCADIA | 0.1079 |
| ELO | 0.0844 |
SOURCE eval_metrics.ece_quantile
RUN exit8-persist-0814-cal
·
MEASUREMENT PROVENANCE
| FIGURE | VALUE | SOURCE |
|---|---|---|
| SEASONS | 5 | SELECT COUNT(DISTINCT season) FROM plays |
| GAMES SCORED | 1,424 | completed games carrying play-by-play rows (games JOIN plays, home_score NOT NULL) |
| SIGNALS | 97 | registry.load_production_families() |
| PLAYS INGESTED | 247,284 | SELECT COUNT(*) FROM plays |
| PLAY-BY-PLAY COVERAGE | 100.0% | completed games carrying play-by-play rows, over completed games in the schedule |
| WEATHER FORECAST ROWS | 0 | SELECT COUNT(*) FROM weather WHERE kind = 'forecast' |
DISCLOSURES
-
promotion of 'v1-logreg-ma-2026.08c' (variant=market_aware) is BLOCKED: it is bit-identical to 'v1-logreg-mf-2026.08c' (variant=market_free, status=production) — same probability function for every possible input row (probability_fingerprint 5c984a4821f90444…, 55 effective terms). MEASURED on the artifact: all 11 'market'-family feature(s) carry coefficient exactly 0.0 (e.g. mkt_books_n, mkt_home_prob_close, mkt_home_prob_close_missing), so that family contributes nothing to the logit. Two production artifacts that are the same function are ONE probability stream under two names: the Arcadia Score would count it twice, agreement.min_families would be satisfied by a duplicate, and the dispersion over it is identically 0 — full agreement credit on no evidence. UPSTREAM (eval.market.clv_status): available=True, feed=current — 37660 pre-kickoff snapshot rows across 272 games from 25 capture instant(s); feed current — 37660 pre-kickoff odds snapshot(s), newest 2026-08-15T18:12:00Z is 1.2h old (stale after 20.0h) This block is about the CONDITION (no independent signal), not about a variant name: it clears itself the moment the two artifacts are genuinely different functions. Retrain once real pre-kickoff odds snapshots exist, or archive one of them.
data/models/registry.json + registry.probability_fingerprint -
STILL PROVISIONAL, AND THE ANCHORS ARE DELIBERATELY NOT FITTED (EXP-006, rewritten and ratified 2026-08-13). CORRECTION OF A FALSE PUBLISHED STATEMENT: this field previously claimed 'on the 2025 pregame slate the tier populations are Pass 52.7% and Elite 0.0%, against AC-07's 25-40% Pass / <=3% Elite'. THAT WAS STALE. It was measured in the sigma_ref = 0.06 era, where sigma_ref sat below 98.2% of measured bootstrap sigmas, so certainty was 0 almost everywhere, the multiplicative composite was 0, and EVERY score collapsed to 0 — which is why everything was Pass and nothing was Elite. sigma_ref was refitted 2026-08-10 and the claim has not described the shipped system since. MEASURED TRUTH on the untouched 2025 season (n=250, as-of gated, production artifact v1-logreg-mf-2026.08c, p_cal = identity, uniform-raw agreement footing): Pass 33.6% and Elite 2.0% — BOTH AC-07 CRITERIA ARE IN BAND with the hand-reasoned placeholders. WHY NO REFIT: a refit is measured to make a passing criterion FAIL (in-sample on 2025, the most favourable case a refit can get, it moves Elite 2.0% -> 3.2%, out of band), and the recipe cannot clear the bound by construction — anchor_map pins ordinate 0.95 at the q=0.97 x-anchor, rounding is floor(1000*A+0.5) so A=0.95 scores exactly 950, the Elite floor is 950 INCLUSIVE, therefore Elite population == #{raw >= q97} ~ 0.03n + 1, i.e. >= 3% at every n this project will have. A refit that worsens a criterion is not a fix. THE FLAG THEREFORE STAYS true BECAUSE THE X-VALUES REMAIN UNFITTED, not because work is outstanding: 'we measured that fitting makes it worse' is a good reason to keep placeholders and a bad reason to call them fitted. BINDING VALIDATION FAILED AND IS DISCLOSED, NOT REPAIRED (PLAN.md Part 7: a failed validation is a finding, not a re-tune): on untouched 2025 the monotone tier chain inverts at Very High 0.806 > Elite 0.800 (deficit +0.006, z=+0.03, on n=5 Elite rows — one game), and Sep-vs-Dec parity fails on the stratified CMH statistic at z=+2.83 with NO individual tier binding-failing. Persisted to score_tier_stats; see docs/experiments/EXP-006-anchor-refit.md. ALSO DISCLOSED: agreement.families still spell names ending '_cal', but under EXIT 6's NO CALIBRATOR verdict (extended to score.py by operator ruling 2026-08-13, see score.UNIFORM_RAW_FOOTING) NO family carries a calibrator — the suffix is historical and family_sources records the true per-game provenance. certainty.sigma_ref is FITTED (see certainty.sigma_ref_note) but was fitted against v1-logreg-mf-2026.08 while production is now v1-logreg-mf-2026.08c; that pairing is carried, not re-derived. GOVERNED BY EXP-009, which re-derives sigma_ref on UNIFORM RAW FOOTING (pre-registered under CONVENTIONS RULE 5 at docs/experiments/EXP-009-sigma-ref-raw-footing.md, falsification conditions declared before results, HARD STOP 2026-09-09 - closed before the first in-season edition). THIS IS THE THIRD ARTICULATION OF ONE DEBT, not three items, and EXP-009 retires all three at once or fails informatively: (1) MODEL VERSION - sigma_ref was fitted against v1-logreg-mf-2026.08, production is v1-logreg-mf-2026.08c (the pairing above); (2) SCALE OF THE CONSTANT - sigma_ref_note records the 795 bootstrap sigmas were 'passed through that model's Platt calibrator', so the reference sits on the CALIBRATED scale while under EXIT 6 / park ruling R5 the shipped probability is RAW (p_cal == p_raw); (3) SCALE AT SERVE TIME - score._sigma_from_artifact still applies that Platt fit to the 200 draws, bending the MEASUREMENT to match the mis-scaled REFERENCE. (3) is the engine's last live calibrator application, registered CONTAINED_PENDING_EXP-009 - governed and contained, NOT ungoverned - in tests/model/test_calibration_single_dispatch.py; effect while it stands, mean 8.2 score points over 16 games, 0 tier moves. They cannot be split: dropping (3) alone compares a raw-scale sigma against a calibrated-scale reference, and deleting the artifact makes the seam raise ScoreUnavailableError('sigma_uncalibrated') for EVERY game, so both must be re-derived against (1)'s production artifact - a measurement, not an edit, which is why R5 declined to close it by fiat. RULED BY THE OPERATOR 2026-08-15, narrower reading upheld: certainty.sigma_ref_provisional's false is TRUE at source. The CONSTANT is fitted and 2025-validated (certainty.sigma_ref, sigma_ref_note); what is CARRIED is (1)'s model-version pairing, EXP-009's jurisdiction. A reader who checks that flag and stops is told something TRUE; what it omits is the pairing. Nothing in this note licenses flipping that flag - and no hand-flip was available: refit_config() OWNS the field and pops this note (model/score.py:1125, :1131), so both are overwritten by the next refit. The flag is the measurement process's testimony, not an editor's - effective-config attestation in its strongest form. Flag flip, version bump and errata record: each considered, none taken. See docs/plans/engine-2-model-backtest-score.md Item 7.
data/models/score/score-v0.1.json -
Operator ruling (docket G2): the market family is inert at any PREGAME as_of BY DESIGN. Closing lines are kickoff-gated (blueprint §9 trap #1) so mkt_* MUST be null before kickoff — that is the leakage protection working, not a break. Part 5's benchmark reads asof.closing_lines() at a post-kickoff as_of instead of asking this family. If these columns ever populate at a Friday as_of, THAT is the integration failure.
arcadia.features.integrity.DECLARED_EMPTY_EXEMPTIONS -
Operator ruling 2026-08-06 (docket G1): weather is INERT for v1. The DB holds 1,424 observed rows and ZERO forecasts, and asof.py correctly forbids observed weather pregame (it is post-game knowledge, blueprint §9). No purchased historical forecasts and no observed-as-forecast substitution — that would be a leak. The family stays wired so 2026 forecasts accumulate; it becomes learnable at the first retrain with real forecast history. 'Weather inert in v1' must be disclosed wherever model composition is described, the Model screen included. SCOPE (2026-08-10): this ruling is about absent FORECAST history and covers the forecast columns only; roof_closed and surface_turf are structural venue data, ruled on their own causes.
arcadia.features.integrity.DECLARED_EMPTY_EXEMPTIONS -
How far to trust the tier table. The confidence tiers are not in outcome order on the measured season — Very High 0.806452 over Elite 0.800000, a gap of 0.006452. A higher tier with a lower hit rate is a failure of the monotone criterion, and it is disclosed here because it was not remediated.Elite is the higher tier but not the better outcome: Very High hit 0.806452 against Elite's 0.800000 over 250 scored games, a deficit of 0.006452. Tier order and outcome order disagree here; the failure is reported rather than tuned away, and Elite is not claimed to be distinguishable from Very High on this evidence. Moderate is the higher tier but not the better outcome: Pass hit 0.619048 against Moderate's 0.456522 over 250 scored games, a deficit of 0.162526. Tier order and outcome order disagree here; the failure is reported rather than tuned away, and Moderate is not claimed to be distinguishable from Pass on this evidence.
score_tier_stats, pinned to one persisted run via eval_runs · run exit8-persist-0814-tiers-2025 -
How far to trust the tier table. Elite rests on fewer than 19 scored games. Its hit rate is printed with the rest of the table and is not the same kind of measurement as the tiers below it.Elite is validated on 5 of 250 scored games, under the 19-game floor below which a hit rate is too wide to separate a tier from the rest of the table. Read its 0.800000 as an estimate this season cannot pin down, not as a settled rate standing level with the tiers beneath it.
score_tier_stats, pinned to one persisted run via eval_runs · run exit8-persist-0814-tiers-2025
MODEL COMPOSITION
SOURCE registry.load_production_families()
- INJURY ADJUSTMENT 13 SIGNALS
- DATA QUALITY 4 SIGNALS
- MARKET VALUE 11 SIGNALS MARKET INERT IN V1 Operator ruling (docket G2): the market family is inert at any PREGAME as_of BY DESIGN. Closing lines are kickoff-gated (blueprint §9 trap #1) so mkt_* MUST be null before kickoff — that is the leakage protection working, not a break. Part 5's benchmark reads asof.closing_lines() at a post-kickoff as_of instead of asking this family. If these columns ever populate at a Friday as_of, THAT is the integration failure.
- LONG-RUN RATING 6 SIGNALS
- QUARTERBACK EDGE 14 SIGNALS
- REST AND TRAVEL 19 SIGNALS
- TEAM FORM 18 SIGNALS
- WEATHER 12 SIGNALS WEATHER INERT IN V1 Operator ruling 2026-08-06 (docket G1): weather is INERT for v1. The DB holds 1,424 observed rows and ZERO forecasts, and asof.py correctly forbids observed weather pregame (it is post-game knowledge, blueprint §9). No purchased historical forecasts and no observed-as-forecast substitution — that would be a leak. The family stays wired so 2026 forecasts accumulate; it becomes learnable at the first retrain with real forecast history. 'Weather inert in v1' must be disclosed wherever model composition is described, the Model screen included. SCOPE (2026-08-10): this ruling is about absent FORECAST history and covers the forecast columns only; roof_closed and surface_turf are structural venue data, ruled on their own causes.
INERT-FAMILY RULINGS arcadia.features.integrity.DECLARED_EMPTY_EXEMPTIONS
THE ARCADIA SCORE
0–1000 CONFIDENCE INDEX · NEVER A PERCENTAGE
Arcadia Score v0.1. score = 1000 * A(edge^1.0 * certainty^0.6 * agreement^0.6 * dq^1.0). The score is a 0-1000 confidence index and is NEVER a win probability or a percentage; the win probability travels separately (detail.p_cal).
ROUNDING floor(1000 * A(raw) + 0.5), clamped to [0, 1000]
SCORE CONFIGURATION PROVISIONAL
| TIER | FROM | TO |
|---|---|---|
| ELITE | 950 | 1000 |
| VERY HIGH | 850 | 949 |
| HIGH | 750 | 849 |
| STRONG | 650 | 749 |
| MODERATE | 550 | 649 |
| PASS | 0 | 549 |
SOURCE data/models/score/score-v0.1.json
How far to trust the tier table.
- 1 — The top of the chain rests on five games. The Elite tier holds 5 of 250 scored games on the 2025 validation season. Its hit rate is 0.800 with a 95% interval of [0.38, 0.96] — consistent with almost anything. The monotone tier chain fails on a single inversion, Very High 0.806 over Elite 0.800, a gap of 0.006 decided by one game (z = +0.03). We report the failure rather than tuning it away, and we do not claim Elite is distinguishable from Very High on this season's evidence. It is not.
- 2 — The Pass band is an abstention band, and its pooled hit rate is confounded. Games reach Pass for different reasons, and averaging them answers no question. Games abstained for a thin edge measure the tier; games held on data quality carry picks that were never published, so they measure the model. Split by cause on 2025: edge-driven 0.587 (n=63), data-quality-driven 0.714 (n=21). The pooled 0.619 is the mixture of the two and should not be compared with any other tier. The edge-driven stratum runs above Moderate (0.457) by 13 points, which does not clear our significance bar at this sample size (z = +1.35) and would need roughly two more seasons to settle. We record it as an open question, not a finding.
Neither limitation is remediated by moving an anchor, and neither has been. The anchor map is unfitted and declared provisional; a refit is measured to make a passing criterion fail, and is refused.
RULED COPY docs/experiments/EXP-006-anchor-refit.md §11c, ruled for wiring by docs/design/EXIT-12-PARK.md §5.4 (ruling R2)
MEASURED FROM score_tier_stats, pinned to one persisted run via eval_runs
TIER HONESTY · WHAT THE TIERS ACTUALLY DID
| TIER | GAMES | HIT RATE | DISCLOSURE |
|---|---|---|---|
| ELITE | 5 | 0.800000 | Elite is the higher tier but not the better outcome: Very High hit 0.806452 against Elite's 0.800000 over 250 scored games, a deficit of 0.006452. Tier order and outcome order disagree here; the failure is reported rather than tuned away, and Elite is not claimed to be distinguishable from Very High on this evidence. Elite is validated on 5 of 250 scored games, under the 19-game floor below which a hit rate is too wide to separate a tier from the rest of the table. Read its 0.800000 as an estimate this season cannot pin down, not as a settled rate standing level with the tiers beneath it. |
| VERY HIGH | 31 | 0.806452 | |
| HIGH | 39 | 0.743590 | |
| STRONG | 45 | 0.666667 | |
| MODERATE | 46 | 0.456522 | Moderate is the higher tier but not the better outcome: Pass hit 0.619048 against Moderate's 0.456522 over 250 scored games, a deficit of 0.162526. Tier order and outcome order disagree here; the failure is reported rather than tuned away, and Moderate is not claimed to be distinguishable from Pass on this evidence. |
| PASS | 84 | 0.619048 |
SOURCE score_tier_stats, pinned to one persisted run via eval_runs · run exit8-persist-0814-tiers-2025
DATA ON HAND
PLAYS INGESTED 247,284
PLAY-BY-PLAY COVERAGE 100.0% 1,424 of 1,424 completed games
WEATHER FORECAST ROWS 0 against 1,424 observed rows
| DATASET | ROWS | SOURCE |
|---|---|---|
| GAMES ON SCHEDULE | 1,696 | SELECT COUNT(*) FROM games |
| PLAYS | 247,284 | SELECT COUNT(*) FROM plays |
| INJURY REPORT ROWS | 29,149 | SELECT COUNT(*) FROM injuries |
| ODDS SNAPSHOTS | 42,133 | SELECT COUNT(*) FROM odds_history |
| WEATHER OBSERVATIONS | 1,424 | SELECT COUNT(*) FROM weather |
| SNAP-COUNT ROWS | 132,616 | SELECT COUNT(*) FROM snap_counts |
| DEPTH-CHART ROWS | 754,635 | SELECT COUNT(*) FROM depth_charts |
| FEATURE BUILDS | 334 | SELECT COUNT(*) FROM feature_runs |
| PREDICTIONS RECORDED | 358 | SELECT COUNT(*) FROM predictions |
| EDITIONS BUILT | 5 | SELECT COUNT(*) FROM editions |
| EVALUATION RUNS | 23 | SELECT COUNT(*) FROM eval_runs |
FEED FRESHNESS
- DEPTH CHARTS PAST THRESHOLD 120 h · THRESHOLD 48 h
- INJURIES PAST THRESHOLD 551 h · THRESHOLD 36 h
- ODDS AGE UNKNOWN 25 h · THRESHOLD 20 h
- SCHEDULES PAST THRESHOLD 550 h · THRESHOLD 168 h
- WEATHER PAST THRESHOLD 551 h · THRESHOLD 24 h
hours, surfaced so the UI never hard-codes them. These drive the freshness BADGE only -- edition holds are decided by gate_preview, not by ingest age. Exactly ONE of them is derived: 'odds' is the B15 bound taken from the declared merge cadence. The other four are operator-calibration and say so in threshold_sources -- only odds has a declared cadence to enumerate, and inventing derivations for the rest would be inventing the numbers.
PIPELINE RECENCY
- ODDS WITHIN THRESHOLD 1.2 h · THRESHOLD 20 h
- PREDICTIONS INCONCLUSIVE 26.5 h · NO DERIVED THRESHOLD
- EDITIONS PUBLISHED INCONCLUSIVE 4.4 h · NO DERIVED THRESHOLD
- EVAL RUNS INCONCLUSIVE 26.7 h · NO DERIVED THRESHOLD
PREDICTIONS — INCONCLUSIVE -- this is NOT evidence the surface is healthy. The age is measured; no staleness bound is derivable for this subject, because no publishing calendar is installed to enumerate. See arcadia/freshness.py for the derivation method a bound must come from.
EDITIONS PUBLISHED — INCONCLUSIVE -- this is NOT evidence the surface is healthy. The age is measured; no staleness bound is derivable for this subject, because no publishing calendar is installed to enumerate. See arcadia/freshness.py for the derivation method a bound must come from.
EVAL RUNS — INCONCLUSIVE -- this is NOT evidence the surface is healthy. The age is measured; no staleness bound is derivable for this subject, because no publishing calendar is installed to enumerate. See arcadia/freshness.py for the derivation method a bound must come from.
COUNTS ARE NOT RECENCY. A count cannot tell a live pipeline from a stopped one, which is the defect arcadia/freshness.py exists to name; every subject here reports the newest KNOWLEDGE instant it has and its age. A subject with no derived staleness bound reports its age and refuses the verdict rather than reporting green.