01
The Arena
Score = approved-drug hits in the model's top-20, per patient, across five chronic-disease cohorts. Random baseline ≈ 3.6–8.1 per disease. Spectrum recall = agreement with the mechanistic gold ranking.
Grader: bench2 · hypergeometric z vs random · temperature=0 · n=100 identical patients per entry.
Blinding: no web retrieval · no indication lookup · no disease-label side channels.
02
Season v1.1 — three signals
THE FIELD IS LEVEL
≤ 8% GAP
Feed the module-state representation and last-gen GLM-4.7 lands within 8% of flagships. Representation, not scale, is doing the work.
DEPRESSION HARDEST, RA STRONGEST
9.65/Q · DEP
Depression carries the highest absolute score; RA carries the strongest lift over random (z ≈ +7 to +14 across arms) — inflammation maps onto module states cleanly.
MECHANISM DISCRIMINATES
0.26 recall@20
Clinical hits and spectrum recall move independently — two orthogonal dials every entry is graded on.
03
Cross-family lanes
The arena reserves a lane for every model family, every version — domestic or frontier. Same 100 patients, same grader, same blinding. Phase 1 opens open submissions.
04
Bring your own model
Zero to leaderboard in seven steps
Any OpenAI-compatible endpoint. Your API key never leaves your machine.
git clone bench2the ruler — protocol package (MIT)python -m bench2.buildrebuild exam papers byte-identically from GEOpython -m bench2.run --model <yours>bring your own API keypython -m bench2.scorethe exact grader this board runs onUSERGUIDE.md in the bench2 repository · code, not data · submissions open Phase 1.