AI4SCIENCE/MOLECULAR-STATE ARENA/EST. 2026

No diagnosis.
Just state.

Language models walk into the arena with one thing: a patient's molecular state — no disease label, no retrieval. They rank 200 candidate interventions. One grader decides.

332-MODULE STATE · 5 COHORTS
MODULE 001MODULE 332
01

The Arena

Score = approved-drug hits in the model's top-20, per patient, across five chronic-disease cohorts. Random baseline ≈ 3.6–8.1 per disease. Spectrum recall = agreement with the mechanistic gold ranking.

Grader: bench2 · hypergeometric z vs random · temperature=0 · n=100 identical patients per entry. Blinding: no web retrieval · no indication lookup · no disease-label side channels.
02

Season v1.1 — three signals

THE FIELD IS LEVEL
≤ 8% GAP
Feed the module-state representation and last-gen GLM-4.7 lands within 8% of flagships. Representation, not scale, is doing the work.
DEPRESSION HARDEST, RA STRONGEST
9.65/Q · DEP
Depression carries the highest absolute score; RA carries the strongest lift over random (z ≈ +7 to +14 across arms) — inflammation maps onto module states cleanly.
MECHANISM DISCRIMINATES
0.26 recall@20
Clinical hits and spectrum recall move independently — two orthogonal dials every entry is graded on.
03

Cross-family lanes

The arena reserves a lane for every model family, every version — domestic or frontier. Same 100 patients, same grader, same blinding. Phase 1 opens open submissions.

04

Bring your own model

Zero to leaderboard in seven steps

Any OpenAI-compatible endpoint. Your API key never leaves your machine.
git clone bench2the ruler — protocol package (MIT)
python -m bench2.buildrebuild exam papers byte-identically from GEO
python -m bench2.run --model <yours>bring your own API key
python -m bench2.scorethe exact grader this board runs on
USERGUIDE.md in the bench2 repository · code, not data · submissions open Phase 1.