Current leaderboard
A lane is the complete combination of model, agent harness, provider route, and effort. Balanced Review Accuracy gives equal weight to anchored recall and usable specificity. Tier-1 recall and positive-fixture noise ≤ 0.12 are hard gates.
DeepSeek V4.1 Flash via DeepSeek via CLIProxyAPI is deployed and passes the current qualification gates.
Scroll horizontally to see every column.
| Rank† | Model | Balanced score | 95% CI | Status |
|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash | 97.81% | 95.6%–100.0% | Deployed |
Recall | ||||
| 1 | Qwen 3.8 Flash | 95.63% | 92.7%–98.6% | Below gate: Tier-1 miss |
Recall | ||||
| 1 | Grok 4.6 | 95.48% | 92.2%–98.8% | Candidate |
Recall | ||||
| 4 | GLM-5.3-Flash | 94.93% | 91.6%–98.3% | Candidate |
Recall | ||||
| 4 | GLM-5.3-Flash | 94.81% | 91.5%–98.1% | Below gate: Tier-1 miss |
Recall | ||||
| 6 | DeepSeek V4 Flash Vision Exp | 91.66% | 87.7%–95.6% | Below gate: Tier-1 miss |
Recall | ||||
| 6 | GPT-5.6 Terra | 91.61% | 86.7%–96.5% | Below gate: Tier-1 miss |
Recall | ||||
| 6 | GPT-5.6 Terra | 89.95% | 83.8%–96.1% | Candidate |
Recall | ||||
| 6 | GPT-5.6 Sol | 88.41% | 81.1%–95.7% | Candidate |
Recall | ||||
Status · Deployed ships the current release · Candidate passes the same gates but is not shipped · Below gate missed a hard gate; its score stays visible for comparison. Open a lane name for its diagnostics, reproduce command, and raw report.
Bars show Balanced Review Accuracy with 95% confidence whiskers on a 0–100 scale. Orange ranks first, the green mark ships in the current release, and hollow bars missed a hard gate but keep their score for comparison, with the failed gate noted beside the effort. The table remains canonical.
Below the gate
These complete reports missed a Tier-1 defect or exceeded the 0.12 positive-noise envelope. Their balanced scores remain visible for comparison, but they receive no rank.
Scroll horizontally to see every column.
| Rank† | Model | Balanced score | 95% CI | Status |
|---|---|---|---|---|
| — | GPT-5.6 Terra | 90.39% | 85.3%–95.5% | Below gate: Tier-1 miss |
Recall | ||||
| — | GPT-5.6 Luna | 88.43% | 81.6%–95.2% | Below gate: Tier-1 miss |
Recall | ||||
Not run
Unavailable routes are shown as blocked, never converted into a zero score.
- Qwen3.8 Flash Next
qwen3.8-flash-nextOpenCode Go subscriptionNot exposed by the authenticated subscription endpoint or Pi model catalog on 2026-08-31.
Not ranked
Compromised or operationally invalid reports remain visible as evidence but never enter the leaderboard.
- Qwen3.8 Max incomplete run
opencode-go/qwen3.8-maxOpenCode Go subscriptionIncomplete at 218/258 saved draws, including 158 reusable valid results. The authenticated provider returned a monthly usage-limit error on 2026-08-31 and reported a reset in 16 days; paid balance was not enabled. Raw Qwen3.8 Max incomplete run report
- Grok 4.6 initial run
grok-4.6Grok subscriptionVoided after anti-cheat v2 detected structured canary adoption in candidate and final review text. Excluded from every aggregate and rank. Raw Grok 4.6 initial run report
How to read it
Updated 2026-09-10
prompt e62d0889fc704541
fixtures e9923bbc7753a04a
scorer 8bbc6152d8b45a43
Balanced Review Accuracy
The primary score is the arithmetic mean of anchored recall and usable specificity. Invalid model output cannot count as a correct positive or negative result, so each unusable draw is counted once. Point-sorted uncertainty groups are anchored to their highest-scoring lane; lower lanes share that rank while their paired 95% normal interval versus the anchor includes zero. This avoids non-transitive bridge comparisons. The table also shows each lane's 95% interval. A Tier-1 miss or positive-fixture noise above 0.12 overrides the rank: the lane keeps its score for comparison but leaves the ranked table. Verdict match, invalid rate, and speed remain separate diagnostics.
Anchored recall
A defect counts only when the finding matches the expected behavior and the expected file. Missing a Tier-1 defect removes the lane from the ranked table regardless of its average.
False positives and noise
Clean fixtures measure whether a reviewer blocks a change that should pass. Positive-fixture noise counts unrelated blocking findings on buggy changes; recall bought with noise above the production envelope is not qualified.
Integrity
Every published lane includes sealed holdouts, a disposable HOME, a planted canary, full-transcript scanning, and post-run repository mutation checks. Harness, provider, and route labels are operator-attested report metadata, not independently derived by the site generator. Provider failures are operational failures, not proof of model quality.
Reproduce
Read the chronological experiment record, inspect each raw report, or run the evaluation harness.