Desk intelligence · measured consensus
Where the bench divides them.
Unanimous misses
589
every model misses
Unanimous passes
73
every model passes
Split trials
838
at least one pass and miss
Comparable
1,500
4 models · common trial ids
What this evidence can say
55.9% of the suite splits the measured field.
838 trials divide at least two of the 4 qualifying models. 589 are shared misses; 73 are saturated passes. The suite has 891 independent answer-shape cells, so trial counts are descriptive, not independent confidence samples.
Inspect a paired comparisonCapability landscape
4 current full-suite models
Family pass@1 pools the same number of complete models on each trial. Split counts mark where models disagree.
Select · 170 trials
Tool selection
50.7%
120 split trials · browse family
Args · 170 trials
Argument filling
49.4%
132 split trials · browse family
Flows · 150 trials
Multi-step flows
27.0%
109 split trials · browse family
Shell · 136 trials
Terminal skill
15.4%
55 split trials · browse family
Files · 114 trials
File operations
30.0%
78 split trials · browse family
Classify · 125 trials
Classifier judgments
50.8%
110 split trials · browse family
Robust · 125 trials
Robustness
34.0%
107 split trials · browse family
Omarchy · 166 trials
Omarchy control
5.4%
21 split trials · browse family
Arch · 125 trials
Arch system care
3.0%
13 split trials · browse family
Shell+ · 115 trials
Pro shell tools
3.5%
16 split trials · browse family
Keys · 104 trials
Hotkey knowledge
64.9%
77 split trials · browse family
Difficulty gradient
Pass@1 by computed grade across the measured field.
A benchmark that discriminates.
The paired report shows who wins each disputed trial. The trial browser shows the prompt, oracle, reference, null, witness, and measured badge. Follow an item below to audit the underlying task.
Trial-level evidence