Head to head · paired evidence
Two models. Same trials.
Different elicitation settings
The response protocol or decoding settings differ. Read the paired outcomes as a comparison of these recorded runs; the model and elicitation effects are combined.
A pass@1 · paired
18.1%
60 exclusive wins
B pass@1 · paired
41.5%
411 exclusive wins
Shared trials
1500
211 both pass · 818 both miss
McNemar χ²
261.57
signal at p<0.05
Paired verdict
rustlean leads on paired passes.
60 A-only wins · 411 B-only wins · 1500 matched trials. The test uses Yates correction below 25 discordant pairs, matching the Rust engine.
Where they separate
Percentage points · paired within family
Multi-step flows
-61.3 pp
150 paired trials
Hotkey knowledge
-41.3 pp
104 paired trials
Argument filling
-35.3 pp
170 paired trials
Robustness
-33.6 pp
125 paired trials
File operations
-21.9 pp
114 paired trials
Tool selection
-18.2 pp
170 paired trials
Terminal skill
-15.4 pp
136 paired trials
Classifier judgments
-10.4 pp
125 paired trials
Omarchy control
-9.6 pp
166 paired trials
Pro shell tools
-7.0 pp
115 paired trials
Arch system care
+0.0 pp
125 paired trials
Proof, trial by trial