Head to head · paired evidence
Two models. Same trials.
Different elicitation settings
The response protocol or decoding settings differ. Read the paired outcomes as a comparison of these recorded runs; the model and elicitation effects are combined.
A pass@1 · paired
41.5%
295 exclusive wins
B pass@1 · paired
34.3%
188 exclusive wins
Shared trials
1500
327 both pass · 690 both miss
McNemar χ²
23.70
signal at p<0.05
Paired verdict
rustlean leads on paired passes.
295 A-only wins · 188 B-only wins · 1500 matched trials. The test uses Yates correction below 25 discordant pairs, matching the Rust engine.
Where they separate
Percentage points · paired within family
Multi-step flows
+42.0 pp
150 paired trials
Hotkey knowledge
+30.8 pp
104 paired trials
Argument filling
+29.4 pp
170 paired trials
Classifier judgments
-21.6 pp
125 paired trials
File operations
-14.0 pp
114 paired trials
Omarchy control
+6.6 pp
166 paired trials
Arch system care
-6.4 pp
125 paired trials
Tool selection
-5.3 pp
170 paired trials
Terminal skill
+3.7 pp
136 paired trials
Pro shell tools
+3.5 pp
115 paired trials
Robustness
+1.6 pp
125 paired trials
Proof, trial by trial