Head to head · paired evidence
Two models. Same trials.
Different elicitation settings
The response protocol or decoding settings differ. Read the paired outcomes as a comparison of these recorded runs; the model and elicitation effects are combined.
A pass@1 · paired
34.3%
188 exclusive wins
B pass@1 · paired
41.5%
295 exclusive wins
Shared trials
1500
327 both pass · 690 both miss
McNemar χ²
23.70
signal at p<0.05
Paired verdict
rustlean leads on paired passes.
188 A-only wins · 295 B-only wins · 1500 matched trials. The test uses Yates correction below 25 discordant pairs, matching the Rust engine.
Where they separate
Percentage points · paired within family
Multi-step flows
-42.0 pp
150 paired trials
Hotkey knowledge
-30.8 pp
104 paired trials
Argument filling
-29.4 pp
170 paired trials
Classifier judgments
+21.6 pp
125 paired trials
File operations
+14.0 pp
114 paired trials
Omarchy control
-6.6 pp
166 paired trials
Arch system care
+6.4 pp
125 paired trials
Tool selection
+5.3 pp
170 paired trials
Terminal skill
-3.7 pp
136 paired trials
Pro shell tools
-3.5 pp
115 paired trials
Robustness
-1.6 pp
125 paired trials
Proof, trial by trial