Head to head · paired evidence
Two models. Same trials.
Different elicitation settings
The response protocol or decoding settings differ. Read the paired outcomes as a comparison of these recorded runs; the model and elicitation effects are combined.
A pass@1 · paired
27.8%
159 exclusive wins
B pass@1 · paired
41.5%
364 exclusive wins
Shared trials
1500
258 both pass · 719 both miss
McNemar χ²
80.35
signal at p<0.05
Paired verdict
rustlean leads on paired passes.
159 A-only wins · 364 B-only wins · 1500 matched trials. The test uses Yates correction below 25 discordant pairs, matching the Rust engine.
Where they separate
Percentage points · paired within family
Argument filling
-48.2 pp
170 paired trials
Multi-step flows
-34.0 pp
150 paired trials
File operations
-29.8 pp
114 paired trials
Hotkey knowledge
-26.0 pp
104 paired trials
Classifier judgments
+25.6 pp
125 paired trials
Tool selection
-21.8 pp
170 paired trials
Terminal skill
-16.2 pp
136 paired trials
Robustness
+14.4 pp
125 paired trials
Arch system care
+5.6 pp
125 paired trials
Pro shell tools
-3.5 pp
115 paired trials
Omarchy control
-3.0 pp
166 paired trials
Proof, trial by trial