KYBER 0.1.0 · gate GREEN · fp 1adbc0d3 · 1,500 trials

Head to head · paired evidence

Two models. Same trials.

First-answer outcomes align by trial id. Wins count only where one model passes and the other misses; McNemar tests the discordant pairs. Partial overlap is reported as paired coverage, never hidden in two separate denominators.

Different elicitation settings

The response protocol or decoding settings differ. Read the paired outcomes as a comparison of these recorded runs; the model and elicitation effects are combined.

A pass@1 · paired

18.1%

60 exclusive wins

B pass@1 · paired

41.5%

411 exclusive wins

Shared trials

1500

211 both pass · 818 both miss

McNemar χ²

261.57

signal at p<0.05

Paired verdict

rustlean leads on paired passes.

60 A-only wins · 411 B-only wins · 1500 matched trials. The test uses Yates correction below 25 discordant pairs, matching the Rust engine.

Where they separate

Percentage points · paired within family

Where they separate

Multi-step flows

-61.3 pp

150 paired trials

A
0%
B
61%

Hotkey knowledge

-41.3 pp

104 paired trials

A
48%
B
89%

Argument filling

-35.3 pp

170 paired trials

A
42%
B
78%

Robustness

-33.6 pp

125 paired trials

A
6%
B
39%

File operations

-21.9 pp

114 paired trials

A
18%
B
39%

Tool selection

-18.2 pp

170 paired trials

A
41%
B
59%

Terminal skill

-15.4 pp

136 paired trials

A
9%
B
24%

Classifier judgments

-10.4 pp

125 paired trials

A
31%
B
42%

Omarchy control

-9.6 pp

166 paired trials

A
1%
B
10%

Pro shell tools

-7.0 pp

115 paired trials

A
0%
B
7%

Arch system care

+0.0 pp

125 paired trials

A
0%
B
0%

Proof, trial by trial

Discordant trials