KYBER 0.1.0 · gate GREEN · fp 1adbc0d3 · 1,500 trials

Head to head · paired evidence

Two models. Same trials.

First-answer outcomes align by trial id. Wins count only where one model passes and the other misses; McNemar tests the discordant pairs. Partial overlap is reported as paired coverage, never hidden in two separate denominators.

Different elicitation settings

The response protocol or decoding settings differ. Read the paired outcomes as a comparison of these recorded runs; the model and elicitation effects are combined.

A pass@1 · paired

27.8%

159 exclusive wins

B pass@1 · paired

41.5%

364 exclusive wins

Shared trials

1500

258 both pass · 719 both miss

McNemar χ²

80.35

signal at p<0.05

Paired verdict

rustlean leads on paired passes.

159 A-only wins · 364 B-only wins · 1500 matched trials. The test uses Yates correction below 25 discordant pairs, matching the Rust engine.

Where they separate

Percentage points · paired within family

Where they separate

Argument filling

-48.2 pp

170 paired trials

A
29%
B
78%

Multi-step flows

-34.0 pp

150 paired trials

A
27%
B
61%

File operations

-29.8 pp

114 paired trials

A
10%
B
39%

Hotkey knowledge

-26.0 pp

104 paired trials

A
63%
B
89%

Classifier judgments

+25.6 pp

125 paired trials

A
67%
B
42%

Tool selection

-21.8 pp

170 paired trials

A
38%
B
59%

Terminal skill

-16.2 pp

136 paired trials

A
8%
B
24%

Robustness

+14.4 pp

125 paired trials

A
54%
B
39%

Arch system care

+5.6 pp

125 paired trials

A
6%
B
0%

Pro shell tools

-3.5 pp

115 paired trials

A
3%
B
7%

Omarchy control

-3.0 pp

166 paired trials

A
7%
B
10%

Proof, trial by trial

Discordant trials