KYBER 0.1.0 · gate GREEN · fp 1adbc0d3 · 1,500 trials

Desk intelligence · measured consensus

Where the bench divides them.

Agreement reveals saturated and stubborn trials; disagreement reveals the tasks that actually separate local agents. These counts use pass@1 from qualifying, complete provider runs only.

Unanimous misses

589

every model misses

Unanimous passes

73

every model passes

Split trials

838

at least one pass and miss

Comparable

1,500

4 models · common trial ids

What this evidence can say

55.9% of the suite splits the measured field.

838 trials divide at least two of the 4 qualifying models. 589 are shared misses; 73 are saturated passes. The suite has 891 independent answer-shape cells, so trial counts are descriptive, not independent confidence samples.

Inspect a paired comparison

Capability landscape

4 current full-suite models

Capability landscape

Family pass@1 pools the same number of complete models on each trial. Split counts mark where models disagree.

Difficulty gradient

Pass@1 by computed grade across the measured field.

easy · 60727.5%
medium · 70631.5%
hard · 6948.9%
extreme · 11828.4%

A benchmark that discriminates.

The paired report shows who wins each disputed trial. The trial browser shows the prompt, oracle, reference, null, witness, and measured badge. Follow an item below to audit the underlying task.

See model order

Trial-level evidence

Signal and shared misses