KYBER 0.1.0 · gate GREEN · fp 1adbc0d3 · 1,500 trials

Agent bench · engine 0.1.0

1,500 lean trials

120 hand seeds + 1380 generated trials across 11 families. References pass, null cases fail, and shortcuts fail the gate. 4 measured models; open a trial for its result.

Suite

1,500 trials

120 seeds · 1380 generated

Gate

GREEN

1500 trials · fp 1adbc0d3

Witnessed

100%

656/656 byte-reading oracles demand the tool

Oracle cost

22.30 ms

mean eval · 1500 bound rows

Measured

34.3%

lfm2.5-2.6b-fable5-coding-agent-heretic · 1500 trials · pass@1 (+3)

What these numbers can bear

gate: reference passes · null fails · the lazy route fails

Shortcut-resistant
656 / 656
byte-reading oracles that demand the tool, or prove the lazy route fails
Independent cells
891 / 1,500
609 rows repeat an answer shape — confidence intervals use cells, not rows
Grades
607 / 706 / 69 / 118
easy / medium / hard / extreme, scored from steps, checked args, and shell composition
Cheat ceiling
3.2%
no no-knowledge route wins; best blind guess is modal-label
Gate detail
GREEN
0 duplicate prompts · 0 shortcuts passing · 0 null-kind mismatches

Measured · lfm2.5-2.6b-fable5-coding-agent-heretic · 1500 trials

pass@1 34.3% · Wilson [32%, 37%] · score 0.345 · weighted 0.348 · valid-pass 34.3% (515/1500)

clustered 891 answer shapes at 39.2% · Wilson [36%, 42%] — repeated shapes are not independent evidence

mean model 2.07s · mean 253 tokens out · mean heap 16.2 kb

jobs 12 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite

harness fp b561b68a · rules fp bcdc6c16 · seed 7 · source kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl

Measured · minicpm5-1b-claude-opus-fable5-v2-thinking · 1500 trials

pass@1 18.1% · Wilson [16%, 20%] · score 0.184 · weighted 0.169 · valid-pass 18.1% (271/1500)

clustered 891 answer shapes at 21.5% · Wilson [19%, 24%] — repeated shapes are not independent evidence

mean model 1.16s · mean 167 tokens out · mean heap 37.5 kb

jobs 4 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite

harness fp b561b68a · rules fp bcdc6c16 · seed 7 · source kyber_minicpm5-1b-claude-opus-fable5-v2-thinking.jsonl

Measured · minicpm5-2b-claude-fable5-1-thinking-agentic · 1500 trials

pass@1 27.8% · Wilson [26%, 30%] · score 0.296 · weighted 0.283 · valid-pass 27.8% (417/1500)

clustered 891 answer shapes at 25.7% · Wilson [23%, 29%] — repeated shapes are not independent evidence

mean model 696ms · mean 53 tokens out · mean heap 34.4 kb

jobs 4 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite

harness fp b561b68a · rules fp bcdc6c16 · seed 7 · source kyber_minicpm5-2b-claude-fable5-1-thinking-agentic.jsonl

Measured · rustlean · 1500 trials

pass@1 41.5% · Wilson [39%, 44%] · score 0.432 · weighted 0.474 · valid-pass 41.5% (622/1500)

clustered 891 answer shapes at 50.5% · Wilson [47%, 54%] — repeated shapes are not independent evidence

mean model 180ms · mean 32 tokens out · mean heap 16.6 kb

jobs 1 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite

harness fp 98459811 · rules fp bcdc6c16 · seed 7 · source kyber_rustlean.jsonl

Compare measured models

paired McNemar on pass@1 · discordants by family

Each model keeps its own newest row per trial, so these are the exact revisions the badges above came from. A paired test only counts the trials both runs share, and it says so when the difference is inside noise.

kyber compare runs/kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl runs/kyber_minicpm5-1b-claude-opus-fable5-v2-thinking.jsonl
kyber compare runs/kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl runs/kyber_minicpm5-2b-claude-fable5-1-thinking-agentic.jsonl
kyber compare runs/kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl runs/kyber_rustlean.jsonl

By difficulty

easy607

40%

medium706

47%

hard69

5%

extreme118

8%

Design a strike

exact settings · Rust CLI · badges via sync

Auto negotiates JSON Schema, JSON object, then text when the provider rejects a format. Fixed modes stay fixed. Sample size spreads trials across family × difficulty.

Enter the loaded model's exact API id to copy the strike.

kyber strike --model 'MODEL_ID' --jobs 8 --samples 1 --temperature 0 --seed 0 --timeout-s 120 --response-format auto
kyber sync

150 shown · 844 in-process (50 ms) · 656 sandboxed-shell (2 s) · measured 6000

150 rows · page 1 / 4