Agent bench · engine 0.1.0
1,500 lean trials
Suite
1,500 trials
120 seeds · 1380 generated
Gate
GREEN
1500 trials · fp 1adbc0d3
Witnessed
100%
656/656 byte-reading oracles demand the tool
Oracle cost
22.30 ms
mean eval · 1500 bound rows
Measured
34.3%
lfm2.5-2.6b-fable5-coding-agent-heretic · 1500 trials · pass@1 (+3)
What these numbers can bear
gate: reference passes · null fails · the lazy route fails
- Shortcut-resistant
- 656 / 656
- byte-reading oracles that demand the tool, or prove the lazy route fails
- Independent cells
- 891 / 1,500
- 609 rows repeat an answer shape — confidence intervals use cells, not rows
- Grades
- 607 / 706 / 69 / 118
- easy / medium / hard / extreme, scored from steps, checked args, and shell composition
- Cheat ceiling
- 3.2%
- no no-knowledge route wins; best blind guess is modal-label
- Gate detail
- GREEN
- 0 duplicate prompts · 0 shortcuts passing · 0 null-kind mismatches
Measured · lfm2.5-2.6b-fable5-coding-agent-heretic · 1500 trials
pass@1 34.3% · Wilson [32%, 37%] · score 0.345 · weighted 0.348 · valid-pass 34.3% (515/1500)
clustered 891 answer shapes at 39.2% · Wilson [36%, 42%] — repeated shapes are not independent evidence
mean model 2.07s · mean 253 tokens out · mean heap 16.2 kb
jobs 12 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite
harness fp b561b68a · rules fp bcdc6c16 · seed 7 · source kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl
Measured · minicpm5-1b-claude-opus-fable5-v2-thinking · 1500 trials
pass@1 18.1% · Wilson [16%, 20%] · score 0.184 · weighted 0.169 · valid-pass 18.1% (271/1500)
clustered 891 answer shapes at 21.5% · Wilson [19%, 24%] — repeated shapes are not independent evidence
mean model 1.16s · mean 167 tokens out · mean heap 37.5 kb
jobs 4 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite
harness fp b561b68a · rules fp bcdc6c16 · seed 7 · source kyber_minicpm5-1b-claude-opus-fable5-v2-thinking.jsonl
Measured · minicpm5-2b-claude-fable5-1-thinking-agentic · 1500 trials
pass@1 27.8% · Wilson [26%, 30%] · score 0.296 · weighted 0.283 · valid-pass 27.8% (417/1500)
clustered 891 answer shapes at 25.7% · Wilson [23%, 29%] — repeated shapes are not independent evidence
mean model 696ms · mean 53 tokens out · mean heap 34.4 kb
jobs 4 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite
harness fp b561b68a · rules fp bcdc6c16 · seed 7 · source kyber_minicpm5-2b-claude-fable5-1-thinking-agentic.jsonl
Measured · rustlean · 1500 trials
pass@1 41.5% · Wilson [39%, 44%] · score 0.432 · weighted 0.474 · valid-pass 41.5% (622/1500)
clustered 891 answer shapes at 50.5% · Wilson [47%, 54%] — repeated shapes are not independent evidence
mean model 180ms · mean 32 tokens out · mean heap 16.6 kb
jobs 1 · samples 1 · temp 0 · filter full · suite fp 1adbc0d3 ✓ matches current suite
harness fp 98459811 · rules fp bcdc6c16 · seed 7 · source kyber_rustlean.jsonl
Compare measured models
paired McNemar on pass@1 · discordants by family
Each model keeps its own newest row per trial, so these are the exact revisions the badges above came from. A paired test only counts the trials both runs share, and it says so when the difference is inside noise.
kyber compare runs/kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl runs/kyber_minicpm5-1b-claude-opus-fable5-v2-thinking.jsonlkyber compare runs/kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl runs/kyber_minicpm5-2b-claude-fable5-1-thinking-agentic.jsonlkyber compare runs/kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl runs/kyber_rustlean.jsonlBy difficulty
40%
47%
5%
8%
Design a strike
exact settings · Rust CLI · badges via sync
Auto negotiates JSON Schema, JSON object, then text when the provider rejects a format. Fixed modes stay fixed. Sample size spreads trials across family × difficulty.
Enter the loaded model's exact API id to copy the strike.
kyber strike --model 'MODEL_ID' --jobs 8 --samples 1 --temperature 0 --seed 0 --timeout-s 120 --response-format autokyber sync125 shown · 844 in-process (50 ms) · 656 sandboxed-shell (2 s) · measured 6000
125 rows · page 1 / 4