Measured models · report cards
A model is more than one number.
rank 1 · 1500 trials
rustlean
41.5%
rank 2 · 1500 trials
lfm2.5-2.6b-fable5-coding-agent-heretic
34.3%
rank 3 · 1500 trials
minicpm5-2b-claude-fable5-1-thinking-agentic
27.8%
rank 4 · 1500 trials
minicpm5-1b-claude-opus-fable5-v2-thinking
18.1%
Selected model · 1500/1500 current trials
minicpm5-1b-claude-opus-fable5-v2-thinking
kyber_minicpm5-1b-claude-opus-fable5-v2-thinking.jsonl
Pass@1
18.1%
271/1500 · Wilson 16%–20%
Clustered
21.5%
891 answer shapes · Wilson 19%–24%
Grade weighted
0.169
valid outcomes, difficulty weighted
Mean score
0.184
18.1% valid-pass
Capability map
Pass@1 · each family's own denominator
Select
Tool selection
41.2%
70/170 first-answer passes · open trials
Args
Argument filling
42.4%
72/170 first-answer passes · open trials
Flows
Multi-step flows
0.0%
0/150 first-answer passes · open trials
Shell
Terminal skill
8.8%
12/136 first-answer passes · open trials
Files
File operations
17.5%
20/114 first-answer passes · open trials
Classify
Classifier judgments
31.2%
39/125 first-answer passes · open trials
Robust
Robustness
5.6%
7/125 first-answer passes · open trials
Omarchy
Omarchy control
0.6%
1/166 first-answer passes · open trials
Arch
Arch system care
0.0%
0/125 first-answer passes · open trials
Shell+
Pro shell tools
0.0%
0/115 first-answer passes · open trials
Keys
Hotkey knowledge
48.1%
50/104 first-answer passes · open trials
Grade cliff
Pass@1 across difficulty grades assigned by work, args, and shell composition.
Cost and outcomes
Model time
1.16s
Tokens out
167
Peak heap
37.5 KB
Outcome kinds