Measured models · report cards
A model is more than one number.
rank 1 · 1500 trials
rustlean
41.5%
rank 2 · 1500 trials
lfm2.5-2.6b-fable5-coding-agent-heretic
34.3%
rank 3 · 1500 trials
minicpm5-2b-claude-fable5-1-thinking-agentic
27.8%
rank 4 · 1500 trials
minicpm5-1b-claude-opus-fable5-v2-thinking
18.1%
Selected model · 1500/1500 current trials
minicpm5-2b-claude-fable5-1-thinking-agentic
kyber_minicpm5-2b-claude-fable5-1-thinking-agentic.jsonl
Pass@1
27.8%
417/1500 · Wilson 26%–30%
Clustered
25.7%
891 answer shapes · Wilson 23%–29%
Grade weighted
0.283
valid outcomes, difficulty weighted
Mean score
0.296
27.8% valid-pass
Capability map
Pass@1 · each family's own denominator
Select
Tool selection
37.6%
64/170 first-answer passes · open trials
Args
Argument filling
29.4%
50/170 first-answer passes · open trials
Flows
Multi-step flows
27.3%
41/150 first-answer passes · open trials
Shell
Terminal skill
8.1%
11/136 first-answer passes · open trials
Files
File operations
9.6%
11/114 first-answer passes · open trials
Classify
Classifier judgments
67.2%
84/125 first-answer passes · open trials
Robust
Robustness
53.6%
67/125 first-answer passes · open trials
Omarchy
Omarchy control
7.2%
12/166 first-answer passes · open trials
Arch
Arch system care
5.6%
7/125 first-answer passes · open trials
Shell+
Pro shell tools
3.5%
4/115 first-answer passes · open trials
Keys
Hotkey knowledge
63.5%
66/104 first-answer passes · open trials
Grade cliff
Pass@1 across difficulty grades assigned by work, args, and shell composition.
Cost and outcomes
Model time
696ms
Tokens out
53
Peak heap
34.4 KB
Outcome kinds