Measured models · report cards
A model is more than one number.
rank 1 · 1500 trials
rustlean
41.5%
rank 2 · 1500 trials
lfm2.5-2.6b-fable5-coding-agent-heretic
34.3%
rank 3 · 1500 trials
minicpm5-2b-claude-fable5-1-thinking-agentic
27.8%
rank 4 · 1500 trials
minicpm5-1b-claude-opus-fable5-v2-thinking
18.1%
Selected model · 1500/1500 current trials
lfm2.5-2.6b-fable5-coding-agent-heretic
kyber_lfm2_5-2_6b-fable5-coding-agent-heretic.jsonl
Pass@1
34.3%
515/1500 · Wilson 32%–37%
Clustered
39.2%
891 answer shapes · Wilson 36%–42%
Grade weighted
0.348
valid outcomes, difficulty weighted
Mean score
0.345
34.3% valid-pass
Capability map
Pass@1 · each family's own denominator
Select
Tool selection
64.7%
110/170 first-answer passes · open trials
Args
Argument filling
48.2%
82/170 first-answer passes · open trials
Flows
Multi-step flows
19.3%
29/150 first-answer passes · open trials
Shell
Terminal skill
20.6%
28/136 first-answer passes · open trials
Files
File operations
53.5%
61/114 first-answer passes · open trials
Classify
Classifier judgments
63.2%
79/125 first-answer passes · open trials
Robust
Robustness
37.6%
47/125 first-answer passes · open trials
Omarchy
Omarchy control
3.6%
6/166 first-answer passes · open trials
Arch
Arch system care
6.4%
8/125 first-answer passes · open trials
Shell+
Pro shell tools
3.5%
4/115 first-answer passes · open trials
Keys
Hotkey knowledge
58.7%
61/104 first-answer passes · open trials
Grade cliff
Pass@1 across difficulty grades assigned by work, args, and shell composition.
Cost and outcomes
Model time
2.07s
Tokens out
253
Peak heap
16.2 KB
Outcome kinds