
General tool-use + agentic robustness benchmark · engine 0.1.0 Live desk: kyber.xcy.cx · source: github.com/xcycx/kyber
KYBER's question: which small model — light enough for a low-VRAM laptop or low-tier GPU, fast enough to stream ~200+ tokens a second — can most reliably drive an Omarchy Linux machine? The desk ranks how close each light, affordable model gets.
KYBER scores how useful a local language model is as an *agent on an Omarchy Linux machine*: tool selection, argument filling, multi-step flows, terminal/file commands in a sandbox, classifier judgments, and robustness under paraphrase and noise — plus native-shell skill: driving the omarchy CLI, Arch package/service care, pro shell tools (eza/fd/rg/bat/tldr/zoxide), and hotkey knowledge. Shell trials marked oma execute against a deterministic mock machine (crates/kyber/oma/) with fixture state. Shipped: 1500 ultra-fast lean trials (120 hand seeds + 1380 deterministic generated; in-process oracles in milliseconds, shell trials with tight timeouts). KYBER runs suites/kyber/; the desk opens at /, with measured results at /leaderboards, /models, /compare, /insights, and /runs, trial evidence at /kyber, and this handbook at /guide.
Identity is KYBER. Engine version 0.1.0. Auth and the database stay off.
This file is the product handbook. The same text renders in the desk at /guide.
Kyber workflow
cargo run --manifest-path crates/kyber/Cargo.toml -- author # regenerate 1380 trials (xorshift seed 7, deterministic, deduped, witnessed)
cargo run --manifest-path crates/kyber/Cargo.toml -- release # gate --emit + floor + trials-ts + sync, in order
cargo run --manifest-path crates/kyber/Cargo.toml -- floor [--sample N] [--jobs 8] [--family terminal --diff hard] # bounds + latency/tokens/heap rows (thread-pool, ordered)
cargo run --manifest-path crates/kyber/Cargo.toml -- strike --model MODEL_ID [--samples 3] [--seed 7] [--jobs 1] # strike LM Studio (pass@1 + pass@k, seeded, file cache, weighted + valid-pass)
cargo run --manifest-path crates/kyber/Cargo.toml -- stats runs/kyber_MODEL_ID.jsonl # aggregate: pass@1/pass@k/Wilson/clustered/weighted/valid + family + grade tables
cargo run --manifest-path crates/kyber/Cargo.toml -- compare A.jsonl B.jsonl # paired McNemar + discordants by family
cargo run --manifest-path crates/kyber/Cargo.toml -- triage runs/kyber_MODEL.jsonl # why a run failed, ranked by cause (+ --trial ID for full evidence)
cargo run --manifest-path crates/kyber/Cargo.toml -- probe # anti-cheat: score no-knowledge strategies (structural ones must be 0%)
cargo run --manifest-path crates/kyber/Cargo.toml -- show --trial KB-T0001 # print exactly what a model receives
cargo run --manifest-path crates/kyber/Cargo.toml -- version # engine version + oracle rules fingerprint
cargo run --manifest-path crates/kyber/Cargo.toml -- hash --suite suites/kyber/kyber.json # suite fingerprint + dup-prompt audit
cargo run --manifest-path crates/kyber/Cargo.toml -- audit # read-only family coverage + normalized prompt duplicates
cargo run --manifest-path crates/kyber/Cargo.toml -- audit --suite NEXT.json --strict # reject sentence-punctuation/case-only repeats in a future suite
cargo run --manifest-path crates/kyber/Cargo.toml -- sync # provider rows + gate/floor evidence -> desk twins
cargo run --manifest-path crates/kyber/Cargo.toml -- trials-ts --suite suites/kyber/kyber.json --out src/data/kyber-trials.gen.ts
cargo test --manifest-path crates/kyber/Cargo.toml # core unit tests (std-only, zero deps)
npm run test:desk # desk unit tests + taxonomy drift guards
npm run parity:kyber # engine vs desk aggregates, field by field
npm run parity:compare # paired comparison vs Rust McNemar output
npm run parity:command # copied command through shell, Rust, provider, sync and desk- Suite:
suites/kyber/kb_tool01.json(40 select/fill) ·kb_agent01.json(20 flows) ·kb_shell01.json(20 shell/files) ·kb_cls01.json(20 classify/robust) ·kb_oma01.json(20 omarchy/arch/shell/keys) ·kyber_gen01.json(1380 generated) ·kyber.json(1500 aggregate across 11 families). - Core (Rust, std-only, zero dependencies):
crates/kyber—jsoncodec ·modeltrial shape + contract checks ·matchxoracle regex ·oraclesix kinds + prefix credit + process witness ·sandboxallowlist, tracedsh -xexecution, process-group timeout kills ·poolordered thread pool with a streaming sink (--jobs, default 16) ·gateshape + reference-100/null-0 + shortcut-fail + dup-prompt discrimination ·authordeterministic generator (prompt-deduped, null-hardened, witness-deriving, self-calibrating) ·runfloor bounds + LM Studio strikes (minimal HTTP/1.1 client,--samplespass@k,--temperature, harness-fingerprinted response cache, retry + circuit breaker, weighted + valid-pass) ·statsWilson / clustered Wilson / McNemar (on pass@1) / grades ·hashFNV fingerprints ·emitJSONL + TS twins. No Python in the Kyber runtime. - Shortcut resistance: a product-only oracle (
stdout_regex,exit_zero,file_exact) reads bytes, so every such trial must either witness the process (oracle.ran— the tool must appear in the run'ssh -xtrace) or declare ashortcut_casethat the gate proves fails. The author derives witnesses from each reference's own execution trace, so printing the expected answer is never a pass. - Anti-cheat, measured not asserted:
kyber probescores seven no-knowledge strategies against the real oracles, using only what a model can see and *knowing* the oracle kind — a strictly stronger attacker than any model. Four are structural (echo the pattern's witness,true, forge the file the oracle reads, answer{}) and must score exactly zero;releasefails if one does not. The other three are chance (first announced tool, most common label, the option already named in the task text) and are reported as the floor a model has to beat — currently 3.2%. - The interface is stated, the answer never is:
run::render_promptis the one place the user message is built — task, announced tool surface, and a reply-shape hint for that oracle kind. It is used for the request, the cache key, andkyber show, so the three cannot disagree, andno_unwitnessed_trial_states_its_expected_observablefails if any trial states its target without a witness. - Diagnostics:
kyber triageturns an outcome kind into a cause — "1094 failed" becomes "330 sandbox rejection, 286 wrong channel, 171 wrong output, 124 wrong args…" — and--trial IDprints one trial's full evidence. The matcher is bounded bySTEP_LIMITso a pathological pattern fails the gate instead of hanging a run, and the provider path is tested against missingchoices, nullcontent, non-JSON bodies, 500s and dead sockets. - Two fingerprints, two anchors: the suite fingerprint covers the trials, and
oracle_fp(kyber version) covers the sandbox and oracle rules that grade them.syncrefuses strike files whose rules fingerprint has moved, so a harness change cannot quietly re-interpret old evidence. - Provider compatibility: strikes start with a neutral JSON Schema for the documented calls, text, or shell response channel. Explicit format rejection (HTTP 400/422 mentioning
response_formatorjson_schema) falls back to JSON object, then unstructured text; wrong answers never cause negotiation.--response-format auto|json_schema|json_object|textcan pin a format for controlled comparisons; fixed modes never fall back and carry distinct elicitation fingerprints. Native tool calls and text content blocks are supported. Empty completions and malformed native arguments are provider failures. Contracts recordresponse_policy, rows recordresponse_formats, and cache keys isolate endpoint, replica, seed, and elicitation fingerprint./runsdistinguishes legacy protocols; compare elicitation fingerprints before interpreting cross-model differences. - Reproducibility:
--seed Nrides every request (samplekusesseed + k) and is recorded in the contract. A local server is still not bit-reproducible attemperature 0— measured against LM Studio, the same 24 trials flipped 7/24 at--jobs 4, 3/24 with a seed, and 1/24 at--jobs 1 --seed 7, because batching changes kernel reduction order and genuine ties are broken by the hardware. Aggregates are far more stable than single trials; the cache is the real replay mechanism. Cache entries are keyed per sample, so a one-sample run extends to--samples 3without re-paying. - Honest statistics:
pass@1is reported besidepass@k, the cost columns are per-sample means,difficultyis computed from the work a trial demands rather than rolled, and a clustered Wilson interval over answer-shape cells sits next to the per-trial one — repeated shapes are one question, not many. Fractional token means are preserved from the runner through the desk. Missing completion usage is estimated separately for each successful answer; skipped provider samples contribute zero tokens. - Future authoring audit:
kyber auditreports trial, title, exact reference-shape, grade, and oracle-kind counts for every family, then flags low title or answer-shape variety and missing calibrated hard/extreme cases as authoring notes. It also finds prompts that differ only in case, whitespace, or sentence-ending punctuation within one family; shell syntax inside backticks and meaningful path punctuation stay distinct. The published suite currently has one such pair (KB-T0020/KB-T0121); it stays published so existing measured scores remain comparable. Use--stricton a proposed next edition to reject that class of duplicate before publishing, then rerun the full gate and model strikes for the new fingerprint. Shape counts indicate repeated expected payloads, not necessarily identical reasoning tasks; coverage notes are advisory because each family has different semantics. - Desk:
/,/leaderboards, and/runsread the compact, fingerprintedsrc/data/kyber-overview.gen.tssummary emitted bysync./models,/compare,/insights, and/kybersharesrc/data/kyber-analysis.gen.ts: a trial index and eight measured values per model/trial, with explicit nulls for missing rows. The shared decoder retains exact model identity, oracle latency, and fractional costs. Only/kyberloads detailed prompts, references, and oracle bodies. The full strike twin remains an independent parity artifact. Both compact payloads carry the same measured snapshot fingerprint; the trial twin carries the suite fingerprint. Mixed or stale emissions display unavailable evidence and recovery commands. Never hand-edit.gen.ts: regenerate withtrials-ts/sync.npm run parity:kyberchecks every compact row, cost, profile, pair, and consensus against the independent detailed twin for published and sparse synthetic evidence across all 11 families, then checks aggregates against the engine.npm run parity:comparechecks the compact paired outcomes and McNemar againstkyber compare. Shared taxonomy:src/data/kyber.ts(families, grades, oracle kinds, infra kinds, shell allowlist) andsrc/data/kyber-suite.ts(detailed meta, coverage, Wilson, clustered Wilson, validity). - Strike designer:
/kyberexplicitly copies jobs, samples, temperature, seed, timeout, and response policy, plus optional endpoint and stratified slice filters. Choose Bash / zsh or PowerShell to quote opaque API ids correctly. Windows PowerShell cannot reliably preserve embedded double quotes or a trailing backslash in an id; the designer directs those ids to Bash. HTTP endpoints support IPv4, DNS names, and bracketed IPv6. Credentials, query strings, and fragments do not belong in the provider base URL.npm run parity:commandruns the actual copied settings through a local mock provider and checks the JSONL and synced desk costs without changing published runs. - Codec performance: the std-only JSON parser traverses the already validated UTF-8 input without revalidating the remaining string for every character.
cargo run --release --manifest-path crates/kyber/Cargo.toml --example codec_benchmeasures three Unicode/escape string sizes with five samples each and reports medians. These are reproducible observations, not timing gates; deterministic roundtrip, malformed-surrogate, and document-truncation tests guard correctness. - Official top 25:
syncpublishes one newest valid strike contract per model and records its source file. A rank requires engine 0.1.0, current suite and oracle fingerprints, an unfiltered full-suite run without sampling or a limit, and one result for all 1500 trials. Pass@1 orders qualifying runs; grade-weighted pass and model id break ties. The board shows only real qualifying runs, up to 25, and/runsexplains why a synced partial or stale run is excluded./modelsshows capability and cost per run;/comparealigns pass@1 by trial id and distinguishes matching, different, or missing elicitation fingerprints, with links to each run contract. Different protocols combine model and elicitation effects./insightsuses common trials across qualifying models for consensus and disagreement. - Desk ↔ engine: the browser's filters live in the URL (
/kyber?family=terminal,?trial=KB-T0001), so a family card, a palette hit and a pasted link all land on the same view. Measured evidence is aggregated per model —synccan hold several, and blending them would report a pass rate that describes none of them. The desk's TS mirrors of the engine's taxonomies are checked against the Rust source bysrc/lib/kyber-taxonomy.test.ts, and its aggregates againstkyber statsbynpm run parity:kyber. - Scoring: pass 1 / fail 0, plus prefix partial credit on
calls_exact— the longest correct prefix over the longer of the two lists, so a chain that gets 2 of 3 calls right scores 0.667 and one that adds hallucinated extras cannot score 1.0. Oracle-eval latency is budget-enforced (50 ms in-process, 2000 ms shell; over-budget fails asover_budget); providermodel_msis recorded but never budgeted. Thread-local heap high-water marks ride on every row without cross-worker contamination. Rows stream to disk as their position in the trial order settles, so a run that dies at trial 1200 keeps 1200 scored rows. Measured-only: onlystrikerows badge trials; floor bounds never do.