KYBER 0.1.0 · gate GREEN · fp 1adbc0d3 · 1,500 trials

Handbook

How the desk is scored

The same README that ships with the repo — families, oracles, sandbox budgets, JSONL rows, and the desk.
KYBER desk — 1,500 lean trials for small local-LLM agents
KYBER desk — 1,500 lean trials for small local-LLM agents

General tool-use + agentic robustness benchmark · engine 0.1.0 Live desk: kyber.xcy.cx · source: github.com/xcycx/kyber

KYBER's question: which small model — light enough for a low-VRAM laptop or low-tier GPU, fast enough to stream ~200+ tokens a second — can most reliably drive an Omarchy Linux machine? The desk ranks how close each light, affordable model gets.

KYBER scores how useful a local language model is as an *agent on an Omarchy Linux machine*: tool selection, argument filling, multi-step flows, terminal/file commands in a sandbox, classifier judgments, and robustness under paraphrase and noise — plus native-shell skill: driving the omarchy CLI, Arch package/service care, pro shell tools (eza/fd/rg/bat/tldr/zoxide), and hotkey knowledge. Shell trials marked oma execute against a deterministic mock machine (crates/kyber/oma/) with fixture state. Shipped: 1500 ultra-fast lean trials (120 hand seeds + 1380 deterministic generated; in-process oracles in milliseconds, shell trials with tight timeouts). KYBER runs suites/kyber/; the desk opens at /, with measured results at /leaderboards, /models, /compare, /insights, and /runs, trial evidence at /kyber, and this handbook at /guide.

Identity is KYBER. Engine version 0.1.0. Auth and the database stay off.

This file is the product handbook. The same text renders in the desk at /guide.

Kyber workflow

text
cargo run --manifest-path crates/kyber/Cargo.toml -- author                                 # regenerate 1380 trials (xorshift seed 7, deterministic, deduped, witnessed)
cargo run --manifest-path crates/kyber/Cargo.toml -- release   # gate --emit + floor + trials-ts + sync, in order
cargo run --manifest-path crates/kyber/Cargo.toml -- floor [--sample N] [--jobs 8] [--family terminal --diff hard]   # bounds + latency/tokens/heap rows (thread-pool, ordered)
cargo run --manifest-path crates/kyber/Cargo.toml -- strike --model MODEL_ID [--samples 3] [--seed 7] [--jobs 1] # strike LM Studio (pass@1 + pass@k, seeded, file cache, weighted + valid-pass)
cargo run --manifest-path crates/kyber/Cargo.toml -- stats runs/kyber_MODEL_ID.jsonl      # aggregate: pass@1/pass@k/Wilson/clustered/weighted/valid + family + grade tables
cargo run --manifest-path crates/kyber/Cargo.toml -- compare A.jsonl B.jsonl               # paired McNemar + discordants by family
cargo run --manifest-path crates/kyber/Cargo.toml -- triage runs/kyber_MODEL.jsonl          # why a run failed, ranked by cause (+ --trial ID for full evidence)
cargo run --manifest-path crates/kyber/Cargo.toml -- probe                                  # anti-cheat: score no-knowledge strategies (structural ones must be 0%)
cargo run --manifest-path crates/kyber/Cargo.toml -- show --trial KB-T0001                   # print exactly what a model receives
cargo run --manifest-path crates/kyber/Cargo.toml -- version                                # engine version + oracle rules fingerprint
cargo run --manifest-path crates/kyber/Cargo.toml -- hash --suite suites/kyber/kyber.json   # suite fingerprint + dup-prompt audit
cargo run --manifest-path crates/kyber/Cargo.toml -- audit                                 # read-only family coverage + normalized prompt duplicates
cargo run --manifest-path crates/kyber/Cargo.toml -- audit --suite NEXT.json --strict       # reject sentence-punctuation/case-only repeats in a future suite
cargo run --manifest-path crates/kyber/Cargo.toml -- sync                                   # provider rows + gate/floor evidence -> desk twins
cargo run --manifest-path crates/kyber/Cargo.toml -- trials-ts --suite suites/kyber/kyber.json --out src/data/kyber-trials.gen.ts
cargo test --manifest-path crates/kyber/Cargo.toml                                          # core unit tests (std-only, zero deps)
npm run test:desk                                                                           # desk unit tests + taxonomy drift guards
npm run parity:kyber                                                                        # engine vs desk aggregates, field by field
npm run parity:compare                                                                      # paired comparison vs Rust McNemar output
npm run parity:command                                                                      # copied command through shell, Rust, provider, sync and desk
  • Suite: suites/kyber/kb_tool01.json (40 select/fill) · kb_agent01.json (20 flows) · kb_shell01.json (20 shell/files) · kb_cls01.json (20 classify/robust) · kb_oma01.json (20 omarchy/arch/shell/keys) · kyber_gen01.json (1380 generated) · kyber.json (1500 aggregate across 11 families).
  • Core (Rust, std-only, zero dependencies): crates/kyber — json codec · model trial shape + contract checks · matchx oracle regex · oracle six kinds + prefix credit + process witness · sandbox allowlist, traced sh -x execution, process-group timeout kills · pool ordered thread pool with a streaming sink (--jobs, default 16) · gate shape + reference-100/null-0 + shortcut-fail + dup-prompt discrimination · author deterministic generator (prompt-deduped, null-hardened, witness-deriving, self-calibrating) · run floor bounds + LM Studio strikes (minimal HTTP/1.1 client, --samples pass@k, --temperature, harness-fingerprinted response cache, retry + circuit breaker, weighted + valid-pass) · stats Wilson / clustered Wilson / McNemar (on pass@1) / grades · hash FNV fingerprints · emit JSONL + TS twins. No Python in the Kyber runtime.
  • Shortcut resistance: a product-only oracle (stdout_regex, exit_zero, file_exact) reads bytes, so every such trial must either witness the process (oracle.ran — the tool must appear in the run's sh -x trace) or declare a shortcut_case that the gate proves fails. The author derives witnesses from each reference's own execution trace, so printing the expected answer is never a pass.
  • Anti-cheat, measured not asserted: kyber probe scores seven no-knowledge strategies against the real oracles, using only what a model can see and *knowing* the oracle kind — a strictly stronger attacker than any model. Four are structural (echo the pattern's witness, true, forge the file the oracle reads, answer {}) and must score exactly zero; release fails if one does not. The other three are chance (first announced tool, most common label, the option already named in the task text) and are reported as the floor a model has to beat — currently 3.2%.
  • The interface is stated, the answer never is: run::render_prompt is the one place the user message is built — task, announced tool surface, and a reply-shape hint for that oracle kind. It is used for the request, the cache key, and kyber show, so the three cannot disagree, and no_unwitnessed_trial_states_its_expected_observable fails if any trial states its target without a witness.
  • Diagnostics: kyber triage turns an outcome kind into a cause — "1094 failed" becomes "330 sandbox rejection, 286 wrong channel, 171 wrong output, 124 wrong args…" — and --trial ID prints one trial's full evidence. The matcher is bounded by STEP_LIMIT so a pathological pattern fails the gate instead of hanging a run, and the provider path is tested against missing choices, null content, non-JSON bodies, 500s and dead sockets.
  • Two fingerprints, two anchors: the suite fingerprint covers the trials, and oracle_fp (kyber version) covers the sandbox and oracle rules that grade them. sync refuses strike files whose rules fingerprint has moved, so a harness change cannot quietly re-interpret old evidence.
  • Provider compatibility: strikes start with a neutral JSON Schema for the documented calls, text, or shell response channel. Explicit format rejection (HTTP 400/422 mentioning response_format or json_schema) falls back to JSON object, then unstructured text; wrong answers never cause negotiation. --response-format auto|json_schema|json_object|text can pin a format for controlled comparisons; fixed modes never fall back and carry distinct elicitation fingerprints. Native tool calls and text content blocks are supported. Empty completions and malformed native arguments are provider failures. Contracts record response_policy, rows record response_formats, and cache keys isolate endpoint, replica, seed, and elicitation fingerprint. /runs distinguishes legacy protocols; compare elicitation fingerprints before interpreting cross-model differences.
  • Reproducibility: --seed N rides every request (sample k uses seed + k) and is recorded in the contract. A local server is still not bit-reproducible at temperature 0 — measured against LM Studio, the same 24 trials flipped 7/24 at --jobs 4, 3/24 with a seed, and 1/24 at --jobs 1 --seed 7, because batching changes kernel reduction order and genuine ties are broken by the hardware. Aggregates are far more stable than single trials; the cache is the real replay mechanism. Cache entries are keyed per sample, so a one-sample run extends to --samples 3 without re-paying.
  • Honest statistics: pass@1 is reported beside pass@k, the cost columns are per-sample means, difficulty is computed from the work a trial demands rather than rolled, and a clustered Wilson interval over answer-shape cells sits next to the per-trial one — repeated shapes are one question, not many. Fractional token means are preserved from the runner through the desk. Missing completion usage is estimated separately for each successful answer; skipped provider samples contribute zero tokens.
  • Future authoring audit: kyber audit reports trial, title, exact reference-shape, grade, and oracle-kind counts for every family, then flags low title or answer-shape variety and missing calibrated hard/extreme cases as authoring notes. It also finds prompts that differ only in case, whitespace, or sentence-ending punctuation within one family; shell syntax inside backticks and meaningful path punctuation stay distinct. The published suite currently has one such pair (KB-T0020 / KB-T0121); it stays published so existing measured scores remain comparable. Use --strict on a proposed next edition to reject that class of duplicate before publishing, then rerun the full gate and model strikes for the new fingerprint. Shape counts indicate repeated expected payloads, not necessarily identical reasoning tasks; coverage notes are advisory because each family has different semantics.
  • Desk: /, /leaderboards, and /runs read the compact, fingerprinted src/data/kyber-overview.gen.ts summary emitted by sync. /models, /compare, /insights, and /kyber share src/data/kyber-analysis.gen.ts: a trial index and eight measured values per model/trial, with explicit nulls for missing rows. The shared decoder retains exact model identity, oracle latency, and fractional costs. Only /kyber loads detailed prompts, references, and oracle bodies. The full strike twin remains an independent parity artifact. Both compact payloads carry the same measured snapshot fingerprint; the trial twin carries the suite fingerprint. Mixed or stale emissions display unavailable evidence and recovery commands. Never hand-edit .gen.ts: regenerate with trials-ts / sync. npm run parity:kyber checks every compact row, cost, profile, pair, and consensus against the independent detailed twin for published and sparse synthetic evidence across all 11 families, then checks aggregates against the engine. npm run parity:compare checks the compact paired outcomes and McNemar against kyber compare. Shared taxonomy: src/data/kyber.ts (families, grades, oracle kinds, infra kinds, shell allowlist) and src/data/kyber-suite.ts (detailed meta, coverage, Wilson, clustered Wilson, validity).
  • Strike designer: /kyber explicitly copies jobs, samples, temperature, seed, timeout, and response policy, plus optional endpoint and stratified slice filters. Choose Bash / zsh or PowerShell to quote opaque API ids correctly. Windows PowerShell cannot reliably preserve embedded double quotes or a trailing backslash in an id; the designer directs those ids to Bash. HTTP endpoints support IPv4, DNS names, and bracketed IPv6. Credentials, query strings, and fragments do not belong in the provider base URL. npm run parity:command runs the actual copied settings through a local mock provider and checks the JSONL and synced desk costs without changing published runs.
  • Codec performance: the std-only JSON parser traverses the already validated UTF-8 input without revalidating the remaining string for every character. cargo run --release --manifest-path crates/kyber/Cargo.toml --example codec_bench measures three Unicode/escape string sizes with five samples each and reports medians. These are reproducible observations, not timing gates; deterministic roundtrip, malformed-surrogate, and document-truncation tests guard correctness.
  • Official top 25: sync publishes one newest valid strike contract per model and records its source file. A rank requires engine 0.1.0, current suite and oracle fingerprints, an unfiltered full-suite run without sampling or a limit, and one result for all 1500 trials. Pass@1 orders qualifying runs; grade-weighted pass and model id break ties. The board shows only real qualifying runs, up to 25, and /runs explains why a synced partial or stale run is excluded. /models shows capability and cost per run; /compare aligns pass@1 by trial id and distinguishes matching, different, or missing elicitation fingerprints, with links to each run contract. Different protocols combine model and elicitation effects. /insights uses common trials across qualifying models for consensus and disagreement.
  • Desk ↔ engine: the browser's filters live in the URL (/kyber?family=terminal, ?trial=KB-T0001), so a family card, a palette hit and a pasted link all land on the same view. Measured evidence is aggregated per model — sync can hold several, and blending them would report a pass rate that describes none of them. The desk's TS mirrors of the engine's taxonomies are checked against the Rust source by src/lib/kyber-taxonomy.test.ts, and its aggregates against kyber stats by npm run parity:kyber.
  • Scoring: pass 1 / fail 0, plus prefix partial credit on calls_exact — the longest correct prefix over the longer of the two lists, so a chain that gets 2 of 3 calls right scores 0.667 and one that adds hallucinated extras cannot score 1.0. Oracle-eval latency is budget-enforced (50 ms in-process, 2000 ms shell; over-budget fails as over_budget); provider model_ms is recorded but never budgeted. Thread-local heap high-water marks ride on every row without cross-worker contamination. Rows stream to disk as their position in the trial order settles, so a run that dies at trial 1200 keeps 1200 scored rows. Measured-only: only strike rows badge trials; floor bounds never do.