# pki.sgit.ai/bench — MVPs and experiments, and what each does not prove # # Every entry below carries a `does not prove` list. It is mandatory: the # generator refuses to build an entry without one, because a section that # collected demonstrations without their limits would manufacture the false # assurance this site exists to argue against. # # 16 of 17 are built and running. Hub: https://pki.sgit.ai/bench/index.html ## The register — live where https://pki.sgit.ai/registry/index.html code registry/, registry/tools/registry_tool.py since v0.1.26 (last moved v0.1.48) is A static register of agent identities, roles, mandates, grants, acceptances and revocations — eleven records and twenty-three signed statements at constructed public URLs, verifiable with the shipped sgit pki commands. DOES NOT PROVE: - That anything here is trustworthy. Ten of the eleven records are fixtures — private keys published on purpose — so every signature verifies and proves nothing - That the root can be relied on: it is a fixture root, and roots.json says so in its own entry - That enrolment works without a human: the write path is a git commit reviewed by a maintainer, not the account-less lane the pack designs ## The mandate hook — live where https://pki.sgit.ai/packs/grant-and-mandate/enforcement.html code .githooks/pre-push, packs/grant-and-mandate/tools/mandate.py, packs/grant-and-mandate/mandates/ since v0.1.28 (last moved v0.1.29) is A signed mandate compiled into a pre-push hook that git runs — and that refused a real push to dev with error: failed to push some refs. DOES NOT PROVE: - That the mandate has any authority. Its issuer is the fixture root, so anybody could forge it and the hook would enforce the forgery just as diligently - That the constraint is a boundary: it reached tier setting, and --no-verify still gets past it - That it protects a fresh clone — the hook file is committed, the config that activates it is local and does not travel ## Grant measurement — live where https://pki.sgit.ai/packs/grant-and-mandate/library.html code packs/grant-and-mandate/tools/measure.py, packs/grant-and-mandate/library/ since v0.1.27 (last moved v0.1.29) is A tool that generates a grant document for the environment it runs in, and two measured entries — a hosted agent container and a CI runner — that join at the push edge. DOES NOT PROVE: - That the measurement is complete. An agent measuring its own grant reports what it can see; it is a floor, not a census, and says so on its face - Anything about environments nobody has measured — two entries, one agent, and a blind-spot delta needs at least two agents against a common reference - That a hand-assembled entry is as good as a measured one: the gallery caught schema drift in the hand-written entry and none in the tool-generated one ## The building blocks — live where https://pki.sgit.ai/packs/grant-and-mandate/blocks.html code assets/gm-blocks.css, admin/build/gen_blocks.py since v0.1.31 (last moved v0.1.31) is Nine reusable components — tier and evidence badges, cards, the delta block, the grant tree — shipped as a stylesheet and a gallery that renders the real documents. DOES NOT PROVE: - That the components survive contact with a population. They have been exercised against two environments and one mandate, all measured by one agent - That the layouts hold on a phone — the grant tree below 390px has a proposed degradation nobody has tested - That a second consumer will find the contract workable; RiskMandate is committed to consuming it and has not yet ## Map your own case — live where https://pki.sgit.ai/assess/index.html code assess/ since v0.1.16 (last moved v0.1.19) is A visitor assembles their own agent installations as grant trees, sees the gap, and records a decision per gap — storing the choices and never the answers. DOES NOT PROVE: - That the library covers anybody's real estate. Scenario 5 has no tree to point at at all - That the assessment changes what anybody does — it has no backend, so it can measure none of its own success measures - That the acceptor model is sound: it offers a role where the pack's own standard asks for a named person ## The chain room — live where https://pki.sgit.ai/experiments/the-room/index.html code experiments/the-room/, packs/grant-and-mandate/instance-fixture.synthetic.json, admin/build/gen_room.py since v0.1.42 (last moved v0.1.43) is The RiskMandate workflow walked end to end as a playable room — eight stations, the product boundary drawn on the floor, four verbs, and a work item that travels the chain. The left half is real artefacts; the right half is a marked simulation. DOES NOT PROVE: - That the workflow works. The right half is synthetic and says so on every surface: no risk has been derived, priced, accepted or monitored by anybody - That a room gets read where a table gets skimmed — the genre's inherited bet, now two implementations old with zero user tests between them - That the conditions can be monitored for real: 3 of 4 hold by observation, and the fourth — the boundary-tier enforcement point — has never held anywhere in this estate - That the acceptance shown right of the line is RiskMandate's actual product behaviour: the shape is read off their positioning card, not their system ## The table — live where https://pki.sgit.ai/experiments/the-table/index.html code experiments/the-table/, admin/build/gen_table.py since v0.1.43 (last moved v0.1.43) is Actions resolving against grants and mandates, played as cards: six suits, four players including the systems, and the estate's real 26 August incident replayed forward as simulation and backward as audit — with every resolution re-run through the enforcement tool at build time. DOES NOT PROVE: - Coverage. One scenario, four turns, one agent, one control — the mechanics, not the space of plays - That a DOES card is a receipt: nothing is signed by the actor at the time of action. The table shows where receipts would sit, which is not the same as having them - That proposed-action simulation works — playing a hypothetical card against the twin is specified in the brief and deliberately not built here. RETIRED at v0.1.46: it is built, at /simulator/, where the cards are playable against either twin and every outcome is a precomputed verdict of the real tool. A does-not-prove retired by later work is recorded here, not quietly dropped - That the card grammar survives a population: the genre bets are now three implementations deep across two estates, still with zero user tests ## The scenario engine (two worlds) — live where https://pki.sgit.ai/experiments/push-to-github/index.html code experiments/push-to-github/scenario.json, experiments/the-deploy/scenario.json, admin/build/gen_scenario.py, experiments/scenario.css since v0.1.44 (last moved v0.1.44) is One engine, two worlds: Push to GitHub and The Deploy are rendered by the same generator from two scenario.json files, each referencing a measured twin — the engine holds no capabilities of its own, and every card is a twin node wearing scene clothes, with a confidence rung computed from its evidence and a micro-animation per capability kind. DOES NOT PROVE: - That two worlds are many. The memo says tonnes of scenarios; the engine has rendered exactly two, both from twins this estate measured itself - That the animations simulate anything — a travelling dot is a depiction of a capability, not an execution of one; no action is resolved here (that is the table's job) - That the platform library exists: Codex, Lovable and the rest are named in the brief as future scenario.json files, and not one has been written — the fact-based variation catalogue is an argument, not an artefact - That a rung above 2 is reachable: independent evidence exists nowhere in this estate, so the scale's top rung has never been exercised - That an agent reads these pages — the memo's claim that agents would also appreciate the visual representation is untested for both humans and agents ## The control room — live where https://pki.sgit.ai/experiments/the-control-room/index.html code experiments/the-control-room/, admin/build/gen_control.py since v0.1.45 (last moved v0.1.45) is Both scenario worlds on one operator board — mimic diagrams, annunciator tiles whose lamp colour is the tier and nothing else, faceplates on click, and the 26 August incident as a replayable sequence-of-events log with every verdict re-run through the enforcement tool at build time. One more renderer, zero new data. DOES NOT PROVE: - That an operator can run a plant from it. Two units and one recorded incident is a diorama with excellent manners, not a control room under load - That the annunciator scales past twenty tiles a unit, or the log past one incident — the genre solves both (paging, filtering, alarm shelving) and this build implements neither - That anyone reads a mimic faster than a table — the genre bet is now four implementations deep across two estates, still with zero user tests - That REPLAY ever becomes LIVE: a live board needs the registry's write path, monitors feeding facts, and a mandate service — all still stated design, which is why the mode chip is pinned where it is ## The simulator — live where https://pki.sgit.ai/simulator/index.html code simulator/, admin/build/gen_simulator.py since v0.1.46 (last moved v0.1.46) is The first surface here that answers to the visitor rather than replaying this estate's history: play cards against a measured twin and watch what they do. Every outcome is a verdict of the real enforcement tool, a reading of the twin, or UNKNOWN — precomputed at build time, because the browser is not an enforcement point. DOES NOT PROVE: - That the simulation is predictive. It composes measured facts and real verdicts; it cannot model an environment nobody measured, and every outcome carries the date of the measurement behind it - That the hand is the space of plays: eight cards, two worlds, one enforcement tool, and a blast radius that is the twin's own reachability rather than a discovered attack path - That anyone learns more by playing than by reading — the genre bet is now five implementations deep across two estates, still with zero user tests - That any of it is live: nothing is executed, and the ladder on the control room says exactly which four doors would have to open before a board here could claim to describe the present ## The workbench — live where https://pki.sgit.ai/workbench/index.html code workbench/ since v0.1.47 (last moved v0.1.47) is An experimental app: the primitives on a rail — identities, grants (the twin), mandates, facts, actions, a simulator — and every decision hands back an evidence pack, because the decision is disposable and the record is not. DOES NOT PROVE: - That anything is enforced. The simulator decides nothing outside the page; the live decision points remain the pre-push hook (setting) and a boundary that does not exist (N12) — and every pack says so about itself - That the twin matches the environment now — it is a recording; the memo's real-time question needs a re-measurement at decision time, which this app cannot perform and says it cannot - That a verified signature carries authority — the issuer is a fixture until N11 - That evidence-pack/v0 is settled — introduced here, proposed to the pack, adopted by nobody ## The Delta Is Where the Insurance Lives (the insurance book) — live where https://pki.sgit.ai/insurance-book/index.html code insurance-book/ — content/, build.py, build_quotes.py, gen_pages.py, gen_pdf.mjs, shots/ since v0.1.61 (last moved v0.1.61) is The second volume, by the method of the first: ten voice memos and a pivot briefing, filed verbatim before they were read, made into one argument — the gap between what an agent can do and what it was authorised to do is an insurable exposure, and the honest first product is a rating anybody can recompute, not a premium nobody can check. Seventeen chapters in five parts, with the audit's six defects walked in full. DOES NOT PROVE: - Anything the corpus does not prove. The book inherits all twenty of the insurance folder's does-not-prove entries and adds none of its own evidence: nothing described is insurance, nothing described is built, and no external fact has been gathered - That the book's coherence is the corpus's. A book's job is the through-line, and a through-line is a choice — the five-part structure and the title are the writing session's, recorded as such in BRIEF.md - That a second volume means the method scales: same harness, same gates, one writer, zero readers so far - That anyone needed the book. The doctrine documents are shorter, the memos are primary, and the book's claim to exist is the argument between them — which is a claim about readers, tested by none yet ## A Key Means Nothing Alone (the book) — live where https://pki.sgit.ai/book/index.html code book/, book/content/, book/shots/, book/build.py since v0.1.33 (last moved v0.1.36) is One volume explaining what this site built, how it composes with RiskMandate.ai, and what none of it proves — 17 chapters in five parts, 14 figures each taken at the release tag its caption names, and 65 quotations re-read out of their sources on every build. DOES NOT PROVE: - That the estate it describes is trustworthy. The book's own centre of gravity is a register whose ten fixture records prove nothing and whose root is a fixture — a reader who finishes believing otherwise has read a book that failed - That a participant's account can be neutral. The mitigations are real and are not independence: the strongest bias in such an account is not what it says but what it thinks to check, and there is no way for the writer to know what it did not think to run - That the estate is mature enough to deserve a book — two environments, one agent, one mandate, a fixture root, and one outside reader in its entire history, whose single pass produced half the open contradictions in chapter 15 - That any of this is needed. Nobody outside the project has been asked, which the estate's own doctrine appendix rates a Phase I hole rather than a nice-to-have ## Probes, not tables (the capability registry) — live where https://pki.sgit.ai/probes/index.html code probes/ (primitives.json, probes.json, run.py, schema/, profiles/, evidence/, incidents/, reductions.json), admin/build/gen_probes.py since v0.1.69 (last moved v0.1.70) is A registry of capability primitives (verb × object × reach + reversible) and measured grants where a grant claim never travels alone: it points at the probe that established it, the date, the environment and the output, so a challenge is a rerun rather than an argument. A runner that emits findings/v1 in OpenSSF Scorecard's probe/finding shape plus reversibility and tier; seven profiles, two of them measured; grants public, mandates private. DOES NOT PROVE: - Independence. Every evidence file was produced by the environment it describes — the weakest tier the model has, and stated on every file - Completeness. A self-run probe reports what the subject can see; a capability it does not know it has will not appear. A floor, not a census - That a derived profile is true of any instance. Five of the seven are reasoned from what a surface architecturally is; no probe has been run on them - That the primitive set is right: a starting set, wrong at the edges from the first week by its own admission ## What you authorised and never asked for — live where https://pki.sgit.ai/authorised/index.html code authorised/ (index.html, app.js, README.md), assets/authorised.css since v0.1.69 (last moved v0.1.69) is A self-assessment that runs in the browser: name your tools, the grant appears from measurements other people contributed, four questions about work produce the mandate, and the gap is rendered with the irreversible rows first and the reduction on the same screen. The verdict is a statement, not a grade; the tier is on the result; the counts-only tuple is shown and never sent. DOES NOT PROVE: - Anything about the visitor's environment. It matched profiles from answers, or read a file the visitor brought; it saw nothing itself - That the gap is a loss. It is authorised, and most of it is harmless most days; the irreversible rows decide - That the questionnaire is the mandate. Twelve purposes and ten exclusions are coarse; the full elicitation is an agent asking in the person's own words - Calibration. No surprise count exists for any assessment made here; the validity test is stated, not passed ## Which Agent Is It? (the game) — live where https://pki.sgit.ai/guess/index.html code guess/ (engine.js, app.js, graph.js, report.js, index.html, graph.html, report.html, data.html, tree.json, selftest.json), probes/mesh/, admin/build/gen_mesh.py, admin/build/gen_guess.py since v0.1.69 (last moved v0.1.73) is Named for the question a person has (working title guess the agent). A guessing game whose output is a measurement: cheap questions with obvious answers, ordinary decision-tree induction over a belief across the public profiles, a prediction step before the reveal, and the prediction gap with the reduction on one screen. Deterministic, in the browser, the path always shown, nothing sent. DOES NOT PROVE: - Anything about the player's environment. It matched a profile from self-reported answers, the weakest tier available, and produced a hypothesis, not a measurement - That the tree is right. Seven profiles, fifteen questions, probabilities estimated by the author on one day and answered by nobody yet - That a small gap is safety. Predicting the grant correctly does not narrow it - Anything at scale: no tuple has been submitted, and the aggregate is empty - Calibration. Reliabilities and expected answers are the author's estimates until play data exists ## Synthetic readers — specified where https://pki.sgit.ai/packs/map-your-case/readers/index.html code packs/map-your-case/readers/ since v0.1.23 (last moved v0.1.24) is A programme that puts a page in front of an agent which receives pixels and nothing else — no text, no structure, no knowledge of the project — and watches where it goes. DOES NOT PROVE: - That synthetic readers can report preferences. They find defects; a preference from a simulated reader is not evidence and the programme says so - That the findings generalise — one run, one archetype, one page - That the simulation marker survives export, which is the rule most likely to be broken by accident # Putting something on the bench: its own folder and its own code, an origin # (a brief or a dev-pack document), and a does_not_prove list plus gates — # both enforced by admin/build/gen_bench.py. # All content CC BY 4.0.