# Reach Is A Node In A Mesh Rather Than A Rung On A Ladder, Questions Divide Into Ones That Identify And Ones That Measure, And Each Carries A Reliability Because A Wrong Answer Is Not Noise But The Finding: The Prediction Gap Is Collected Throughout Rather Than Once At The End

**version** v0.33.65
**date** 5 September 2026
**from** Human (project lead)
**to** Product, Engineering, Design

**type** Dev brief

*Third of 5 September. The guessing game shipped to yesterday's specification and was read before writing this: seven profiles across fifteen questions, belief updating rather than branch pruning, the next question chosen to split the belief most evenly, a prediction step before the reveal, the gap as the output, deterministic arithmetic in the tab with nothing sent, and the prediction gap kept distinct from the surprise defined elsewhere. The memo asks for an evolution and one of its asides overturns a design decision in that specification, because a user who answers wrongly is not producing noise, they are producing the measurement. Limitation: the shipped page was read rather than played through, and the reliability weights proposed below are a mechanism with no values behind them yet.*

---

## What This Is

An evolution of the guessing game, built on one correction the memo makes and one it implies: **the correction is that a capability's reach is not a rung on a ladder but a node in a mesh, since reading a file system is a property whose meaning depends entirely on which file system, and the local disk, the user's home, a corporate share, a network mount, a cloud store and a container's own layer are six different exposures wearing one verb, and each of those is itself connected to running environments and products so that starting from the local file system and walking outward reaches desktop applications and command line tools rather than being reached by them, which is the estate's own start-anywhere principle and makes the primitive a triple of verb, object and a reach node rather than a scalar; the second structural point is that the graph needs coarse and fine views of the same thing so a reader can ask about file access or about one mount without changing format, which is the fractal test again and is the only way a mesh this size stays legible; the memo's request for funny questions is better than it sounds because a question about whether the product carries a person's name or the vendor's name contains a particular word is enormously discriminating and measures nothing at all, which forces a distinction the shipped game does not make between questions that identify and questions that measure, and only the second kind may contribute to the finding; the aside that overturns yesterday's design is the observation that a user asked whether the thing can reach every credential on their machine may answer no because they do not know, which means every answer is already a prediction and the gap is not one number collected at the end but a set of per-capability disagreements collected throughout, and it also means the engine must not treat a low-confidence answer as strong evidence about which product this is, so each question carries a reliability that weights it heavily for identification and lightly for belief, or the other way round; the detective panel the memo wants is the corpus's own evidence discipline rendered as an interface, showing what the player asserted, what the engine inferred, and what remains possible, in three visually distinct classes that are never mixed, because a hypothesis drawn like a fact is the thing this whole estate exists to prevent; the deterministic engine stays the default and a model earns its place only at three named edges, free text, question generation and narration, none of which is the core; and the tree grows toward things people do not think of as agents, which is where the best guesses live, and it joins the regulatory graphs at the end with the honest caveat that a profile matched from answers is the weakest evidence tier there is, so an obligation surfaced this way is a hypothesis to check rather than a compliance finding.** New contributions: **reach as a node with the mesh it implies, two levels of granularity over one format, the split between identifying and measuring questions, a reliability per question, the prediction gap collected per capability throughout, the three-class inspector, the three edges where a model earns its place, and the regulatory join with its tier caveat.**

## What Shipped

Read from the page on 5 September, and it matches the specification: **seven profiles across fifteen questions**; answers update a belief rather than prune a branch so a wrong answer is recoverable; the next question is the one that splits the remaining belief most evenly; the player commits to a prediction before the reveal; the output is the matched profile, the measured grant, the prediction gap and the path taken; it is deterministic arithmetic over a published tree running in the tab with nothing sent; and it measures prediction gaps rather than surprises, which are two different things and stayed apart.

**The name is right and should not be churned.** Yesterday's rule is to name a thing for the question a visitor has, and "which agent is it" is literally that question. It passes its own test.

## The Correction: Reach Is A Node, Not A Rung

The repository brief proposed a capability as a verb crossed with an object class crossed with a **reach**, and gave reach as a ladder: self, project, host, tenant, world. The memo shows that is wrong, and the example is exact.

> Read a file system is a property. **But which one?**

| The same verb and object | The actual exposure |
|---|---|
| read x file, the container's own layer | Nothing of yours. Ephemeral |
| read x file, the user's home directory | Documents, keys, history, other projects |
| read x file, a corporate share | Other people's work, and the audit question that follows |
| read x file, a network mount | Whatever is mounted today, which is not a fixed set |
| read x file, a cloud store | A different blast radius and a different regulator |

**Those are five different exposures wearing one verb.** A ladder cannot express them because they are not ordered: a corporate share is not more or less than a cloud store, it is elsewhere.

So the primitive becomes:

```
   capability = verb  x  object class  x  REACH NODE      + reversible?

   where a reach node has its own edges:
      fs:user-home        --runs-in-->  env:desktop, env:wsl
                          --exposed-by-->  product:desktop-app, product:cli
      fs:corporate-share  --runs-in-->  env:managed-desktop
                          --governed-by-->  obligation:...
      env:container       --provided-by-->  vendor:...
```

### And that makes the traversal bidirectional

The memo's own test: start at the local file system, walk outward, and you should arrive at the desktop applications and command line tools that touch it. That is the estate's **start anywhere** principle from August, applied here, and it produces a second use of the same data for free.

| Direction | Question it answers | Who asks it |
|---|---|---|
| Product outward | What does this thing reach? | The player, in the game |
| **Reach inward** | **Which products touch my corporate share?** | An organisation, and it is a better question |

**The inward direction is the more valuable one and the game currently cannot answer it.** Building the mesh rather than a ladder gets it without extra data.

## Two Levels Of Granularity, One Format

The memo asks for more or less granularity in the graph. The rule is the fractal test the estate already holds: **zooming must not change the format.**

```
   coarse   capability:file-access          one node
   fine     read x file x fs:user-home      one node, same shape, same edges
```

A coarse node is a named set of fine ones, addressable and traversable identically. **A coarse view that is a different object from a fine view is a hierarchy, and it will need special cases at every boundary.** The game shows coarse by default and lets a curious player descend, and the same data serves both.

## Two Classes Of Question, And Only One Of Them Measures

The memo asks for funny questions and gives good ones: does the name contain a particular word, is it named after a person, is the vendor's name a word rather than a person's. These are excellent and they force a distinction the shipped game does not make.

| Class | Example | Identifies? | Measures? |
|---|---|---|---|
| **Discriminating** | Does the product have a person's name? Do you use it in a terminal? | **Strongly** | **Not at all** |
| **Eliciting** | Can it read every credential on your machine? Can it send email as you? | Weakly | **Strongly** |

**A question about a product's name splits the profile space beautifully and tells us nothing about anybody's exposure.** That is fine, and it must be labelled, because only eliciting answers may contribute to the prediction gap. Mixing them would let a player who guessed the vendor's naming convention appear to understand their grant.

Two design rules follow. **Open with discriminators**, because they are cheap, fast and funny, and they get a player through the first third with the belief already concentrating. **Close with elicitors**, because by then the candidate set is small enough that a disagreement is informative about that specific product rather than about the population.

And the humour is not decoration. A game that is enjoyable is completed, and the whole instrument depends on completion.

## Reliability Per Question, Which Is The Real Correction

The memo's sharpest aside:

> Does the thing have access to all the credentials on your computer? The user might say no, **but this might be because the user is not aware of it.**

That breaks an assumption in the shipped engine, which treats every answer as evidence about which profile this is. **A user who wrongly says no gets moved toward the wrong profile by their own ignorance**, which is precisely the ignorance the game exists to measure.

The fix is one field per question:

```
   question:
     text
     class:        discriminating | eliciting
     reliability:  0.0 .. 1.0     how likely a player is to know the true answer
     weight_identify:  derived from reliability
     weight_believe:   derived from (1 - reliability)
```

| Question | Reliability | Used for |
|---|---|---|
| Do you use it in a terminal? | **High.** You know where you type | Identification, strongly |
| Does its name contain a particular word? | **High** | Identification, strongly |
| Can it read every credential on your machine? | **Low.** Almost nobody knows | **Belief, strongly.** Identification, barely |
| Can it send email as you? | Medium | Both, moderately |

**A low-reliability answer barely moves the profile belief and fully counts toward the prediction gap.** That is the whole correction and it makes the engine robust to exactly the user error the memo describes, while turning that error into the product.

## The Prediction Gap Is Collected Throughout

Which produces the second consequence, and it supersedes yesterday's design.

The shipped game asks for one prediction before one reveal. **But every eliciting answer is already a prediction**, so the gap is not a single number at the end; it is a set of per-capability disagreements accumulated as the player goes.

```
   for each eliciting question answered:
      player said        no
      matched profile    yes, at reach fs:user-home, irreversible
      -> a disagreement, on a named capability, with its reach and reversibility

   the gap  =  the disagreement set
   the number  =  what fraction of the true grant the player knew about
```

Three things that buys, none available from a single end-of-game prediction.

**The gap is specific.** Not "you underestimated" but "you did not know it reaches your home directory", which is actionable in a way a percentage is not.

**It is weighted by what matters.** A disagreement on an irreversible capability counts differently from one on a reversible one, which is the reversibility field earning its place.

**And it can be shown at the moment it happens** rather than saved for the end, which is better for the player and, per the corpus's own finding that high threat without efficacy produces denial, must arrive with the fix beside it.

**Keep the end-of-game prediction as well.** It measures something different: whether a player's overall sense of scale is right, which is not the sum of their per-question errors.

## The Detective Panel: Three Classes, Never Mixed

The memo wants the engine's inner workings visible, as facts, hypotheses, current thinking and possibilities. That is this estate's evidence discipline rendered as an interface, and it should be built with the same strictness.

| Class | What it is | How it must look |
|---|---|---|
| **Asserted** | The player said this | One visual treatment. **Never inferred, always theirs** |
| **Inferred** | The engine derived it from the matched profile's measured grant | A second treatment, clearly weaker |
| **Possible** | Still consistent with the belief, not yet decided | A third, faint |

**A hypothesis drawn like a fact is the thing this whole estate exists to prevent**, so the panel is a demonstration of the product's own argument rather than a debugging aid that happens to be pretty. Which also means it should be on by default at a low level of detail rather than hidden behind a toggle.

Below the three classes, the belief itself: the candidate profiles with their current probabilities, the question that produced the largest change, and what the next question is expected to resolve. **A player who can see why it asked what it asked trusts the guess**, and a player who cannot is playing with a fortune teller.

## Where A Model Earns Its Place

The memo asks again whether a deterministic formula can carry this. It can, and the shipped game proves it. A model earns three specific edges and nothing else.

| Edge | Why a model | Why not the core |
|---|---|---|
| **Free text** | "It is a bit like that but with a plugin" is not answerable by a tree | The tree's failure to place somebody is a recorded gap; a model's guess is not |
| **Question generation** | Proposing new discriminators as the profile set grows is drudgery a model does well | A generated question still needs a reliability and a class, which are judgements |
| **Narration** | Explaining a path in words, on request | The path is already the explanation |

**The rule stands: computation where the answer is checkable, a model where it is not, and never a model in front of an arithmetic step.** Third statement of the same rule in three days, which is enough to consider it settled.

## Growing The Tree Toward Things Nobody Calls An Agent

The memo makes a point worth pulling to the front: the set should include things that are not agentic, or do not feel it, but are.

**The most valuable guess this game can make is one the player did not think was in scope.** A browser extension with broad host permissions, a continuous integration runner with repository credentials, a mail plugin with a token, a scheduled job with a service account. Each has a grant, most have no mandate written anywhere, and none of them is what somebody pictures when they read the word agent.

So the profile set grows along two axes, and the second is the interesting one:

| Axis | Examples |
|---|---|
| More agents | The next assistants, the next coding tools |
| **More things with grants** | Extensions, runners, plugins, service accounts, scheduled jobs |

And an entry to the game that says think of an agent will never surface the second. **A later variant should ask the player to think of a tool they have connected to something**, which is a wider net over the same mesh.

## The Regulatory Join, With Its Caveat

The memo wants to reach obligations: whether a matched thing is in scope for particular regimes and what its compliance level is. The mesh makes that a traversal rather than a feature, because a reach node can carry an edge to an obligation and an article graph already exists in the estate with resolvable identifiers.

```
   product -> capability -> reach node -> governed-by -> article node (identifier, hash)
```

**And the caveat is not optional.** A profile matched from a player's answers is the weakest evidence tier this estate has: it is self-reported, about a product rather than a deployment, and inferred rather than measured. So what comes back is **an obligation worth checking**, never a compliance finding, and the wording has to carry that or the game becomes exactly the kind of unearned assurance the conformance work refuses to give.

The honest arc, which is the same arc as the assessment site: **the game raises the question in two minutes, a probe answers it in ten, and a conformance assertion with a named acceptor answers it properly in a quarter.**

## Build Order

| # | Step | Produces |
|---|---|---|
| 1 | Reach as a node, with the mesh and both traversal directions | The primitive corrected, and the which-products-touch-my-share question |
| 2 | Coarse and fine as the same shape | Legibility without a second format |
| 3 | Question class and reliability fields on the existing fifteen | The engine stops being fooled by ignorance |
| 4 | Per-capability disagreements collected throughout | **The gap becomes specific and actionable** |
| 5 | The three-class inspector, on by default | The estate's own argument, demonstrated |
| 6 | Funny discriminators, a dozen of them | Completion, which the instrument depends on |
| 7 | Non-agentic profiles | The guesses that surprise people most |
| 8 | The obligation edges, with the tier caveat in the copy | The compliance question, honestly framed |
| 9 | The model at the three named edges | Coverage for the unusual |

**Steps 3 and 4 are the substance.** Everything else is reach or polish.

## What This Must Not Do

- **Not let a discriminating answer count toward the gap.** Guessing a naming convention is not understanding a grant.
- **Not treat a low-reliability answer as strong evidence about the product.** That is the engine being fooled by the ignorance it exists to measure.
- **Not draw a hypothesis like a fact.** Three classes, three treatments, never mixed.
- **Not present a matched profile as a measurement.** It is inferred from self-reported answers, which is the weakest tier available.
- **Not surface an obligation as a compliance finding.** It is a question worth asking.
- **Not put a model in front of the arithmetic.**
- **Not show a gap without the fix on the same screen.**

## What This Does Not Try To Be

- **Not a measurement.** It matches a profile from answers; the probes measure.
- **Not a compliance assessment.** It can point at an article and never at a verdict.
- **Not a complete mesh.** Reach nodes will be wrong at the edges from the first week, which is what the correction path is for.
- **Not a replacement for the end-of-game prediction.** That measures a sense of scale, which per-question errors do not sum to.
- **Not calibrated.** The reliability values are a mechanism with no numbers behind them until somebody plays it enough to fit them.

## Honest Tensions

| Tension | Note |
|---------|------|
| Reach as a mesh | It is what the world is like and it is much harder to author than five rungs |
| Two question classes | It keeps the finding honest and it means half the questions contribute nothing to the measurement |
| Reliability per question | It fixes the ignorance problem and every value is a guess until there is play data to fit |
| Per-capability disagreements | It makes the gap actionable and it makes the reveal longer and more confronting |
| The inspector on by default | It demonstrates the argument and it makes a light game look like an instrument |
| Reaching for obligations | It is the most commercially interesting edge and the one most likely to be over-read by somebody in a hurry |

## Open Questions

| Question | Notes |
|----------|-------|
| How are reliability values obtained? | Fitted from play data once submissions accumulate, and guessed until then, which should be said on the page |
| How many reach nodes before the mesh is unauthorable? | Six file-system reaches already, and network and identity will be worse |
| Does the coarse view need its own priors? | Probably, and a coarse belief that disagrees with the fine one would be a bug worth catching |
| What is the entry line for non-agentic things? | Think of an agent will not surface a runner, and think of a tool you connected to something might |
| Should the disagreement set be exportable? | It is the beginning of a findings file, which would join the game to the probes properly |
| Who authors a new funny question's reliability? | It is a judgement, and a generated question without one should not ship |

## Relationship To Previous Briefs

| Date | Document | Relationship |
|---|---|---|
| 4 Sep | `v0.33.64__dev-brief__guess-the-agent-the-tree-guesses-your-grant-and-the-product-is-the-gap-between-what-you-predicted-and-what-it-found.md` | The specification this evolves, and whose single end-of-game prediction is superseded by per-capability collection |
| 4 Sep | `v0.33.64__dev-brief__the-grant-mandate-repo-ships-probes-not-tables-grants-are-measured-per-tool-and-contributed-mandates-are-elicited-locally-and-never-leave.md` | **Corrected here**: reach is a node in a mesh rather than a rung on a ladder |
| 5 Sep | `v0.33.65__research-brief__one-measurement-five-artefacts-the-interchange-schemas-already-exist-as-standards-and-every-stage-must-produce-a-finished-thing.md` | The progression this game is stage one of, and the findings shape a disagreement set would export into |
| 4 Sep | `v0.33.64__strategy-brief__name-the-question-not-the-concept-a-surprise-is-an-action-outside-the-grant-so-the-assessment-is-falsifiable-and-the-scan-cannot-run-from-outside.md` | The naming rule the game's own name passes, and the surprise it must not be confused with |
| 4 Sep | `v0.33.64__dev-brief__the-conformance-layer-specified-and-shipped-the-same-day-two-edges-kept-apart-unevidenced-as-the-default-and-62-of-1126-crosswalks-resolved.md` | The article identifiers an obligation edge would resolve to, and the tier discipline the caveat inherits |
| 9 Aug | `v0.33.57__arch-brief__sg-send-index-is-not-a-source-caching-nodes-are-prunable-start-anywhere.md` | Start anywhere, which is why the mesh must traverse inward as well as outward |

---

## Key Claims

| # | Claim |
|---|-------|
| 1 | Reach is a node in a mesh rather than a rung on a ladder, because six file-system reaches are elsewhere from each other rather than above or below |
| 2 | A reach node carries its own edges to running environments, products and obligations, which makes the traversal bidirectional |
| 3 | Walking inward answers which products touch my corporate share, which is a better question than the game currently asks |
| 4 | Coarse and fine views must be the same shape, or every boundary needs a special case |
| 5 | Questions divide into ones that identify and ones that measure, and only the second may contribute to the finding |
| 6 | A question about a product's name discriminates beautifully and measures nothing, which is fine once it is labelled |
| 7 | Each question carries a reliability, because a player who does not know the answer would otherwise be moved toward the wrong profile by their own ignorance |
| 8 | A low-reliability answer barely moves the profile belief and fully counts toward the gap, which turns the user's error into the product |
| 9 | Every eliciting answer is already a prediction, so the gap is a set of per-capability disagreements collected throughout rather than one number at the end |
| 10 | The inspector shows asserted, inferred and possible in three treatments that are never mixed, because a hypothesis drawn like a fact is what this estate exists to prevent |
| 11 | A model earns three edges, free text, question generation and narration, and never sits in front of an arithmetic step |
| 12 | The most valuable guesses are things nobody calls an agent, and an obligation reached this way is a question worth asking rather than a compliance finding |

## Sources

- The Which Agent Is It page on pki.sgit.ai, read on 5 September 2026: seven profiles across fifteen questions, belief updating rather than pruning, the next question chosen to split the belief most evenly, the prediction step and the gap as output, the deterministic in-tab statement that nothing is sent, and the distinction kept between a prediction gap and a surprise. https://pki.sgit.ai/guess/index.html

---

This document is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).
