# The Precedents Lead With A Grade And Put The Evidence Below It, They Scale From Individual To Corporate By Selling The Same Thing Larger Rather Than Something Different, The Ones That Reach Engineering Ship A Documented JSON Output And A Way To Run It Where The Thing Lives, And The Probe And Finding Vocabulary We Were About To Coin Already Exists

**version** v0.33.64
**date** 4 September 2026
**from** Human (project lead)
**to** Product, Design, Engineering, Ambassador

**type** Research brief

*Seventh of 4 September. The memo asks for research into the free assessment services this product is modelled on, with attention to pacing, inputs, outputs, the free and paid boundary, the individual to corporate ladder, and schemas for import and export. Six services were examined on 4 September and their published pricing, documentation and vocabulary read rather than recalled. Two findings change decisions taken in earlier briefs today, one of them a vocabulary we were about to invent and which already exists in a widely adopted project. Limitation: the services were read from their own published pages, so their descriptions of themselves are the source, and no product was operated end to end.*

---

## What This Is

Research into the services this one is modelled on, and what transfers: **six free assessment services were examined and their shared shape is more consistent than their subjects, since almost all of them take one small input, return a single legible verdict above the fold with the evidence beneath it, require no account for the basic case, and produce a result the visitor wants to show somebody, which is the growth engine rather than a nicety; the pacing lesson is uniform and worth copying exactly, that the grade arrives first and the reasoning second, because a visitor who has to read to find out whether they are in trouble leaves before they find out; the individual to corporate ladder is the most instructive finding, because the service that does it best does not sell a different product to organisations, it sells the same product larger, sized by how many things you are watching and how fast you may ask, with the free tier genuinely useful for one person and the enterprise tier being the same searches without limits, and that shape fits this product exactly since an individual measures one agent and an organisation measures four hundred; the services that reach engineering teams rather than stopping at a landing page all do two things, they publish a documented machine-readable output and they offer a way to run the assessment where the thing being assessed lives rather than only from a server, and the one closest to our own constraint publishes structured findings in place of a score so that consumers can apply their own policy engine, which is the same argument made yesterday against a levels ladder that multiplies coverage by tier; the vocabulary finding is the one to act on immediately, because the probe and finding words proposed here yesterday already exist in that project with definitions close enough to adopt, a probe being an individual heuristic about a distinct behaviour the subject may or may not be doing and a finding being the result of one probe, so the anchoring rule this corpus settled in August applies and we should use theirs rather than mint our own; and the certification ladder from the national scheme supplies the missing commercial rung, since it sells the same questionnaire twice, once self-assessed and once audited by a third party, which is the evidence tier written as a price list and is the cleanest available precedent for charging for independence rather than for information.** New contributions: **the six services compared on the axes the memo asked about, the grade-first pacing rule, the ladder that sells the same thing larger, the two properties that separate services that reach engineering from those that do not, the vocabulary already in use that we should adopt, and the self-assessed against audited tier as a price list.**

## The Services Examined

| Service | Input | Output | Account | Corporate path |
|---|---|---|---|---|
| **Network exposure probe** (Gibson Research) | Nothing. It probes the visitor's address | A verdict per port and a plain-language summary | No | None. It stayed a public good |
| **Transport security grader** (SSL Labs) | A hostname | **A letter grade**, then a long report | No | An API and a documented assessment methodology |
| **HTTP header grader** (securityheaders.com) | A URL | **A letter grade**, then per-header advice | No | A related paid monitoring product |
| **Website observatory** (Mozilla) | A URL | A score and per-test results, with the scoring rules published | No | Open source scoring |
| **Breach lookup** (Have I Been Pwned) | An email address | Yes or no, and the breach list | No, for the basic case | **A sized subscription ladder** |
| **Repository health** (OpenSSF Scorecard) | A repository | A score, and **structured findings** | No | **Runs in the pipeline, JSON output** |

## Finding One: The Grade Comes First

Every service that succeeded at reaching non-specialists leads with a single legible verdict and puts the reasoning beneath it. A letter, a score, a yes or no.

**A visitor who has to read in order to discover whether they are in trouble will leave before discovering it.** That is the pacing rule and it is not a stylistic preference; it is what separates the services people sent to their colleagues from the ones that stayed among specialists.

Three properties travel with it.

**The verdict is comparable to itself over time.** A letter grade can improve, and improvement is why somebody returns.

**The evidence is immediately below, not behind a click.** The services that hid the reasoning were trusted less, and the ones that published their scoring rules outright were trusted most.

**And the verdict is shareable without embarrassment being the only motive.** A good result is worth posting, which matters more than it sounds: a service where only bad news is shareable spreads through fear and stalls.

For this product that argues for a verdict on the gap that a person can be pleased with, not only alarmed by, and against a score that can only ever look bad on a first visit.

## Finding Two: The Ladder Sells The Same Thing Larger

The clearest individual-to-corporate ladder among the six does not sell organisations a different product. It sells the same searches at greater scale, priced on two dimensions.

| Tier | Shape |
|---|---|
| Free | Your own address, notifications, and monitoring for a domain with a small number of affected addresses |
| Paid, entry | The same searches, sized by requests per minute and number of domains watched |
| Paid, higher | The same again at larger rates and counts, plus a few capabilities aimed at people doing it for others |
| Enterprise | The same with no rate limits, deployable under the customer's own name, with callbacks |

**Nothing on that ladder is a feature an individual cannot have. It is a quantity an individual does not need.** That is a much easier ladder to explain and to defend than a feature matrix, and it avoids the failure where the free tier is deliberately crippled and the product is resented.

It transfers directly: **an individual maps one agent, a team maps a dozen, an organisation maps hundreds and wants them watched.** The unit is agents watched and profiles tracked, and the free tier should be genuinely sufficient for one person mapping their own tools, because the hypothesis being tested is about individuals as well as companies.

## Finding Three: Two Properties Separate Reach From Reputation

Several of these services are respected and stop at their own landing page. The ones that became part of how engineering teams work share two properties, and neither is about the assessment itself.

**A documented machine-readable output.** A JSON result with a stated shape, which somebody can store, diff and act on.

**A way to run it where the thing lives.** Either a command-line tool or a pipeline step, so the assessment runs against the real thing rather than against whatever a server can reach from outside.

The second is the one that matters most here, because this product has no choice: agent exposure is not visible from outside, so running where the thing lives is not an enhancement, it is the only available mode. **What the precedents show is that this is a strength rather than a compromise**, since the services with that property are the ones that ended up in other people's build systems.

The corollary is that the import and export schema is not a later feature. **It is the thing that decides whether this reaches teams at all**, and it should exist before the website does.

## Finding Four: The Vocabulary Already Exists, So Use It

The strongest single finding, and it corrects a decision taken in a brief written earlier today.

The repository brief proposed that the unit of contribution should be a probe, with each grant claim carrying the command that establishes it. That vocabulary already exists in the repository health project, with definitions close enough to adopt outright:

> A **probe** is "an individual heuristic, which provides information about a distinct behavior a project under analysis may or may not be doing."
>
> A **finding** is "the result of an individual probe. Each finding describes the outcome of a particular probe, and optionally, a location in the repo where this behavior was observed."

**Their findings are returned as JSON carrying the probe name, a description, a message, an outcome, and optional remediation guidance.** That is a working schema for exactly the object the repository brief specified, published, in use, and battle-tested at scale.

This corpus settled in August that you anchor to published identifiers rather than minting new ones, and that an anchor resolves identity and never authority. The same rule applies to vocabulary. **Adopt probe and finding with their definitions, cite the source, and diverge only where our subject genuinely differs**, which it does in at least two places worth naming: our probes carry a reversibility property because recoverability decides insurability, and our findings carry an evidence tier because some grants can only be read from documentation.

## Finding Five: Their Argument Against Scores Is Our Argument Too

The same project made a change worth reading closely, because it is the argument made here yesterday, arrived at independently.

They moved from returning only a score to returning structured findings, on the reasoning that the score "combines multiple data points" with limited transparency, while structured results give "granular information about a project's security practices" that consumers can "tailor to their needs" and "act on using a policy engine of their choice."

That is precisely yesterday's objection to a maturity level that multiplies coverage by evidence tier: **a single number cannot say which of its inputs moved.** Finding a large open project reaching the same conclusion from a different subject is the strongest available evidence that the objection is structural rather than a preference.

**It also settles the tension in finding one.** Lead with a grade, because that is what makes a service legible and shareable. Return structured findings, because that is what makes it useful. The two are not in conflict as long as **the grade is derived from the findings and the findings are what travels.**

## Finding Six: Independence Is The Thing Worth Charging For

The national certification scheme examined sells the same requirements twice: once as a self-assessed questionnaire, and once with a third party verifying the same claims on the customer's systems. The second costs materially more and is the one buyers ask for when they are answering somebody else's question rather than their own.

**That is the evidence tier written as a price list**, and it is the cleanest precedent for this product's commercial boundary:

| What is sold | Tier |
|---|---|
| The questions, the primitives, the guidance, and your own answers | **Free.** Self-reported |
| Somebody running the probes in your environment and vouching for the result | **Paid.** Independently observed |

That is a better articulation of the paid boundary than the one written earlier today. **The paid thing is not access to information and not customisation; it is independence.** Customisation is real work and sells too, but independence is the thing a customer cannot produce for themselves at any price, and it is exactly what the evidence-tier model says is worth more.

## What Transfers, And What Does Not

| Property | Transfers? |
|---|---|
| One small input, verdict first, evidence below | **Yes.** Copy exactly |
| No account for the basic case | **Yes** |
| A shareable result | **Yes**, and it should be pleasant to share when good |
| Published scoring rules | **Yes.** The most trusted services publish theirs |
| Sized ladder rather than a feature matrix | **Yes** |
| Structured findings under the grade | **Yes**, with reversibility and tier added |
| Probing the visitor from a server | **No.** There is nothing to probe from outside |
| A result computed entirely on our infrastructure | **No**, and its absence is the privacy story |
| A letter grade as the primary artefact | **Probably not.** A grade implies a scale somebody calibrated, and nobody has calibrated one for this |

**That last row is the one to argue about.** The precedents say lead with a verdict, and this estate's own position is that an uncalibrated number invites false confidence. The reconciliation is a verdict that is a statement rather than a score: the count of capabilities held and never asked for, with the irreversible ones named, which is legible in one line, honest, comparable to itself over time, and needs no scale nobody has validated.

## What To Build, In The Order The Research Suggests

| # | Step | Why this order |
|---|---|---|
| 1 | **The findings schema**, using probe and finding as defined by the existing project, plus reversibility and tier | It decides whether this reaches teams, and it must exist before the site |
| 2 | **The command-line runner** that produces those findings where the agent lives | The only mode available, and the property that gets a tool adopted |
| 3 | **The verdict**: a one-line statement, not a grade, computed from the findings | The pacing rule, honestly satisfied |
| 4 | **The site** that explains, accepts a findings file, and renders the verdict and the evidence | The landing page comes fourth, not first |
| 5 | **The shareable artefact**, pleasant when the news is good | The growth engine |
| 6 | **The sized ladder**: agents watched, profiles tracked, and nothing withheld that an individual needs | The commercial shape, copied |
| 7 | **The independent run**, as the first genuinely paid thing | Independence is what cannot be self-produced |

## What This Does Not Try To Be

- **Not a competitive analysis.** None of these services is a competitor; they are precedents for a shape.
- **Not an endorsement of any scoring method.** The scoring is the part least transferable, and the section above says why.
- **Not a complete survey.** Six services, read from their own published pages on one date.
- **Not a pricing proposal.** The ladder's shape is copied; the numbers are somebody else's and not ours.
- **Not a claim to have operated these products.** Their documentation is the source.

## Honest Tensions

| Tension | Note |
|---------|------|
| Lead with a verdict, refuse a grade | It keeps the pacing and gives up the single most shareable artefact these services have |
| Adopting another project's vocabulary | It avoids minting a fifth private term and ties our schema to somebody else's decisions |
| The schema before the website | It is the right order for reaching teams and the wrong order for showing anybody anything |
| A ladder sized by quantity | It is honest and it means an individual with many agents pays like a small company |
| Independence as the paid thing | It is the most defensible boundary and it is the least scalable service we could sell |
| Six services read from their own pages | It is what was available and every one of them is describing itself |

## Open Questions

| Question | Notes |
|----------|-------|
| Is a one-line verdict shareable enough? | The grade is what spread these services, and refusing one costs something real |
| How closely should the findings schema follow the existing one? | Close enough to reuse tooling, and our reversibility and tier fields are genuine divergences |
| What is the unit on the sized ladder? | Agents watched and profiles tracked are the candidates, and they behave differently for a large estate |
| Who performs an independent run? | It is the paid tier and it is people, which is the part of the model that does not scale |
| Does publishing the scoring rules invite gaming? | The most trusted precedents publish theirs anyway, which suggests the answer is yes and it is worth it |
| Should the free tier include the machine-readable output? | Withholding it would be the classic crippled free tier, and including it makes the paid tier purely about independence and scale |

## Relationship To Previous Briefs

| Date | Document | Relationship |
|---|---|---|
| 4 Sep | `v0.33.64__dev-brief__the-grant-mandate-repo-ships-probes-not-tables-grants-are-measured-per-tool-and-contributed-mandates-are-elicited-locally-and-never-leave.md` | **Corrected here**: the probe and finding vocabulary it proposed already exists and should be adopted rather than minted |
| 4 Sep | `v0.33.64__strategy-brief__name-the-question-not-the-concept-a-surprise-is-an-action-outside-the-grant-so-the-assessment-is-falsifiable-and-the-scan-cannot-run-from-outside.md` | The site this researches, and the paid boundary sharpened here from customisation to independence |
| 4 Sep | `v0.33.64__arch-brief__the-shipped-levels-ladder-blends-coverage-with-tier-and-evidence-mode-is-the-axis-that-decides-the-insurance-instrument.md` | The argument against collapsing a number, which a large open project reached independently |
| 3 Sep | `v0.33.63__strategy-brief__the-level-problem-we-built-an-endgame-solution-and-every-audience-is-on-level-one-each-levels-solution-creates-the-next-levels-problem-and-concepts-are-introduced-in-order-of-attrition.md` | The ladder the sized tiers map onto |
| 9 Aug | `v0.33.57__arch-brief__sg-send-enrichment-and-shared-anchors-research-paid-once-wikidata-is-the-concept-layer.md` | Anchor to published identifiers rather than minting new ones, applied here to vocabulary |
| 20 Aug | `v0.33.61__dev-brief__user-section-is-a-conformance-test-store-the-choices-not-the-answers-high-threat-without-efficacy-produces-denial.md` | Why the verdict must be shareable when the news is good, not only when it is bad |

---

## Key Claims

| # | Claim |
|---|-------|
| 1 | Every service that reached non-specialists leads with a single legible verdict and puts the evidence immediately below it |
| 2 | A visitor who must read to discover whether they are in trouble leaves before discovering it |
| 3 | A result must be pleasant to share when the news is good, or the service spreads only through fear and stalls |
| 4 | The clearest individual-to-corporate ladder sells the same product larger, sized by scale rather than by withheld features |
| 5 | Nothing on that ladder is a feature an individual cannot have; it is a quantity an individual does not need |
| 6 | The services that reached engineering teams publish a documented machine-readable output and run where the thing lives |
| 7 | Running where the thing lives is our only available mode, and the precedents show it is a strength rather than a compromise |
| 8 | The import and export schema decides whether this reaches teams, so it exists before the website |
| 9 | The probe and finding vocabulary already exists with usable definitions and a published JSON shape, and should be adopted rather than minted |
| 10 | A large open project replaced a score with structured findings for the same reason we objected to a collapsed level, arrived at independently |
| 11 | Lead with a verdict and return structured findings: the grade is derived from the findings and the findings are what travels |
| 12 | Independence rather than information or customisation is the thing a customer cannot produce for themselves, and it is the honest paid boundary |

## Sources

- ShieldsUP, Gibson Research: a free network exposure probe requiring no input from the visitor beyond visiting. https://www.grc.com/shieldsup and https://en.wikipedia.org/wiki/ShieldsUP
- Have I Been Pwned subscription tiers, read on 4 September 2026 for the individual-to-organisation ladder: a free tier covering an individual's own address and a small domain, then plans sized by requests per minute and domains monitored, then an enterprise tier described as the public plans without rate limits and with white-label deployment. https://haveibeenpwned.com/Subscription and https://haveibeenpwned.com/API/v3
- OpenSSF Scorecard structured results, read on 4 September 2026 for the probe and finding definitions quoted above, the JSON shape carrying probe name, description, message and outcome, and the stated reasoning for exposing findings in place of a score so consumers can apply their own policy engine. https://openssf.org/blog/2024/04/17/beyond-scores-with-openssf-scorecard-granular-structured-results-for-custom-policy-enforcement/ and https://github.com/ossf/scorecard
- Mozilla Observatory, examined for the published scoring rules and the score-plus-per-test output pattern. https://developer.mozilla.org/en-US/observatory
- The Cyber Essentials certification ladder, read on 4 September 2026 for the self-assessed against independently audited tiers of the same requirements. https://iasme.co.uk/articles/cyber-essentials-and-cyber-essentials-plus-what-is-the-difference/

---

This document is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).
