# The Screenshot Boundary Is The Instrument And The Patience Budget Must Come From Outside The Model: Synthetic Readers Find Defects Rather Than Preferences, And The Service In Part Two Is One Prompt Away From The Tool This Corpus Ruled Out On 9 August

**version** v0.33.61
**date** 20 August 2026
**from** Human (project lead)
**to** Product, Design, Architecture, the pki.sgit.ai site agent

**type** Dev brief

*Fifteenth of 20 August. The two simulation rules are quoted verbatim from the repository rather than paraphrased, because this memo proposes something one of them forbids and the wording of the rule decides how narrow the exception can be. The rendering capability named here is an existing service in this estate rather than something to build. Limitation: no run has been performed, so the methodological corrections are argued from how the mechanism works rather than from a failed exercise, and the calibration record that would make part two a product does not exist yet.*

---

## What This Is

The tabletop programme the memo asks for, three methodological corrections that decide whether it measures anything, and a rule it collides with: **the memo proposes two agents working side by side, one rendering the site and taking screenshots and one acting as a persona that sees only those screenshots and says where it would click, with the persona's patience, understanding and available time modelled rather than assumed infinite, the whole run exported with screenshots and reactions, published on the site, and ideally performed against mockups before anything is coded, with a second part that packages persona simulation as a service across projects; the rendering half already exists in this estate as a browser-automation service that navigates, runs scripted sequences and returns screenshots, so that agent is a caller rather than a build; the strict separation the memo describes is the most important design decision in it and should be defended rather than optimised away, because a persona agent able to read the page source would understand the page better than any human could, so the screenshot boundary is the instrument and not a limitation of it; two methodological errors would silently void the whole exercise, the first being that the same model must not both generate the page and react to it, since a page invented per turn is written to be reacted to and the run then measures self-consistency rather than a design, and the second being that a patience budget generated by the same model that generates the confusion will always cohere with it, so patience has to be set exogenously as a fixed number of screens or minutes decided before the run, which turns abandonment into a measured event rather than a story; the honest limit on the method is that synthetic readers reliably find defects and cannot report preferences, meaning they will catch a page that does not answer its own question, a term used before definition, a dead end or a missing next step, and cannot tell you which of two clear designs real people prefer, which maps exactly onto the previous brief's split, so synthetic readers should clear the levels and humans should judge the variants; and the memo proposes deriving personas from named real people, which is precisely what the second of two rules set on 9 August forbids, being raised now for the third time, so the exception has to be narrow and testable, with properties sampled from several people, an archetype recognisable as none of them, no names anywhere, and the line drawn at comprehension rather than decision, which is also the line the part-two service must be built around rather than promise, because the same machinery pointed at what would this executive approve is the tool that rule exists to prevent.** It is the fifteenth document of 20 August (cross-ref: the v0.33.57 executive view brief, the v0.33.61 levels and variants brief, the v0.33.61 user section brief, the v0.33.58 hub specification brief, and the v0.33.59 comparison pages brief). New contributions: **the rendering half located as an existing service, the screenshot boundary named as the instrument, the fixed-artefact requirement, the exogenous patience budget, defects separated from preferences, the third raising of the individual-simulation rule with a narrow testable exception, and comprehension against decision as the boundary the service needs in its design.**

## The Rendering Half Already Exists

The memo describes one agent's job in detail. The project lead: **"you can literally run playwright and take a screenshot."**

**That agent is a caller rather than a build.** This estate already has a browser-automation service reachable over HTTP that navigates to a page, runs scripted multi-step sequences, returns screenshots and rendered documents, extracts page content, probes one page from several angles, and holds a stateful session across turns. Everything the rendering agent needs is on that surface.

So the work in part one is not the mechanism. **It is the protocol between the two agents and the discipline of the run**, which is the rest of this brief.

One consequence worth taking now rather than later: because that service can also extract page text and structure, **the rendering agent must be configured not to pass it on.** That is the subject of the next section and it is a configuration decision rather than a capability gap.

## The Screenshot Boundary Is The Instrument, Not A Limitation

The memo specifies the separation almost in passing. The project lead: **"the second agent literally just processes screenshots and then gives the answers. And if he needs to move the mouse to click something, he should say, click on here, click there."**

**That constraint is the entire validity of the exercise, and somebody will try to improve it away.** The reasoning is short. A persona agent given the page's text or structure reads it perfectly: no scanning, no missed heading, no ambiguity about what is a button, no visual hierarchy to misjudge. **It would understand the page better than any human being could**, and every finding would be optimistic.

So the rule to write down before the first run:

> **The persona agent receives pixels and nothing else.** No page text, no structure, no source, no accessibility tree, no filenames, no knowledge of what the page was trying to achieve.

Two smaller consequences follow and both are easy to breach accidentally.

**The persona agent must not be told what the page is for.** A reader arriving at a page does not know its purpose. If the persona is briefed with the intent, it will find the intent, and the most common real failure, which is a page whose purpose is not apparent, becomes undetectable.

**And the click instruction should be spatial rather than semantic.** Click the thing at the top right, not click the assessment button, because naming the element proves the persona already recognised it and skips the step where a real person does not.

## The Page Under Test Must Be A Fixed Artefact

The memo offers three modes for the rendering agent and one of them voids the exercise. The project lead: **"there's a mode or couple modes where the agent acting like the web server actually renders the website because you can literally run playwright and take a screenshot, or he runs a simulation, or even just creates a temp mockup and shares it."**

**The first is a test. The third is not.**

If the same model both invents the page and reacts to it, the run measures the model's self-consistency. A page generated in response to a persona is written, however unconsciously, to be reacted to by that persona, and the reaction will be favourable because both come from the same place.

| Mode | What it measures |
|---|---|
| Render a real page from real markup | **The design** |
| Render a fixed mockup authored in advance, in a file, unchanged during the run | **The design.** This is the before-it-is-built case and it is legitimate |
| Generate a mockup during the run | **The model agreeing with itself** |

So the requirement is not that the page be built. **It is that the page be fixed, authored before the run, and identical across every persona in the round.** A hand-written HTML mockup committed to a file satisfies this completely, which is what makes the memo's best idea, testing before coding, available.

## The Patience Budget Has To Come From Outside The Model

The memo's most valuable addition, and it needs one correction to work. The project lead: **"it's very important that this also has a sort of a stop moment where we should kind of measure the emotional state and the status of understanding of patience that the person has. So the point here is not to create an agent that has infinite patience."**

**Modelling patience is right. Generating it is circular.** If the same model produces both the confusion and the reaction to it, the two will cohere: a persona that has been rendered confused will also report losing patience, not because impatience was measured but because that is the consistent story. What comes out is a narrative, and it will be a plausible one, which is worse than an obviously wrong one.

**So the budget is set from outside, in advance, and spent by mechanical rule:**

| Property | Set how | Spent how |
|---|---|---|
| Screens | A fixed number, per persona, before the run | One per screenshot delivered |
| Time | A fixed number of minutes | Estimated per screen by a fixed reading rate, not by the model |
| Clicks | A fixed number | One per instruction |
| **Comprehension** | Not budgeted | **Asked at each stop, and recorded** |

Then abandonment is an event rather than an opinion: **the budget ran out before the persona could state what the page told it.** That is a measurement, it is comparable across variants, and no part of it is generated.

And the comprehension question at each stop should be fixed wording across every persona and every variant, for the same reason. **Ask what would you do now, and what did that page tell you**, in those words, every time. A varying question produces varying answers that look like findings.

## Synthetic Readers Find Defects, Not Preferences

The honest limit on the whole method, and stating it is what makes the published output credible.

The memo frames the output as synthetic traffic that shows how good the site is. The project lead: **"fundamentally, what we are doing here is creating synthetic traffic. We're creating synthetic users that will allow us to understand, you know, how good the website is."**

**A synthetic reader is a reliable detector of defects and an unreliable reporter of preferences.**

| What it will find | Why it can |
|---|---|
| A page that does not answer the question it poses | The failure is internal to the page |
| A term used before it is defined | Detectable from the sequence alone |
| A dead end, or a step with no next action | Structural |
| Two pages that contradict each other | Comparison, not taste |
| A screen whose purpose is not apparent | Provided the persona was not briefed |

| What it will not find | Why not |
|---|---|
| Which of two clear designs people prefer | Its preferences are the model's, not a population's |
| Whether a tone lands as confident or arrogant | Same |
| Whether the result feels alarming enough to act on | The previous brief's outcome measure, unavailable here |

That maps exactly onto the split made in the previous brief. **Synthetic readers clear the levels, which are about completeness. Humans judge the variants, which are about rendering.** That is also the cheaper order: defects are removed for free before any human time is spent, and the small number of real users you can actually recruit are spent on the only question they can answer.

## Rule One Applies Hardest Precisely Because You Want To Publish

Two rules were set on 9 August and described as non-negotiable rather than defaults. The first, verbatim:

> **A simulated acceptance must never be confusable with a real one.** Different storage, different rendering, and an indelible marker that survives export.

The memo wants the runs exported as documents and published on the site as assets, including historical ones. **Export is exactly the moment a marker is lost**, and a published transcript of a persona saying it understood the page is one screenshot away from being quoted as a user saying it.

So the marker has to survive the export format rather than live in the interface: in the filename, in the document header, in the running header of every page, and beside each quoted reaction. **A synthetic run should be difficult to quote misleadingly even by somebody trying.** That is a higher bar than labelling it once at the top, and it is the right one for material intended to be shared.

## Rule Two Is Raised For The Third Time, And The Memo Proposes What It Forbids

The second rule, verbatim:

> **Simulate the role, not the named individual.** Modelling how a chief financial officer generally responds is a training aid. Modelling how a specific named person will respond, and tuning a presentation against that model, is building a tool for routing around a colleague. The line is between preparing for a conversation and pre-empting a person.

The memo proposes the opposite. The project lead: **"I'll basically provide, for example, a couple of individuals that I know, and then you can create a persona from that basically real life person, and then we should basically simulate what that person is."**

**This is the third raising of the same rule.** It was set on 9 August, and that brief records that the corpus had already flagged it once before, when the field demo work concluded that a compliance agent is persona simulation rather than a reviewer. A rule that keeps being proposed against is either wrong or badly explained, so it is worth saying what it is actually protecting.

**The rule's harm is asymmetric power, and this use does not have it.** Rule two exists to stop somebody modelling a gatekeeper in order to get past them. A usability persona is a reader you are trying to serve. Different relation, different harm, and the rule should not be applied mechanically to a case it was not written for.

**But the memo's proposal reaches further than usability needs**, and the phrase that matters is tuning a presentation against that model, which is exactly what a variant programme does. So the exception has to be narrow and it has to be testable:

| Allowed | Not allowed |
|---|---|
| Sampling properties from several real people | Building a portrait of one |
| Composing an archetype that maps to no individual | A persona whose source could recognise themselves |
| Recording the property list | Recording who it came from |
| Publishing the archetype | Naming, identifying, or describing anybody |

And the test, which is checkable by somebody other than the author:

> **If the person it came from, or a colleague of theirs, would recognise them in it, it is a portrait rather than an archetype.**

There is also a practical argument that points the same way, and it is worth more than the rule. **A portrait of one person is a fixed point and generalises to nobody**, so a finding from it is about that person. An archetype composed from several is parameterised, and the memo has already written the parameters: technical level, time available, motivation, prior interest. **The property list is the archetype. The named individual is only where the properties were sampled from**, and confusing the two loses the ability to vary them independently, which is the whole point of having personas at all.

## Comprehension, Not Decision, Is The Line The Service Needs In Its Design

The second half of the memo, and it inherits everything above. The project lead: **"here's a particular user. Here's a particular persona, particular type of exec, particular type of user. What would would they get from here? What would they understand? Where would it add value? Where is the confusion?"**

Every question in that list is about comprehension, and every one is legitimate. **The same machinery, given one different question, becomes the tool rule two exists to prevent.**

| The question asked | What the product is |
|---|---|
| What would this reader understand? | **A usability instrument** |
| Where would they be confused? | **A usability instrument** |
| What would this reader value? | Marketing research, and defensible |
| **What would this person approve?** | **The thing the rule forbids** |
| **How should I present this so they say yes?** | **Worse, and it is the natural next feature request** |

**The difference is the question, not the technology**, which means it cannot be handled in terms of service or in guidance. It has to be a property of the product: the interface asks about understanding and confusion, it does not accept a question about approval, and a simulated verdict is never rendered as a decision.

That also connects to the first rule. **A simulated acceptance must never be confusable with a real one**, and a service whose output is a persona's verdict on a document is producing exactly the object that rule is about.

And the thing that would make it a product rather than a plausible opinion generator is the one the memo already names. The project lead: **"we'll calibrate this with the real ones, a couple of rinse and repeat."** That is right, and it should be the headline rather than the closing thought: **publish the calibration record, including the misses.** How often did the synthetic verdict match the real reader's, on what sample, on what date. Without it the output is unfalsifiable, and this estate's own comparison discipline says a dated re-runnable test is evidence while an assertion from a participant is marketing.

## Testing Before Building Is The Strongest Part And Is Nearly Thrown Away

The memo puts this near the end and it should be the reason to do the programme at all. The project lead: **"this is even more cool when we do some of these before it has been built... so we can learn from it before we actually do a lot of coding."**

**That is the highest-value use and it is the cheapest**, because a fixed hand-written mockup satisfies the fixed-artefact requirement completely, and a defect found before implementation costs a file edit rather than a refactor.

It also matches a position this corpus already took. On 14 August the argument for publishing a design before building it was that a design published before it is built is the strongest available demonstration of the working method the site already documents. **A published tabletop against an unbuilt page is that same move one layer deeper**, and it is unusual enough to be worth doing for its own sake.

## Publishing The Runs Means Publishing The Failures

The memo wants everything on the site. The project lead: **"everything here should be published on a website because these are all valid assets to have, even valid historical assets to have."**

Agreed, with the discipline the site already carries. **A participant publishing synthetic users' verdicts on the participant's own pages is marketing unless the failures are published too**, and the comparison rules from 16 August apply unchanged: state who produced it, publish the method before the findings, and publish where it loses.

Concretely, the runs worth publishing most are the ones where the persona ran out of budget without understanding the page, because those are the ones that changed the design. A section of successful runs is an advertisement. A section that shows three abandoned runs, the change made, and the run that then succeeded, is the working method the site exists to demonstrate.

## What This Does Not Try To Be

- **Not a new mechanism.** The rendering half is an existing service in this estate and the persona machinery was specified on 9 August.
- **Not a replacement for real users.** Synthetic readers find defects, and preferences need people.
- **Not an exemption from rule two.** A narrow, testable exception is proposed and the rule stands.
- **Not a measurement of emotion.** The budget is set outside the model; only its exhaustion is measured.
- **Not a service specification.** Part two is scoped by the question it may accept, and the rest is not settled here.

## Honest Tensions

| Tension | Note |
|---------|------|
| The screenshot boundary | It is what makes the exercise valid and it makes the persona agent worse at the task than the tooling allows, which will feel like leaving capability unused |
| Exogenous patience budgets | They make abandonment measurable and they are arbitrary numbers somebody has to choose, and the choice will drive the results |
| Archetypes rather than portraits | It honours the rule and it discards the specificity that made real people attractive as sources |
| Synthetic readers as the first pass | It is cheap and correct and it will be quoted as though it were user research by somebody who was not in the room |
| Publishing the runs | It demonstrates the method and it hands competitors a catalogue of what confused people about your own pages |
| A service one question from a banned tool | The boundary is real and it depends on refusing a feature customers will ask for by name |

## Open Questions

| Question | Notes |
|----------|-------|
| Who sets the budgets, and from what? | They drive every result and there is no basis for the first set beyond judgement |
| How is the marker made to survive export? | Rule one turns on this and export formats are where markers are lost |
| What is the fixed comprehension question? | It must not vary, and nobody has written it |
| How many archetypes before the property list is stable? | Composing from too few is a portrait with extra steps |
| Where does the calibration record live? | Without it part two is unfalsifiable, and it is the least interesting thing to build |
| Does the persona agent get memory across runs? | A returning reader is a different test from a first-time one, and conflating them would be easy |

## Relationship To Previous Briefs

| Date | Document | Relationship |
|---|---|---|
| 9 Aug | `v0.33.57__dev-brief__sg-send-executive-view-is-compiled-framing-is-governed-grc-connects-simulation-finds-gaps.md` | The two simulation rules, quoted here, one of which this memo proposes against for the third time |
| 20 Aug | `v0.33.61__dev-brief__levels-and-variants-are-two-axes-persona-generator-specified-in-august-and-the-worked-instance-is-this-session.md` | Levels against variants, which decides that synthetic readers clear the first and humans judge the second |
| 20 Aug | `v0.33.61__dev-brief__user-section-is-a-conformance-test-store-the-choices-not-the-answers-high-threat-without-efficacy-produces-denial.md` | The outcome measure a synthetic reader cannot produce, and the pages these runs would test |
| 14 Aug | `v0.33.58__dev-brief__sgit-specification-for-the-hub-briefing-pack.md` | Publishing a design before it is built as a demonstration of the method, which this extends to testing it |
| 16 Aug | `v0.33.59__strategy-brief__sgit-comparison-pages-as-reproducible-tests-privileges-is-the-missing-column.md` | Dated, re-runnable, method before findings, which the calibration record must follow |
| 2 Aug | `v0.33.55__arch-brief__sg-send-end-to-end-worked-example-article-26-5-creditworthiness-agent-fact-to-board.md` | Unanswered questions as the primary output, which is what an abandoned run produces |

---

## Key Claims

| # | Claim |
|---|-------|
| 1 | The rendering agent is a caller of an existing browser-automation service rather than something to build |
| 2 | The persona agent must receive pixels only, because page text would let it read better than any human |
| 3 | It must also not be told the page's purpose, or the commonest real failure becomes undetectable |
| 4 | A page generated during the run measures the model agreeing with itself, so the artefact must be fixed in advance |
| 5 | A hand-written mockup satisfies that, which is what makes testing before building available |
| 6 | Patience generated by the same model that generates the confusion will always cohere with it |
| 7 | So the budget is set exogenously and abandonment becomes a measured event rather than a narrative |
| 8 | Synthetic readers reliably find defects and cannot report preferences |
| 9 | So they clear the levels and humans judge the variants, which is also the cheaper order |
| 10 | Rule one applies hardest here because the runs are to be published, and export is where markers are lost |
| 11 | Rule two is raised for the third time, and the exception must be testable: recognisable by its source means portrait, not archetype |
| 12 | The part two service is separated from the tool that rule forbids by the question it accepts, so the boundary belongs in the product |

---

## Sources

- The two simulation rules set on 9 August, that a simulated acceptance must never be confusable with a real one and requires different storage, different rendering and an indelible marker that survives export, and that simulation should model the role rather than the named individual because modelling how a specific person will respond and tuning a presentation against that model is building a tool for routing around a colleague, together with the note that the corpus had already flagged the same line once before: the project repository, cloned and searched on 20 August 2026, with the document named in the relationship table above
- The browser-automation service available in this estate, which navigates to a page, runs scripted multi-step sequences, returns screenshots and rendered documents, extracts page content and structure, probes one page from several angles and holds a stateful session: the skill documentation available to this session on 20 August 2026, held with the project record

---

This document is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).
