Methodology exhibit · reference page
How this tool decided.
Every classification the Diagnostic produces is computed by the rules on this page — published in full, versioned, and running as deterministic code you can read in your browser's devtools. As shipped today (v1), no model appears anywhere in this tool; even the narrative memo is assembled from templates. The architecture reserves exactly one place a model could ever sit — below the boundary — and if we fill it, this page changes first, as a published event. This page is the complete rubric; nothing about the verdict is withheld.
Why the verdict is code, not a model
Before building this tool, we assessed it with its own rubric. The obvious version — hand your answers to a language model, let it write an assessment — classifies badly: rendering judgment on a stranger's workflow is judgment-dense work (W4), checking any single output requires an expert read (V3), and a confidently wrong verdict published under our name is an irreversible, reputation-scale action. Run those classes through the autonomy lookup below and the answer comes back: that system doesn't get to operate autonomously in public. So the load-bearing layer is code — and in v1 we went further than the rubric demanded: even the memo is assembled deterministically from templates keyed to your classifications. The one place a model could ever sit is below the boundary, narrating results it cannot change. The tool is our own advice, taken.
The full account, including what this decision cost:Field note 01 — The Diagnostic had to pass its own test.
Fig. 01 — the boundary
Structured intake (16 questions)
│
▼
┌───────────────────────────────┐ deterministic, versioned code
│ SCORING ENGINE │ same answers → same result, always
│ legibility · exceptions · │ auditable in devtools — deliberately
│ verification · risk · W-class│
└───────────────┬───────────────┘
│ computed JSON (locked)
▼
┌───────────────────────────────┐ deterministic templates (v1)
│ NARRATIVE MEMO │ assembled from computed results only
│ (no model in v1) │ ── reserved slot: if a model ever
└───────────────────────────────┘ writes this memo, it sits HERE,
below the boundary — announced,
never slipped inThe memo is convenience; the verdict is code — and in v1, so is the memo. Free-text answers are length-capped, score-inert, and treated as untrusted input: they may be echoed in the memo's framing, never in the verdict.
What each question maps to
| Section | Question | Feeds |
|---|---|---|
| A · Workflow | Description, frequency, who performs it | Narrative only — never scored |
| B · Legibility | Documentation quality, replaceability, deviation frequency | Legibility grade |
| C · Exceptions | Share needing special handling, known vs. new types | Exception mass |
| D · Verification | How you'd catch a wrong instance; check-vs-produce time | Verification class (V0–V4) |
| E · Consequence | Reversibility; who's affected | Risk class (I–IV) |
| F · Judgment & regime | Expert disagreement frequency; regulatory context | W-class, oversight floor |
| G · Data | Where information lives | K-class flag, integration note |
Taxonomy definitions live in the Field Guide's enterprise taxonomies index. The tables below are the Diagnostic's actual decision rules, and where the two differ, these govern.
Fig. 02 — risk class: reversibility × blast radius
| Internal record | One client | Many clients | Reg / legal / financial | |
|---|---|---|---|---|
| Fully reversible | I | I | II | II |
| Reversible w/ effort | I | II | II | III |
| Partially irreversible | II | III | III | IV |
| Irreversible | III | III | IV | IV |
Your E1 answer selects the row; E2 selects the column. The result page shows you exactly which cell you landed in, and why.
Fig. 03 — max autonomy tier: risk class × verification class
| V0–V1 | V2 | V3 | V4 | |
|---|---|---|---|---|
| Risk I | A4 · Monitored | A3 · Sampled | A2 · Exception-gated | A1 · Gated |
| Risk II | A3 · Sampled | A3 · Sampled | A1 · Gated | A1 · Gated |
| Risk III | A2 · Exception-gated | A2 · Exception-gated | A1 · Gated | A0 · Advisory |
| Risk IV | A1 · Gated | A1 · Gated | A0 · Advisory | A0 · Advisory |
The lookup gives the base tier. Three caps then apply, in order:
- Low legibility caps at A2 — and the recommendation leads with making the process legible before automating it. You cannot safely delegate what you cannot describe.
- Regulated or fiduciary context caps at A2 (unless Risk I) and sets the oversight floor to exception-gated with documentary provenance.
- Judgment-dense work (W4) reframes the recommendation as agent prepares, human decides — whatever the tier arithmetic says.
Rules that keep it honest
Inconsistencies are flagged, not resolved
If your answers contradict each other — errors are "caught by a simple check" but checking is "harder than doing" — the engine bumps the verification class conservatively and tells you which answer it privileged. Nothing is reconciled silently.
Some classes are out of reach, on purpose
Sixteen questions cannot honestly assign the extreme complexity classes (W1, W5). The tool never claims them.
"Don’t automate this yet" is a real output
High risk + expensive verification + low legibility produces exactly that verdict. A diagnostic that flatters every workflow isn’t a diagnostic.
Your session is yours
Answers aren’t stored unless you request the emailed PDF. You can download your full result — inputs, rubric version, outputs — as JSON. That’s decision provenance, demonstrated.
No model, and we’ll tell you if that changes
In v1 the memo is templated, not model-written. If a model-drafted memo ever ships, it ships below the boundary, and this page is updated first — as a published event.
The rubric is versioned
Every result carries its rubric version (currently v1). When these tables change, that’s a published editorial event, not a silent update.
The limitation, stated plainly
Self-report measures the documented process. The paid assessment measures the practiced one. The gap between them is legibility debt — and it's usually where the answer changes.
What's simulated vs. production
Nothing here is simulated — the diagnostic genuinely computes what it claims from what it's given, using the tables above, with no model in the loop. What it cannot do is observe practice. That's the paid assessment's entire job, and the reason the two products exist as a pair rather than one tool with a price toggle.