← WORKFLOW DIAGNOSTIC

Methodology exhibit · reference page

How this tool decided.

Every classification the Diagnostic produces is computed by the rules on this page — published in full, versioned, and running as deterministic code you can read in your browser's devtools. As shipped today (v1), no model appears anywhere in this tool; even the narrative memo is assembled from templates. The architecture reserves exactly one place a model could ever sit — below the boundary — and if we fill it, this page changes first, as a published event. This page is the complete rubric; nothing about the verdict is withheld.

Why the verdict is code, not a model

Before building this tool, we assessed it with its own rubric. The obvious version — hand your answers to a language model, let it write an assessment — classifies badly: rendering judgment on a stranger's workflow is judgment-dense work (W4), checking any single output requires an expert read (V3), and a confidently wrong verdict published under our name is an irreversible, reputation-scale action. Run those classes through the autonomy lookup below and the answer comes back: that system doesn't get to operate autonomously in public. So the load-bearing layer is code — and in v1 we went further than the rubric demanded: even the memo is assembled deterministically from templates keyed to your classifications. The one place a model could ever sit is below the boundary, narrating results it cannot change. The tool is our own advice, taken.

The full account, including what this decision cost:Field note 01 — The Diagnostic had to pass its own test.

Fig. 01 — the boundary

Structured intake (16 questions)
        │
        ▼
┌───────────────────────────────┐   deterministic, versioned code
│  SCORING ENGINE               │   same answers → same result, always
│  legibility · exceptions ·    │   auditable in devtools — deliberately
│  verification · risk · W-class│
└───────────────┬───────────────┘
                │  computed JSON (locked)
                ▼
┌───────────────────────────────┐   deterministic templates (v1)
│  NARRATIVE MEMO               │   assembled from computed results only
│  (no model in v1)             │   ── reserved slot: if a model ever
└───────────────────────────────┘      writes this memo, it sits HERE,
                                        below the boundary — announced,
                                        never slipped in

The memo is convenience; the verdict is code — and in v1, so is the memo. Free-text answers are length-capped, score-inert, and treated as untrusted input: they may be echoed in the memo's framing, never in the verdict.

What each question maps to

SectionQuestionFeeds
A · WorkflowDescription, frequency, who performs itNarrative only — never scored
B · LegibilityDocumentation quality, replaceability, deviation frequencyLegibility grade
C · ExceptionsShare needing special handling, known vs. new typesException mass
D · VerificationHow you'd catch a wrong instance; check-vs-produce timeVerification class (V0–V4)
E · ConsequenceReversibility; who's affectedRisk class (I–IV)
F · Judgment & regimeExpert disagreement frequency; regulatory contextW-class, oversight floor
G · DataWhere information livesK-class flag, integration note

Taxonomy definitions live in the Field Guide's enterprise taxonomies index. The tables below are the Diagnostic's actual decision rules, and where the two differ, these govern.

Fig. 02 — risk class: reversibility × blast radius

Internal recordOne clientMany clientsReg / legal / financial
Fully reversibleIIIIII
Reversible w/ effortIIIIIIII
Partially irreversibleIIIIIIIIIV
IrreversibleIIIIIIIVIV

Your E1 answer selects the row; E2 selects the column. The result page shows you exactly which cell you landed in, and why.

Fig. 03 — max autonomy tier: risk class × verification class

V0–V1V2V3V4
Risk IA4 · MonitoredA3 · SampledA2 · Exception-gatedA1 · Gated
Risk IIA3 · SampledA3 · SampledA1 · GatedA1 · Gated
Risk IIIA2 · Exception-gatedA2 · Exception-gatedA1 · GatedA0 · Advisory
Risk IVA1 · GatedA1 · GatedA0 · AdvisoryA0 · Advisory

The lookup gives the base tier. Three caps then apply, in order:

  1. Low legibility caps at A2 — and the recommendation leads with making the process legible before automating it. You cannot safely delegate what you cannot describe.
  2. Regulated or fiduciary context caps at A2 (unless Risk I) and sets the oversight floor to exception-gated with documentary provenance.
  3. Judgment-dense work (W4) reframes the recommendation as agent prepares, human decides — whatever the tier arithmetic says.

Rules that keep it honest

Inconsistencies are flagged, not resolved

If your answers contradict each other — errors are "caught by a simple check" but checking is "harder than doing" — the engine bumps the verification class conservatively and tells you which answer it privileged. Nothing is reconciled silently.

Some classes are out of reach, on purpose

Sixteen questions cannot honestly assign the extreme complexity classes (W1, W5). The tool never claims them.

"Don’t automate this yet" is a real output

High risk + expensive verification + low legibility produces exactly that verdict. A diagnostic that flatters every workflow isn’t a diagnostic.

Your session is yours

Answers aren’t stored unless you request the emailed PDF. You can download your full result — inputs, rubric version, outputs — as JSON. That’s decision provenance, demonstrated.

No model, and we’ll tell you if that changes

In v1 the memo is templated, not model-written. If a model-drafted memo ever ships, it ships below the boundary, and this page is updated first — as a published event.

The rubric is versioned

Every result carries its rubric version (currently v1). When these tables change, that’s a published editorial event, not a silent update.

The limitation, stated plainly

Self-report measures the documented process. The paid assessment measures the practiced one. The gap between them is legibility debt — and it's usually where the answer changes.

What's simulated vs. production

Nothing here is simulated — the diagnostic genuinely computes what it claims from what it's given, using the tables above, with no model in the loop. What it cannot do is observe practice. That's the paid assessment's entire job, and the reason the two products exist as a pair rather than one tool with a price toggle.