@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/CHANGELOG.md +351 -0
  2. package/LICENSE +21 -0
  3. package/README.md +163 -2
  4. package/VERSION +1 -0
  5. package/adapters/_template/adapter.tsv +16 -0
  6. package/adapters/_template/detect.sh +10 -0
  7. package/adapters/_template/emit.sh +5 -0
  8. package/adapters/_template/verify.sh +4 -0
  9. package/adapters/claude/adapter.tsv +8 -0
  10. package/adapters/claude/detect.sh +8 -0
  11. package/adapters/claude/verify.sh +47 -0
  12. package/adapters/codex/adapter.tsv +12 -0
  13. package/adapters/codex/detect.sh +9 -0
  14. package/adapters/codex/verify.sh +45 -0
  15. package/adapters/copilot/adapter.tsv +10 -0
  16. package/adapters/copilot/detect.sh +8 -0
  17. package/adapters/copilot/verify.sh +45 -0
  18. package/adapters/cursor/adapter.tsv +11 -0
  19. package/adapters/cursor/detect.sh +10 -0
  20. package/adapters/cursor/verify.sh +45 -0
  21. package/adapters/gemini/adapter.tsv +15 -0
  22. package/adapters/gemini/detect.sh +11 -0
  23. package/adapters/gemini/verify.sh +49 -0
  24. package/adapters/hermes/adapter.tsv +9 -0
  25. package/adapters/hermes/detect.sh +8 -0
  26. package/adapters/hermes/verify.sh +27 -0
  27. package/adapters/opencode/adapter.tsv +14 -0
  28. package/adapters/opencode/detect.sh +9 -0
  29. package/adapters/opencode/verify.sh +45 -0
  30. package/automations/README.md +53 -0
  31. package/automations/bugreporter-intake.sh +145 -0
  32. package/automations/drift-audit.sh +139 -0
  33. package/automations/report.schema.tsv +10 -0
  34. package/bans/README.md +82 -0
  35. package/bans/grep-ban.sh +84 -0
  36. package/bans/layer-check.sh +57 -0
  37. package/bin/goblin +119 -0
  38. package/bin/goblin-audit +145 -0
  39. package/bin/goblin-bans +178 -0
  40. package/bin/goblin-doctor +233 -0
  41. package/bin/goblin-emit +482 -0
  42. package/bin/goblin-init +519 -0
  43. package/bin/goblin-install +720 -0
  44. package/bin/goblin-lib.sh +289 -0
  45. package/bin/goblin-model +105 -0
  46. package/bin/goblin-upgrade +572 -0
  47. package/bin/goblin-verify +2798 -0
  48. package/bin/goblin.js +48 -0
  49. package/docs/ADOPTION.md +168 -0
  50. package/docs/CI.md +187 -0
  51. package/docs/CONTRACTS.md +197 -0
  52. package/docs/DESIGN.md +92 -0
  53. package/docs/ENFORCEMENT.md +225 -0
  54. package/docs/FLOWS.md +164 -0
  55. package/docs/GUARDRAILS.md +126 -0
  56. package/docs/GUIDE.md +610 -0
  57. package/docs/INTEGRATION.md +92 -0
  58. package/docs/LIMITS.md +591 -0
  59. package/docs/LOOP.md +165 -0
  60. package/docs/RE-PLAYBOOK.md +183 -0
  61. package/docs/RISKS.md +70 -0
  62. package/docs/ROLES.md +105 -0
  63. package/manifest/bans.tsv +9 -0
  64. package/manifest/classes.tsv +61 -0
  65. package/manifest/enforcement.tsv +88 -0
  66. package/manifest/glossary.tsv +25 -0
  67. package/manifest/playbooks.tsv +16 -0
  68. package/package.json +36 -4
  69. package/presets/A-shipped-software.yaml +48 -0
  70. package/presets/B-service-config.yaml +40 -0
  71. package/presets/C-game.yaml +38 -0
  72. package/presets/D-knowledge.yaml +41 -0
  73. package/presets/E-fleet-config.yaml +42 -0
  74. package/presets/F-electron.yaml +67 -0
  75. package/roles.yaml +54 -0
  76. package/skills/goblin-bootstrap/SKILL.md +51 -0
  77. package/skills/goblin-bugfix/SKILL.md +26 -0
  78. package/skills/goblin-bugreporter/SKILL.md +52 -0
  79. package/skills/goblin-drift-audit/SKILL.md +43 -0
  80. package/skills/goblin-eval/SKILL.md +68 -0
  81. package/skills/goblin-feature/SKILL.md +26 -0
  82. package/skills/goblin-feature-map/SKILL.md +140 -0
  83. package/skills/goblin-handoff/SKILL.md +28 -0
  84. package/skills/goblin-investigation/SKILL.md +26 -0
  85. package/skills/goblin-judge/SKILL.md +74 -0
  86. package/skills/goblin-loop/SKILL.md +88 -0
  87. package/skills/goblin-mode/SKILL.md +70 -0
  88. package/skills/goblin-overnight/SKILL.md +42 -0
  89. package/skills/goblin-pr-gate/SKILL.md +42 -0
  90. package/skills/goblin-re-mobile/SKILL.md +51 -0
  91. package/skills/goblin-refactor/SKILL.md +23 -0
  92. package/skills/goblin-sweep/SKILL.md +23 -0
  93. package/skills/goblin-tdd-repro/SKILL.md +27 -0
  94. package/skills/goblin-verify-author/SKILL.md +50 -0
  95. package/skills/practice/SKILL.md +37 -0
  96. package/templates/AGENTS.md.tmpl +23 -0
  97. package/templates/HANDOFF.md.tmpl +43 -0
  98. package/templates/SPEC.md.tmpl +34 -0
  99. package/templates/audit-waiver.tsv.tmpl +10 -0
  100. package/templates/boundary-waivers.tmpl +8 -0
  101. package/templates/checks/assert.mjs.tmpl +60 -0
  102. package/templates/checks/gate.sh.tmpl +29 -0
  103. package/templates/ci/goblin-gate.yml.tmpl +46 -0
  104. package/templates/goblin.yaml.tmpl +138 -0
  105. package/templates/install-hooks.allowlist.tmpl +9 -0
  106. package/templates/loop/decisions.tsv.tmpl +1 -0
  107. package/templates/loop/predicate.tmpl +16 -0
  108. package/templates/report.yaml.tmpl +16 -0
package/docs/LIMITS.md ADDED
@@ -0,0 +1,591 @@
1
+ # Limits — where this is weaker, and what is unproven
2
+
3
+ Honest accounting. Nothing here is a hedge for a defect that could be fixed; each is either a
4
+ deliberate trade or an unfilled gap.
5
+
6
+ ## Weaker than the Cursor harness it learns from
7
+
8
+ 1. **No live-drive lane is shipped.** A harness that launches and drives the real application is
9
+ what several of the source harness's playbooks depend on, and it is inert without a plugin
10
+ that is not shipped. `P6` can *author* a project-local driver; goblin-stack cannot ship one,
11
+ because that is per-project work. **The "live lane is the floor" rule is stated and
12
+ unfilled.**
13
+ 2. **No cloud agents.** There is no per-agent computer here, so nothing that needs a machine of
14
+ its own.
15
+ 3. **No agent graph and no verdict-ledger daemon.** The board replaces both at a lower
16
+ resolution: `parents` expresses ordering, not data flow, and sibling cards cannot see each
17
+ other.
18
+ 4. **Fifteen playbooks against twenty-three.** The cuts in `docs/FLOWS.md` are deliberate and each
19
+ is argued, but real coverage is lost: performance hillclimbing, pixel parity, trace
20
+ forensics, stack landing, worktree hygiene.
21
+ 5. **No swarm or arena fan-out.** The read-versus-write axis says that is correct for this work
22
+ mix; it is still a capability the other harness has and this does not.
23
+
24
+ ## Weaker than the status quo it is meant to improve
25
+
26
+ 6. **The biggest measured defect is out of reach.** The orchestrator's routing text says a bare
27
+ subagent spawn reaches the specialist profiles; it does not. That is the highest-cost defect
28
+ on the box, and goblin-stack cannot fix it from inside a project repo. It ships corrected
29
+ text in `docs/INTEGRATION.md` and lints its own artifacts (`MD-03`); the fleet-side edit is
30
+ escalated.
31
+ 7. **The staleness of a fleet-config repo is detectable and not fixable here** — the E-class gate
32
+ notices it, and the underlying job bug belongs to another repository.
33
+ 8. **A new surface to maintain.** Per-repo vendored `.goblin/` plus `.hermes/skills/` means
34
+ upgrade debt in every adopted repo, plus one more command pair to learn. The counter is that
35
+ the alternative — profile copies — already failed.
36
+ 9. **Ten rows are labelled `advisory`, and nine of them are prose with no check at all:**
37
+ `HP-04`, `HS-03`, `CM-02`, `MD-03`, `PG-04`, `DOC-01`, `DOC-02`, `SC-09` and `JG-03` (G2, the
38
+ judge lane that has never returned a non-`done` verdict). One (`MD-02`) is
39
+ advisory-labelled but still reports its state as `ADV`. Each is counted and capped, but a
40
+ counted rule is still not an enforced one, and **the cap is a policy, not a proof.**
41
+ 10. **It does not reduce the profile-skill surface.** It only stops that surface from growing
42
+ with project procedures; the existing divergence is a separate cleanup.
43
+ 11. **`HS-02` cannot prove anything until a round lands.** With no pinned pre-change commit it is
44
+ skipped with a reason. That is honest, and it means a brand-new install has **no** REPLAY
45
+ evidence at all. The shipped `checks/assert.mjs` is a scaffold: it asserts something true by
46
+ construction, so `HS-02` correctly reports it as a check that proves nothing until it is
47
+ replaced with real probes.
48
+ 12. **`HP-03` proves a date exists, not that a number is fresh.** A HANDOFF can carry yesterday's
49
+ number with today's date and pass. Since V1 the row is anchored on the gate names
50
+ `.goblin/goblin.yaml` DECLARES and skips the template's own example sentence, so a real gate
51
+ line can no longer lose its `measured <date>` in silence (G8-2 measured the old row passing
52
+ exactly that); but nothing re-measures the number, and a gate-bearing line that names no
53
+ declared gate and carries no gate-shaped keyword is still unseen.
54
+ 13. **`PG-04` is documentation, and `PG-05` is a text reading — re-declared at W4 rather than
55
+ called a gate.** `PG-04` stays `advisory` because the forge is unobservable from inside a
56
+ repo: a protected branch whose only admin is the person pushing protects nothing, so a
57
+ required check armed under the sole admin's own identity binds nobody — the second half of
58
+ that sentence is measured (Harvey is the sole admin of every repo he owns), and changing it
59
+ is a forge/account change, not a repository change. `PG-05` **used to be a heuristic**: it
60
+ counted "every step is guarded", so a single **job-level** `if:`, a job with no step, and the
61
+ measured real shape (one *unguarded* step deciding whether the guarded gate step runs) all
62
+ passed it while GitHub reported Success. It now refuses a conditional job, a conditional
63
+ step, a job that declares no step and a workflow with no `jobs:` — see `docs/CI.md` §1 for
64
+ why a skipped job is a green light. It is still a **text reading**: no YAML parser
65
+ (`docs/CONTRACTS.md` allows none), so a flow-style `jobs: {…}` mapping is refused rather than
66
+ parsed, a `#` inside a quoted string is read as a comment, and a conditional step that is
67
+ genuinely safe is indistinguishable from the trap. It cannot see branch protection, the
68
+ required-check list, or whether the job ever ran — `PG-04` is the row that says so. `PG-06`
69
+ is the other half and has its own limit: it proves the declared gate is **invoked** in a file
70
+ under `.github/workflows/`, never that the forge marks that job required, never that it is
71
+ the job the forge waits on, and never that the workflow can fail.
72
+ 14. **`CM-02` cannot be enforced.** A backtick lost to command substitution leaves no trace a
73
+ later check can read.
74
+
75
+ ## Unproven at the level that matters
76
+
77
+ 15. **The controlled evidence now exists, and its counter-finding is about this repo (corrected
78
+ 2026-09-25).** This entry used to say that **no controlled study of skill efficacy exists**, so
79
+ that nothing here claimed the installed skills change agent behaviour. That is out of date:
80
+ **SkillsBench 1.1** (benchmark published 2026-06-16) reports that **curated Skills raise the
81
+ mean task-resolution rate from 33.9% to 50.5% — +16.6 points, a 25.5% normalized gain — across
82
+ 87 tasks, 8 domains and 18 model–harness configurations**, with every one of the 18
83
+ configurations higher with Skills (configuration-level gains from +4.1 to +25.7 points).
84
+ **The counter-finding matters more here than the headline:** in the same benchmark's
85
+ **self-generated condition — the agent authors its own Skills before solving — all three tested
86
+ configurations landed BELOW their no-Skills baseline** (−8.1, −11.3 and −11.5 points), while
87
+ curated Skills stayed above it. **goblin-stack installs agent-authored skills**, so the lower
88
+ row of that result is a warning about its own output, not someone else's: a generated skill
89
+ accepted after a skim is a different proposition from a written one. It is why `P6` hands the
90
+ generated skill to `P12` before any `verified:` date advances (`docs/RISKS.md` K15).
91
+ **`P12` is still the mechanism to find out, and it has still never been run**: the record format
92
+ is now specified (`skills/goblin-eval/SKILL.md`) and **no row reads a lane** (#31).
93
+ Pin: `https://www.skillsbench.ai/blogs/skillsbench-1-1`, sha256 of the retrieved page
94
+ `d812bb7c2da702cc67556eb5f46b9e93faaca19d3545376544866e940126ef15`, fetched 2026-09-25; the
95
+ eleven-token ban and the judge procedure are argued in
96
+ `https://ai.engineer/talks/0vphxNt4wyk-don-t-ship-skills-without-evals`. **Not registered**:
97
+ the source registry is `manifest/sources.tsv` (G9's artifact) and that file does not exist in
98
+ this repo — measured, `git ls-files manifest/` names five files and none of them is
99
+ `sources.tsv` — so the registry row is composed in the W2 report for G9's patch instead of
100
+ being written into a file this card does not own.
101
+ 16. **Every number in the design is a file read or a third-party published number; none of it is
102
+ a measurement of this harness under load.** The harness was built and its verifier was shown
103
+ RED under a deliberate break; that is a statement about the mechanism, not about outcomes.
104
+ 17. **One spec deviation, recorded rather than hidden.** The design spec's literal HANDOFF check
105
+ is `grep -q "$(git rev-parse --short HEAD)" HANDOFF.md`, which **can never pass** — committing
106
+ the HANDOFF moves HEAD, so the file can only ever name an ancestor. `HP-05` is therefore
107
+ mechanised as *the HANDOFF names a commit that exists in this repo and is an ancestor of
108
+ HEAD*, which still catches the defect the rule exists for (an artifact that names no commit
109
+ at all). The deviation and its reason are in the row's own `if_not_why` column.
110
+ 18. **`.goblin/installed.json` is not signed, so nothing here proves it was not rewritten.** Every
111
+ drift check — `IN-02`, `SK-02`, and `HS-01`'s hash of the harness dir — reads its expected
112
+ hash out of that one file, and that file is the one file no check protects. Measured: append a
113
+ byte to `.goblin/bin/goblin-verify`, rewrite its recorded hash in `installed.json`, commit, and
114
+ the run is **fully GREEN** (`43 passed, 0 failed`, exit 0). One edit defeats three rows at
115
+ once, and it is the cheapest way to fake a green run. Doing better needs an anchor the target
116
+ cannot edit — a signature, or a hash held outside the repo — and goblin-stack has no such
117
+ trust root: the source checkout is not guaranteed to exist at verify time, and any value
118
+ stored in the tree is editable by the same hand. It is therefore **recorded here and printed
119
+ in the "cannot see" footer on every run**, not claimed away.
120
+ 19. **No automation has ever run here.** Cost per run, the respawn guards under a nightly
121
+ producer, and whether `researcher` is the right reporter profile are all unmeasured; the
122
+ first watched run of A-02 is what produces those numbers. The producer's own ceiling is a
123
+ **run count**, not a dollar figure — no config key holds the spend cap, so no automation can
124
+ read it, and none pretends to.
125
+ 20. **The dedup key is a dedup, not a mutex.** The board's lookup runs before the write
126
+ transaction, so a concurrent create can insert twice and the next lookup stabilises on the
127
+ newest. The key stops a duplicate storm; it does not make one impossible.
128
+ 21. **`AU-02`'s normalisation is only measured on synthetic reports.** Two differently-typed
129
+ copies of one symptom give one key and a different symptom gives another, but whether a real
130
+ report set normalises well enough is unknown. Its failure mode is a duplicate card, never a
131
+ lost report.
132
+ 22. **No audit has ever run against a real registry here.** Everything `SC-07` does was measured
133
+ against a canned npm-audit report (`tests/t-audit.sh`), so the parse is proven, the *policy*
134
+ is not: whether the recorded waiver set matches the real advisory set is unknown until the
135
+ first real `goblin-audit`. Its skip-with-a-reason behaviour on a repo with no record is what
136
+ keeps that honest rather than silent.
137
+ 23. **The audit record is read by field name, not by a JSON parser.** `goblin-audit` recognises
138
+ npm's `vulnerabilities` / `via` shape (`source`, `name`, `url`, `range`) with awk and REFUSES
139
+ (exit 5) anything it cannot parse, rather than writing an empty record that `SC-07` would read
140
+ as clean. A different audit tool is therefore a refusal, not a silent pass.
141
+ 24. **The freshness rows need a GNU `date -d`.** `SC-07` parses the record's date that way; on a
142
+ host without it the row FAILS with the reason rather than assuming the record is fresh.
143
+ 25. **`SC-04` reads one statement, not one program.** A cookie write spread over three lines (or
144
+ assembled through a helper) is not seen, and the row says so in its own cell.
145
+ 26. **The last advisory slot is an open decision, not a rule.** Measured (V1): the advisory rows
146
+ are `HP-04`, `HS-03`, `CM-02`, `MD-02`, `MD-03`, `PG-04`, `DOC-01`, `DOC-02`, `SC-09` — 9 at a
147
+ ceiling of 10 — so **exactly one slot was free**, and `SK-03` reported that arithmetic at the
148
+ time (`advisory 9 of ceiling 10 (1 free slot)`; **corrected 2026-09-25 (AB3):** the run prints
149
+ `advisory 10 of ceiling 10 (0 free slots: the next advisory row FAILs)`, W3's `JG-03` having
150
+ taken the slot — the two notes below carry the chronology). Two planned cards each wanted the
151
+ slot: G1's `FM-03` (the feature map) and G2's `JG-03` (the judge agent). **Nothing in this repo
152
+ chooses between them**, and V1 deliberately spent nothing. The cap is a count, not a strict
153
+ bound: measured, 10 advisory rows at a ceiling of 10 **pass**, and the 11th FAILs (10 at a
154
+ ceiling of 9 FAILs). So the tenth row is allowed; the eleventh is not. Whoever lands second
155
+ brings a real command.
156
+ **Decided 2026-09-25 (W2):** G1's `FM-03` does **not** take the slot. The feature map ships
157
+ `FM-01` and `FM-02` as real commands, and the one thing they cannot check — whether the map
158
+ lists every feature — is recorded as #30 instead of as a counted row; the slot is left free for
159
+ G2's `JG-03`, whose judge is the mechanism `P12` actually needs. Measured after W2, `SK-03`
160
+ still reads `advisory 9 of ceiling 10 (1 free slot)`.
161
+ **Spent 2026-09-25 (W3):** G2 landed and `JG-03` took it. The slot is now **full** — measured
162
+ `advisory 10 of ceiling 10 (0 free slots: the next advisory row FAILs)` — so the next author who
163
+ wants an advisory row must raise `advisory_ceiling` in the same change and write down why,
164
+ rather than discovering the cap from a red run. This is not a rule change; it is the arithmetic
165
+ the ceiling was always meant to force into the open.
166
+ 30. **The feature map is an inventory, and nothing checks that it is complete.** `FM-01` checks the
167
+ index against the feature files that exist and the four-H2 entry contract; `FM-02` is a tripwire
168
+ over `entry_paths:`. **Neither can see a feature nobody wrote down** — that needs semantic
169
+ judgement over the app, which no command here performs. Three smaller blind spots are named in
170
+ the rows' own why-cells and repeated here: `FM-02` searches for the token under `source_root:`
171
+ and therefore **reads a vendored copy, a lockfile or a build artifact as "still resolves"**; it
172
+ sees a *file* change and not a *behaviour* change, so a refactor that leaves the route alone
173
+ reports stale-and-wrong and a behaviour change in a file the token does not appear in is missed;
174
+ and a `verified:` date is a **claim the row cannot test** — nothing distinguishes a feature
175
+ driven that day from a date typed that day. The first of those is why `source_root:` should name
176
+ the source tree, not the repo root, and the third is why the upkeep pass in
177
+ `skills/goblin-feature-map/SKILL.md` requires the date to advance only for a feature someone
178
+ actually drove.
179
+ 31. **The P6↔P12 loop is wired as a contract with no runner.** `P6` now hands a generated
180
+ verification skill to `P12` and `verified:` does not advance until an eval record exists; the
181
+ record's shape, the eleven-token ban, the cheap-checks-first ladder, the merge rule and the pass
182
+ condition (every seeded defect detected, the control's number at zero, every correction RED
183
+ before GREEN) are all specified in `skills/goblin-eval/SKILL.md`. **No row reads a lane, and
184
+ nothing executes an eval** — measured after W2, `manifest/enforcement.tsv` has no `EV-*` row,
185
+ and `P12` has still never been run. The runner and the record checks (`EV-01`..`EV-04` in G1's
186
+ design) are deferred to a follow-up card, deliberately and in the open, rather than half-built:
187
+ a row that reads a record nobody writes would pass vacuously and look like enforcement.
188
+ 32. **The judge is a language model grading prose, and the record proves a handle exists — never
189
+ that the handle supports the verdict.** `JG-01` FAILs a `done` verdict whose evidence resolves
190
+ to nothing: `sha:<hex>` must be a commit in `git rev-list --all`, `file:<path>` a path under the
191
+ root, `sha256:<hex>` the digest of a file under `.goblin/loop/`. All three say *exists*. A judge
192
+ may cite a real commit that has nothing to do with the claim and pass, and a `cmd:<command>`
193
+ token resolves **nothing on purpose** — the command ran, its output is not in the record, and a
194
+ verdict resting on it is the self-report the row refuses. Two smaller blind spots are stated in
195
+ the row's own why-cell and repeated here: the row cannot see **which lane returned a verdict**
196
+ (no file in a repo observes the profile that ran — `MD-03`), and **an unresolved judge lane is
197
+ an `ADV`, not a failure**, because a repo cannot choose the fleet's routing (`docs/ROLES.md`,
198
+ "the measured caveat"): measured on this box, the judge lane resolves to no profile at all.
199
+ Finally, the `JG-03` counter-measure is **policy, not a check** — one known-red control verdict
200
+ per wave, recorded in `docs/LOOP.md`: a lane that has judged twice cannot be called always-yes,
201
+ and the history that would show a bad lane lives across cards and repos.
202
+ 33. **`LP-04` measures a changed evidence pointer, which is a proxy for progress — not progress.**
203
+ Three consecutive verdict rows with an identical non-empty pointer and a result that is not
204
+ `predicate:green` is a FAIL naming the row numbers, and that is the strongest thing a repo can
205
+ read without running the loop. A loop that edits a file each turn to keep the pointer moving is
206
+ not caught, and **nothing in Hermes detects a lack of progress either**: measured in
207
+ `hermes_cli/goals.py`, `run_kanban_goal_loop` carries no progress state at all — its whole
208
+ state is `last_response`, `turns_used` and `nudged_to_finalize`, so a loop that returns
209
+ `continue` for the same reason nineteen times spends nineteen turns and then blocks
210
+ (`hermes_cli/goals.py:1636-1638`, `:1689-1696`). So the budget is the backstop, and the budget
211
+ has its own blind spots: `LP-03` proves the declared budget is a positive integer at or under
212
+ `loop_max_turns_ceiling` and that the record holds no more verdict rows than the budget — it
213
+ cannot see whether the budget is **affordable**, and cost is not a field the record holds
214
+ (neither turns nor tokens nor the per-turn auxiliary judge call). Three further "not yets" are
215
+ structural rather than measurable: `LP-01` proves a recorded first run exists and that its
216
+ timestamp is at or before the first log row — **not that the command ran and not that the
217
+ `exit=` value was measured rather than typed** (`HP-03`'s defect, one artifact over); `LP-02`
218
+ proves the predicate file still hashes to its pin — **not that the predicate is the right one,
219
+ and not who edited it**; and a predicate that was **vacuously true from the start** (a `grep -c`
220
+ against a renamed directory) passes `LP-01` and `LP-02` and ends the loop green on nothing.
221
+ `LP-05` makes a write-up mandatory and therefore visible — nothing can make it true.
222
+
223
+ ## What the harness refuses to do
224
+
225
+ It does not claim a green run means the work is right. `goblin-verify` asserts that the installed
226
+ files are the files on disk, that every rule with a command still passes, and that the
227
+ untestable remainder is counted and capped — and it prints, on every single run, what it cannot
228
+ see: the five upstream blind spots, plus the ban lane's own (the unsigned ban table #28, the
229
+ text-probe gap #27, and a ban that is invisible until verify runs, `V3-1`), plus the judge/loop
230
+ lane's (#32: a handle that exists is not a handle that supports the verdict; #33: a changed
231
+ pointer is a proxy for progress, not progress), plus the CI lane's (#34: a file is not a gate —
232
+ the required-check list, the bypass switch and the push identity are forge state; #35: the
233
+ Electron perf number is a host gate, and the ratchet deliberately carries a different metric).
234
+
235
+ 27. **The ban probes are text probes, not ASTs.** `BN-01`, `BN-02`, `BN-03` and `BN-05` are
236
+ `grep` over source under `bash`/`grep`/`awk` only — the dependency contract in
237
+ `docs/CONTRACTS.md` allows no parser and no `npm`. So a `: any` inside a string or a comment
238
+ is reported, `Record<string, any>` (no leading colon) is missed, and BN-05 does not resolve
239
+ module aliases or dynamic imports. The AST-grade form of the same bans (BN-01..BN-04 in
240
+ `G5.md` §C.2) needs ESLint and `dependency-cruiser`; that is why **G5's `BN-04` (the nine
241
+ named unnecessary-effect patterns) is NOT shipped** — it cannot be mechanised without a
242
+ parser, and a ban that cannot go red is worse than advisory, so it is recorded here rather
243
+ than as a row that would cost the last advisory slot. A text probe with a stated
244
+ false-positive set is still a gate: each BN row is shown RED under its own violation and
245
+ GREEN when it is removed (`tests/t-verify-red.sh`).
246
+ 28. **The ban table's integrity rides on `IN-02`, and `IN-02` rides on an unsigned record.** A
247
+ project could empty `manifest/bans.tsv` (or edit a `detect` command) and the ban gate would
248
+ pass vacuously — `BN-00` fails closed on an *empty* table, but it cannot see a table whose
249
+ rows were weakened, because `.goblin/manifest/bans.tsv` is hashed by `IN-02` and
250
+ `.goblin/installed.json` is not signed (`docs/LIMITS.md` #18). The ban list is not
251
+ tamper-proof; it is as strong as the record every drift check trusts. **Measured (W5-12):** with
252
+ `BN-01`'s `detect` cell set to `true`, `BN-00` and `BN-01` both PASS and the only row that
253
+ changes verdict is `IN-02`'s drift check — the `W5-12` control in `tests/t-verify-red.sh` pins
254
+ exactly that, and nothing else can. Z1's verdict is to **record this, not fix it**: the `detect`
255
+ cell is executable content, and a row able to judge whether another row's command *means*
256
+ something would be an `eval` over the matrix, which the same card rules out. The gap is now
257
+ stated, measured and controlled one row over, which is the most this table can do without
258
+ becoming an interpreter.
259
+ 29. **The perf ceiling must EQUAL the recorded baseline, so a budget with headroom is not
260
+ expressible.** `PF-01` FAILs when `ratchet.ceiling` and `perf.baseline_value` disagree
261
+ (G8-6b), which is what stops a one-line ceiling raise from passing while printing the
262
+ contradiction. The cost is real and the row cannot see it: a project that wants the ratchet
263
+ to allow, say, 10% growth over the measured baseline cannot write `ceiling: 44000` beside
264
+ `baseline_value: 40000` — the row reads that as a disagreement. Headroom is expressed by
265
+ re-anchoring BOTH (measure on a pinned commit, then set `baseline_value` and `ceiling` to the
266
+ new number together, and record it as an operator action), which is the deliberate re-anchor
267
+ the row's own why-cell names. It is a bound on the budget's shape, not a claim that the
268
+ budget is the right one — and it still never re-measures.
269
+ 34. **The CI lane reads files, and a file is not a gate.** The shipped workflow
270
+ (`templates/ci/goblin-gate.yml.tmpl` → `.github/workflows/goblin-gate.yml`) has no `if:` at any
271
+ level, but nothing in a repository can make GitHub **require** it: the required-check list,
272
+ the bypass switch and the push identity are forge state (four settings, `docs/CI.md` §1).
273
+ `PG-06` proves the whole declared gate set is invoked in a file under `.github/workflows/` and
274
+ cannot prove that file is the one the forge waits on; `PG-05`'s reader has no parser, so a
275
+ flow-style `jobs: {…}` mapping is refused, and a `#` inside a quoted string truncates the line
276
+ it is on. Two further measured gaps: the template's job is `ubuntu-latest` with no cache, so a
277
+ repo whose gate needs a display, a licence, a GPU or a signed-in session cannot use it at all
278
+ (that is a **host** gate — the class-C rule, restated for class F), and a private repo's
279
+ Actions minutes are billed to the account (2,000/month free). Measured ground truth at W4:
280
+ **one** first-party workflow exists in the whole estate and it self-skips; five of the six
281
+ repos with a remote have none.
282
+ 35. **The Electron perf number is a host gate, and the ratchet carries a different metric.**
283
+ `presets/F-electron.yaml` declares `main_thread_busy_pct` as `perf_host_gate:` and uses
284
+ `app_bundle_bytes` for `ratchet:` — a **deliberate deviation** from G6 §B.3, which put the FPS
285
+ number in the ratchet. The instrument that produces it (CDP `Performance.getMetrics`, or
286
+ `app.getAppMetrics()[i].cpu.percentCPUUsage` inside a real Electron) needs Playwright or
287
+ Electron plus a GUI, and a shipped rule may use nothing but bash/git/awk/sed/grep/python3
288
+ (`docs/CONTRACTS.md`) — so `ratchet.cmd` pointing at the probe would make a fresh install
289
+ **born RED**, which is the one thing the install path must not produce. The probe belongs to
290
+ the project; a project whose CI needs npm runs it in **its own** workflow. What the number
291
+ means has a limit of its own, and it is measured: this box is an LXC with no display and no
292
+ system Chromium, so the G6 sweep came from a *bundled headless* Chromium with no compositor
293
+ and no vsync. Relative comparisons on one machine are meaningful (which is why
294
+ `main_thread_busy_pct` works at all, and why frame time does not — p50 stayed flat at 16.70 ms
295
+ while the main thread went from 1.8 % to 54.5 % busy, `docs/CI.md` §3.3); an absolute FPS
296
+ claim is not. Three further Electron failure modes are **recorded, not mechanised**, and
297
+ `docs/CI.md` §4 says why: the dependency-graph boundary check, `ipcMain` sender validation,
298
+ and fuses at package time.
299
+ 36. **A ban's exemption reaches the probe through its environment, so a custom probe can ignore it.**
300
+ `bans_exempt:` and the inline `// BAN-OK(<id>): <reason>` are filtered *before* the exit code is
301
+ chosen, because a filter applied to a probe's stdout afterwards cannot change a verdict — that
302
+ was W5-1, and the old code made every exemption a permanent RED. The engine therefore exports
303
+ `GOBLIN_BANS_ID` and `GOBLIN_BANS_EXEMPT` and the two shipped probes honour them. A project's
304
+ **own** `detect` command that ignores the variables keeps the old behaviour: a violation inside
305
+ an exempted path stays RED. That direction is **closed**, never open, and it is the trade this
306
+ choice makes. The engine's own stdout is no longer filtered at all, so a probe that ignores the
307
+ contract prints the exempted lines it reported — loud, and still a FAIL.
308
+ 37. **`FM-02` refuses the harness, not every non-source file.** The search skips `.git/`,
309
+ `.goblin/`, `.hermes/`, the declared `harness_dir` and the map's own directory, so a stub map
310
+ whose token occurs only in the install no longer resolves (W5-4: `entry_paths: [export]` used to
311
+ "resolve" to `./.goblin/bin/goblin-verify`). It does **not** exclude the target's own `docs/`,
312
+ `tests/` or build output: a token that occurs only there still reads as resolved, and excluding
313
+ them would be a guess about a layout goblin-stack does not know. Nor is there a frequency bound
314
+ — a common token ("export", "main") is satisfied by the first of hundreds of files.
315
+ 38. **The judge lane's model family is reported, never enforced — and on this box it is not even
316
+ mapped yet.** `JG-02` proves the declared profile *names* are disjoint; W5-6 measured that a
317
+ profile mapped to the author's own model passed it. `MD-02` now resolves the judge lane and
318
+ compares its model with the code lane's, so the state is loud — but it stays `advisory`
319
+ (`return 2`, never a FAIL): goblin-stack cannot choose the fleet's models, and a repo-local file
320
+ cannot observe *which* model a lane actually ran. Measured on this box at X1: Harvey's
321
+ `fleet-model.yaml` names **no `judge:` profile at all**, so `MD-02` reports the judge lane
322
+ **unresolved** and prints the one-line remedy, while every profile it *does* name — nine of
323
+ them, the named per-fleet role profiles this setup routes its lanes through —
324
+ resolves to one model under an active promotion. So on this box the honest reading is: the judge
325
+ lane is not mapped, and the moment it is mapped it will be the author's own family. The
326
+ comparison is exact model equality, not a version-stripped "family": two spellings of the same
327
+ family that differ only in a suffix would read as different.
328
+ 39. **`LP-02` cannot tell a weaker predicate from a re-scope. The close-and-reopen is recorded, not
329
+ prevented.** A loop could archive its bar under `closed-<date>/`, write a weaker one, re-pin,
330
+ and pass `LP-02` and `LP-05` with nothing in the record (W5-7, measured). The row now requires
331
+ every archive to hold its predicate **and** the pin it was closed under, and the live
332
+ `predicate.sha256` to name the archived digest on a `previous:` line — so a **silent** relaxation
333
+ is caught and a legitimate re-scope costs one line. Whether the new predicate is weaker, and who
334
+ edited it, is not decidable from a digest, and a predicate that calls a script elsewhere is
335
+ pinned only at its call site. The residual is the same one `LP-01` carries one file over: the
336
+ record is evidence, not proof that the loop stopped for the right reason.
337
+ 40. **A fresh install into a repo with no commits is born RED, and X1 did not change that.** Measured
338
+ at X1: `goblin-install --class A` into a `git init` with zero commits, then `goblin-verify`, gives
339
+ `37 passed, 6 failed, 11 advisory, 24 skipped`, exit 1 — `HP-05`, `SP-02`, `GT-02`, `CM-01`,
340
+ `CM-03` and `PT-02` all read a HEAD that does not exist yet. The install never creates the seed
341
+ commit (`bin/goblin-install` writes files and stops), and it must not: a tool that commits into
342
+ Harvey's repo on first contact is the overreach `docs/CONTRACTS.md` rules out. Carried from W5
343
+ §5's `G8-9` rather than fixed here — it is a **sequencing** limit, not a hole in a row: one commit
344
+ clears all six, and `tests/t-verify-green.sh` seeds one before it installs.
345
+ 41. **Eleven of Y1 §7's eighteen documented mechanisms still carry no control of their own, and one
346
+ exemption is by design.** The census, stated separably so a reader can check the arithmetic:
347
+ **18 listed · 2 fixed · 2 closed · 3 controlled · 11 recorded.** Y1 §7 listed eighteen
348
+ mechanisms; Z1 **fixed** the two load-bearing ones (17, the unrendered token — Z1-3; 18,
349
+ `replay.cmd` — Z1-4, which now have controls and therefore sit *outside* the "carry no control"
350
+ set) and **closed** two more (4, the ban engine's `exit 2` paths; 5, `--list` — three assertions
351
+ in `tests/t-verify-red.sh`, RED against a deliberately broken copy of `bin/goblin-bans`, which
352
+ is PR-03's second branch: those mechanisms *worked*, they were just unguarded); AA1
353
+ **controlled** three (1, 2 and 6 — the ban-engine cluster, below). The remaining **eleven** are
354
+ **recorded here rather than mechanised**. Each was measured WORKING, so an assertion would pin
355
+ behaviour that already holds, and this box does not have eleven controls' worth of budget; the
356
+ cost of the gap is exactly the shape of both regressions in this repo's history — a mechanism
357
+ documented, and nothing asserting it.
358
+ - items 1, 2, 6 — `bans_exempt:` on the layer probe, the engine→probe
359
+ `GOBLIN_BANS_ID`/`GOBLIN_BANS_EXEMPT` contract, and its segment alignment: **CONTROLLED at
360
+ 0.4.2 (AA1)** — seven controls in `tests/t-verify-red.sh`, in both directions each (the
361
+ layer probe's exempt path PASSes and the same crossing import outside it FAILs; a prefix
362
+ covers its own subtree but does not swallow `src/okay`; a prefix written with a trailing
363
+ slash exempts nothing; and a project's own probe that exits 0 only when both environment
364
+ variables arrive). This was the cluster Z1 named as the one with a real engine underneath,
365
+ and `grep -c GOBLIN_BANS_EXEMPT tests/` was **0** before it.
366
+ - item 3 — a project's own probe that ignores those variables fails closed: the mechanism is
367
+ this file's #36, where the FAIL is measured.
368
+ - item 7 — `BAN-OK` must sit on the offending line: the marker's *effect* is controlled, the
369
+ on-this-line clause is not (a marker on another line does not clear the violation).
370
+ - items 8, 9 — `SC-05`'s `.goblin/boundary-waivers` and `SC-08`'s
371
+ `.goblin/install-hooks.allowlist`: both files are copied and both FAIL directions are
372
+ controlled; neither *honoured* direction is. (`SC-08`'s reader is the row the minified
373
+ lockfile defeated — Z2-2; that was a shape, not this item, and it is fixed and controlled.)
374
+ - item 10 — `SC-03` clause 2, `sec_build_output`: a literal in `dist/assets` FAILs; only the
375
+ clause-1 path is controlled.
376
+ - items 11, 12, 15 — `scaffold_checks:`'s SKIP branch, `perf_host_gate` (#35) and
377
+ `templates/loop/*.tmpl`: declared no-ops (SKIP, a host gate, and templates `docs/LOOP.md`
378
+ says nothing installs). A control here would assert that nothing happens.
379
+ - item 13 — `docs/CI.md`'s flow-style `jobs: {…}` refusal: fails closed already, so a control
380
+ would pin the strict direction of a check that cannot pass vacantly.
381
+ - item 14 — `bin/goblin-model`: works (`code` → the resolved lane, unknown role → exit 2);
382
+ only its *absence* from an install is asserted.
383
+ - item 16 — `docs/LOOP.md` §6's "one known-red control verdict per wave, recorded in this
384
+ file": a prose obligation, not a command. It is honoured in the write-ups (or not) and no
385
+ check can read a wave.
386
+ **W5-10 is answered here too.** The tenant strings `PT-01` forbids do not reach `docs/`, because
387
+ `docs/` is never installed: the row's directory list is `skills manifest bin templates presets
388
+ .goblin .hermes`, so a string under `docs/` is source-tree prose that no operator's repo ever
389
+ receives. That is the whole reason, it is deliberate, and the row is not weakened by it — Y1
390
+ agreed, and Z1 leaves it. This is the sentence Z1 added so the reason is visible in the shipped
391
+ artifact rather than only in the wave's own notes.
392
+
393
+ 42. **The doc walk cannot read a path broken at the directory/name boundary.**
394
+ `tests/t-doc-promises.sh` reads a command directory only as `.goblin/bin/` or `bin/` **with its
395
+ trailing slash**; when a line ends on the bare directory (`.goblin/bin`, `bin`) and the slash
396
+ leads the next line (`/goblin-doctor`) or is dropped (`goblin-doctor`), the directory and its
397
+ continuation are both invisible — the first fragment stops before the name the grammar needs, the
398
+ second is not preceded by `bin/`, so neither is a token and neither is asserted. Measured
399
+ 2026-09-26 (AB7): a plant of that form in `docs/CI.md` is reported **`PASS`, rc 0**, mentioning
400
+ the plant **zero** times, and a line that ends on the bare directory with nothing after it is
401
+ silent too, because the `dangling()` guard needs the trailing slash as well. **0 live instances**
402
+ — measured: `grep -rnE '\.goblin/bin$|bin/$' $(git ls-files)` → **no output, exit 1**. **What it
403
+ costs:** a false path written in that form — the exact defect this file's whole species is named
404
+ for — passes the watch silently, so a future document could hand a reader a command that does not
405
+ exist and nothing in the walk would say so. **Ticketed, not gated:** reading the form is a change
406
+ to the tokeniser inside `tests/t-doc-promises.sh`, not to a document, and it is a separate card;
407
+ this entry records the boundary so a green run cannot imply a coverage it does not have. The
408
+ control's own header now names the form as uncovered instead of claiming a sensitivity it does
409
+ not have, and the slash's load-bearing role is stated where the dangling rule is.
410
+
411
+ 43. **The engine footer names its judge, and the statement is unsigned — one forgery covers N
412
+ repos.** W1's engine split prints `cli_sha256` and `enforcement_tsv_sha256` on every run
413
+ (and `installed.json` gains an optional `engine:` block recording the same pair), so a repo
414
+ can state WHICH engine judged it. Nothing signs either hash: the footer is printf output of
415
+ the very binary it names, and the `engine:` block is a JSON stanza inside the same unsigned
416
+ record `docs/LIMITS.md` #18 already covers. An edited engine, or a hand-written block
417
+ claiming a mode that was never resolved, prints whatever it likes — and because the
418
+ statement now rides in every repo's run, one forgery propagates to every repo that trusts
419
+ it, which is strictly worse than #18's per-repo record. The old defences still hold and are
420
+ what the footer must not be read to replace: `--source` pins the engine by flag, a vendored
421
+ engine keeps beating the global one in the resolution chain, and IN-02 still hashes the
422
+ repo-local bytes. What is NOT fixed (G6, the no-signing non-goal): no trust root, no
423
+ detached signature, no third-party attestation. Measured with W1: tamper one byte of the
424
+ vendored manifest and the footer's `enforcement_tsv_sha256` moves — the footer is a real
425
+ fingerprint of what ran, it is just not PROOF of it.
426
+
427
+ ## Verdicts recorded, not built (the note-8 questions)
428
+
429
+ Two library questions were investigated to a verdict and **deliberately built nothing here**, so a
430
+ later session does not re-derive them. The measurements and sources are in
431
+ `goblin-stack-research/G6.md` Part C; this is the durable half.
432
+
433
+ **Pretext (`chenglou/pretext`) — USE, narrowly, and not as a runtime dependency.** Verified from
434
+ the source of truth: a pure JS/TS library for multiline text measurement and layout, MIT, that
435
+ *"side-steps the need for DOM measurements (e.g. `getBoundingClientRect`, `offsetHeight`), which
436
+ trigger layout reflow"*, using the browser's own font engine as ground truth. Author confirmed
437
+ (Cheng Lou, `_npmUser: chenglou`); the README credits Sebastian Markbage's earlier `text-layout`,
438
+ whose repository now says *"This project is archived. The ideas here evolved into Pretext"*.
439
+ Two verified corrections to the shipped skill: the skill pins `@chenglou/pretext@0.0.6` while npm's
440
+ latest is **`0.0.9`**, and *"15KB zero-dependency"* describes the **runtime**, not the install
441
+ (published tarball `unpackedSize: 887142` across 69 files, `sideEffects: false`, subpaths `.` and
442
+ `./rich-inline`). The concrete use is a **dev-time / harness-time label-fit check** for
443
+ `diagram-studio` and `goblin-ui` — "does this string fit this box at this font", without a layout
444
+ read, as a `checks/*.mjs` assertion with a REPLAY — in exactly the repos that today hand-roll that
445
+ arithmetic (26 `getBoundingClientRect()` calls, no `measureText` call, no label-overflow probe).
446
+ **Do not** put it in the runtime bundle: the app is offline by design, and the skill's own stack
447
+ table imports it through a CDN, which contradicts that. The second use is one blog demo, not
448
+ infrastructure.
449
+
450
+ **mise (`jdx/mise`) — NOT NOW, with a named trigger that flips it.** Verified: mise-en-place, MIT,
451
+ macOS/Linux/Windows, one CLI that declares tool versions, environment variables and commands in
452
+ `mise.toml` and uses them in the shell, the editor and CI; the polyglot successor to
453
+ asdf/nvm/pyenv, and `.tool-versions` already works; installed with `curl https://mise.run | sh`.
454
+ The premise for adopting it did not survive measurement: the projects do **not** run differing Node
455
+ versions, and nothing declares a version at all — 0 `.nvmrc`/`.node-version`/`.tool-versions` files
456
+ under `~/projects`, `engines.node` in exactly **one** first-party manifest, and a single Node on
457
+ the box (v22.23.1, reached through `~/.local/bin`, which a shell profile puts first). The one real
458
+ divergence is a *package-manager* split, which mise cannot resolve. Adopting it now would add a
459
+ second version source beside the Node Hermes bundles. **Adopt when either becomes true:** two
460
+ projects need different Node majors, or the first Electron app lands (Electron ships its own
461
+ Node/Chromium, so the host Node stops mattering for the shell and starts mattering for the build
462
+ tooling). **The cheap thing to do meanwhile:** declare the version that already exists — one
463
+ `engines.node` line where it is missing, and a `measured <date>` gate line in the HANDOFF naming
464
+ the Node the gates ran under.
465
+
466
+ 44. **The version statement is a convention, not an enforcement row — the sync is test-side and
467
+ the engine itself never checks it.** W2 closed the measured hole PLAN-V1 §4.4 recorded (no
468
+ test read any of the five `GOBLIN_*_VERSION` constants in `bin/`): `tests/t-version-sync.sh`
469
+ now pins the count of those constants at 5, asserts every constant and
470
+ `package.json.version` equals `VERSION`, and asserts `goblin --version` — through both the
471
+ bash CLI and the node shim — prints `VERSION` byte-for-byte. What remains open: the sync
472
+ lives only in this repo's test suite, which a target repo never runs — and when the npm
473
+ tarball exists (W3/W5 packaging, planned to ship no `tests/`, deliberately), a published
474
+ package's `goblin.js` and its bash payload could
475
+ drift apart with nothing in the shipped artifact noticing. The engine has no self-check row
476
+ that reads its own `--version` against a manifest record, and adding one would make the
477
+ version a rule — which is a real option, not done here. Until then the guarantee is
478
+ development-time only: green in this checkout, unverifiable in the wild.
479
+
480
+ 45. **The migration is crash-safe by sequence, not by journal — a kill mid-upgrade leaves a shape
481
+ only three of which are detected.** `goblin upgrade` runs eight steps across two commits, and
482
+ the spec's refusal conditions name the shapes a crashed run can leave: record says global but
483
+ the 18 files are still here (R8 catches it), the record lost its engine: block entirely (R9
484
+ catches it), commits half-landed (the preflight's own state resolution). What is NOT caught:
485
+ a crash between the engine landing (step 3) and commit A (step 4) leaves `~/.goblin/engine`
486
+ written but the repo untouched — harmless, invisible, and never re-verified (the next upgrade
487
+ run compares hashes and refuses R6 if the payload has since changed, so the stale engine can
488
+ sit there being wrong until someone looks). And the shadowing footer (W3 §3) reads only the
489
+ vendored-payload shape; a repo whose engine_dir points at an engine that no longer exists
490
+ FAILs the resolution chain loudly (exit 2), but a repo pointing at an engine whose bytes have
491
+ silently changed since migration day runs green on the drifted table — the footer's hashes
492
+ name what ran, nothing compares them to migration day's. **Ticketed, not gated:** a
493
+ migration-day hash pin in the record compared per-run is the real fix and is a row-shaped
494
+ change (a new IN clause), not a W3 patch. Measured: R6 refuses the wrong-engine reuse, R8/R9
495
+ catch the two crash shapes they name, and the between-steps engine write is unwatched.
496
+
497
+ 46. **The emitted platform config is unsigned and hand-editable — DRIFT is detected per run,
498
+ never prevented.** `goblin emit` writes skills byte-copies and one delimited block in the
499
+ platform's context file, recorded in `~/.goblin-stack/emissions.tsv` with pre-image hashes
500
+ (so `--uninstall` restores byte-exactly), but nothing signs what it writes: an edit to an
501
+ emitted `SKILL.md` or to the bytes inside the `goblin-stack:begin/end` block makes the
502
+ platform's `verify.sh` oracle report DRIFT on the next `goblin doctor` run — and that is
503
+ all it does. There is no lock, no signature, and no write protection on any emitted file;
504
+ a platform (or the user) can change them between two doctor runs and nothing notices
505
+ until someone runs one. The same holds for `--unshadow`: it removes only project copies
506
+ whose hash equals the source payload and refuses-and-names any that differ, so a real
507
+ local edit survives, but nothing reconciles it either. Measured: a one-byte tamper in an
508
+ emitted `SKILL.md` and a stale marker VERSION in the context block each report DRIFT (exit
509
+ 1) on the next run, and uninstall refuses to delete a recorded file whose bytes no longer
510
+ match its post-image (R6).
511
+
512
+ 47. **The four W4b adapter conventions are documented shapes, not run-probed installs — and
513
+ two of the seven platforms cannot block commands outright.** cursor and codex are not
514
+ installed on the build machine, so their rows pin the official docs (read 2026-09-29),
515
+ not a live CLI; a platform changing its layout invalidates the adapter silently until a
516
+ doctor DRIFT names it. codex and gemini report cap_command_blocking `partial` — codex
517
+ disables skills via `~/.codex/config.toml` `[[skills.config]]`, gemini only narrows via
518
+ approval modes — so an emitted skill is *available* there even when the operator would
519
+ forbid it; the doctor prints the codex hooks caveat (sessionStart only). And gemini's
520
+ id cell was measured to be exactly its platform name (the `gem_ini` typo shipped in one
521
+ W4b build and the doctor's DRIFT caught it — the schema check works).
522
+
523
+ 48. **`SP-01`'s rule text once overclaimed; it now states the check's actual scope.** The text
524
+ used to read "The current round has a SPEC" while the check is
525
+ `ls ./*-SPEC.md >/dev/null 2>&1`, which passes on ANY spec-shaped file at the repo root. The
526
+ measured consequence: a fresh class-A install ships the scaffold's `ROUND-000-SPEC.md`, and
527
+ `SP-01` PASSes on that scaffold alone — no round is open, and the row is green anyway. A stale
528
+ spec from a finished round keeps satisfying the row exactly as well as a live one, because the
529
+ check has no notion of "current". **Fixed in W6 (text, not check):** the row now reads
530
+ "A *-SPEC.md file exists at the repo root (any round, not the current one - round-scoping
531
+ arrives with the W6 staged chain)" — manifest and `docs/ENFORCEMENT.md` re-rendered in the
532
+ same commit, the check cell untouched. What it costs: the SPEC-exists signal in a run summary is
533
+ weaker than a round-scoped claim would suggest — read it as "a `*-SPEC.md` file is present", not "this round's spec is
534
+ here". Measured: `goblin-verify` on a fresh probe install reports `PASS SP-01 (ls
535
+ ./*-SPEC.md >/dev/null 2>&1)` with `ROUND-000-SPEC.md` the only file the glob sees, and the
536
+ same PASS after renaming it to a non-round name. **Ticketed, not gated:** "current round" is a stage-order notion, and the W6 staged workflow chain (SC-01:
537
+ a SPEC committed before the changes it governs) is what gives the word meaning; the staged
538
+ chain's stage-order rows are what will make round-scoping REAL — the reword names the
539
+ boundary until then.
540
+
541
+ 49. **The engine footer prints to captured stdout, and a script parsing `goblin-verify` output
542
+ must expect it.** The two-line footer (`engine: mode=… cli_sha256=… enforcement_tsv_sha256=…`)
543
+ is unconditional `printf` output — there is no isatty guard, by design, so a captured run still
544
+ names the engine that judged it (that is the point of #43's statement). The cost is parser
545
+ friction: a script that treats every stdout line as a verdict line trips on the banner and
546
+ footer lines, which are not `PASS`/`FAIL` rows. The contract is the EXIT CODE, never the line
547
+ set: 0 green, 1 red, 2 refuse — parse the exit code, or filter to `^[A-Z]{2}-[0-9]{2}` before
548
+ reading lines. Measured: `./.goblin/bin/goblin-verify > out.txt` on a probe install puts
549
+ `engine: mode=vendored cli_sha256=… enforcement_tsv_sha256=…` in the captured file
550
+ (`tests/t-engine-dir.sh` itself asserts the footer IN captured output, so removing it or
551
+ tty-gating it would break the engine's own suite). Recorded so the pitfall is findable from
552
+ this file; no change to the engine is implied or wanted.
553
+
554
+ 50. **`CM-01` reads only the most recent commit — historical commits with a wrong identity pass
555
+ unseen. Current stated scope: the row gates the identity of HEAD at the moment of the run,
556
+ and nothing older; that is the row's whole claim, by design.** The check is
557
+ `test "$(git log -1 --format='%ae')"` against the configured
558
+ `owner_email`, so it gates the identity of HEAD at the moment of the run and nothing older.
559
+ Measured: a probe history `owner → wrong@old.co → owner` reports `--only CM-01` clean (exit 0)
560
+ with the wrong-identity commit sitting one below HEAD. That is consistent with the
561
+ gate-at-the-moment design — every row judges the tree and history as they stand when verify
562
+ runs, and retrofitting an identity sweep over `git log --all` is a policy change, not a bug
563
+ fix — but it should be named: the row's green means "the latest commit carries the owner
564
+ identity", never "no commit in this repo's history carries an ambient one". A wrong-identity
565
+ commit that has since been followed by correct ones is invisible to every run. Recorded as a
566
+ boundary; no row change implied. `docs/ENFORCEMENT.md`'s CM-01 row carries the same one-line
567
+ scope statement.
568
+
569
+ 51. **`gob init` writes its three added values into `goblin.yaml` itself, not through a
570
+ template.** Install renders `.goblin/goblin.yaml` from `templates/goblin.yaml.tmpl` and
571
+ owns it (`put_once`: never rewritten after the first install) — and install correctly
572
+ carries no `--branch/--email/--gate` flags, because those keys are the project's to edit.
573
+ The wizard therefore `sed`-patches `branch:` and `owner_email:` and rewrites the first
574
+ gate's `cmd:` line right after install renders the file, fail-closed: the gate must read
575
+ back through the engine's own `g_yaml_gates` or the run stops. The cost is a second
576
+ writer for exactly those three lines: a hand-customised comment placement survives, but a
577
+ future template change to those lines' shapes (renamed keys, a multi-gate default) must
578
+ be mirrored in `bin/goblin-init`. The wizard never rewrites anything outside the declared
579
+ keys, and a re-run after a hand edit of another line leaves that line alone. Recorded as
580
+ a boundary; the alternative — teaching install three wizard-only flags — would put wizard
581
+ vocabulary into the installer's contract for no gain.
582
+
583
+ 52. **`scope:source` and `scope:target` name WHOSE burden a row carries — the framework's or
584
+ the adopting repo's.** `scope:source` is goblin-stack's own proof burden: those rows
585
+ (`PR-01`..`PR-05`) are the framework testing ITSELF while it is being developed — they run
586
+ in this repo, under `tests/run-tests.sh`, and a dev of goblin-stack is the one who owes the
587
+ run. `scope:target` is the adopting repo's proof burden: those rows run in an installed
588
+ repo via `goblin-verify`, and the repo's owner owes the run. Same matrix, two creditors:
589
+ a source row can never fail a user's repo, and a target row can never substitute for the
590
+ framework's own suite. Recorded as a definition; `docs/ENFORCEMENT.md`'s scope paragraph
591
+ carries the same sentence for the reader who arrives there first.