@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +351 -0
- package/LICENSE +21 -0
- package/README.md +217 -2
- package/VERSION +1 -0
- package/adapters/_template/adapter.tsv +16 -0
- package/adapters/_template/detect.sh +10 -0
- package/adapters/_template/emit.sh +5 -0
- package/adapters/_template/verify.sh +4 -0
- package/adapters/claude/adapter.tsv +8 -0
- package/adapters/claude/detect.sh +8 -0
- package/adapters/claude/verify.sh +47 -0
- package/adapters/codex/adapter.tsv +12 -0
- package/adapters/codex/detect.sh +9 -0
- package/adapters/codex/verify.sh +45 -0
- package/adapters/copilot/adapter.tsv +10 -0
- package/adapters/copilot/detect.sh +8 -0
- package/adapters/copilot/verify.sh +45 -0
- package/adapters/cursor/adapter.tsv +11 -0
- package/adapters/cursor/detect.sh +10 -0
- package/adapters/cursor/verify.sh +45 -0
- package/adapters/gemini/adapter.tsv +15 -0
- package/adapters/gemini/detect.sh +11 -0
- package/adapters/gemini/verify.sh +49 -0
- package/adapters/hermes/adapter.tsv +9 -0
- package/adapters/hermes/detect.sh +8 -0
- package/adapters/hermes/verify.sh +27 -0
- package/adapters/opencode/adapter.tsv +14 -0
- package/adapters/opencode/detect.sh +9 -0
- package/adapters/opencode/verify.sh +45 -0
- package/automations/README.md +53 -0
- package/automations/bugreporter-intake.sh +145 -0
- package/automations/drift-audit.sh +139 -0
- package/automations/report.schema.tsv +10 -0
- package/bans/README.md +82 -0
- package/bans/grep-ban.sh +84 -0
- package/bans/layer-check.sh +57 -0
- package/bin/goblin +119 -0
- package/bin/goblin-audit +145 -0
- package/bin/goblin-bans +178 -0
- package/bin/goblin-doctor +233 -0
- package/bin/goblin-emit +484 -0
- package/bin/goblin-init +519 -0
- package/bin/goblin-install +720 -0
- package/bin/goblin-lib.sh +289 -0
- package/bin/goblin-model +105 -0
- package/bin/goblin-upgrade +572 -0
- package/bin/goblin-verify +2798 -0
- package/bin/goblin.js +103 -0
- package/docs/ADOPTION.md +168 -0
- package/docs/CI.md +187 -0
- package/docs/CONTRACTS.md +197 -0
- package/docs/DESIGN.md +92 -0
- package/docs/ENFORCEMENT.md +225 -0
- package/docs/FLOWS.md +164 -0
- package/docs/GUARDRAILS.md +126 -0
- package/docs/GUIDE.md +610 -0
- package/docs/INTEGRATION.md +92 -0
- package/docs/LIMITS.md +591 -0
- package/docs/LOOP.md +165 -0
- package/docs/RE-PLAYBOOK.md +183 -0
- package/docs/RISKS.md +70 -0
- package/docs/ROLES.md +105 -0
- package/manifest/bans.tsv +9 -0
- package/manifest/classes.tsv +61 -0
- package/manifest/enforcement.tsv +88 -0
- package/manifest/glossary.tsv +25 -0
- package/manifest/playbooks.tsv +16 -0
- package/package.json +37 -4
- package/presets/A-shipped-software.yaml +48 -0
- package/presets/B-service-config.yaml +40 -0
- package/presets/C-game.yaml +38 -0
- package/presets/D-knowledge.yaml +41 -0
- package/presets/E-fleet-config.yaml +42 -0
- package/presets/F-electron.yaml +67 -0
- package/roles.yaml +54 -0
- package/skills/goblin-bootstrap/SKILL.md +51 -0
- package/skills/goblin-bugfix/SKILL.md +26 -0
- package/skills/goblin-bugreporter/SKILL.md +52 -0
- package/skills/goblin-drift-audit/SKILL.md +43 -0
- package/skills/goblin-eval/SKILL.md +68 -0
- package/skills/goblin-feature/SKILL.md +26 -0
- package/skills/goblin-feature-map/SKILL.md +140 -0
- package/skills/goblin-handoff/SKILL.md +28 -0
- package/skills/goblin-investigation/SKILL.md +26 -0
- package/skills/goblin-judge/SKILL.md +74 -0
- package/skills/goblin-loop/SKILL.md +88 -0
- package/skills/goblin-mode/SKILL.md +70 -0
- package/skills/goblin-overnight/SKILL.md +42 -0
- package/skills/goblin-pr-gate/SKILL.md +42 -0
- package/skills/goblin-re-mobile/SKILL.md +51 -0
- package/skills/goblin-refactor/SKILL.md +23 -0
- package/skills/goblin-sweep/SKILL.md +23 -0
- package/skills/goblin-tdd-repro/SKILL.md +27 -0
- package/skills/goblin-verify-author/SKILL.md +50 -0
- package/skills/practice/SKILL.md +37 -0
- package/templates/AGENTS.md.tmpl +23 -0
- package/templates/HANDOFF.md.tmpl +43 -0
- package/templates/SPEC.md.tmpl +34 -0
- package/templates/audit-waiver.tsv.tmpl +10 -0
- package/templates/boundary-waivers.tmpl +8 -0
- package/templates/checks/assert.mjs.tmpl +60 -0
- package/templates/checks/gate.sh.tmpl +29 -0
- package/templates/ci/goblin-gate.yml.tmpl +46 -0
- package/templates/goblin.yaml.tmpl +138 -0
- package/templates/install-hooks.allowlist.tmpl +9 -0
- package/templates/loop/decisions.tsv.tmpl +1 -0
- package/templates/loop/predicate.tmpl +16 -0
- package/templates/report.yaml.tmpl +16 -0
package/docs/LIMITS.md
ADDED
|
@@ -0,0 +1,591 @@
|
|
|
1
|
+
# Limits — where this is weaker, and what is unproven
|
|
2
|
+
|
|
3
|
+
Honest accounting. Nothing here is a hedge for a defect that could be fixed; each is either a
|
|
4
|
+
deliberate trade or an unfilled gap.
|
|
5
|
+
|
|
6
|
+
## Weaker than the Cursor harness it learns from
|
|
7
|
+
|
|
8
|
+
1. **No live-drive lane is shipped.** A harness that launches and drives the real application is
|
|
9
|
+
what several of the source harness's playbooks depend on, and it is inert without a plugin
|
|
10
|
+
that is not shipped. `P6` can *author* a project-local driver; goblin-stack cannot ship one,
|
|
11
|
+
because that is per-project work. **The "live lane is the floor" rule is stated and
|
|
12
|
+
unfilled.**
|
|
13
|
+
2. **No cloud agents.** There is no per-agent computer here, so nothing that needs a machine of
|
|
14
|
+
its own.
|
|
15
|
+
3. **No agent graph and no verdict-ledger daemon.** The board replaces both at a lower
|
|
16
|
+
resolution: `parents` expresses ordering, not data flow, and sibling cards cannot see each
|
|
17
|
+
other.
|
|
18
|
+
4. **Fifteen playbooks against twenty-three.** The cuts in `docs/FLOWS.md` are deliberate and each
|
|
19
|
+
is argued, but real coverage is lost: performance hillclimbing, pixel parity, trace
|
|
20
|
+
forensics, stack landing, worktree hygiene.
|
|
21
|
+
5. **No swarm or arena fan-out.** The read-versus-write axis says that is correct for this work
|
|
22
|
+
mix; it is still a capability the other harness has and this does not.
|
|
23
|
+
|
|
24
|
+
## Weaker than the status quo it is meant to improve
|
|
25
|
+
|
|
26
|
+
6. **The biggest measured defect is out of reach.** The orchestrator's routing text says a bare
|
|
27
|
+
subagent spawn reaches the specialist profiles; it does not. That is the highest-cost defect
|
|
28
|
+
on the box, and goblin-stack cannot fix it from inside a project repo. It ships corrected
|
|
29
|
+
text in `docs/INTEGRATION.md` and lints its own artifacts (`MD-03`); the fleet-side edit is
|
|
30
|
+
escalated.
|
|
31
|
+
7. **The staleness of a fleet-config repo is detectable and not fixable here** — the E-class gate
|
|
32
|
+
notices it, and the underlying job bug belongs to another repository.
|
|
33
|
+
8. **A new surface to maintain.** Per-repo vendored `.goblin/` plus `.hermes/skills/` means
|
|
34
|
+
upgrade debt in every adopted repo, plus one more command pair to learn. The counter is that
|
|
35
|
+
the alternative — profile copies — already failed.
|
|
36
|
+
9. **Ten rows are labelled `advisory`, and nine of them are prose with no check at all:**
|
|
37
|
+
`HP-04`, `HS-03`, `CM-02`, `MD-03`, `PG-04`, `DOC-01`, `DOC-02`, `SC-09` and `JG-03` (G2, the
|
|
38
|
+
judge lane that has never returned a non-`done` verdict). One (`MD-02`) is
|
|
39
|
+
advisory-labelled but still reports its state as `ADV`. Each is counted and capped, but a
|
|
40
|
+
counted rule is still not an enforced one, and **the cap is a policy, not a proof.**
|
|
41
|
+
10. **It does not reduce the profile-skill surface.** It only stops that surface from growing
|
|
42
|
+
with project procedures; the existing divergence is a separate cleanup.
|
|
43
|
+
11. **`HS-02` cannot prove anything until a round lands.** With no pinned pre-change commit it is
|
|
44
|
+
skipped with a reason. That is honest, and it means a brand-new install has **no** REPLAY
|
|
45
|
+
evidence at all. The shipped `checks/assert.mjs` is a scaffold: it asserts something true by
|
|
46
|
+
construction, so `HS-02` correctly reports it as a check that proves nothing until it is
|
|
47
|
+
replaced with real probes.
|
|
48
|
+
12. **`HP-03` proves a date exists, not that a number is fresh.** A HANDOFF can carry yesterday's
|
|
49
|
+
number with today's date and pass. Since V1 the row is anchored on the gate names
|
|
50
|
+
`.goblin/goblin.yaml` DECLARES and skips the template's own example sentence, so a real gate
|
|
51
|
+
line can no longer lose its `measured <date>` in silence (G8-2 measured the old row passing
|
|
52
|
+
exactly that); but nothing re-measures the number, and a gate-bearing line that names no
|
|
53
|
+
declared gate and carries no gate-shaped keyword is still unseen.
|
|
54
|
+
13. **`PG-04` is documentation, and `PG-05` is a text reading — re-declared at W4 rather than
|
|
55
|
+
called a gate.** `PG-04` stays `advisory` because the forge is unobservable from inside a
|
|
56
|
+
repo: a protected branch whose only admin is the person pushing protects nothing, so a
|
|
57
|
+
required check armed under the sole admin's own identity binds nobody — the second half of
|
|
58
|
+
that sentence is measured (Harvey is the sole admin of every repo he owns), and changing it
|
|
59
|
+
is a forge/account change, not a repository change. `PG-05` **used to be a heuristic**: it
|
|
60
|
+
counted "every step is guarded", so a single **job-level** `if:`, a job with no step, and the
|
|
61
|
+
measured real shape (one *unguarded* step deciding whether the guarded gate step runs) all
|
|
62
|
+
passed it while GitHub reported Success. It now refuses a conditional job, a conditional
|
|
63
|
+
step, a job that declares no step and a workflow with no `jobs:` — see `docs/CI.md` §1 for
|
|
64
|
+
why a skipped job is a green light. It is still a **text reading**: no YAML parser
|
|
65
|
+
(`docs/CONTRACTS.md` allows none), so a flow-style `jobs: {…}` mapping is refused rather than
|
|
66
|
+
parsed, a `#` inside a quoted string is read as a comment, and a conditional step that is
|
|
67
|
+
genuinely safe is indistinguishable from the trap. It cannot see branch protection, the
|
|
68
|
+
required-check list, or whether the job ever ran — `PG-04` is the row that says so. `PG-06`
|
|
69
|
+
is the other half and has its own limit: it proves the declared gate is **invoked** in a file
|
|
70
|
+
under `.github/workflows/`, never that the forge marks that job required, never that it is
|
|
71
|
+
the job the forge waits on, and never that the workflow can fail.
|
|
72
|
+
14. **`CM-02` cannot be enforced.** A backtick lost to command substitution leaves no trace a
|
|
73
|
+
later check can read.
|
|
74
|
+
|
|
75
|
+
## Unproven at the level that matters
|
|
76
|
+
|
|
77
|
+
15. **The controlled evidence now exists, and its counter-finding is about this repo (corrected
|
|
78
|
+
2026-09-25).** This entry used to say that **no controlled study of skill efficacy exists**, so
|
|
79
|
+
that nothing here claimed the installed skills change agent behaviour. That is out of date:
|
|
80
|
+
**SkillsBench 1.1** (benchmark published 2026-06-16) reports that **curated Skills raise the
|
|
81
|
+
mean task-resolution rate from 33.9% to 50.5% — +16.6 points, a 25.5% normalized gain — across
|
|
82
|
+
87 tasks, 8 domains and 18 model–harness configurations**, with every one of the 18
|
|
83
|
+
configurations higher with Skills (configuration-level gains from +4.1 to +25.7 points).
|
|
84
|
+
**The counter-finding matters more here than the headline:** in the same benchmark's
|
|
85
|
+
**self-generated condition — the agent authors its own Skills before solving — all three tested
|
|
86
|
+
configurations landed BELOW their no-Skills baseline** (−8.1, −11.3 and −11.5 points), while
|
|
87
|
+
curated Skills stayed above it. **goblin-stack installs agent-authored skills**, so the lower
|
|
88
|
+
row of that result is a warning about its own output, not someone else's: a generated skill
|
|
89
|
+
accepted after a skim is a different proposition from a written one. It is why `P6` hands the
|
|
90
|
+
generated skill to `P12` before any `verified:` date advances (`docs/RISKS.md` K15).
|
|
91
|
+
**`P12` is still the mechanism to find out, and it has still never been run**: the record format
|
|
92
|
+
is now specified (`skills/goblin-eval/SKILL.md`) and **no row reads a lane** (#31).
|
|
93
|
+
Pin: `https://www.skillsbench.ai/blogs/skillsbench-1-1`, sha256 of the retrieved page
|
|
94
|
+
`d812bb7c2da702cc67556eb5f46b9e93faaca19d3545376544866e940126ef15`, fetched 2026-09-25; the
|
|
95
|
+
eleven-token ban and the judge procedure are argued in
|
|
96
|
+
`https://ai.engineer/talks/0vphxNt4wyk-don-t-ship-skills-without-evals`. **Not registered**:
|
|
97
|
+
the source registry is `manifest/sources.tsv` (G9's artifact) and that file does not exist in
|
|
98
|
+
this repo — measured, `git ls-files manifest/` names five files and none of them is
|
|
99
|
+
`sources.tsv` — so the registry row is composed in the W2 report for G9's patch instead of
|
|
100
|
+
being written into a file this card does not own.
|
|
101
|
+
16. **Every number in the design is a file read or a third-party published number; none of it is
|
|
102
|
+
a measurement of this harness under load.** The harness was built and its verifier was shown
|
|
103
|
+
RED under a deliberate break; that is a statement about the mechanism, not about outcomes.
|
|
104
|
+
17. **One spec deviation, recorded rather than hidden.** The design spec's literal HANDOFF check
|
|
105
|
+
is `grep -q "$(git rev-parse --short HEAD)" HANDOFF.md`, which **can never pass** — committing
|
|
106
|
+
the HANDOFF moves HEAD, so the file can only ever name an ancestor. `HP-05` is therefore
|
|
107
|
+
mechanised as *the HANDOFF names a commit that exists in this repo and is an ancestor of
|
|
108
|
+
HEAD*, which still catches the defect the rule exists for (an artifact that names no commit
|
|
109
|
+
at all). The deviation and its reason are in the row's own `if_not_why` column.
|
|
110
|
+
18. **`.goblin/installed.json` is not signed, so nothing here proves it was not rewritten.** Every
|
|
111
|
+
drift check — `IN-02`, `SK-02`, and `HS-01`'s hash of the harness dir — reads its expected
|
|
112
|
+
hash out of that one file, and that file is the one file no check protects. Measured: append a
|
|
113
|
+
byte to `.goblin/bin/goblin-verify`, rewrite its recorded hash in `installed.json`, commit, and
|
|
114
|
+
the run is **fully GREEN** (`43 passed, 0 failed`, exit 0). One edit defeats three rows at
|
|
115
|
+
once, and it is the cheapest way to fake a green run. Doing better needs an anchor the target
|
|
116
|
+
cannot edit — a signature, or a hash held outside the repo — and goblin-stack has no such
|
|
117
|
+
trust root: the source checkout is not guaranteed to exist at verify time, and any value
|
|
118
|
+
stored in the tree is editable by the same hand. It is therefore **recorded here and printed
|
|
119
|
+
in the "cannot see" footer on every run**, not claimed away.
|
|
120
|
+
19. **No automation has ever run here.** Cost per run, the respawn guards under a nightly
|
|
121
|
+
producer, and whether `researcher` is the right reporter profile are all unmeasured; the
|
|
122
|
+
first watched run of A-02 is what produces those numbers. The producer's own ceiling is a
|
|
123
|
+
**run count**, not a dollar figure — no config key holds the spend cap, so no automation can
|
|
124
|
+
read it, and none pretends to.
|
|
125
|
+
20. **The dedup key is a dedup, not a mutex.** The board's lookup runs before the write
|
|
126
|
+
transaction, so a concurrent create can insert twice and the next lookup stabilises on the
|
|
127
|
+
newest. The key stops a duplicate storm; it does not make one impossible.
|
|
128
|
+
21. **`AU-02`'s normalisation is only measured on synthetic reports.** Two differently-typed
|
|
129
|
+
copies of one symptom give one key and a different symptom gives another, but whether a real
|
|
130
|
+
report set normalises well enough is unknown. Its failure mode is a duplicate card, never a
|
|
131
|
+
lost report.
|
|
132
|
+
22. **No audit has ever run against a real registry here.** Everything `SC-07` does was measured
|
|
133
|
+
against a canned npm-audit report (`tests/t-audit.sh`), so the parse is proven, the *policy*
|
|
134
|
+
is not: whether the recorded waiver set matches the real advisory set is unknown until the
|
|
135
|
+
first real `goblin-audit`. Its skip-with-a-reason behaviour on a repo with no record is what
|
|
136
|
+
keeps that honest rather than silent.
|
|
137
|
+
23. **The audit record is read by field name, not by a JSON parser.** `goblin-audit` recognises
|
|
138
|
+
npm's `vulnerabilities` / `via` shape (`source`, `name`, `url`, `range`) with awk and REFUSES
|
|
139
|
+
(exit 5) anything it cannot parse, rather than writing an empty record that `SC-07` would read
|
|
140
|
+
as clean. A different audit tool is therefore a refusal, not a silent pass.
|
|
141
|
+
24. **The freshness rows need a GNU `date -d`.** `SC-07` parses the record's date that way; on a
|
|
142
|
+
host without it the row FAILS with the reason rather than assuming the record is fresh.
|
|
143
|
+
25. **`SC-04` reads one statement, not one program.** A cookie write spread over three lines (or
|
|
144
|
+
assembled through a helper) is not seen, and the row says so in its own cell.
|
|
145
|
+
26. **The last advisory slot is an open decision, not a rule.** Measured (V1): the advisory rows
|
|
146
|
+
are `HP-04`, `HS-03`, `CM-02`, `MD-02`, `MD-03`, `PG-04`, `DOC-01`, `DOC-02`, `SC-09` — 9 at a
|
|
147
|
+
ceiling of 10 — so **exactly one slot was free**, and `SK-03` reported that arithmetic at the
|
|
148
|
+
time (`advisory 9 of ceiling 10 (1 free slot)`; **corrected 2026-09-25 (AB3):** the run prints
|
|
149
|
+
`advisory 10 of ceiling 10 (0 free slots: the next advisory row FAILs)`, W3's `JG-03` having
|
|
150
|
+
taken the slot — the two notes below carry the chronology). Two planned cards each wanted the
|
|
151
|
+
slot: G1's `FM-03` (the feature map) and G2's `JG-03` (the judge agent). **Nothing in this repo
|
|
152
|
+
chooses between them**, and V1 deliberately spent nothing. The cap is a count, not a strict
|
|
153
|
+
bound: measured, 10 advisory rows at a ceiling of 10 **pass**, and the 11th FAILs (10 at a
|
|
154
|
+
ceiling of 9 FAILs). So the tenth row is allowed; the eleventh is not. Whoever lands second
|
|
155
|
+
brings a real command.
|
|
156
|
+
**Decided 2026-09-25 (W2):** G1's `FM-03` does **not** take the slot. The feature map ships
|
|
157
|
+
`FM-01` and `FM-02` as real commands, and the one thing they cannot check — whether the map
|
|
158
|
+
lists every feature — is recorded as #30 instead of as a counted row; the slot is left free for
|
|
159
|
+
G2's `JG-03`, whose judge is the mechanism `P12` actually needs. Measured after W2, `SK-03`
|
|
160
|
+
still reads `advisory 9 of ceiling 10 (1 free slot)`.
|
|
161
|
+
**Spent 2026-09-25 (W3):** G2 landed and `JG-03` took it. The slot is now **full** — measured
|
|
162
|
+
`advisory 10 of ceiling 10 (0 free slots: the next advisory row FAILs)` — so the next author who
|
|
163
|
+
wants an advisory row must raise `advisory_ceiling` in the same change and write down why,
|
|
164
|
+
rather than discovering the cap from a red run. This is not a rule change; it is the arithmetic
|
|
165
|
+
the ceiling was always meant to force into the open.
|
|
166
|
+
30. **The feature map is an inventory, and nothing checks that it is complete.** `FM-01` checks the
|
|
167
|
+
index against the feature files that exist and the four-H2 entry contract; `FM-02` is a tripwire
|
|
168
|
+
over `entry_paths:`. **Neither can see a feature nobody wrote down** — that needs semantic
|
|
169
|
+
judgement over the app, which no command here performs. Three smaller blind spots are named in
|
|
170
|
+
the rows' own why-cells and repeated here: `FM-02` searches for the token under `source_root:`
|
|
171
|
+
and therefore **reads a vendored copy, a lockfile or a build artifact as "still resolves"**; it
|
|
172
|
+
sees a *file* change and not a *behaviour* change, so a refactor that leaves the route alone
|
|
173
|
+
reports stale-and-wrong and a behaviour change in a file the token does not appear in is missed;
|
|
174
|
+
and a `verified:` date is a **claim the row cannot test** — nothing distinguishes a feature
|
|
175
|
+
driven that day from a date typed that day. The first of those is why `source_root:` should name
|
|
176
|
+
the source tree, not the repo root, and the third is why the upkeep pass in
|
|
177
|
+
`skills/goblin-feature-map/SKILL.md` requires the date to advance only for a feature someone
|
|
178
|
+
actually drove.
|
|
179
|
+
31. **The P6↔P12 loop is wired as a contract with no runner.** `P6` now hands a generated
|
|
180
|
+
verification skill to `P12` and `verified:` does not advance until an eval record exists; the
|
|
181
|
+
record's shape, the eleven-token ban, the cheap-checks-first ladder, the merge rule and the pass
|
|
182
|
+
condition (every seeded defect detected, the control's number at zero, every correction RED
|
|
183
|
+
before GREEN) are all specified in `skills/goblin-eval/SKILL.md`. **No row reads a lane, and
|
|
184
|
+
nothing executes an eval** — measured after W2, `manifest/enforcement.tsv` has no `EV-*` row,
|
|
185
|
+
and `P12` has still never been run. The runner and the record checks (`EV-01`..`EV-04` in G1's
|
|
186
|
+
design) are deferred to a follow-up card, deliberately and in the open, rather than half-built:
|
|
187
|
+
a row that reads a record nobody writes would pass vacuously and look like enforcement.
|
|
188
|
+
32. **The judge is a language model grading prose, and the record proves a handle exists — never
|
|
189
|
+
that the handle supports the verdict.** `JG-01` FAILs a `done` verdict whose evidence resolves
|
|
190
|
+
to nothing: `sha:<hex>` must be a commit in `git rev-list --all`, `file:<path>` a path under the
|
|
191
|
+
root, `sha256:<hex>` the digest of a file under `.goblin/loop/`. All three say *exists*. A judge
|
|
192
|
+
may cite a real commit that has nothing to do with the claim and pass, and a `cmd:<command>`
|
|
193
|
+
token resolves **nothing on purpose** — the command ran, its output is not in the record, and a
|
|
194
|
+
verdict resting on it is the self-report the row refuses. Two smaller blind spots are stated in
|
|
195
|
+
the row's own why-cell and repeated here: the row cannot see **which lane returned a verdict**
|
|
196
|
+
(no file in a repo observes the profile that ran — `MD-03`), and **an unresolved judge lane is
|
|
197
|
+
an `ADV`, not a failure**, because a repo cannot choose the fleet's routing (`docs/ROLES.md`,
|
|
198
|
+
"the measured caveat"): measured on this box, the judge lane resolves to no profile at all.
|
|
199
|
+
Finally, the `JG-03` counter-measure is **policy, not a check** — one known-red control verdict
|
|
200
|
+
per wave, recorded in `docs/LOOP.md`: a lane that has judged twice cannot be called always-yes,
|
|
201
|
+
and the history that would show a bad lane lives across cards and repos.
|
|
202
|
+
33. **`LP-04` measures a changed evidence pointer, which is a proxy for progress — not progress.**
|
|
203
|
+
Three consecutive verdict rows with an identical non-empty pointer and a result that is not
|
|
204
|
+
`predicate:green` is a FAIL naming the row numbers, and that is the strongest thing a repo can
|
|
205
|
+
read without running the loop. A loop that edits a file each turn to keep the pointer moving is
|
|
206
|
+
not caught, and **nothing in Hermes detects a lack of progress either**: measured in
|
|
207
|
+
`hermes_cli/goals.py`, `run_kanban_goal_loop` carries no progress state at all — its whole
|
|
208
|
+
state is `last_response`, `turns_used` and `nudged_to_finalize`, so a loop that returns
|
|
209
|
+
`continue` for the same reason nineteen times spends nineteen turns and then blocks
|
|
210
|
+
(`hermes_cli/goals.py:1636-1638`, `:1689-1696`). So the budget is the backstop, and the budget
|
|
211
|
+
has its own blind spots: `LP-03` proves the declared budget is a positive integer at or under
|
|
212
|
+
`loop_max_turns_ceiling` and that the record holds no more verdict rows than the budget — it
|
|
213
|
+
cannot see whether the budget is **affordable**, and cost is not a field the record holds
|
|
214
|
+
(neither turns nor tokens nor the per-turn auxiliary judge call). Three further "not yets" are
|
|
215
|
+
structural rather than measurable: `LP-01` proves a recorded first run exists and that its
|
|
216
|
+
timestamp is at or before the first log row — **not that the command ran and not that the
|
|
217
|
+
`exit=` value was measured rather than typed** (`HP-03`'s defect, one artifact over); `LP-02`
|
|
218
|
+
proves the predicate file still hashes to its pin — **not that the predicate is the right one,
|
|
219
|
+
and not who edited it**; and a predicate that was **vacuously true from the start** (a `grep -c`
|
|
220
|
+
against a renamed directory) passes `LP-01` and `LP-02` and ends the loop green on nothing.
|
|
221
|
+
`LP-05` makes a write-up mandatory and therefore visible — nothing can make it true.
|
|
222
|
+
|
|
223
|
+
## What the harness refuses to do
|
|
224
|
+
|
|
225
|
+
It does not claim a green run means the work is right. `goblin-verify` asserts that the installed
|
|
226
|
+
files are the files on disk, that every rule with a command still passes, and that the
|
|
227
|
+
untestable remainder is counted and capped — and it prints, on every single run, what it cannot
|
|
228
|
+
see: the five upstream blind spots, plus the ban lane's own (the unsigned ban table #28, the
|
|
229
|
+
text-probe gap #27, and a ban that is invisible until verify runs, `V3-1`), plus the judge/loop
|
|
230
|
+
lane's (#32: a handle that exists is not a handle that supports the verdict; #33: a changed
|
|
231
|
+
pointer is a proxy for progress, not progress), plus the CI lane's (#34: a file is not a gate —
|
|
232
|
+
the required-check list, the bypass switch and the push identity are forge state; #35: the
|
|
233
|
+
Electron perf number is a host gate, and the ratchet deliberately carries a different metric).
|
|
234
|
+
|
|
235
|
+
27. **The ban probes are text probes, not ASTs.** `BN-01`, `BN-02`, `BN-03` and `BN-05` are
|
|
236
|
+
`grep` over source under `bash`/`grep`/`awk` only — the dependency contract in
|
|
237
|
+
`docs/CONTRACTS.md` allows no parser and no `npm`. So a `: any` inside a string or a comment
|
|
238
|
+
is reported, `Record<string, any>` (no leading colon) is missed, and BN-05 does not resolve
|
|
239
|
+
module aliases or dynamic imports. The AST-grade form of the same bans (BN-01..BN-04 in
|
|
240
|
+
`G5.md` §C.2) needs ESLint and `dependency-cruiser`; that is why **G5's `BN-04` (the nine
|
|
241
|
+
named unnecessary-effect patterns) is NOT shipped** — it cannot be mechanised without a
|
|
242
|
+
parser, and a ban that cannot go red is worse than advisory, so it is recorded here rather
|
|
243
|
+
than as a row that would cost the last advisory slot. A text probe with a stated
|
|
244
|
+
false-positive set is still a gate: each BN row is shown RED under its own violation and
|
|
245
|
+
GREEN when it is removed (`tests/t-verify-red.sh`).
|
|
246
|
+
28. **The ban table's integrity rides on `IN-02`, and `IN-02` rides on an unsigned record.** A
|
|
247
|
+
project could empty `manifest/bans.tsv` (or edit a `detect` command) and the ban gate would
|
|
248
|
+
pass vacuously — `BN-00` fails closed on an *empty* table, but it cannot see a table whose
|
|
249
|
+
rows were weakened, because `.goblin/manifest/bans.tsv` is hashed by `IN-02` and
|
|
250
|
+
`.goblin/installed.json` is not signed (`docs/LIMITS.md` #18). The ban list is not
|
|
251
|
+
tamper-proof; it is as strong as the record every drift check trusts. **Measured (W5-12):** with
|
|
252
|
+
`BN-01`'s `detect` cell set to `true`, `BN-00` and `BN-01` both PASS and the only row that
|
|
253
|
+
changes verdict is `IN-02`'s drift check — the `W5-12` control in `tests/t-verify-red.sh` pins
|
|
254
|
+
exactly that, and nothing else can. Z1's verdict is to **record this, not fix it**: the `detect`
|
|
255
|
+
cell is executable content, and a row able to judge whether another row's command *means*
|
|
256
|
+
something would be an `eval` over the matrix, which the same card rules out. The gap is now
|
|
257
|
+
stated, measured and controlled one row over, which is the most this table can do without
|
|
258
|
+
becoming an interpreter.
|
|
259
|
+
29. **The perf ceiling must EQUAL the recorded baseline, so a budget with headroom is not
|
|
260
|
+
expressible.** `PF-01` FAILs when `ratchet.ceiling` and `perf.baseline_value` disagree
|
|
261
|
+
(G8-6b), which is what stops a one-line ceiling raise from passing while printing the
|
|
262
|
+
contradiction. The cost is real and the row cannot see it: a project that wants the ratchet
|
|
263
|
+
to allow, say, 10% growth over the measured baseline cannot write `ceiling: 44000` beside
|
|
264
|
+
`baseline_value: 40000` — the row reads that as a disagreement. Headroom is expressed by
|
|
265
|
+
re-anchoring BOTH (measure on a pinned commit, then set `baseline_value` and `ceiling` to the
|
|
266
|
+
new number together, and record it as an operator action), which is the deliberate re-anchor
|
|
267
|
+
the row's own why-cell names. It is a bound on the budget's shape, not a claim that the
|
|
268
|
+
budget is the right one — and it still never re-measures.
|
|
269
|
+
34. **The CI lane reads files, and a file is not a gate.** The shipped workflow
|
|
270
|
+
(`templates/ci/goblin-gate.yml.tmpl` → `.github/workflows/goblin-gate.yml`) has no `if:` at any
|
|
271
|
+
level, but nothing in a repository can make GitHub **require** it: the required-check list,
|
|
272
|
+
the bypass switch and the push identity are forge state (four settings, `docs/CI.md` §1).
|
|
273
|
+
`PG-06` proves the whole declared gate set is invoked in a file under `.github/workflows/` and
|
|
274
|
+
cannot prove that file is the one the forge waits on; `PG-05`'s reader has no parser, so a
|
|
275
|
+
flow-style `jobs: {…}` mapping is refused, and a `#` inside a quoted string truncates the line
|
|
276
|
+
it is on. Two further measured gaps: the template's job is `ubuntu-latest` with no cache, so a
|
|
277
|
+
repo whose gate needs a display, a licence, a GPU or a signed-in session cannot use it at all
|
|
278
|
+
(that is a **host** gate — the class-C rule, restated for class F), and a private repo's
|
|
279
|
+
Actions minutes are billed to the account (2,000/month free). Measured ground truth at W4:
|
|
280
|
+
**one** first-party workflow exists in the whole estate and it self-skips; five of the six
|
|
281
|
+
repos with a remote have none.
|
|
282
|
+
35. **The Electron perf number is a host gate, and the ratchet carries a different metric.**
|
|
283
|
+
`presets/F-electron.yaml` declares `main_thread_busy_pct` as `perf_host_gate:` and uses
|
|
284
|
+
`app_bundle_bytes` for `ratchet:` — a **deliberate deviation** from G6 §B.3, which put the FPS
|
|
285
|
+
number in the ratchet. The instrument that produces it (CDP `Performance.getMetrics`, or
|
|
286
|
+
`app.getAppMetrics()[i].cpu.percentCPUUsage` inside a real Electron) needs Playwright or
|
|
287
|
+
Electron plus a GUI, and a shipped rule may use nothing but bash/git/awk/sed/grep/python3
|
|
288
|
+
(`docs/CONTRACTS.md`) — so `ratchet.cmd` pointing at the probe would make a fresh install
|
|
289
|
+
**born RED**, which is the one thing the install path must not produce. The probe belongs to
|
|
290
|
+
the project; a project whose CI needs npm runs it in **its own** workflow. What the number
|
|
291
|
+
means has a limit of its own, and it is measured: this box is an LXC with no display and no
|
|
292
|
+
system Chromium, so the G6 sweep came from a *bundled headless* Chromium with no compositor
|
|
293
|
+
and no vsync. Relative comparisons on one machine are meaningful (which is why
|
|
294
|
+
`main_thread_busy_pct` works at all, and why frame time does not — p50 stayed flat at 16.70 ms
|
|
295
|
+
while the main thread went from 1.8 % to 54.5 % busy, `docs/CI.md` §3.3); an absolute FPS
|
|
296
|
+
claim is not. Three further Electron failure modes are **recorded, not mechanised**, and
|
|
297
|
+
`docs/CI.md` §4 says why: the dependency-graph boundary check, `ipcMain` sender validation,
|
|
298
|
+
and fuses at package time.
|
|
299
|
+
36. **A ban's exemption reaches the probe through its environment, so a custom probe can ignore it.**
|
|
300
|
+
`bans_exempt:` and the inline `// BAN-OK(<id>): <reason>` are filtered *before* the exit code is
|
|
301
|
+
chosen, because a filter applied to a probe's stdout afterwards cannot change a verdict — that
|
|
302
|
+
was W5-1, and the old code made every exemption a permanent RED. The engine therefore exports
|
|
303
|
+
`GOBLIN_BANS_ID` and `GOBLIN_BANS_EXEMPT` and the two shipped probes honour them. A project's
|
|
304
|
+
**own** `detect` command that ignores the variables keeps the old behaviour: a violation inside
|
|
305
|
+
an exempted path stays RED. That direction is **closed**, never open, and it is the trade this
|
|
306
|
+
choice makes. The engine's own stdout is no longer filtered at all, so a probe that ignores the
|
|
307
|
+
contract prints the exempted lines it reported — loud, and still a FAIL.
|
|
308
|
+
37. **`FM-02` refuses the harness, not every non-source file.** The search skips `.git/`,
|
|
309
|
+
`.goblin/`, `.hermes/`, the declared `harness_dir` and the map's own directory, so a stub map
|
|
310
|
+
whose token occurs only in the install no longer resolves (W5-4: `entry_paths: [export]` used to
|
|
311
|
+
"resolve" to `./.goblin/bin/goblin-verify`). It does **not** exclude the target's own `docs/`,
|
|
312
|
+
`tests/` or build output: a token that occurs only there still reads as resolved, and excluding
|
|
313
|
+
them would be a guess about a layout goblin-stack does not know. Nor is there a frequency bound
|
|
314
|
+
— a common token ("export", "main") is satisfied by the first of hundreds of files.
|
|
315
|
+
38. **The judge lane's model family is reported, never enforced — and on this box it is not even
|
|
316
|
+
mapped yet.** `JG-02` proves the declared profile *names* are disjoint; W5-6 measured that a
|
|
317
|
+
profile mapped to the author's own model passed it. `MD-02` now resolves the judge lane and
|
|
318
|
+
compares its model with the code lane's, so the state is loud — but it stays `advisory`
|
|
319
|
+
(`return 2`, never a FAIL): goblin-stack cannot choose the fleet's models, and a repo-local file
|
|
320
|
+
cannot observe *which* model a lane actually ran. Measured on this box at X1: Harvey's
|
|
321
|
+
`fleet-model.yaml` names **no `judge:` profile at all**, so `MD-02` reports the judge lane
|
|
322
|
+
**unresolved** and prints the one-line remedy, while every profile it *does* name — nine of
|
|
323
|
+
them, the named per-fleet role profiles this setup routes its lanes through —
|
|
324
|
+
resolves to one model under an active promotion. So on this box the honest reading is: the judge
|
|
325
|
+
lane is not mapped, and the moment it is mapped it will be the author's own family. The
|
|
326
|
+
comparison is exact model equality, not a version-stripped "family": two spellings of the same
|
|
327
|
+
family that differ only in a suffix would read as different.
|
|
328
|
+
39. **`LP-02` cannot tell a weaker predicate from a re-scope. The close-and-reopen is recorded, not
|
|
329
|
+
prevented.** A loop could archive its bar under `closed-<date>/`, write a weaker one, re-pin,
|
|
330
|
+
and pass `LP-02` and `LP-05` with nothing in the record (W5-7, measured). The row now requires
|
|
331
|
+
every archive to hold its predicate **and** the pin it was closed under, and the live
|
|
332
|
+
`predicate.sha256` to name the archived digest on a `previous:` line — so a **silent** relaxation
|
|
333
|
+
is caught and a legitimate re-scope costs one line. Whether the new predicate is weaker, and who
|
|
334
|
+
edited it, is not decidable from a digest, and a predicate that calls a script elsewhere is
|
|
335
|
+
pinned only at its call site. The residual is the same one `LP-01` carries one file over: the
|
|
336
|
+
record is evidence, not proof that the loop stopped for the right reason.
|
|
337
|
+
40. **A fresh install into a repo with no commits is born RED, and X1 did not change that.** Measured
|
|
338
|
+
at X1: `goblin-install --class A` into a `git init` with zero commits, then `goblin-verify`, gives
|
|
339
|
+
`37 passed, 6 failed, 11 advisory, 24 skipped`, exit 1 — `HP-05`, `SP-02`, `GT-02`, `CM-01`,
|
|
340
|
+
`CM-03` and `PT-02` all read a HEAD that does not exist yet. The install never creates the seed
|
|
341
|
+
commit (`bin/goblin-install` writes files and stops), and it must not: a tool that commits into
|
|
342
|
+
Harvey's repo on first contact is the overreach `docs/CONTRACTS.md` rules out. Carried from W5
|
|
343
|
+
§5's `G8-9` rather than fixed here — it is a **sequencing** limit, not a hole in a row: one commit
|
|
344
|
+
clears all six, and `tests/t-verify-green.sh` seeds one before it installs.
|
|
345
|
+
41. **Eleven of Y1 §7's eighteen documented mechanisms still carry no control of their own, and one
|
|
346
|
+
exemption is by design.** The census, stated separably so a reader can check the arithmetic:
|
|
347
|
+
**18 listed · 2 fixed · 2 closed · 3 controlled · 11 recorded.** Y1 §7 listed eighteen
|
|
348
|
+
mechanisms; Z1 **fixed** the two load-bearing ones (17, the unrendered token — Z1-3; 18,
|
|
349
|
+
`replay.cmd` — Z1-4, which now have controls and therefore sit *outside* the "carry no control"
|
|
350
|
+
set) and **closed** two more (4, the ban engine's `exit 2` paths; 5, `--list` — three assertions
|
|
351
|
+
in `tests/t-verify-red.sh`, RED against a deliberately broken copy of `bin/goblin-bans`, which
|
|
352
|
+
is PR-03's second branch: those mechanisms *worked*, they were just unguarded); AA1
|
|
353
|
+
**controlled** three (1, 2 and 6 — the ban-engine cluster, below). The remaining **eleven** are
|
|
354
|
+
**recorded here rather than mechanised**. Each was measured WORKING, so an assertion would pin
|
|
355
|
+
behaviour that already holds, and this box does not have eleven controls' worth of budget; the
|
|
356
|
+
cost of the gap is exactly the shape of both regressions in this repo's history — a mechanism
|
|
357
|
+
documented, and nothing asserting it.
|
|
358
|
+
- items 1, 2, 6 — `bans_exempt:` on the layer probe, the engine→probe
|
|
359
|
+
`GOBLIN_BANS_ID`/`GOBLIN_BANS_EXEMPT` contract, and its segment alignment: **CONTROLLED at
|
|
360
|
+
0.4.2 (AA1)** — seven controls in `tests/t-verify-red.sh`, in both directions each (the
|
|
361
|
+
layer probe's exempt path PASSes and the same crossing import outside it FAILs; a prefix
|
|
362
|
+
covers its own subtree but does not swallow `src/okay`; a prefix written with a trailing
|
|
363
|
+
slash exempts nothing; and a project's own probe that exits 0 only when both environment
|
|
364
|
+
variables arrive). This was the cluster Z1 named as the one with a real engine underneath,
|
|
365
|
+
and `grep -c GOBLIN_BANS_EXEMPT tests/` was **0** before it.
|
|
366
|
+
- item 3 — a project's own probe that ignores those variables fails closed: the mechanism is
|
|
367
|
+
this file's #36, where the FAIL is measured.
|
|
368
|
+
- item 7 — `BAN-OK` must sit on the offending line: the marker's *effect* is controlled, the
|
|
369
|
+
on-this-line clause is not (a marker on another line does not clear the violation).
|
|
370
|
+
- items 8, 9 — `SC-05`'s `.goblin/boundary-waivers` and `SC-08`'s
|
|
371
|
+
`.goblin/install-hooks.allowlist`: both files are copied and both FAIL directions are
|
|
372
|
+
controlled; neither *honoured* direction is. (`SC-08`'s reader is the row the minified
|
|
373
|
+
lockfile defeated — Z2-2; that was a shape, not this item, and it is fixed and controlled.)
|
|
374
|
+
- item 10 — `SC-03` clause 2, `sec_build_output`: a literal in `dist/assets` FAILs; only the
|
|
375
|
+
clause-1 path is controlled.
|
|
376
|
+
- items 11, 12, 15 — `scaffold_checks:`'s SKIP branch, `perf_host_gate` (#35) and
|
|
377
|
+
`templates/loop/*.tmpl`: declared no-ops (SKIP, a host gate, and templates `docs/LOOP.md`
|
|
378
|
+
says nothing installs). A control here would assert that nothing happens.
|
|
379
|
+
- item 13 — `docs/CI.md`'s flow-style `jobs: {…}` refusal: fails closed already, so a control
|
|
380
|
+
would pin the strict direction of a check that cannot pass vacantly.
|
|
381
|
+
- item 14 — `bin/goblin-model`: works (`code` → the resolved lane, unknown role → exit 2);
|
|
382
|
+
only its *absence* from an install is asserted.
|
|
383
|
+
- item 16 — `docs/LOOP.md` §6's "one known-red control verdict per wave, recorded in this
|
|
384
|
+
file": a prose obligation, not a command. It is honoured in the write-ups (or not) and no
|
|
385
|
+
check can read a wave.
|
|
386
|
+
**W5-10 is answered here too.** The tenant strings `PT-01` forbids do not reach `docs/`, because
|
|
387
|
+
`docs/` is never installed: the row's directory list is `skills manifest bin templates presets
|
|
388
|
+
.goblin .hermes`, so a string under `docs/` is source-tree prose that no operator's repo ever
|
|
389
|
+
receives. That is the whole reason, it is deliberate, and the row is not weakened by it — Y1
|
|
390
|
+
agreed, and Z1 leaves it. This is the sentence Z1 added so the reason is visible in the shipped
|
|
391
|
+
artifact rather than only in the wave's own notes.
|
|
392
|
+
|
|
393
|
+
42. **The doc walk cannot read a path broken at the directory/name boundary.**
|
|
394
|
+
`tests/t-doc-promises.sh` reads a command directory only as `.goblin/bin/` or `bin/` **with its
|
|
395
|
+
trailing slash**; when a line ends on the bare directory (`.goblin/bin`, `bin`) and the slash
|
|
396
|
+
leads the next line (`/goblin-doctor`) or is dropped (`goblin-doctor`), the directory and its
|
|
397
|
+
continuation are both invisible — the first fragment stops before the name the grammar needs, the
|
|
398
|
+
second is not preceded by `bin/`, so neither is a token and neither is asserted. Measured
|
|
399
|
+
2026-09-26 (AB7): a plant of that form in `docs/CI.md` is reported **`PASS`, rc 0**, mentioning
|
|
400
|
+
the plant **zero** times, and a line that ends on the bare directory with nothing after it is
|
|
401
|
+
silent too, because the `dangling()` guard needs the trailing slash as well. **0 live instances**
|
|
402
|
+
— measured: `grep -rnE '\.goblin/bin$|bin/$' $(git ls-files)` → **no output, exit 1**. **What it
|
|
403
|
+
costs:** a false path written in that form — the exact defect this file's whole species is named
|
|
404
|
+
for — passes the watch silently, so a future document could hand a reader a command that does not
|
|
405
|
+
exist and nothing in the walk would say so. **Ticketed, not gated:** reading the form is a change
|
|
406
|
+
to the tokeniser inside `tests/t-doc-promises.sh`, not to a document, and it is a separate card;
|
|
407
|
+
this entry records the boundary so a green run cannot imply a coverage it does not have. The
|
|
408
|
+
control's own header now names the form as uncovered instead of claiming a sensitivity it does
|
|
409
|
+
not have, and the slash's load-bearing role is stated where the dangling rule is.
|
|
410
|
+
|
|
411
|
+
43. **The engine footer names its judge, and the statement is unsigned — one forgery covers N
|
|
412
|
+
repos.** W1's engine split prints `cli_sha256` and `enforcement_tsv_sha256` on every run
|
|
413
|
+
(and `installed.json` gains an optional `engine:` block recording the same pair), so a repo
|
|
414
|
+
can state WHICH engine judged it. Nothing signs either hash: the footer is printf output of
|
|
415
|
+
the very binary it names, and the `engine:` block is a JSON stanza inside the same unsigned
|
|
416
|
+
record `docs/LIMITS.md` #18 already covers. An edited engine, or a hand-written block
|
|
417
|
+
claiming a mode that was never resolved, prints whatever it likes — and because the
|
|
418
|
+
statement now rides in every repo's run, one forgery propagates to every repo that trusts
|
|
419
|
+
it, which is strictly worse than #18's per-repo record. The old defences still hold and are
|
|
420
|
+
what the footer must not be read to replace: `--source` pins the engine by flag, a vendored
|
|
421
|
+
engine keeps beating the global one in the resolution chain, and IN-02 still hashes the
|
|
422
|
+
repo-local bytes. What is NOT fixed (G6, the no-signing non-goal): no trust root, no
|
|
423
|
+
detached signature, no third-party attestation. Measured with W1: tamper one byte of the
|
|
424
|
+
vendored manifest and the footer's `enforcement_tsv_sha256` moves — the footer is a real
|
|
425
|
+
fingerprint of what ran, it is just not PROOF of it.
|
|
426
|
+
|
|
427
|
+
## Verdicts recorded, not built (the note-8 questions)
|
|
428
|
+
|
|
429
|
+
Two library questions were investigated to a verdict and **deliberately built nothing here**, so a
|
|
430
|
+
later session does not re-derive them. The measurements and sources are in
|
|
431
|
+
`goblin-stack-research/G6.md` Part C; this is the durable half.
|
|
432
|
+
|
|
433
|
+
**Pretext (`chenglou/pretext`) — USE, narrowly, and not as a runtime dependency.** Verified from
|
|
434
|
+
the source of truth: a pure JS/TS library for multiline text measurement and layout, MIT, that
|
|
435
|
+
*"side-steps the need for DOM measurements (e.g. `getBoundingClientRect`, `offsetHeight`), which
|
|
436
|
+
trigger layout reflow"*, using the browser's own font engine as ground truth. Author confirmed
|
|
437
|
+
(Cheng Lou, `_npmUser: chenglou`); the README credits Sebastian Markbage's earlier `text-layout`,
|
|
438
|
+
whose repository now says *"This project is archived. The ideas here evolved into Pretext"*.
|
|
439
|
+
Two verified corrections to the shipped skill: the skill pins `@chenglou/pretext@0.0.6` while npm's
|
|
440
|
+
latest is **`0.0.9`**, and *"15KB zero-dependency"* describes the **runtime**, not the install
|
|
441
|
+
(published tarball `unpackedSize: 887142` across 69 files, `sideEffects: false`, subpaths `.` and
|
|
442
|
+
`./rich-inline`). The concrete use is a **dev-time / harness-time label-fit check** for
|
|
443
|
+
`diagram-studio` and `goblin-ui` — "does this string fit this box at this font", without a layout
|
|
444
|
+
read, as a `checks/*.mjs` assertion with a REPLAY — in exactly the repos that today hand-roll that
|
|
445
|
+
arithmetic (26 `getBoundingClientRect()` calls, no `measureText` call, no label-overflow probe).
|
|
446
|
+
**Do not** put it in the runtime bundle: the app is offline by design, and the skill's own stack
|
|
447
|
+
table imports it through a CDN, which contradicts that. The second use is one blog demo, not
|
|
448
|
+
infrastructure.
|
|
449
|
+
|
|
450
|
+
**mise (`jdx/mise`) — NOT NOW, with a named trigger that flips it.** Verified: mise-en-place, MIT,
|
|
451
|
+
macOS/Linux/Windows, one CLI that declares tool versions, environment variables and commands in
|
|
452
|
+
`mise.toml` and uses them in the shell, the editor and CI; the polyglot successor to
|
|
453
|
+
asdf/nvm/pyenv, and `.tool-versions` already works; installed with `curl https://mise.run | sh`.
|
|
454
|
+
The premise for adopting it did not survive measurement: the projects do **not** run differing Node
|
|
455
|
+
versions, and nothing declares a version at all — 0 `.nvmrc`/`.node-version`/`.tool-versions` files
|
|
456
|
+
under `~/projects`, `engines.node` in exactly **one** first-party manifest, and a single Node on
|
|
457
|
+
the box (v22.23.1, reached through `~/.local/bin`, which a shell profile puts first). The one real
|
|
458
|
+
divergence is a *package-manager* split, which mise cannot resolve. Adopting it now would add a
|
|
459
|
+
second version source beside the Node Hermes bundles. **Adopt when either becomes true:** two
|
|
460
|
+
projects need different Node majors, or the first Electron app lands (Electron ships its own
|
|
461
|
+
Node/Chromium, so the host Node stops mattering for the shell and starts mattering for the build
|
|
462
|
+
tooling). **The cheap thing to do meanwhile:** declare the version that already exists — one
|
|
463
|
+
`engines.node` line where it is missing, and a `measured <date>` gate line in the HANDOFF naming
|
|
464
|
+
the Node the gates ran under.
|
|
465
|
+
|
|
466
|
+
44. **The version statement is a convention, not an enforcement row — the sync is test-side and
|
|
467
|
+
the engine itself never checks it.** W2 closed the measured hole PLAN-V1 §4.4 recorded (no
|
|
468
|
+
test read any of the five `GOBLIN_*_VERSION` constants in `bin/`): `tests/t-version-sync.sh`
|
|
469
|
+
now pins the count of those constants at 5, asserts every constant and
|
|
470
|
+
`package.json.version` equals `VERSION`, and asserts `goblin --version` — through both the
|
|
471
|
+
bash CLI and the node shim — prints `VERSION` byte-for-byte. What remains open: the sync
|
|
472
|
+
lives only in this repo's test suite, which a target repo never runs — and when the npm
|
|
473
|
+
tarball exists (W3/W5 packaging, planned to ship no `tests/`, deliberately), a published
|
|
474
|
+
package's `goblin.js` and its bash payload could
|
|
475
|
+
drift apart with nothing in the shipped artifact noticing. The engine has no self-check row
|
|
476
|
+
that reads its own `--version` against a manifest record, and adding one would make the
|
|
477
|
+
version a rule — which is a real option, not done here. Until then the guarantee is
|
|
478
|
+
development-time only: green in this checkout, unverifiable in the wild.
|
|
479
|
+
|
|
480
|
+
45. **The migration is crash-safe by sequence, not by journal — a kill mid-upgrade leaves a shape
|
|
481
|
+
only three of which are detected.** `goblin upgrade` runs eight steps across two commits, and
|
|
482
|
+
the spec's refusal conditions name the shapes a crashed run can leave: record says global but
|
|
483
|
+
the 18 files are still here (R8 catches it), the record lost its engine: block entirely (R9
|
|
484
|
+
catches it), commits half-landed (the preflight's own state resolution). What is NOT caught:
|
|
485
|
+
a crash between the engine landing (step 3) and commit A (step 4) leaves `~/.goblin/engine`
|
|
486
|
+
written but the repo untouched — harmless, invisible, and never re-verified (the next upgrade
|
|
487
|
+
run compares hashes and refuses R6 if the payload has since changed, so the stale engine can
|
|
488
|
+
sit there being wrong until someone looks). And the shadowing footer (W3 §3) reads only the
|
|
489
|
+
vendored-payload shape; a repo whose engine_dir points at an engine that no longer exists
|
|
490
|
+
FAILs the resolution chain loudly (exit 2), but a repo pointing at an engine whose bytes have
|
|
491
|
+
silently changed since migration day runs green on the drifted table — the footer's hashes
|
|
492
|
+
name what ran, nothing compares them to migration day's. **Ticketed, not gated:** a
|
|
493
|
+
migration-day hash pin in the record compared per-run is the real fix and is a row-shaped
|
|
494
|
+
change (a new IN clause), not a W3 patch. Measured: R6 refuses the wrong-engine reuse, R8/R9
|
|
495
|
+
catch the two crash shapes they name, and the between-steps engine write is unwatched.
|
|
496
|
+
|
|
497
|
+
46. **The emitted platform config is unsigned and hand-editable — DRIFT is detected per run,
|
|
498
|
+
never prevented.** `goblin emit` writes skills byte-copies and one delimited block in the
|
|
499
|
+
platform's context file, recorded in `~/.goblin-stack/emissions.tsv` with pre-image hashes
|
|
500
|
+
(so `--uninstall` restores byte-exactly), but nothing signs what it writes: an edit to an
|
|
501
|
+
emitted `SKILL.md` or to the bytes inside the `goblin-stack:begin/end` block makes the
|
|
502
|
+
platform's `verify.sh` oracle report DRIFT on the next `goblin doctor` run — and that is
|
|
503
|
+
all it does. There is no lock, no signature, and no write protection on any emitted file;
|
|
504
|
+
a platform (or the user) can change them between two doctor runs and nothing notices
|
|
505
|
+
until someone runs one. The same holds for `--unshadow`: it removes only project copies
|
|
506
|
+
whose hash equals the source payload and refuses-and-names any that differ, so a real
|
|
507
|
+
local edit survives, but nothing reconciles it either. Measured: a one-byte tamper in an
|
|
508
|
+
emitted `SKILL.md` and a stale marker VERSION in the context block each report DRIFT (exit
|
|
509
|
+
1) on the next run, and uninstall refuses to delete a recorded file whose bytes no longer
|
|
510
|
+
match its post-image (R6).
|
|
511
|
+
|
|
512
|
+
47. **The four W4b adapter conventions are documented shapes, not run-probed installs — and
|
|
513
|
+
two of the seven platforms cannot block commands outright.** cursor and codex are not
|
|
514
|
+
installed on the build machine, so their rows pin the official docs (read 2026-09-29),
|
|
515
|
+
not a live CLI; a platform changing its layout invalidates the adapter silently until a
|
|
516
|
+
doctor DRIFT names it. codex and gemini report cap_command_blocking `partial` — codex
|
|
517
|
+
disables skills via `~/.codex/config.toml` `[[skills.config]]`, gemini only narrows via
|
|
518
|
+
approval modes — so an emitted skill is *available* there even when the operator would
|
|
519
|
+
forbid it; the doctor prints the codex hooks caveat (sessionStart only). And gemini's
|
|
520
|
+
id cell was measured to be exactly its platform name (the `gem_ini` typo shipped in one
|
|
521
|
+
W4b build and the doctor's DRIFT caught it — the schema check works).
|
|
522
|
+
|
|
523
|
+
48. **`SP-01`'s rule text once overclaimed; it now states the check's actual scope.** The text
|
|
524
|
+
used to read "The current round has a SPEC" while the check is
|
|
525
|
+
`ls ./*-SPEC.md >/dev/null 2>&1`, which passes on ANY spec-shaped file at the repo root. The
|
|
526
|
+
measured consequence: a fresh class-A install ships the scaffold's `ROUND-000-SPEC.md`, and
|
|
527
|
+
`SP-01` PASSes on that scaffold alone — no round is open, and the row is green anyway. A stale
|
|
528
|
+
spec from a finished round keeps satisfying the row exactly as well as a live one, because the
|
|
529
|
+
check has no notion of "current". **Fixed in W6 (text, not check):** the row now reads
|
|
530
|
+
"A *-SPEC.md file exists at the repo root (any round, not the current one - round-scoping
|
|
531
|
+
arrives with the W6 staged chain)" — manifest and `docs/ENFORCEMENT.md` re-rendered in the
|
|
532
|
+
same commit, the check cell untouched. What it costs: the SPEC-exists signal in a run summary is
|
|
533
|
+
weaker than a round-scoped claim would suggest — read it as "a `*-SPEC.md` file is present", not "this round's spec is
|
|
534
|
+
here". Measured: `goblin-verify` on a fresh probe install reports `PASS SP-01 (ls
|
|
535
|
+
./*-SPEC.md >/dev/null 2>&1)` with `ROUND-000-SPEC.md` the only file the glob sees, and the
|
|
536
|
+
same PASS after renaming it to a non-round name. **Ticketed, not gated:** "current round" is a stage-order notion, and the W6 staged workflow chain (SC-01:
|
|
537
|
+
a SPEC committed before the changes it governs) is what gives the word meaning; the staged
|
|
538
|
+
chain's stage-order rows are what will make round-scoping REAL — the reword names the
|
|
539
|
+
boundary until then.
|
|
540
|
+
|
|
541
|
+
49. **The engine footer prints to captured stdout, and a script parsing `goblin-verify` output
|
|
542
|
+
must expect it.** The two-line footer (`engine: mode=… cli_sha256=… enforcement_tsv_sha256=…`)
|
|
543
|
+
is unconditional `printf` output — there is no isatty guard, by design, so a captured run still
|
|
544
|
+
names the engine that judged it (that is the point of #43's statement). The cost is parser
|
|
545
|
+
friction: a script that treats every stdout line as a verdict line trips on the banner and
|
|
546
|
+
footer lines, which are not `PASS`/`FAIL` rows. The contract is the EXIT CODE, never the line
|
|
547
|
+
set: 0 green, 1 red, 2 refuse — parse the exit code, or filter to `^[A-Z]{2}-[0-9]{2}` before
|
|
548
|
+
reading lines. Measured: `./.goblin/bin/goblin-verify > out.txt` on a probe install puts
|
|
549
|
+
`engine: mode=vendored cli_sha256=… enforcement_tsv_sha256=…` in the captured file
|
|
550
|
+
(`tests/t-engine-dir.sh` itself asserts the footer IN captured output, so removing it or
|
|
551
|
+
tty-gating it would break the engine's own suite). Recorded so the pitfall is findable from
|
|
552
|
+
this file; no change to the engine is implied or wanted.
|
|
553
|
+
|
|
554
|
+
50. **`CM-01` reads only the most recent commit — historical commits with a wrong identity pass
|
|
555
|
+
unseen. Current stated scope: the row gates the identity of HEAD at the moment of the run,
|
|
556
|
+
and nothing older; that is the row's whole claim, by design.** The check is
|
|
557
|
+
`test "$(git log -1 --format='%ae')"` against the configured
|
|
558
|
+
`owner_email`, so it gates the identity of HEAD at the moment of the run and nothing older.
|
|
559
|
+
Measured: a probe history `owner → wrong@old.co → owner` reports `--only CM-01` clean (exit 0)
|
|
560
|
+
with the wrong-identity commit sitting one below HEAD. That is consistent with the
|
|
561
|
+
gate-at-the-moment design — every row judges the tree and history as they stand when verify
|
|
562
|
+
runs, and retrofitting an identity sweep over `git log --all` is a policy change, not a bug
|
|
563
|
+
fix — but it should be named: the row's green means "the latest commit carries the owner
|
|
564
|
+
identity", never "no commit in this repo's history carries an ambient one". A wrong-identity
|
|
565
|
+
commit that has since been followed by correct ones is invisible to every run. Recorded as a
|
|
566
|
+
boundary; no row change implied. `docs/ENFORCEMENT.md`'s CM-01 row carries the same one-line
|
|
567
|
+
scope statement.
|
|
568
|
+
|
|
569
|
+
51. **`gob init` writes its three added values into `goblin.yaml` itself, not through a
|
|
570
|
+
template.** Install renders `.goblin/goblin.yaml` from `templates/goblin.yaml.tmpl` and
|
|
571
|
+
owns it (`put_once`: never rewritten after the first install) — and install correctly
|
|
572
|
+
carries no `--branch/--email/--gate` flags, because those keys are the project's to edit.
|
|
573
|
+
The wizard therefore `sed`-patches `branch:` and `owner_email:` and rewrites the first
|
|
574
|
+
gate's `cmd:` line right after install renders the file, fail-closed: the gate must read
|
|
575
|
+
back through the engine's own `g_yaml_gates` or the run stops. The cost is a second
|
|
576
|
+
writer for exactly those three lines: a hand-customised comment placement survives, but a
|
|
577
|
+
future template change to those lines' shapes (renamed keys, a multi-gate default) must
|
|
578
|
+
be mirrored in `bin/goblin-init`. The wizard never rewrites anything outside the declared
|
|
579
|
+
keys, and a re-run after a hand edit of another line leaves that line alone. Recorded as
|
|
580
|
+
a boundary; the alternative — teaching install three wizard-only flags — would put wizard
|
|
581
|
+
vocabulary into the installer's contract for no gain.
|
|
582
|
+
|
|
583
|
+
52. **`scope:source` and `scope:target` name WHOSE burden a row carries — the framework's or
|
|
584
|
+
the adopting repo's.** `scope:source` is goblin-stack's own proof burden: those rows
|
|
585
|
+
(`PR-01`..`PR-05`) are the framework testing ITSELF while it is being developed — they run
|
|
586
|
+
in this repo, under `tests/run-tests.sh`, and a dev of goblin-stack is the one who owes the
|
|
587
|
+
run. `scope:target` is the adopting repo's proof burden: those rows run in an installed
|
|
588
|
+
repo via `goblin-verify`, and the repo's owner owes the run. Same matrix, two creditors:
|
|
589
|
+
a source row can never fail a user's repo, and a target row can never substitute for the
|
|
590
|
+
framework's own suite. Recorded as a definition; `docs/ENFORCEMENT.md`'s scope paragraph
|
|
591
|
+
carries the same sentence for the reader who arrives there first.
|