@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +351 -0
- package/LICENSE +21 -0
- package/README.md +217 -2
- package/VERSION +1 -0
- package/adapters/_template/adapter.tsv +16 -0
- package/adapters/_template/detect.sh +10 -0
- package/adapters/_template/emit.sh +5 -0
- package/adapters/_template/verify.sh +4 -0
- package/adapters/claude/adapter.tsv +8 -0
- package/adapters/claude/detect.sh +8 -0
- package/adapters/claude/verify.sh +47 -0
- package/adapters/codex/adapter.tsv +12 -0
- package/adapters/codex/detect.sh +9 -0
- package/adapters/codex/verify.sh +45 -0
- package/adapters/copilot/adapter.tsv +10 -0
- package/adapters/copilot/detect.sh +8 -0
- package/adapters/copilot/verify.sh +45 -0
- package/adapters/cursor/adapter.tsv +11 -0
- package/adapters/cursor/detect.sh +10 -0
- package/adapters/cursor/verify.sh +45 -0
- package/adapters/gemini/adapter.tsv +15 -0
- package/adapters/gemini/detect.sh +11 -0
- package/adapters/gemini/verify.sh +49 -0
- package/adapters/hermes/adapter.tsv +9 -0
- package/adapters/hermes/detect.sh +8 -0
- package/adapters/hermes/verify.sh +27 -0
- package/adapters/opencode/adapter.tsv +14 -0
- package/adapters/opencode/detect.sh +9 -0
- package/adapters/opencode/verify.sh +45 -0
- package/automations/README.md +53 -0
- package/automations/bugreporter-intake.sh +145 -0
- package/automations/drift-audit.sh +139 -0
- package/automations/report.schema.tsv +10 -0
- package/bans/README.md +82 -0
- package/bans/grep-ban.sh +84 -0
- package/bans/layer-check.sh +57 -0
- package/bin/goblin +119 -0
- package/bin/goblin-audit +145 -0
- package/bin/goblin-bans +178 -0
- package/bin/goblin-doctor +233 -0
- package/bin/goblin-emit +484 -0
- package/bin/goblin-init +519 -0
- package/bin/goblin-install +720 -0
- package/bin/goblin-lib.sh +289 -0
- package/bin/goblin-model +105 -0
- package/bin/goblin-upgrade +572 -0
- package/bin/goblin-verify +2798 -0
- package/bin/goblin.js +103 -0
- package/docs/ADOPTION.md +168 -0
- package/docs/CI.md +187 -0
- package/docs/CONTRACTS.md +197 -0
- package/docs/DESIGN.md +92 -0
- package/docs/ENFORCEMENT.md +225 -0
- package/docs/FLOWS.md +164 -0
- package/docs/GUARDRAILS.md +126 -0
- package/docs/GUIDE.md +610 -0
- package/docs/INTEGRATION.md +92 -0
- package/docs/LIMITS.md +591 -0
- package/docs/LOOP.md +165 -0
- package/docs/RE-PLAYBOOK.md +183 -0
- package/docs/RISKS.md +70 -0
- package/docs/ROLES.md +105 -0
- package/manifest/bans.tsv +9 -0
- package/manifest/classes.tsv +61 -0
- package/manifest/enforcement.tsv +88 -0
- package/manifest/glossary.tsv +25 -0
- package/manifest/playbooks.tsv +16 -0
- package/package.json +37 -4
- package/presets/A-shipped-software.yaml +48 -0
- package/presets/B-service-config.yaml +40 -0
- package/presets/C-game.yaml +38 -0
- package/presets/D-knowledge.yaml +41 -0
- package/presets/E-fleet-config.yaml +42 -0
- package/presets/F-electron.yaml +67 -0
- package/roles.yaml +54 -0
- package/skills/goblin-bootstrap/SKILL.md +51 -0
- package/skills/goblin-bugfix/SKILL.md +26 -0
- package/skills/goblin-bugreporter/SKILL.md +52 -0
- package/skills/goblin-drift-audit/SKILL.md +43 -0
- package/skills/goblin-eval/SKILL.md +68 -0
- package/skills/goblin-feature/SKILL.md +26 -0
- package/skills/goblin-feature-map/SKILL.md +140 -0
- package/skills/goblin-handoff/SKILL.md +28 -0
- package/skills/goblin-investigation/SKILL.md +26 -0
- package/skills/goblin-judge/SKILL.md +74 -0
- package/skills/goblin-loop/SKILL.md +88 -0
- package/skills/goblin-mode/SKILL.md +70 -0
- package/skills/goblin-overnight/SKILL.md +42 -0
- package/skills/goblin-pr-gate/SKILL.md +42 -0
- package/skills/goblin-re-mobile/SKILL.md +51 -0
- package/skills/goblin-refactor/SKILL.md +23 -0
- package/skills/goblin-sweep/SKILL.md +23 -0
- package/skills/goblin-tdd-repro/SKILL.md +27 -0
- package/skills/goblin-verify-author/SKILL.md +50 -0
- package/skills/practice/SKILL.md +37 -0
- package/templates/AGENTS.md.tmpl +23 -0
- package/templates/HANDOFF.md.tmpl +43 -0
- package/templates/SPEC.md.tmpl +34 -0
- package/templates/audit-waiver.tsv.tmpl +10 -0
- package/templates/boundary-waivers.tmpl +8 -0
- package/templates/checks/assert.mjs.tmpl +60 -0
- package/templates/checks/gate.sh.tmpl +29 -0
- package/templates/ci/goblin-gate.yml.tmpl +46 -0
- package/templates/goblin.yaml.tmpl +138 -0
- package/templates/install-hooks.allowlist.tmpl +9 -0
- package/templates/loop/decisions.tsv.tmpl +1 -0
- package/templates/loop/predicate.tmpl +16 -0
- package/templates/report.yaml.tmpl +16 -0
package/docs/DESIGN.md
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Design
|
|
2
|
+
|
|
3
|
+
## The thesis
|
|
4
|
+
|
|
5
|
+
goblin-stack is a small portable repository that installs three things into any target repo:
|
|
6
|
+
|
|
7
|
+
1. an **executable rule manifest** — every rule carries a runnable check or is explicitly
|
|
8
|
+
counted as advisory;
|
|
9
|
+
2. a set of **Hermes project-local skills** that carry the flows;
|
|
10
|
+
3. an **installer and a verifier** that prove the first two are still true.
|
|
11
|
+
|
|
12
|
+
It is not a rules document. A rules document has no enforcement, and the published evidence is
|
|
13
|
+
that an LLM-generated context file costs tokens and can reduce resolution. A rule that cannot
|
|
14
|
+
be checked is counted and capped instead of asserted.
|
|
15
|
+
|
|
16
|
+
It is not a port of any Cursor harness. The primitives that carry those harnesses —
|
|
17
|
+
`subagent_type`, `/loop`, `/goal`, `environment: "cloud"`, `readonly`, a vendored plugin
|
|
18
|
+
directory — do not exist in Hermes. What transfers is the *shape*: an executable matrix, a
|
|
19
|
+
bounded panel, a pinned-SHA verdict. What does not transfer is cut, with the reason recorded in
|
|
20
|
+
`docs/FLOWS.md`.
|
|
21
|
+
|
|
22
|
+
## Three load-bearing decisions
|
|
23
|
+
|
|
24
|
+
### D1 — the enforcement matrix is an executable artifact, not a document
|
|
25
|
+
|
|
26
|
+
`manifest/enforcement.tsv` has one row per rule. The `check` column holds a real command, the
|
|
27
|
+
literal `advisory`, or the marker `goblin-verify --only <ID>` for a check that needs more than
|
|
28
|
+
one shell line. A self-check row (IN-03) fails the manifest when a row has neither, and SK-03
|
|
29
|
+
caps the advisory count. **The number of unenforceable rules is itself a gate.**
|
|
30
|
+
|
|
31
|
+
### D2 — skills install as *project-local* skills, never into a profile
|
|
32
|
+
|
|
33
|
+
Hermes discovers skills at `<project-root>/.hermes/skills/` and gives the project tier the
|
|
34
|
+
highest precedence — above `~/.hermes/skills/` and above any profile copy. Project dirs are
|
|
35
|
+
treated as repo-owned, so autonomous skill maintenance never rewrites them, and they are
|
|
36
|
+
versioned with the code they govern.
|
|
37
|
+
|
|
38
|
+
This is the design that removes a whole class of rot instead of managing it. Two living copies
|
|
39
|
+
of the same rule with no equality check drift; one copy plus a recorded hash does not. The
|
|
40
|
+
installer puts the skills in the repo, and `SK-02` hashes what it installed.
|
|
41
|
+
|
|
42
|
+
### D3 — the gate is chosen by the change, not by the repo, and every review names its SHA
|
|
43
|
+
|
|
44
|
+
The stakes ladder S0–S4 becomes the `goblin-pr-gate` playbook. The one-line upgrade that
|
|
45
|
+
applies at every tier — including a direct push — is that a review note names the SHA it
|
|
46
|
+
reviewed: `reviews/<slug>-<head7>.md` with `head:`, `base:`, `patch-id:`, `stakes:`,
|
|
47
|
+
`checks-run:`. A verdict is a claim about an artifact, not about a moment, and `PG-03`
|
|
48
|
+
re-checks the patch-id so a new head voids it.
|
|
49
|
+
|
|
50
|
+
## The composition decision: the house style is referenced, never copied
|
|
51
|
+
|
|
52
|
+
The split is by **kind**, and it is checkable.
|
|
53
|
+
|
|
54
|
+
- The **standard this repo references** owns the house style: the HANDOFF shape and the
|
|
55
|
+
stale-sentence rule, the SPEC lifecycle, the harness house style and the pinned-commit
|
|
56
|
+
REPLAY, the gate vocabulary, commit discipline, the delegation tiers, data safety, the
|
|
57
|
+
documentation duty, adopt-don't-replace.
|
|
58
|
+
- **goblin-stack owns the mechanism**: which rule is enforced by what, how it is installed, how
|
|
59
|
+
it is verified, which class a project is, which flow applies, which role runs a flow.
|
|
60
|
+
|
|
61
|
+
goblin-stack carries **no copy** of the standard's text. `.goblin/goblin.yaml` holds
|
|
62
|
+
`practice:` and `practice_sha256:`; `goblin-verify` re-checks the hash, so a silently edited
|
|
63
|
+
standard is visible rather than assumed. An edit that *is* intended is re-pinned by one explicit,
|
|
64
|
+
printed command (`goblin-install --target <dir> --re-pin`), never automatically — the pin exists
|
|
65
|
+
to catch a silent edit, so it may not update itself. If no standard is configured, the checks that
|
|
66
|
+
depend on it report advisory, never a failure — that is what makes the repo portable.
|
|
67
|
+
|
|
68
|
+
## Rejected alternatives
|
|
69
|
+
|
|
70
|
+
| Rejected | Why |
|
|
71
|
+
|---|---|
|
|
72
|
+
| A file-for-file port of a Cursor harness | Most of its playbooks are inert without a plugin that is not shipped, and its primitives do not exist here. |
|
|
73
|
+
| A "goblin-stack rules" document | The #1 anti-pattern: a rules doc with no enforcement is a measured net cost. |
|
|
74
|
+
| Copying the referenced standard into this repo | It is the strongest artifact on the box and must not be weakened; a second copy is exactly the duplication that already rots. |
|
|
75
|
+
| Replacing the referenced standard | Would discard the pinned-commit REPLAY rule, the stale-sentence rule, the data-safety rule and adopt-don't-replace — four things no imported source has. |
|
|
76
|
+
| Per-project bespoke harnesses | The *gate vocabulary* differs by class, not the harness. Six class presets plus a real off switch. |
|
|
77
|
+
| Fan-out by default, auto-merge, or an unattended hillclimb | The axis is read-versus-write, parallel lanes cost about N times the tokens, and no reviewed source ships unconditional auto-merge. |
|
|
78
|
+
| A plugin or marketplace package | There is no marketplace here. The portable unit is a `SKILL.md` plus a bash installer. |
|
|
79
|
+
| Shipping no workflow at all (the pre-W4 position) | Measured: `PG-05` cannot bite without a workflow to read, and CI and `goblin-verify` were free to report different truths about one SHA. Reversed at W4 as an **architecture change**, which is why the part is in `manifest/classes.tsv` and the installer step is explicit rather than a loose template file. |
|
|
80
|
+
|
|
81
|
+
## What it deliberately does not do
|
|
82
|
+
|
|
83
|
+
See `docs/LIMITS.md` and the non-goals in `docs/RISKS.md`. The short version: it does not choose
|
|
84
|
+
models, does not write the vault, does not replace any project's existing gate, writes nothing
|
|
85
|
+
outside its target, and does not pretend the prose rules are enforced.
|
|
86
|
+
|
|
87
|
+
**One amendment, at W4.** It used to say *"there is no CI workflow"*. It now writes **at most one**
|
|
88
|
+
workflow, into its own target, only for a class that requires or permits the `ci-gate` part
|
|
89
|
+
(`.github/workflows/goblin-gate.yml`), never overwriting a file it did not write. What it cannot do
|
|
90
|
+
is make that file a **gate**: the required-check list, the bypass switch and the push identity are
|
|
91
|
+
forge state, so `docs/CI.md` §1 writes them down and `PG-04` stays advisory. The workflow runs the
|
|
92
|
+
repo's own declared gate set through `goblin-verify`, so it adds no second source of truth.
|
|
@@ -0,0 +1,225 @@
|
|
|
1
|
+
# Enforcement matrix
|
|
2
|
+
|
|
3
|
+
`manifest/enforcement.tsv` is the matrix: **one row per rule**. The `check` column holds a
|
|
4
|
+
real command that `goblin-verify` runs, or the literal `advisory`, or the marker
|
|
5
|
+
`goblin-verify --only <ID>` for a check that needs more than one shell line (those are
|
|
6
|
+
builtins in `bin/goblin-verify`, so the row stays self-describing and the matrix stays the
|
|
7
|
+
single source of truth).
|
|
8
|
+
|
|
9
|
+
Some `if_not_why` cells carry internal workstream tags — `W1`..`W5` — recording which redesign
|
|
10
|
+
of this toolkit last touched that rule's check (W1 the global-engine split, W2 the version
|
|
11
|
+
sync and npm shim, W3 the upgrade migration, W4 the platform adapters, W5 the publish prep).
|
|
12
|
+
They are historical provenance only: the rule text and its check are authoritative, not the
|
|
13
|
+
tag. `D1`..`D6` are the design decisions in `docs/DESIGN.md`; `Z1`-style identifiers are
|
|
14
|
+
independent verification pass IDs.
|
|
15
|
+
|
|
16
|
+
`scope` is `target` (runs in an installed repo via `goblin-verify`) or `source` (runs in
|
|
17
|
+
this repo via `tests/run-tests.sh`). The two scopes name WHOSE proof burden a row carries:
|
|
18
|
+
`scope:source` is the framework's own burden when developing goblin-stack (a dev runs
|
|
19
|
+
`run-tests.sh`), `scope:target` is the adopting repo's burden (the user runs `goblin-verify`).
|
|
20
|
+
The source-scope rows (`PR-01`..`PR-05`) therefore never run
|
|
21
|
+
in an installed repo — they execute only in this repo's own test suite. `enforced_by` is one of
|
|
22
|
+
five values, and the enum is
|
|
23
|
+
closed and now ENFORCED (`IN-03`'s third clause, Z1-5): `script`, `lint`, `gate`, `advisory`,
|
|
24
|
+
plus `test` for a source-scope row, whose check is a script under `tests/` run by
|
|
25
|
+
`tests/run-tests.sh`.
|
|
26
|
+
|
|
27
|
+
Measured shape of this table: **87 rows** - 82 target, 5 source; advisory 10, gate 25, lint 28, script 20, test 4.
|
|
28
|
+
|
|
29
|
+
## The rows
|
|
30
|
+
|
|
31
|
+
| id | scope | enforced by | rule | check | if it cannot be enforced, why |
|
|
32
|
+
|---|---|---|---|---|---|
|
|
33
|
+
| `IN-01` | target | script | The install exists and records its version + every file's hash. | goblin-verify --only IN-01 | — (W1: the check is a builtin so the engine.mode=global clause can run — a global, declaration-only repo has no install record and SKIPs with `global engine mode — no per-repo install record`; the vendored clauses are the old one-liner: the record exists and names its version) |
|
|
34
|
+
| `IN-02` | target | script | Every installed file still matches its recorded hash. | `goblin-verify --only IN-02` (builtin) | — |
|
|
35
|
+
| `IN-03` | target | script | The verifier's own manifest is complete: every rule has a check or is advisory. | `goblin-verify --only IN-03` (builtin) | — (this row is the reason the matrix cannot rot; the second clause is D6's shape in general: a row that carries no check must be labelled advisory, or it claims verification it does not perform. The third clause is Z1-5: `enforced_by` is documented as a closed enum in docs/ENFORCEMENT.md and was read by NOTHING, so a typo in that cell changed nothing - `script\|lint\|gate\|advisory` plus `test`, the source-scope value whose check is tests/run-tests.sh. W1: the check runs against whichever manifest the engine actually resolved (the chain in `bin/goblin-verify`), no longer the hardcoded per-repo path — the rule's meaning is untouched, only the path input follows the engine) |
|
|
36
|
+
| `IN-04` | target | script | No file goblin-stack did not create has been overwritten. | `goblin-verify --only IN-04` (builtin) | Detects a file the installer recorded as pre-existing (a `refused` entry) that has since vanished, or that is listed as installed anyway. The second clause is an internal-consistency guard: with correct code a refused path is never written, so it fires only if the installer regresses. The negative control exercises the vanished branch. |
|
|
37
|
+
| `HP-01` | target | gate | HANDOFF.md exists at the root. | `test -f HANDOFF.md` | — |
|
|
38
|
+
| `HP-02` | target | lint | The HANDOFF carries its five required sections. | `for h in 'START HERE' 'STATE\|STATUS' 'GATES?' 'NEXT STEPS\|NEXT' 'NOT VERIFIED\|UNVERIFIED\|NOT PROVEN\|UNPROVEN\|PENDING[^.]*DEVICE TEST'; do grep -qiE "^#{2,3}[[:space:]]+[^[:alpha:]]*($h)\b" HANDOFF.md \|\| { echo "missing section: $h"; exit 1; }; done` | Partial: proves a heading exists whose first word names the slot, not that the section's content is complete. The slot may use the repo's own vocabulary (the model repo's `### NEXT` passes; `Status`, `Gate`, `Unverified`, `Pending <x> device test` are accepted); a heading that merely contains the word (`## BOARD STATE`) does not satisfy `State`. |
|
|
39
|
+
| `HP-03` | target | lint | Every gate number is a measurement with a date, never a copy. | `goblin-verify --only HP-03` (builtin) | Builtin, anchored on the gate NAMES `.goblin/goblin.yaml` declares (GT-01's source of truth) rather than on a hardcoded keyword list: the shipped gates are named `commit` and `todo_ceiling`, neither of which the old list matched, so on a fresh install the only line it could see was the template's own example sentence, and a real gate line could lose its `date` and the row stayed GREEN (G8-2). A line whose text says "example of the required form" is template prose, not a gate number, and is skipped, so the template cannot satisfy the row. Partial: it proves a dated line inside the `Gates` section exists and that every gate-bearing line there carries a date - not that the number was re-measured that day, and not a gate-bearing line that names no declared gate and carries no gate-shaped keyword (so a line the config does not declare and the keyword list does not recognise is unseen). Historical gate lines outside that section are exempt by design - PROJECT-PRACTICE section 1's stale-sentence rule requires them to be kept. |
|
|
40
|
+
| `HP-04` | target | advisory | A stale sentence is corrected in place with a dated parenthetical, never deleted. | `advisory` | Detecting a silent deletion needs semantic judgement; a diff heuristic (>=5 removed non-empty lines with no 'corrected' addition) is too noisy to gate on. goblin-verify --only HP-04 prints the heuristic as a warning only. |
|
|
41
|
+
| `HP-05` | target | gate | The HANDOFF names the HEAD it describes. | `goblin-verify --only HP-05` (builtin) | Deviation from the design spec, with the reason: the spec's literal check is `grep -q "$(git rev-parse --short HEAD)" HANDOFF.md`, which can never pass - committing the HANDOFF moves HEAD, so the file can only name a commit that is now an ancestor. The mechanised form is therefore 'the HANDOFF names a commit that exists in this repo AND is an ancestor of HEAD', which still catches the defect it exists for (a review/handoff artifact that names no commit at all). |
|
|
42
|
+
| `SP-01` | target | gate | A *-SPEC.md file exists at the repo root (any round, not the current one - round-scoping arrives with the W6 staged chain). | `ls ./*-SPEC.md >/dev/null 2>&1` | Skipped when class: D and disabled: [spec]. |
|
|
43
|
+
| `SP-02` | target | script | The SPEC is committed, not left untracked. | `test -z "$(git ls-files --others --exclude-standard -- '*-SPEC.md')"` | — |
|
|
44
|
+
| `SP-03` | target | lint | Every AC: item is checkable without a human. | `awk '/^[[:space:]]*[-*][[:space:]]/ && /AC[0-9]*:/ && !/\|==\|===\|exit\|<\|>/ {print; bad=1} END{exit bad}' ./*-SPEC.md` | Partial: structural only - a checkable-looking bullet can still be unfalsifiable. The matcher is `AC[0-9]*:` because the shipped template writes `- AC1: ...`; the literal `/AC:/` only saw a bullet that spelled the label without a number. |
|
|
45
|
+
| `GT-01` | target | gate | The gate set is declared, never inferred from the stack. | `goblin-verify --only GT-01` (builtin) | Builtin: every `- name:` under `gates:` is a DECLARED gate and each one must carry a `cmd:` - a gate whose `cmd:` was deleted, blanked or re-indented FAILs here instead of vanishing from the count (G8-3, the condition of G8's own 9/10 sentence). The count is of declarations, not of runnable pairs. It cannot see whether a declared command is the RIGHT gate for the project: it proves a command exists, not that it is meaningful. |
|
|
46
|
+
| `GT-02` | target | gate | Every declared gate runs and exits 0. | `goblin-verify --only GT-02` (builtin) | — |
|
|
47
|
+
| `GT-03` | target | script | The round reports one line of measured numbers. | `test -f .goblin/last-gate-line && [ .goblin/last-gate-line -nt "$(git rev-parse --git-dir)/logs/HEAD" ]` | The freshness reference is HEAD's reflog (`.git/logs/HEAD`), which every HEAD movement rewrites - a commit in an attached or a detached worktree included, and independently of whether the refs are packed. The clause it replaced read `.git/HEAD`, a file a commit never rewrites (only branch operations do), so a commit landing after the measured line left the row GREEN while the round had moved on (D2, measured: `.git/HEAD`'s mtime unchanged across a real commit, `--only GT-03` exit 0). Two limits remain, recorded rather than hidden: a repo with the reflog disabled (`core.logAllRefUpdates=false`) has no reference to compare against, and `-nt` against a missing path is true, so the clause passes vacuously and the row then proves only that a measured line EXISTS (a repo with no commit yet is the same case); and the clause reads ANY HEAD movement as staleness, so a checkout, a branch rename or a reset FAILs it until the next gate run rewrites the line. That second behaviour is what makes a full run self-freshening: `GT-02` writes the line earlier in the same pass, so the row asserts that the line in front of you came from THIS run. `tests/t-gt03-freshness.sh` is the control, in both directions. |
|
|
48
|
+
| `GT-04` | target | gate | A ratchet is declared with a ceiling. | `goblin-verify --only GT-04` (builtin) | — |
|
|
49
|
+
| `GT-05` | target | gate | The ratchet has not risen. | `goblin-verify --only GT-05` (builtin) | — |
|
|
50
|
+
| `HS-01` | target | lint | Asserting harnesses follow the house shape. | `goblin-verify --only HS-01` (builtin) | A declared harness_dir that is absent FAILS when the class scaffolds one (config `scaffold_checks: yes`, classes A and C); a class that ships no harness dir (B/D/E) SKIPs with a reason. Keying the skip off the path alone let one config line switch this row and HS-02 off. A report utility in the same dir is counted and reported separately rather than failing the run. |
|
|
51
|
+
| `HS-02` | target | gate | A check green on both trees proves nothing - the REPLAY must show RED pre-change. | `goblin-verify --only HS-02` (builtin) | The declared `replay.cmd` is EXECUTED with `{name}` replaced by each harness's name, in the pre-change worktree, with `replay.env=<commit>` set (Z1-4). Two clauses that used to be unheard: a command that interpolates no `{name}` FAILs (it cannot be running the harness it names, so nothing was replayed), and a command that cannot be executed at all - exit 126 or 127 - FAILs rather than counting as a RED harness. The harness's name is substituted SHELL-QUOTED (`printf %q`), because the name is part of the command text the shell parses; unquoted, a name carrying `;` or `#` reached the shell as syntax (Z2-3, the G8-1 surface class), and the control for it carries a metacharacter-bearing file name. What it cannot see: a command that runs *something else* under the harness's name and exits non-zero, and whether the harness tests the right path rather than merely failing on this tree. |
|
|
52
|
+
| `HS-03` | target | advisory | Source probes read text with comments blanked first. | `advisory` | Recognising 'this probe reads source text' is semantic; a grep for the blanking helper produces false FAILs on harnesses that do not probe source. |
|
|
53
|
+
| `CM-01` | target | gate | Commits carry the owner identity, not an ambient one. Current scope: this gates the identity of HEAD (the last commit) at the moment of the run - earlier commits by other authors are not scanned. | `test "$(git log -1 --format='%ae')" = "$(grep '^owner_email:' .goblin/goblin.yaml \| cut -d' ' -f2)"` | — |
|
|
54
|
+
| `CM-02` | target | advisory | The commit message was written to a file, not passed with -m. | `advisory` | A backtick lost to command substitution leaves no trace a later check can read. Reported as a heuristic (unbalanced backticks in a body) only. |
|
|
55
|
+
| `CM-03` | target | script | Commit-as-you-go: the working tree is not carrying a dead run's work. | `goblin-verify --only CM-03` (builtin) | — |
|
|
56
|
+
| `MD-01` | target | lint | No model name is hardcoded in any reusable rule. | `for d in skills manifest bin templates presets .goblin .hermes; do [ -d "$d" ] \|\| continue; grep -rniE '(d[e]epseek\|cl[a]ude\|g[p]t-[0-9]\|gr[o]k\|g[e]mini\|g[l]m-[0-9]\|k[i]mi)[a-z0-9.:_-]*' "$d" && exit 1; done; exit 0` | — (the pattern is written with character classes so this row cannot match itself; tests/t-verify-red.sh proves it still catches a real model name. SCOPE, stated so a repo-wide grep is not re-reported as a hole (G8-10): this row reads the seven directories an install writes rules into - skills/ manifest/ bin/ templates/ presets/ .goblin/ .hermes/. It does NOT read tests/, docs/, or the source automations/ directory; tests/ is where a control must be free to WRITE the thing it guards, and the source directories are covered by MD-01's own body in tests/run-tests.sh, which scans skills manifest bin templates presets automations. Both control strings are assembled at run time, so a repo-wide grep over the whole tree is 0 and the row's own scope is 0 by construction.) |
|
|
57
|
+
| `MD-02` | target | advisory | The review lane is a different model family from the code lane. | `goblin-verify --only MD-02` (builtin) | goblin-stack cannot choose the fleet's models; today's map resolves both roles to the same family. Reported as ADV, never gated. W5-6: the JUDGE lane's resolved model is compared with the code lane's too, because `JG-02` proves only that the declared profile NAMES are disjoint; it is still ADV, never gated, and the measured state of this box is recorded in `docs/LIMITS.md` #38. |
|
|
58
|
+
| `MD-03` | target | advisory | Role-pinned fan-out goes through kanban, not a model-less subagent spawn. | `advisory` | It is a fleet-runtime property: no repo-local file can observe which tool created a worker. Enforced at board level, described in docs/INTEGRATION.md. |
|
|
59
|
+
| `PG-01` | target | gate | The reviewed artifact is named by SHA, and that SHA exists. | `goblin-verify --only PG-01` (builtin) | — |
|
|
60
|
+
| `PG-02` | target | script | The gate is chosen by the change, not the repo, and the tier's evidence exists. | `goblin-verify --only PG-02` (builtin) | — |
|
|
61
|
+
| `PG-03` | target | gate | A new head voids the verdict. | `goblin-verify --only PG-03` (builtin) | — |
|
|
62
|
+
| `PG-04` | target | advisory | Never bypass what the forge enforces. | `advisory` | Needs the forge: GitHub's restrictions do not apply to admins, and a sole-admin repo has nobody the gate binds. Not observable from the repo. |
|
|
63
|
+
| `PG-05` | target | lint | No required check that self-skips. | `goblin-verify --only PG-05` | Text, not a YAML parser. Three clauses per job: the job declares at least one `run:`/`uses:` step (else the check runs nothing); the JOB carries no `if:` (else the whole required check self-skips); and no STEP carries an `if:` (else that step - possibly the gate step - self-skips). The old body flagged only 'every step guarded', which PASSED the real shape: in the estate's one existing workflow a deliberately UNGUARDED credential step decides whether the guarded compile step runs, so `guarded < steps` and the row reported PASS on the workflow it exists to catch (G8 section 3, re-measured V3-4). Deliberately strict, and the strictness is the point: GitHub reports a SKIPPED job as Success even when it is a required check (docs S1/S2), so a conditional step is a step that can green-light a commit whose gate never ran. False positives it cannot avoid: a `#` inside a quoted string is read as a comment, a flow-style (`jobs: {...}`) mapping is refused rather than parsed, and a conditional step that is genuinely safe is indistinguishable from the trap - the remedy is to move the condition into the declared command, or into a second job that is not the required check. It cannot see branch-protection state, the required-check list, or whether the workflow ever ran (`PG-04`), and a repo with no workflow passes with the count printed on the line. |
|
|
64
|
+
| `PG-06` | target | gate | The gate CI runs is the gate the project declares. | `goblin-verify --only PG-06` | The declared set comes from `g_yaml_gates`, the reader `GT-01` uses, so a gate cannot vanish from the comparison in silence (G8-3). A workflow passes when the whole declared set is RUN: the verifier with no `--only` (a full run executes every declared gate through `GT-02`), or each declared gate command verbatim. It reads TEXT with comments blanked first, and only `run:` payloads and block-scalar bodies count - a command sitting in a `name:`, `env:` or `with:` value is dropped, which is what stops a comment or a label from satisfying the row. What it cannot see: that the forge marks that job a REQUIRED check, that the job is the one the forge waits on, or that the workflow can fail at all (`PG-05`). A repo with no workflow reports SKIP with that reason instead of a vacuous pass. The lane's own enforcement is the TARGET repo's, not goblin-stack's (`docs/LIMITS.md` #34): goblin-stack can prove the file invokes the gate and cannot make the forge run it, so this row is a text reading of the target's CI and never a claim that CI gated the SHA. |
|
|
65
|
+
| `DS-01` | target | script | Runtime data is not test fixture: a gate run must not write it. | `goblin-verify --only DS-01` (builtin) | — |
|
|
66
|
+
| `DS-02` | target | gate | Snapshot before, verify after. | `goblin-verify --only DS-02` (builtin) | — |
|
|
67
|
+
| `DOC-01` | target | advisory | A significant change updates the docs that teach it. | `advisory` | 'Significant' is a judgement; a diff-size heuristic fails on the cases that matter. |
|
|
68
|
+
| `DOC-02` | target | advisory | System-level changes are recorded wherever the project's standard says they live. | `advisory` | Where the recording lives may be owned by a stricter external rule than goblin-stack may add; the repo can only state it. |
|
|
69
|
+
| `SK-01` | target | lint | Every shipped skill has name + description frontmatter. | goblin-verify --only SK-01 | W1: the check is a builtin so the engine.mode=global clause can run - in global mode the procedure tier is emitted per platform (not carried in this repo) and the row SKIPs with that reason instead of passing vacuously on a repo with no skills (§2.5 names SK-01 alongside SK-02/SK-04). In vendored mode it is exactly the old loop: every SKILL.md must open with frontmatter carrying both name and description. |
|
|
70
|
+
| `SK-02` | target | script | The installed skills match their recorded hashes (no drift). | `goblin-verify --only SK-02` (builtin) | — |
|
|
71
|
+
| `SK-03` | target | script | A rule with no mechanism is labelled advisory, and the advisory count is reported. | `goblin-verify --only SK-03` (builtin) | — |
|
|
72
|
+
| `SK-04` | target | lint | Every shipped skill says what it cannot see. | goblin-verify --only SK-04 | Partial: proves the section exists, not that what it says is complete or true - the limit every prose rule carries. Every shipped skill already carries it, so the row is GREEN on a fresh install and RED only under a real violation. W1: the check is a builtin so the engine.mode=global clause can run - in global mode the procedure tier is emitted per platform (not carried in this repo) and the row SKIPs with that reason. |
|
|
73
|
+
| `PT-01` | target | lint | No tenant-specific string inside a reusable rule. | `for d in skills manifest bin templates presets .goblin .hermes; do [ -d "$d" ] \|\| continue; grep -rniE --exclude=goblin.yaml --exclude=installed.json '(h[a]rvey\|tech-g[o]blin\|/h[o]me/[a-z]+\|g[o]blin-ui\|op[e]n-door\|sup[r]eme\|bb[t]ech\|c[l]v)' "$d" && exit 1; done; exit 0` | — (the same directory list MD-01 uses: the rules an install actually writes live in `.goblin/` and `.hermes/`, not in the source layout. Two documented exceptions: `.goblin/goblin.yaml`, which holds `models_file:`/`practice:` - per-machine config, not a rule - and `.goblin/installed.json`, which since W1 records the machine's absolute `engine_dir` in its `engine:` block - both per-machine facts, not rules - each excluded by name. The pattern is written with character classes so this row cannot match itself.) |
|
|
74
|
+
| `PT-02` | target | gate | The default branch is declared, not assumed. | `goblin-verify --only PT-02` (builtin) | — |
|
|
75
|
+
| `CL-01` | target | script | Every part the class requires is present, and every part it forbids is absent. | `goblin-verify --only CL-01` (builtin) | — |
|
|
76
|
+
| `CL-02` | target | script | An archive: true project verifies GREEN without a HANDOFF or gates. | `goblin-verify --only CL-02` (builtin) | Falsifiable: FAILs when `archive:` is not `true`/`false`, and when the config's value disagrees with the one the install recorded in `.goblin/installed.json` (so the waiver cannot be flipped on by hand). It cannot observe the *effect* of the waiver on the other rows without re-entering the runner. |
|
|
77
|
+
| `SC-01` | target | lint | No secret file is tracked. | `n=$(git ls-files \| grep -iE '(^\|/)\.env\|\.pem$\|\.key$' \| grep -vcE '\.(example\|sample\|template)$'); printf '%s tracked secret file(s)\n' "$n"; [ "$n" = 0 ]` | Partial: it sees tracked PATHS, never contents - a secret pasted into a tracked file is invisible here, and the pattern is a name family, so a credential inside `config.ts` is missed by construction. |
|
|
78
|
+
| `SC-02` | target | gate | The ignore rules cover the whole secret family. | `goblin-verify --only SC-02` | Builtin, and behavioural: clause 1 reads `.gitignore`; clause 2 asks git's own matcher (`git check-ignore`) for `.env`, `.env.local` and `.env.production` one path at a time, so a rule that looks right but does not match still fails. It cannot see a secret already in git history, or one committed under a name the family does not cover. SKIPs with a reason when there is no `.gitignore` and no `package.json`. |
|
|
79
|
+
| `SC-03` | target | lint | No client-visible name is secret-shaped, and no build output carries a secret literal. | `goblin-verify --only SC-03` | Builtin, two clauses: the `NEXT_PUBLIC_*_(SECRET\|TOKEN\|KEY\|PASSWORD\|PRIVATE)` name pattern over source, and known secret prefixes over the declared build output. Partial: a prefix pattern cannot see a secret that does not look like one, and clause 2 says "no build output to scan" on the line rather than skipping silently. The source scan excludes `.goblin/` and `.hermes/`, so it cannot match the row text that describes it. |
|
|
80
|
+
| `SC-04` | target | lint | Every cookie write carries its flags. | `goblin-verify --only SC-04` | Builtin, same-statement only: a write spread over three lines is not seen, and the row says so. It REPORTS that a JS-written cookie is readable by any script rather than failing on it - that is a design fact, not a bug - so the pass is about the flags, never about the choice. |
|
|
81
|
+
| `SC-05` | target | gate | Every write route validates its input, or is waived. | `goblin-verify --only SC-05` | Builtin: it proves a validator is CALLED (`safeParse\|zod\|valibot\|yup\|ajv\|superstruct\|validate(`), never that the schema is right - a schema that accepts everything passes. `.goblin/boundary-waivers` is the escape hatch, and the waived count is printed, so a silent pile-up is visible. |
|
|
82
|
+
| `SC-06` | target | gate | A lockfile exists, and the repo tracks it. | `goblin-verify --only SC-06` | Builtin: presence, then `git ls-files --error-unmatch`. It cannot see that the lockfile is STALE relative to `package.json` - resolving that needs the package manager, which is a deliberate network-shaped step, not a check. SKIPs with a reason when there is no `package.json`. |
|
|
83
|
+
| `SC-07` | target | gate | The dependency audit record is present, dated, fresh, and clean-or-waived. | `goblin-verify --only SC-07` | Builtin, OFFLINE by construction: it reads the record and never the network. The record comes from `.goblin/bin/goblin-audit`, run once, deliberately; the waiver count is printed on the gate line so the debt is loud even when the row passes. It cannot see an advisory the registry did not know on the day the record was taken. SKIPs with a reason when no record exists yet. |
|
|
84
|
+
| `SC-08` | target | gate | No dependency runs an install-time script that is not on the allowlist. | `goblin-verify --only SC-08` | Builtin over `package-lock.json`'s `hasInstallScript`: a pnpm/yarn lockfile has no such field, so those repos get a SKIP with that reason rather than a vacuous pass, and a MINIFIED (one-line) lockfile is read like a pretty-printed one - the text is normalised into the pretty shape before the line-anchored reader sees it, so a hook inside a one-line lock FAILs instead of being reported as `0 install hook(s)` (Z2-2: before that, a one-line lock PASSED vacuously). A `lockfileVersion` 1 lock records its hooks under `dependencies`, not under `node_modules/` keys, so the line-anchored reader enumerates nothing there either and prints the same `0 install hook(s), 0 allowlisted` PASS over a surface it did not read (measured at 0.4.2; npm <= 6 records no `hasInstallScript` at all, so for a genuine v1 lock this is blindness rather than a wrong answer). Lowest-value row of the ten - regression detection - and the first to cut if the matrix gets heavy. |
|
|
85
|
+
| `SC-09` | target | advisory | Auth is applied consistently across sibling routes. | `advisory` | Prose on purpose: "consistently" is a semantic judgement about a private surface no repo here has yet. Counted (9 of ceiling 10) so the matrix cannot quietly grow prose. |
|
|
86
|
+
| `PF-01` | target | lint | The perf baseline names the commit it measured. | `goblin-verify --only PF-01` | Builtin: the metric must equal `ratchet.name` so the budget and the measurement cannot silently disagree, the value must be numeric, the date must exist, the baseline commit must exist AND be an ancestor of HEAD (`HP-05`'s mechanic, reused rather than re-derived), and `ratchet.ceiling` must equal `perf.baseline_value` - otherwise a one-line ceiling raise passes while the row prints the contradiction, which is `I raised the budget and never measured again` (G8-6b). It cannot see whether the metric is the right one for the product, and it never re-measures: re-anchoring is a deliberate operator action. SKIPs with a reason when the class declares no metric, or none has been recorded yet. |
|
|
87
|
+
| `AU-01` | target | lint | An automation's producer is deterministic and network-free. | goblin-verify --only AU-01 | Partial: proves that no line of a producer begins with a network or forge verb, not that the script is otherwise deterministic. W1: in global mode the first producer glob resolves under the engine dir (the resolution chain); the repo-local globs stay. New clause: a repo with NO producer anywhere (repo or engine) SKIPs with `no automation producer found (repo or engine)` instead of FAILing — the old FAIL was a born-RED artifact of the glob; a repo that declares its own automations but has no producer still FAILs. |
|
|
88
|
+
| `AU-02` | target | script | A report's dedup key is a function of content only - no date, no run id. | `goblin-verify --only AU-02` (builtin) | Builtin: it recomputes the key from the report's own `repo` and `symptom` and requires the recorded `dedup_key` to equal it, then refuses a key carrying a date. Skipped with a reason when the repo holds no report - nothing to dedup. It cannot see whether two reports should have been one: a normalisation that merges two genuinely different symptoms is a duplicate card, not a lost report. |
|
|
89
|
+
| `AU-03` | target | gate | A reporter run leaves the tree and the harness untouched. | `goblin-verify --only AU-03` (builtin) | Builtin: asserts a clean working tree, and - when HEAD is a reporter commit - that no path under the declared harness_dir appears in it. Skipped with a reason when the repo holds no reports/ - no reporter has run here. It cannot see a reporter that edited the tree and committed the edit as part of the report. |
|
|
90
|
+
| `AU-04` | target | lint | An automation's skill declares its own write surface. | goblin-verify --only AU-04 | Partial: proves every installed automation skill carries a `## Write surface` section. W1: the check is a builtin so the engine.mode=global clause can run - in global mode the procedure tier is emitted per platform (not carried in this repo) and the row SKIPs with that reason. |
|
|
91
|
+
| `BN-00` | target | script | Every ban has an enforcement row, every ban row names a replacement, and the ban table is not empty. | `goblin-verify --only BN-00` (builtin) | — (this row is the reason the ban list cannot decay into prose: IN-03's shape applied to bans.tsv, and it agrees in both directions) |
|
|
92
|
+
| `BN-01` | target | lint | No `any` in application TypeScript. | `goblin-verify --only BN-01` (builtin) | Text probe, not an AST: a `: any` inside a string or a comment is reported, and `Record<string, any>` (no leading colon) is missed. The AST form needs a parser the no-npm contract (docs/CONTRACTS.md) forbids (docs/LIMITS.md #27). SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
93
|
+
| `BN-02` | target | lint | No `@ts-ignore` / `@ts-expect-error` suppressions. | `goblin-verify --only BN-02` (builtin) | Text probe: it sees the directive wherever it appears, including inside a string, and cannot tell a suppression hiding a real error from one on a line that would compile anyway. SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
94
|
+
| `BN-03` | target | lint | No direct network call from a component. | `goblin-verify --only BN-03` (builtin) | Text probe over the declared component globs: it stops the call and cannot tell whether a data layer was written or the call merely moved into a helper. SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
95
|
+
| `BN-05` | target | lint | No import across a declared layer boundary. | `goblin-verify --only BN-05` (builtin) | Reads the `layers:` list; an empty list SKIPs with a reason, never a vacuous pass. It matches an import path naming the target directory's last segment - module aliases and dynamic imports are not seen. SKIPs when no file matches its globs. |
|
|
96
|
+
| `BN-06` | target | lint | No renderer with Node access (`nodeIntegration: true`). | `goblin-verify --only BN-06` | Text probe over the ban table's globs, the same mechanism as BN-01..BN-05: it sees `nodeIntegration: true` wherever it appears, including inside a string, and cannot see a webPreferences object built at run time or spread in from another module. The STRONGER form is a runtime measurement - the renderer prints `process.contextIsolated` and `process.sandboxed` and the check requires true/true - and that needs a real Electron process, which the dependency contract (docs/CONTRACTS.md) does not allow a shipped rule to launch: it is the project's host gate (docs/LIMITS.md #34). SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
97
|
+
| `BN-07` | target | lint | No renderer with context isolation or the process sandbox turned off. | `goblin-verify --only BN-07` | One probe for two properties because Electron's own documentation makes them one: disabling `contextIsolation` "also disables process sandboxing", so a repo that has turned either off has lost both. Text probe, with the same false-positive set as BN-06. SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
98
|
+
| `BN-08` | target | lint | No dangerous webPreferences. | `goblin-verify --only BN-08` | Four one-line patterns from Electron's own security checklist (`webSecurity: false`, `allowRunningInsecureContent: true`, `enableBlinkFeatures`, `<webview allowpopups>`). Text probe: `enableBlinkFeatures` is banned by name rather than by value, so the string is reported even in a comment. SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
99
|
+
| `BN-09` | target | lint | No synchronous IPC and no `@electron/remote`. | `goblin-verify --only BN-09` | The banned-list shape the wave's note 9 asks for, applied to Electron: `sendSync(` and `@electron/remote` block the renderer's own thread, which is the freeze the class exists to prevent. Text probe - it sees the call site, not the call graph, so a wrapper around `sendSync` in a file the globs do not match is missed. SKIPs when the ban is not in `bans:` or its globs match no file. |
|
|
100
|
+
| `FM-01` | target | lint | Every feature file is indexed from the map README, declares its slug and at least one entry path, and carries the four-H2 entry contract. | `goblin-verify --only FM-01` (builtin) | SKIPs (exit 3) when feature_map: is empty - a fresh install has no map and must not be born RED (the D8 shape). When a map IS declared: the README must exist, every features/*.md must be linked from it in the (./<slug>.md) form and every relative .md link must resolve, each feature file's `feature:` must equal its filename stem, it must declare >=1 `entry_paths:`, and its H2s must be exactly Sub-features / How to get to it (user POV) / Driving it with <harness> / Gotchas, in that order. Partial: the README's own H2s are prose this row does not read, and 'the map lists every user-facing feature' is not mechanically checkable - that is docs/LIMITS.md #30, not a row. |
|
|
101
|
+
| `FM-02` | target | lint | Every entry point a feature declares still resolves in source, and no entry path changed after the map was verified. | `goblin-verify --only FM-02` (builtin) | SKIPs (exit 3) when feature_map: is empty. A tripwire, not a proof. The token is searched under source_root with occurrences under the map's own directory excluded - without that exclusion the map's own entry-path list satisfies the search and the row could never go RED. Freshness is `git log -1 --format=%cs` on the resolved file against the feature's `verified:` date, and git sees a FILE change, not a behaviour change: the row can be RED-when-stale and never GREEN-means-fresh. A token that also occurs in a vendored copy or a build artifact is read as resolved, and a `verified:` date is itself a claim the row cannot test (docs/LIMITS.md #30). W5-4: the search skips the harness's own directories (`.goblin/`, `.hermes/`, the declared `harness_dir`) and the map's own directory, so a stub map whose token occurs only in the install no longer resolves; a token that occurs only in the target's own `docs/`, `tests/` or build output still does (docs/LIMITS.md #37). Z1-6: the resolved file must be TRACKED (`git ls-files --error-unmatch`) before its date is compared - an untracked file used to make the freshness clause skip in silence, so a map could claim `verified: 2020-01-01` over source that was never committed. |
|
|
102
|
+
| `VA-01` | target | gate | The generated verification skill's doctor command runs and exits 0. | `goblin-verify --only VA-01` (builtin) | SKIPs (exit 3) when verify_doctor: is empty (the replay.commit: "" shape). Runs the DECLARED command exactly as GT-02 runs a declared gate, and never a string read out of file content (the v0.2 blocker). Closes P6's stated-but-unenforced clause 'a generated skill that was never executed is a draft': the doctor is the smallest executable proof that the skill's own instructions still run - and it proves only that, never that the doctor tests the right path. |
|
|
103
|
+
| `JG-01` | target | script | A judge verdict is recorded as a row whose evidence resolves. | `goblin-verify --only JG-01` (builtin) | Builtin: it proves the named handle EXISTS - a commit `git rev-list --all` knows, a path under the root, or a `sha256:` matching a file under `.goblin/loop/` - never that the handle SUPPORTS the verdict beside it: a judge may cite a real commit that has nothing to do with the claim. It also cannot see whether the verdict was reached by the declared judge lane. A row with fewer columns than the header is a FAIL, not a silent skip. SKIPs (exit 3) when there is no loop record: nothing to check is reported as a reason, never as a pass. |
|
|
104
|
+
| `JG-02` | target | gate | The judge lane is disjoint from the author lane. | `goblin-verify --only JG-02` (builtin) | Builtin: it proves the DECLARED lane sets are disjoint (profile names), not that a judge ran on one - no repo-local file observes which profile ran (the blindness `MD-03` records), and family equality is not this row's business (`MD-02` owns it). An unresolved lane is an ADV carrying the one-line remedy, never a FAIL: a box with no mapping file must not be failed for the fleet's routing. Both sentinels count as unresolved - `?` (the installed path) and `unknown` (the checkout path) - because a check that greps for one passes on the other. |
|
|
105
|
+
| `JG-03` | target | advisory | A judge lane that has never returned a non-`done` verdict is escalated. | `advisory` | A lane that has returned two verdicts cannot be called always-yes, and a repo sees only its own rows; the history that would show a bad lane lives across cards and repos. The counter-measure is policy - one known-red control verdict per wave, recorded in docs/LOOP.md - and it is counted, not enforced. |
|
|
106
|
+
| `LP-01` | target | script | The exit predicate is a command, and it ran before iteration 1. | `goblin-verify --only LP-01` (builtin) | Builtin: it proves the predicate is exactly ONE command and that a recorded first run exists whose timestamp is at or before the first log row - not that the command was actually run, and not that the `exit=` value was measured rather than typed (`HP-03`'s defect, one artifact over). It deliberately does NOT re-run the predicate: a stop condition is legitimately red before the loop finishes, so re-running it would fail a repo whose loop is working, and PROJECT-PRACTICE section 7 forbids a probe that writes. Timestamps compare on their `YYYY-MM-DDThh:mm:ss` prefix, so records written in different UTC offsets compare as text. |
|
|
107
|
+
| `LP-02` | target | script | The predicate is pinned at loop start and is never relaxed. | `goblin-verify --only LP-02` (builtin) | Builtin: it proves the predicate file on disk still hashes to the recorded pin - never that the predicate is the RIGHT one, and never who edited it. The digest covers the predicate file alone: a loop that quietly relaxes a predicate living somewhere else is not seen. The remedy is printed with the failure and is never automatic, the `IN-02` `--re-pin` shape. W5-7: a close-and-reopen is auditable - every `closed-<date>/` archive must hold its predicate AND the pin it was closed under, and the live pin must name the archived digest on a `previous:` line - which catches a SILENT relaxation, never a WEAKER one (nothing in bash judges that: `docs/LIMITS.md` #39). |
|
|
108
|
+
| `LP-03` | target | gate | The loop declares a budget, and the record never exceeds it. | `goblin-verify --only LP-03` (builtin) | Builtin: it proves the declared budget is a positive integer at or under the configured `loop_max_turns_ceiling`, and that the record holds no more verdict rows than the budget - not that the budget is affordable. Cost is not a field the record carries: a turn budget bounds turns, and the auxiliary judge call per turn is unpriced (`docs/LIMITS.md` #32). A worker cannot edit its own card body, so a CARD predicate is already protected; this row is the file-predicate half. |
|
|
109
|
+
| `LP-04` | target | gate | A loop making no progress stops instead of thrashing. | `goblin-verify --only LP-04` (builtin) | Builtin: it measures a CHANGED EVIDENCE POINTER, which is a proxy for progress, not progress. Three consecutive verdict rows with an identical non-empty pointer and a non-`predicate:green` result is a FAIL that names the row numbers. A loop that edits a file each turn to keep the pointer moving is not caught, which is why LP-05 and the budget sit beside it. Nothing in the Hermes kanban goal loop detects a lack of progress at all - `run_kanban_goal_loop` carries no progress state. |
|
|
110
|
+
| `LP-05` | target | script | A loop that ended without its predicate green carries a written-up reason. | `goblin-verify --only LP-05` (builtin) | Builtin: the three-non-blank-line threshold is arbitrary and stated as such - it proves a write-up EXISTS, the way `SP-03` proves a bullet LOOKS checkable. Nothing can make the write-up true, and the row cannot see whether it names the right dead end or whether the loop stopped too early. A loop whose last row IS `predicate:green` needs no write-up and the row passes without reading one. |
|
|
111
|
+
| `RC-01` | target | gate | No file in the shipping tree matches a reference-corpus hash. | `goblin-verify --only RC-01` | Four clauses: (1) `reference_manifest:` empty -> SKIP with that reason (the `FM-01`/`VA-01` shape - not born RED); (2) declared but missing or unparseable -> FAIL, never a SKIP; (3) any file under the declared `security: build_output:` whose sha256 appears in an entry's sha256 -> FAIL, naming path + entry; (4) the manifest itself must not sit inside the build output, at the declared path or as a byte-identical copy - it is a listing of every corpus hash, so shipping it ships the corpus's shape. Cannot see: a re-encoded/resized/recoloured asset (level 2), copied text in a shipped string (level 3), or a manifest authored weak - that is `RC-02`. |
|
|
112
|
+
| `RC-02` | target | lint | The reference manifest is shaped so `RC-01` cannot pass vacuously. | `goblin-verify --only RC-02` | Schema `reference-manifest/1`; `generated_from` non-empty; `reference_app.package`/`.version`/`apk_sha256` (64 hex); `entries` non-empty; every entry carries `path`, 64-hex `sha256`, numeric `bytes`; `entry_count` equals `len(entries)`. SKIPs with `RC-01`'s reason when the key is empty. Exists because `RC-01` alone carries the `bans.tsv` weakness: a weakened input passes the check it feeds (`LIMITS.md` #28). It cannot see whether the entries are the *right* hashes - that is as strong as tamper hashing, and `installed.json` is unsigned (`LIMITS.md` #18). |
|
|
113
|
+
| `RC-03` | target | script | A lab repo tracks no extracted byte: scripts, notes, manifests, docs only. | `goblin-verify --only RC-03` | Two clauses: every tracked path falls under a declared allowlist (`scripts/`, `notes/`, `manifests/`, root docs - the harness's own `.goblin/` and `.hermes/` files and the `.gitignore` block it appends are the install, not the lab's content, and are excluded the way `FM-02` excludes them); and no tracked file's sha256 equals any hash in a `manifests/*.sha256`. The manifest itself lists paths and hashes, which the lab repo's own README explicitly permits - the check is about *bytes*, and it must not read the manifest as a violation. SKIPs with a reason when no `manifests/` exists. Cannot see a payload renamed and re-encoded - which is why the allowlist is a second, independent trap. |
|
|
114
|
+
| `RC-04` | target | lint | The acquisition record exists, and the manifest names the target and its source. | `goblin-verify --only RC-04` | The header block must name a target slug + version, a source store, and a checksum (md5 or sha256 hex); the manifest must hold a `*.apk` row, the header's named apk must be that row, and a 64-hex token in the header, if present, must equal the row's sha256 (a header carrying only the md5 passes - the comparison is conditional in the engine). Weakest of the four: it proves a record EXISTS, not that the number came from the store (`HP-03`'s defect; the `JG-01` shape). It is the first to cut if the matrix gets heavy. |
|
|
115
|
+
| `PR-01` | source | test | The installer never writes outside its target. | `tests/run-tests.sh` | — |
|
|
116
|
+
| `PR-02` | source | test | A second install is a no-op, and an upgrade reports created/updated/unchanged. | `tests/run-tests.sh` | — |
|
|
117
|
+
| `PR-03` | source | test | Every target-scope check goes RED under its own violation. | `tests/run-tests.sh` | — (the negative control the verifier re-runs) |
|
|
118
|
+
| `PR-04` | source | lint | The repo is portable: no personal path in any reusable rule. | `tests/run-tests.sh (the PT-01 body over the source tree, plus tests/)` | — |
|
|
119
|
+
| `PR-05` | source | test | The automation producer is silent when there is nothing to report. | `tests/run-tests.sh` | — (the mutation is the control: the same producer, on the same fixture, with one installed file edited, must go from an empty stdout to a record and exit 1. A producer that stays quiet after the mutation is not silent, it is broken.) |
|
|
120
|
+
|
|
121
|
+
## Advisory rows, named
|
|
122
|
+
|
|
123
|
+
10 of the 82 rows are labelled `advisory`. 9 of them carry no executable check at all
|
|
124
|
+
(they are prose the matrix refuses to pretend about); 1 is advisory-labelled but still
|
|
125
|
+
reports its state.
|
|
126
|
+
|
|
127
|
+
- **HP-04** (no check at all) - A stale sentence is corrected in place with a dated parenthetical, never deleted.
|
|
128
|
+
- **HS-03** (no check at all) - Source probes read text with comments blanked first.
|
|
129
|
+
- **CM-02** (no check at all) - The commit message was written to a file, not passed with -m.
|
|
130
|
+
- **MD-02** (reports state as ADV) - The review lane is a different model family from the code lane.
|
|
131
|
+
- **MD-03** (no check at all) - Role-pinned fan-out goes through kanban, not a model-less subagent spawn.
|
|
132
|
+
- **PG-04** (no check at all) - Never bypass what the forge enforces.
|
|
133
|
+
- **DOC-01** (no check at all) - A significant change updates the docs that teach it.
|
|
134
|
+
- **DOC-02** (no check at all) - System-level changes are recorded wherever the project's standard says they live.
|
|
135
|
+
- **SC-09** (no check at all) - Auth is applied consistently across sibling routes. Added when the missing entry was measured: the count said 9 and this list held 8.
|
|
136
|
+
- **JG-03** (no check at all) - A judge lane that has never returned a non-`done` verdict is escalated. The history that would show a bad lane lives across cards and repos, so the counter-measure is policy - one known-red control verdict per wave, recorded in `docs/LOOP.md` - not a command.
|
|
137
|
+
|
|
138
|
+
`advisory_ceiling` (default 10) caps that count: **SK-03 fails the run when the advisory
|
|
139
|
+
count exceeds it.** The number of unenforceable rules is itself a gate. Measured at `43f7f69`:
|
|
140
|
+
`advisory 9 of ceiling 10`; measured after W3's judge/loop rows:
|
|
141
|
+
`advisory 10 of ceiling 10`.
|
|
142
|
+
|
|
143
|
+
**The budget, stated so the next rule author does not have to work it out (V1/G8-5).** The
|
|
144
|
+
count is 10 of a ceiling of 10, so **no advisory slot is free**: the next advisory row FAILs
|
|
145
|
+
`SK-03` unless the ceiling is raised in the same change, with the reason written down. `SK-03`
|
|
146
|
+
prints the arithmetic on every run (`advisory 10 of ceiling 10 (0 free slots: the next advisory
|
|
147
|
+
row FAILs)`), and a ceiling that is not a number is a FAIL rather than a silent ADVISORY.
|
|
148
|
+
|
|
149
|
+
Two planned rows each wanted the last slot - **`FM-03`** (G1, the feature map) and **`JG-03`**
|
|
150
|
+
(G2, the judge agent). **Decided 2026-09-25 (W2):** G1's `FM-03` does **not** take it. The
|
|
151
|
+
feature map ships `FM-01` and `FM-02` as real commands, and the completeness claim `FM-03` would
|
|
152
|
+
have carried is recorded in `docs/LIMITS.md` #30 instead of as a counted row, so the slot stayed
|
|
153
|
+
free for G2's `JG-03`. The decision is recorded in `docs/LIMITS.md` #26; W2 measured `advisory 9
|
|
154
|
+
of ceiling 10 (1 free slot)`. **Spent 2026-09-25 (W3):** G2's `JG-03` took it, measured `advisory
|
|
155
|
+
10 of ceiling 10 (0 free slots)`.
|
|
156
|
+
|
|
157
|
+
`SK-03`'s count is the count the run itself uses: a row is advisory if its `check` cell says so
|
|
158
|
+
OR its `enforced_by` cell does.
|
|
159
|
+
|
|
160
|
+
## The class matrix
|
|
161
|
+
|
|
162
|
+
`manifest/classes.tsv` is the same idea applied to the parts a project must have.
|
|
163
|
+
`R` = required, `O` = optional (installed, reported), `-` = off, and **off is enforced**:
|
|
164
|
+
the installer records every `-` part in `disabled:`, so its rows report `SKIP (opt-out)`
|
|
165
|
+
instead of silently passing, and `CL-01` fails if a forbidden part's artifact exists.
|
|
166
|
+
|
|
167
|
+
| part | A | B | C | D | E | F |
|
|
168
|
+
|---|---|---|---|---|---|---|
|
|
169
|
+
| handoff | R | R | R | R | R | R |
|
|
170
|
+
| spec | R | R | R | - | R | R |
|
|
171
|
+
| gate | R | R | R | O | R | R |
|
|
172
|
+
| replay | R | - | R | - | O | R |
|
|
173
|
+
| ratchet | R | O | O | - | O | R |
|
|
174
|
+
| pr-gate | O | - | O | - | O | O |
|
|
175
|
+
| review-panel | O | - | R | - | O | O |
|
|
176
|
+
| playbooks | R | R | R | R | R | R |
|
|
177
|
+
| tokens | O | - | - | - | - | O |
|
|
178
|
+
| ci-gate | R | - | O | - | O | R |
|
|
179
|
+
|
|
180
|
+
Class **F** is the desktop shell: it needs a part no other class has — a **host gate**, a number
|
|
181
|
+
measured on a machine with a display — and forbids nothing the others allow except a renderer that
|
|
182
|
+
reaches Node directly, which its ban list catches. `ci-gate` is the one part added at W4: when it
|
|
183
|
+
is required or optional the installer renders `templates/ci/goblin-gate.yml.tmpl` into
|
|
184
|
+
`.github/workflows/goblin-gate.yml`, and when it is `-` the artifact must be **absent** (which is
|
|
185
|
+
why `CL-01` keys off that exact path, not the `.github/` directory — a repo is still allowed CI of
|
|
186
|
+
its own). `docs/CI.md` is the contract for what that file does and does not make true.
|
|
187
|
+
|
|
188
|
+
## The ban list (G5)
|
|
189
|
+
|
|
190
|
+
`BN-00`..`BN-09` are not ordinary rows: they read `.goblin/manifest/bans.tsv`, a table whose
|
|
191
|
+
every row carries a real command. A ban with no mechanism is a wish, so `manifest/bans.tsv`
|
|
192
|
+
holds `id`, the ban, the globs, the `detect` command, the replacement code, the escape hatch,
|
|
193
|
+
the reviewer and the source — and `BN-00` fails the whole list if any ban has no enforcement
|
|
194
|
+
row, if any row names no replacement, or if the table is empty. `bin/goblin-bans` is the engine
|
|
195
|
+
(`--only <id>` for one, `--list` for the table); each ban's `detect` exits 0 when the tree is
|
|
196
|
+
clean, 1 when the ban is violated, 3 when it cannot be read (no file matches its globs, no
|
|
197
|
+
`layers:` declared) and 2 when it could not run at all — a 2 FAILS, never passes.
|
|
198
|
+
|
|
199
|
+
Which bans apply is the config's `bans:` list, not the class: an unlisted ban SKIPs with that
|
|
200
|
+
reason (class A and C turn on `BN-01 BN-02 BN-05`; B and D turn on none; E turns on `BN-02`).
|
|
201
|
+
`bans_exempt:` records narrow, reviewed exceptions and `layers:` is what `BN-05` reads. An
|
|
202
|
+
exception reaches the **probe**, not the engine's stdout: the engine exports `GOBLIN_BANS_ID` and
|
|
203
|
+
`GOBLIN_BANS_EXEMPT`, the probe drops the exempted hits *before* it chooses its exit code, and an
|
|
204
|
+
inline `// BAN-OK(<id>): <reason>` on the offending line clears that one line. Filtering stdout
|
|
205
|
+
after the probe had already decided was decorative and made every exemption a permanent RED
|
|
206
|
+
(W5-1, measured `rc 0` at `72490f0` → `rc 1` at `7fec08f`); the trade is that a project's **own**
|
|
207
|
+
probe must honour the two variables or the exception fails **closed** (`docs/LIMITS.md` #36). The
|
|
208
|
+
probes are text probes — `grep`, no `npm`, no AST — so their false-positive sets are stated on
|
|
209
|
+
each row and in `docs/LIMITS.md` #27, and G5's `BN-04` is deliberately unshipped for the same
|
|
210
|
+
reason: it cannot be mechanised without a parser, and a ban that cannot go red is worse than an
|
|
211
|
+
advisory.
|
|
212
|
+
|
|
213
|
+
## How a rule is added
|
|
214
|
+
|
|
215
|
+
A rule enters only by adding a row. If the author cannot write a command, the row's `check`
|
|
216
|
+
is `advisory` and the rule's cost is visible in the summary line. There is no third option,
|
|
217
|
+
and `IN-03` plus `SK-03` are what make that true: a row with neither a command nor the
|
|
218
|
+
`advisory` label fails the manifest before any check runs (exit 3), and the advisory count
|
|
219
|
+
fails the run when it exceeds the ceiling.
|
|
220
|
+
|
|
221
|
+
## What the matrix cannot do
|
|
222
|
+
|
|
223
|
+
A counted rule is still not an enforced one. The cap is a policy, not a proof, and nine of
|
|
224
|
+
these rows are prose. That is stated here rather than implied away.
|
|
225
|
+
|
package/docs/FLOWS.md
ADDED
|
@@ -0,0 +1,164 @@
|
|
|
1
|
+
# The flow catalogue - 15 playbooks
|
|
2
|
+
|
|
3
|
+
`manifest/playbooks.tsv` is the machine-readable form; this is the prose. Every playbook has
|
|
4
|
+
the same six fields, and `verification` is always a *measurable* step that also names what it
|
|
5
|
+
cannot see. The router that picks one is the `goblin-mode` skill.
|
|
6
|
+
|
|
7
|
+
## P1 - `goblin-investigation`
|
|
8
|
+
|
|
9
|
+
- **When:** a read-only question, or "why is this happening"
|
|
10
|
+
- **Steps:** 1 name the question as a falsifiable claim<br>- 2 read the code paths, cite file:line<br>- 3 if the answer is observable by running something, run it instead of asking<br>- 4 write the answer with its evidence
|
|
11
|
+
- **Verification:** every claim carries a file:line or a command+output; no code changes; a claim you cannot source is marked [unverified]
|
|
12
|
+
- **Profiles:** any
|
|
13
|
+
- **Role:** investigate
|
|
14
|
+
|
|
15
|
+
## P2 - `goblin-bugfix`
|
|
16
|
+
|
|
17
|
+
- **When:** a reported defect
|
|
18
|
+
- **Steps:** 1 reproduce it yourself<br>- 2 state the root cause with a measurement, never 'should be'<br>- 3 fix the root, not the symptom<br>- 4 prove absence on the same surface, with a negative control
|
|
19
|
+
- **Verification:** the repro fails before and passes after, on the same command; unit tests show branch behaviour, not bug absence
|
|
20
|
+
- **Profiles:** coder
|
|
21
|
+
- **Role:** code
|
|
22
|
+
|
|
23
|
+
## P3 - `goblin-feature`
|
|
24
|
+
|
|
25
|
+
- **When:** new behaviour
|
|
26
|
+
- **Steps:** 1 SPEC first (measured root cause + AC: list)<br>- 2 name the data shape before the code<br>- 3 land it in units that each end checkable<br>- 4 write the SHA it landed at into the review note
|
|
27
|
+
- **Verification:** every AC: item has a checkable assertion; the gate line is measured; the SHA is named
|
|
28
|
+
- **Profiles:** architect -> coder
|
|
29
|
+
- **Role:** judgment -> code
|
|
30
|
+
|
|
31
|
+
## P4 - `goblin-refactor`
|
|
32
|
+
|
|
33
|
+
- **When:** a behaviour-preserving reshape
|
|
34
|
+
- **Steps:** 1 pin the contract first (characterization test / snapshot / equivalence harness)<br>- 2 shape only, no behaviour<br>- 3 delete the legacy path in the same change
|
|
35
|
+
- **Verification:** the pin is a real assertion run before and after; a type check and lint are not a pin
|
|
36
|
+
- **Profiles:** coder
|
|
37
|
+
- **Role:** code
|
|
38
|
+
|
|
39
|
+
## P5 - `goblin-tdd-repro`
|
|
40
|
+
|
|
41
|
+
- **When:** a defect where a regression test is cheap
|
|
42
|
+
- **Steps:** 1 write the failing test<br>- 2 confirm it fails for the intended reason<br>- 3 smallest production fix<br>- 4 revert the fix -> the test MUST fail -> restore
|
|
43
|
+
- **Verification:** the RED-again step is captured; prefer no new test over a bad test; the skip path is explicit, never silent
|
|
44
|
+
- **Profiles:** coder
|
|
45
|
+
- **Role:** code
|
|
46
|
+
|
|
47
|
+
## P6 - `goblin-verify-author`
|
|
48
|
+
|
|
49
|
+
- **When:** a project has no live lane, or its gates drift
|
|
50
|
+
- **Steps:** 1 read the repo, not the user, for entry points<br>- 2 write the gate/check set into the harness dir with the house harness shape<br>- 3 execute it once end to end<br>- 4 add the REPLAY block<br>- 5 seed the feature map and declare `feature_map:`/`source_root:` (`goblin-feature-map`)<br>- 6 hand the generated skill to P12: `verified:` does not advance until the eval record exists
|
|
51
|
+
- **Verification:** a generated skill that was never executed is a draft, not a deliverable; the harness prints PASS/FAIL and exits non-zero on failure; the REPLAY shows RED pre-change; the map's index and four-H2 entry contract hold (`FM-01`), every declared entry path still resolves (`FM-02`), and the declared `verify_doctor:` exits 0 (`VA-01`)
|
|
52
|
+
- **Profiles:** architect
|
|
53
|
+
- **Role:** judgment
|
|
54
|
+
|
|
55
|
+
## P7 - `goblin-pr-gate`
|
|
56
|
+
|
|
57
|
+
- **When:** anything that should be reviewed before it lands
|
|
58
|
+
- **Steps:** 1 classify stakes S0-S4<br>- 2 run the gate set at the candidate SHA and record the numbers<br>- 3 write reviews/<slug>-<head7>.md with head/base/patch-id/stakes/checks-run<br>- 4 evaluate the panel rule for S3+<br>- 5 at S3+ the foreman is `role-judge`: N lane verdicts are opinions, and one judge turns them into a decision<br>- 6 re-check the patch-id before landing
|
|
59
|
+
- **Verification:** the patch-id of base..head still matches the recorded one; the review note names a SHA that exists in git rev-list; for S2+ a check ran on that SHA; at S3+ the deciding lane is disjoint from the author's (`JG-02`)
|
|
60
|
+
- **Profiles:** reviewer (+ architect for S3)
|
|
61
|
+
- **Role:** review-panel
|
|
62
|
+
|
|
63
|
+
## P8 - `goblin-bootstrap`
|
|
64
|
+
|
|
65
|
+
- **When:** adopting goblin-stack in a repo, or starting one
|
|
66
|
+
- **Steps:** 1 classify the project (A-F)<br>- 2 goblin-install --class <x><br>- 3 goblin-verify GREEN<br>- 4 fix .gitignore BEFORE any git init<br>- 5 first HANDOFF, first SPEC, first check script
|
|
67
|
+
- **Verification:** goblin-verify exits 0 and the created-file list matches installed.json; a repo with no gate declares one and records its first measured numbers
|
|
68
|
+
- **Profiles:** architect
|
|
69
|
+
- **Role:** judgment
|
|
70
|
+
|
|
71
|
+
## P9 - `goblin-handoff`
|
|
72
|
+
|
|
73
|
+
- **When:** ending a session, or picking up another's
|
|
74
|
+
- **Steps:** 1 commit uncommitted edits as one internally-consistent wip: commit<br>- 2 write intent / verified state / next steps / what is NOT verified<br>- 3 stale sentences get a dated parenthetical, never deletion<br>- 4 on pickup: verify inherited claims against the artifact
|
|
75
|
+
- **Verification:** every gate number in the HANDOFF carries 'measured <date>'; the named HEAD matches git rev-parse --short HEAD; the NOT-verified section is non-empty or says 'nothing outstanding'
|
|
76
|
+
- **Profiles:** any
|
|
77
|
+
- **Role:** judgment
|
|
78
|
+
|
|
79
|
+
## P10 - `goblin-overnight`
|
|
80
|
+
|
|
81
|
+
- **When:** an unattended run over a predicate
|
|
82
|
+
- **Steps:** 1 the exit condition is a checkable predicate written before iteration 1<br>- 2 it never gets relaxed<br>- 3 an escape hatch: a genuine dead end writes up why and stops<br>- 4 the morning audit reads the Attention section first
|
|
83
|
+
- **Verification:** the predicate is a command, and its first run is recorded before iteration 1 (`LP-01`); the predicate is pinned and never relaxed (`LP-02`); `goal_max_turns` is set and at or under `loop_max_turns_ceiling` (`LP-03`); no three consecutive rows share an evidence pointer without reaching `predicate:green` (`LP-04`); a run that ends without its predicate green carries `.goblin/loop/stuck.md` naming it (`LP-05`); every landed change has a P7 verdict row; a verdict that says `done` cites a handle the repo resolves (`JG-01`) and comes from a lane disjoint from the author's (`JG-02`)
|
|
84
|
+
- **Profiles:** default + coder
|
|
85
|
+
- **Role:** code + judge
|
|
86
|
+
|
|
87
|
+
## P11 - `goblin-sweep`
|
|
88
|
+
|
|
89
|
+
- **When:** the same change or question across projects
|
|
90
|
+
- **Steps:** 1 enumerate targets with a shell glob, not a memory<br>- 2 classify each (A-F); an archive project is skipped, not processed<br>- 3 one card per project, parents=[sweep]<br>- 4 collect one line per project: what changed / what was refused / what is unfindable
|
|
91
|
+
- **Verification:** the per-project line carries the command it ran; the sweep report states its own coverage (n of m projects, and names the skipped ones)
|
|
92
|
+
- **Profiles:** default
|
|
93
|
+
- **Role:** synthesis
|
|
94
|
+
|
|
95
|
+
## P12 - `goblin-eval`
|
|
96
|
+
|
|
97
|
+
- **When:** a skill or prompt changed, and you want to know if it did anything
|
|
98
|
+
- **Steps:** 1 candidate and control run in sanitized directories<br>- 2 no eval/test/judge/rubric token anywhere the candidate sees<br>- 3 grade the chain from the transcript (which files it actually opened), never self-report<br>- 4 the judge runs on a different model family<br>- 5 write the record to `evals/<slug>/` (prompt, rubric, manifest.tsv, transcripts/, verdict.md)<br>- 6 a change must improve the evaluated cases or add new evaluations; a generated verification skill passes only on a measured sensitivity
|
|
99
|
+
- **Verification:** the judge's verdict is reproducible from the transcripts; candidates never learn other candidates exist; the record exists and every lane names a transcript file that exists; a generated verification skill is verified only by a record, with its sensitivity printed beside the control's number
|
|
100
|
+
- **Profiles:** researcher
|
|
101
|
+
- **Role:** synthesis
|
|
102
|
+
|
|
103
|
+
## P13 - `goblin-bugreporter`
|
|
104
|
+
|
|
105
|
+
- **When:** an event delivered a report — a bug report file, a chat message turned into one, a webhook
|
|
106
|
+
- **Steps:** 1 validate intake (the six required keys; a missing key is a refusal card with no assignee)<br>- 2 freeze `repo` + `revision` as immutable<br>- 3 reproduce (R1 a failing command, then R2 the REPLAY, then R3 a real-UI drive)<br>- 4 write `reports/<slug>/repro.md` with the `pre`/`post` table<br>- 5 create the fix card only on `reproduced`<br>- 6 complete its own card with the verdict
|
|
107
|
+
- **Verification:** `repro.md` carries a command, a revision that exists in `git rev-list`, and a RED `pre` row; the fix card exists iff the verdict is `reproduced`; `git status --porcelain` is empty after the run (`AU-02`, `AU-03`)
|
|
108
|
+
- **Profiles:** researcher
|
|
109
|
+
- **Role:** investigate
|
|
110
|
+
|
|
111
|
+
## P14 - `goblin-drift-audit`
|
|
112
|
+
|
|
113
|
+
- **When:** a recorded claim disagrees with the artifact (drift)
|
|
114
|
+
- **Steps:** 1 enumerate targets by glob, never by memory<br>- 2 compute the drift record per target<br>- 3 print nothing when clean<br>- 4 one card per drifting repo<br>- 5 report `n of m`, naming the skipped targets
|
|
115
|
+
- **Verification:** the summary names each repo and the command it ran; a clean run prints nothing; skipped targets are named; a capped run prints its own line, so "silent because clean" and "silent because capped" are never confused
|
|
116
|
+
- **Profiles:** architect
|
|
117
|
+
- **Role:** judgment
|
|
118
|
+
|
|
119
|
+
**Why a 13th and 14th playbook, rather than folding these into P1/P9.** Every other playbook is
|
|
120
|
+
entered by a *human or orchestrator request* and delivers a *change*. These two are entered by an
|
|
121
|
+
*event* and deliver a *card*. `P10` covers "an unattended run over a predicate" — a run over a
|
|
122
|
+
*condition*, not a run *started by* a condition. Stretching P1 would lose the intake gate;
|
|
123
|
+
stretching P9 would lose the reproduce-first gate. The producer half is
|
|
124
|
+
`automations/drift-audit.sh` (no agent at all) and `automations/bugreporter-intake.sh`; the
|
|
125
|
+
three-part model is in `automations/README.md`.
|
|
126
|
+
|
|
127
|
+
## P15 - `goblin-re-mobile`
|
|
128
|
+
|
|
129
|
+
- **When:** one shipped Android build must be understood as facts for study, with a reproducible, hash-manifested corpus
|
|
130
|
+
- **Steps:** 1 S0 preflight: the sandbox exists and is the one the fences describe<br>- 2 S1 acquire, S2 verify provenance against the published hash<br>- 3 S3 triage: the engine verdict, cheapest test first<br>- 4 S4 static decompile, S5 carve the containers, S6 manifest the corpus<br>- 5 S7 dossier: facts and numbers, never expression<br>- 6 S8 is deferred by design; S9 retention/teardown
|
|
131
|
+
- **Verification:** the corpus manifest verifies `sha256sum -c` where the corpus lives; `RC-01`, `RC-02` and `RC-03` return the exits their rows define (an exact hash inside the build output, a weak manifest and a tracked payload each fail the build); every negative control NC-1..NC-6 was shown RED and then restored
|
|
132
|
+
- **Profiles:** coder
|
|
133
|
+
- **Role:** code
|
|
134
|
+
|
|
135
|
+
**The step list above is the summary, not the procedure**, and the four `RC-` rows are what make
|
|
136
|
+
its verification column measurable rather than aspirational: `RC-01` is the build-time gate over
|
|
137
|
+
the declared `security: build_output:`, `RC-02` is the vacuous-pass guard on the manifest's own
|
|
138
|
+
shape, `RC-03` is the quarantine rule over the lab repo's tracked tree, and `RC-04` is the
|
|
139
|
+
acquisition record. The procedure's non-negotiable fences (an owned build only, one dedicated
|
|
140
|
+
sandbox, the quarantine, nothing extracted entering a repo) are stated with what enforces each.
|
|
141
|
+
|
|
142
|
+
## The cuts - pstack ships 23, this ships 15 (the two automations included)
|
|
143
|
+
|
|
144
|
+
Each cut has a reason, and a cut is recorded rather than deleted silently.
|
|
145
|
+
|
|
146
|
+
| cut | verdict | reason |
|
|
147
|
+
|---|---|---|
|
|
148
|
+
| `perf-issue` | folded into P2 | Its mechanism is "baseline trace, post-fix trace, diff the artifacts" - a step of a bug fix, not a playbook. No project here has a perf-target loop. |
|
|
149
|
+
| `hillclimb` | cut, deferred | Needs a frozen harness with proven sensitivity plus a loop primitive; the kanban's `goal_mode` already provides re-entry with a budget cap. |
|
|
150
|
+
| `runtime-forensics`, `trace-forensics` | merged into P2 | The transferable step is "capture a real artifact, then inject instrumentation into the running process". The library already ships the runners (`node-inspect-debugger`, `python-debugpy`), so a playbook would restate an existing skill. |
|
|
151
|
+
| `prototype` | merged into P1 step 3 | The valuable half is the classifier: a question whose answer is observable by running something is not the human's to answer. |
|
|
152
|
+
| `visual-parity` | cut | Needs a baseline screenshot harness; no project has a pixel-parity migration target. |
|
|
153
|
+
| `authoring-a-skill` | cut | `hermes-agent-skill-authoring` is live in the global library and is a Hermes-native duplicate. |
|
|
154
|
+
| `autopilot-stack` | cut | Nothing to stack: measured 0 branches and 0 PR merges across the repos this was designed for. |
|
|
155
|
+
| `autopilot-full`, `orchestrate` | merged into P10 + P11 + kanban | The fleet-programme half is the board's job (`parents`, `goal_mode`, `request_review`). The patch-id and different-family verifier rules they contain are kept as PG-03 and MD-02. |
|
|
156
|
+
| `multi-phase-plan` | merged into P3 + P8 | The standard already owns the SPEC lifecycle; duplicating it violates the one-owner rule. The machine-checkable half is kept as SP-01/02/03. |
|
|
157
|
+
| `worktree-cleanup` | cut | macOS/Xcode-specific by inspection (`xcrun simctl`, `DerivedData`). |
|
|
158
|
+
| `opening-a-pr` | merged into P7 | It is the terminal step of every other playbook; as its own playbook it would be a step file with one caller. Its PR-body schema becomes P7's review-note schema. |
|
|
159
|
+
|
|
160
|
+
**Deliberate non-imports:** the `swarm` and `arena` fan-out shapes (the axis is read-vs-write,
|
|
161
|
+
and five parallel lanes cost about five times the tokens), the cloud-agent lane (no per-agent
|
|
162
|
+
computer here), and the Slack automation lane (the *shape* transfers, and P11 plus cron is
|
|
163
|
+
that shape).
|
|
164
|
+
|