@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/CHANGELOG.md +351 -0
  2. package/LICENSE +21 -0
  3. package/README.md +217 -2
  4. package/VERSION +1 -0
  5. package/adapters/_template/adapter.tsv +16 -0
  6. package/adapters/_template/detect.sh +10 -0
  7. package/adapters/_template/emit.sh +5 -0
  8. package/adapters/_template/verify.sh +4 -0
  9. package/adapters/claude/adapter.tsv +8 -0
  10. package/adapters/claude/detect.sh +8 -0
  11. package/adapters/claude/verify.sh +47 -0
  12. package/adapters/codex/adapter.tsv +12 -0
  13. package/adapters/codex/detect.sh +9 -0
  14. package/adapters/codex/verify.sh +45 -0
  15. package/adapters/copilot/adapter.tsv +10 -0
  16. package/adapters/copilot/detect.sh +8 -0
  17. package/adapters/copilot/verify.sh +45 -0
  18. package/adapters/cursor/adapter.tsv +11 -0
  19. package/adapters/cursor/detect.sh +10 -0
  20. package/adapters/cursor/verify.sh +45 -0
  21. package/adapters/gemini/adapter.tsv +15 -0
  22. package/adapters/gemini/detect.sh +11 -0
  23. package/adapters/gemini/verify.sh +49 -0
  24. package/adapters/hermes/adapter.tsv +9 -0
  25. package/adapters/hermes/detect.sh +8 -0
  26. package/adapters/hermes/verify.sh +27 -0
  27. package/adapters/opencode/adapter.tsv +14 -0
  28. package/adapters/opencode/detect.sh +9 -0
  29. package/adapters/opencode/verify.sh +45 -0
  30. package/automations/README.md +53 -0
  31. package/automations/bugreporter-intake.sh +145 -0
  32. package/automations/drift-audit.sh +139 -0
  33. package/automations/report.schema.tsv +10 -0
  34. package/bans/README.md +82 -0
  35. package/bans/grep-ban.sh +84 -0
  36. package/bans/layer-check.sh +57 -0
  37. package/bin/goblin +119 -0
  38. package/bin/goblin-audit +145 -0
  39. package/bin/goblin-bans +178 -0
  40. package/bin/goblin-doctor +233 -0
  41. package/bin/goblin-emit +484 -0
  42. package/bin/goblin-init +519 -0
  43. package/bin/goblin-install +720 -0
  44. package/bin/goblin-lib.sh +289 -0
  45. package/bin/goblin-model +105 -0
  46. package/bin/goblin-upgrade +572 -0
  47. package/bin/goblin-verify +2798 -0
  48. package/bin/goblin.js +103 -0
  49. package/docs/ADOPTION.md +168 -0
  50. package/docs/CI.md +187 -0
  51. package/docs/CONTRACTS.md +197 -0
  52. package/docs/DESIGN.md +92 -0
  53. package/docs/ENFORCEMENT.md +225 -0
  54. package/docs/FLOWS.md +164 -0
  55. package/docs/GUARDRAILS.md +126 -0
  56. package/docs/GUIDE.md +610 -0
  57. package/docs/INTEGRATION.md +92 -0
  58. package/docs/LIMITS.md +591 -0
  59. package/docs/LOOP.md +165 -0
  60. package/docs/RE-PLAYBOOK.md +183 -0
  61. package/docs/RISKS.md +70 -0
  62. package/docs/ROLES.md +105 -0
  63. package/manifest/bans.tsv +9 -0
  64. package/manifest/classes.tsv +61 -0
  65. package/manifest/enforcement.tsv +88 -0
  66. package/manifest/glossary.tsv +25 -0
  67. package/manifest/playbooks.tsv +16 -0
  68. package/package.json +37 -4
  69. package/presets/A-shipped-software.yaml +48 -0
  70. package/presets/B-service-config.yaml +40 -0
  71. package/presets/C-game.yaml +38 -0
  72. package/presets/D-knowledge.yaml +41 -0
  73. package/presets/E-fleet-config.yaml +42 -0
  74. package/presets/F-electron.yaml +67 -0
  75. package/roles.yaml +54 -0
  76. package/skills/goblin-bootstrap/SKILL.md +51 -0
  77. package/skills/goblin-bugfix/SKILL.md +26 -0
  78. package/skills/goblin-bugreporter/SKILL.md +52 -0
  79. package/skills/goblin-drift-audit/SKILL.md +43 -0
  80. package/skills/goblin-eval/SKILL.md +68 -0
  81. package/skills/goblin-feature/SKILL.md +26 -0
  82. package/skills/goblin-feature-map/SKILL.md +140 -0
  83. package/skills/goblin-handoff/SKILL.md +28 -0
  84. package/skills/goblin-investigation/SKILL.md +26 -0
  85. package/skills/goblin-judge/SKILL.md +74 -0
  86. package/skills/goblin-loop/SKILL.md +88 -0
  87. package/skills/goblin-mode/SKILL.md +70 -0
  88. package/skills/goblin-overnight/SKILL.md +42 -0
  89. package/skills/goblin-pr-gate/SKILL.md +42 -0
  90. package/skills/goblin-re-mobile/SKILL.md +51 -0
  91. package/skills/goblin-refactor/SKILL.md +23 -0
  92. package/skills/goblin-sweep/SKILL.md +23 -0
  93. package/skills/goblin-tdd-repro/SKILL.md +27 -0
  94. package/skills/goblin-verify-author/SKILL.md +50 -0
  95. package/skills/practice/SKILL.md +37 -0
  96. package/templates/AGENTS.md.tmpl +23 -0
  97. package/templates/HANDOFF.md.tmpl +43 -0
  98. package/templates/SPEC.md.tmpl +34 -0
  99. package/templates/audit-waiver.tsv.tmpl +10 -0
  100. package/templates/boundary-waivers.tmpl +8 -0
  101. package/templates/checks/assert.mjs.tmpl +60 -0
  102. package/templates/checks/gate.sh.tmpl +29 -0
  103. package/templates/ci/goblin-gate.yml.tmpl +46 -0
  104. package/templates/goblin.yaml.tmpl +138 -0
  105. package/templates/install-hooks.allowlist.tmpl +9 -0
  106. package/templates/loop/decisions.tsv.tmpl +1 -0
  107. package/templates/loop/predicate.tmpl +16 -0
  108. package/templates/report.yaml.tmpl +16 -0
@@ -0,0 +1,51 @@
1
+ ---
2
+ name: goblin-bootstrap
3
+ description: P8: adopt goblin-stack in a repo - classify, install, verify, then the first round.
4
+ ---
5
+
6
+ # goblin-bootstrap (P8)
7
+
8
+ Use when adopting goblin-stack in a repo, or starting one.
9
+
10
+ 1. **Classify the project A-F.** The class selects which parts are required, optional or off;
11
+ it is not a stringency level. F is a desktop shell: it adds the electron bans and a host gate.
12
+ 2. **`goblin-install --target <dir> --class <x>`**
13
+ 3. **`goblin-verify`** — a class-A install verifies green: `43 passed, 0 failed, 11
14
+ advisory, 28 skipped`, exit 0, once `HANDOFF.md` names a commit that exists; before that edit the
15
+ scaffold's `0000000` placeholder is `HP-05`'s one expected day-one red (`42 passed, 1 failed`).
16
+ Twenty-eight rows skip with a reason, and the reason matters: `HS-02`
17
+ (no pinned pre-change commit yet, so the REPLAY is not provable), `AU-02`/`AU-03` (no report
18
+ has been filed in this repo), `SC-06`/`SC-07`/`SC-08` (no dependency manifest, no lockfile, no
19
+ audit record), `PF-01` (no perf baseline measured yet), `BN-01`/`BN-02`/`BN-05` (the ban table is
20
+ installed but this fresh repo has no `src/` for a ban to read) and `BN-03` with the four electron
21
+ bans `BN-06`..`BN-09` (not in this class's `bans: [BN-01, BN-02, BN-05]`, so they skip as
22
+ *not enabled* rather than as *unread*),
23
+ `FM-01`/`FM-02`/`VA-01` (no feature map and no declared `verify_doctor:` yet), `RC-01`..`RC-04`
24
+ (no reference corpus declared: `reference_manifest:` ships empty and there is no lab
25
+ `manifests/`), and `JG-01` with
26
+ `LP-01`..`LP-05` (no `.goblin/loop/` record, because no loop has run in this repo yet). Each is
27
+ a *not yet*, not a pass.
28
+ The class's required parts
29
+ that only a round can produce (a first review, a real gate) pass *vacuously*, and that list
30
+ is the repo's first-step list, not a defect.
31
+ 4. **Fix `.gitignore` BEFORE any `git init`.** A credentials file already in the tree is
32
+ committed by the first `git add -A` and is then in history forever.
33
+ 5. **First HANDOFF, first SPEC, first check script** — in that order, each independently
34
+ useful.
35
+ 6. **Record the first perf number, and the commit you measured it on.** Run `ratchet.cmd` on a
36
+ pinned commit, write the value into `ratchet.ceiling` and the same number into
37
+ `perf.baseline_value`/`perf.baseline_commit`/`perf.measured`. Until you do, `PF-01` skips with
38
+ "no perf baseline recorded yet" — a green run that proves nothing about the number you ship.
39
+
40
+ ## Verification
41
+
42
+ - `goblin-verify` exit 0, and the created-file list matches `.goblin/installed.json`.
43
+ - A repo with no gate declares one and records its first measured numbers.
44
+ - `.gitignore` is verified before the first commit, not after.
45
+ - The perf baseline exists, and it names a commit that exists (`PF-01`), rather than the ratchet
46
+ ceiling having been raised by hand.
47
+
48
+ ## What this cannot see
49
+
50
+ Whether the class you chose matches how the project actually ships. Re-classify when it does
51
+ not; the installer records the class and `CL-01` checks the parts against it.
@@ -0,0 +1,26 @@
1
+ ---
2
+ name: goblin-bugfix
3
+ description: P2: reproduce a defect, fix its measured root cause, prove absence with a control.
4
+ ---
5
+
6
+ # goblin-bugfix (P2)
7
+
8
+ 1. **Reproduce it yourself**, with the exact command, before touching anything. A bug you did
9
+ not reproduce is a bug you will not fix.
10
+ 2. **State the root cause with a measurement.** Never "should be" and never "probably". Paste
11
+ the command and its output that shows the cause.
12
+ 3. **Fix the root, not the symptom.** A guard added at the call site of a wrong value is a
13
+ symptom fix; the wrong value is the bug.
14
+ 4. **Prove absence on the same surface, with a negative control.** The repro command fails
15
+ before and passes after. Then revert the fix, confirm the repro fails again, restore.
16
+
17
+ ## Verification
18
+
19
+ - The repro fails before and passes after, on the *same* command.
20
+ - A unit test shows branch behaviour, not bug absence. Do not report a green suite as proof
21
+ the bug is gone; report the repro.
22
+ - If the defect also exists in a sibling project, say so and do not silently fix it there.
23
+
24
+ ## What this cannot see
25
+
26
+ Bugs whose only symptom is a human noticing something looks wrong.
@@ -0,0 +1,52 @@
1
+ ---
2
+ name: goblin-bugreporter
3
+ description: P13: intake a bug report, reproduce it at a named revision, and hand off a fix card only on proof.
4
+ ---
5
+
6
+ # goblin-bugreporter (P13)
7
+
8
+ Use when a report has arrived as an **event** — a report file under `reports/<slug>/`, a chat
9
+ message turned into one, or a webhook payload. You are the reporter. You never fix.
10
+
11
+ 1. **Validate the intake.** `reports/<slug>/report.yaml` must carry the six required keys
12
+ (`repo`, `symptom`, `expected`, `observed`, `repro_steps`, `revision`) and its `dedup_key`.
13
+ Run the validator rather than eyeballing the file:
14
+
15
+ bash .goblin/automations/bugreporter-intake.sh <slug>
16
+
17
+ A missing key is a **refusal**, not a guess: create the same card with **no `--assignee`**,
18
+ so the dispatcher buckets it `skipped_unassigned` and no agent acts on it, and name the
19
+ missing field in a comment. The human sees it; nothing is silently lost or invented.
20
+ 2. **Freeze the coordinates.** `repo` and `revision` are immutable for the run. If `revision`
21
+ is not a commit that resolves (`git rev-parse --verify <rev>^{commit}`), the report is a
22
+ refusal: a repro at a revision nobody can check is not a repro.
23
+ 3. **Reproduce, in this order.**
24
+ - **R1** — a failing command with its output, at the named revision. The best form.
25
+ - **R2** — the REPLAY form: `git show <revision>:<file>` into a temp dir, run the probe
26
+ against it, require RED. Never a probe that is green on both trees.
27
+ - **R3** — a real-UI drive with captured evidence. Admissible only with a configured
28
+ control adapter, and **never** accepted as harness proof; it is evidence for a human.
29
+ 4. **Write `reports/<slug>/repro.md`** as a two-row table, not a story: `command:`/`revision:`,
30
+ a `pre:` row that is RED with the pasted output, and the `post:` row. The `pre` row is the
31
+ negative control and it is not optional — a single green run reproduces nothing.
32
+ 5. **Create the fix card only on `reproduced`.** `--assignee coder`, `--skill goblin-bugfix`,
33
+ `--parent <this card>`, the report's dedup key. The child waits in `todo` until this card is
34
+ `done`, so the gate is a dispatcher-enforced dependency rather than a promise.
35
+ 6. **Complete your own card with the verdict** in the completion metadata:
36
+ `reproduced` | `could-not-reproduce` | `blocked`. `could-not-reproduce` is a complete,
37
+ successful run — not a failure to work around. `blocked` names the missing capability.
38
+
39
+ ## Write surface
40
+
41
+ `reports/**`, and the card this lineage owns (`kanban_create` for the fix card). Nothing else:
42
+ no `git commit`, no source edit, no edit under the harness dir, no vault write. `AU-03` asserts
43
+ the working tree is clean after a run, so a reporter that edited the tree fails its own gate.
44
+
45
+ ## What this cannot see
46
+
47
+ Whether the report is *true*. The gate proves the symptom reproduces at a revision; it cannot
48
+ prove the reporter's description of intent, severity or cause is right, and it cannot see a fix
49
+ that fixed the wrong thing. It cannot see a defect whose only detector is a human eye, anything
50
+ needing a device, a network or a running production instance, or a defect that reproduces only
51
+ on a revision that no longer builds. In those cases the verdict is `could-not-reproduce`, and
52
+ that is the whole answer.
@@ -0,0 +1,43 @@
1
+ ---
2
+ name: goblin-drift-audit
3
+ description: P14: repair a recorded claim that no longer agrees with the artifact, one card per drifting repo.
4
+ ---
5
+
6
+ # goblin-drift-audit (P14)
7
+
8
+ Use when a drift card exists. The producer (`automations/drift-audit.sh`) compared a recorded
9
+ claim against an artifact, disagreed, and filed **one card per drifting repo**. You repair the
10
+ disagreement; you do not re-derive it from memory.
11
+
12
+ 1. **Read the record, not your memory.** The card body carries the repo, the check id and the
13
+ detail line the producer computed. Re-run that exact check yourself and confirm it is still
14
+ RED:
15
+
16
+ <repo>/.goblin/bin/goblin-verify --only <check-id>
17
+
18
+ A drift that is no longer reproducible is a stale card, and closing it with that fact is a
19
+ complete run.
20
+ 2. **Classify it.** A **real regression** (the artifact changed and the claim did not) | a
21
+ **stale-by-design claim** (the claim is historical, and `PROJECT-PRACTICE` section 1 requires
22
+ it kept, so it gets a dated parenthetical) | **fixture drift** (the recorded value is right
23
+ and the check is wrong).
24
+ 3. **Repair the right one.** A hash drift after an intended edit of the referenced standard is
25
+ a deliberate `--re-pin`, never a hand-edited hash. A missing installed file is a re-install.
26
+ An absent gate line is a gate run. **Never edit a harness to make a number green** — that is
27
+ the one repair that must never happen from here, and it is the repair an automated actor is
28
+ most likely to reach for.
29
+ 4. **Prove it.** Re-run the same `--only <check-id>` and show GREEN, and name the command and
30
+ the revision in the completion metadata.
31
+
32
+ ## Write surface
33
+
34
+ Only the drifting repo named by the card, only the repair the card names — the config, the
35
+ ignore file, the pin, or the install. No harness edit, no vault write, no push.
36
+
37
+ ## What this cannot see
38
+
39
+ Whether the *claim* is the right one to keep. It compares a recorded value with an artifact, so
40
+ it cannot detect drift in something nobody recorded: an unwritten assumption, a stale sentence
41
+ with no number in it, or a document that was never generated from the artifact. It also cannot
42
+ see whether the producer's ceiling stopped a real finding from being filed — that is why a
43
+ capped run prints its own line.
@@ -0,0 +1,68 @@
1
+ ---
2
+ name: goblin-eval
3
+ description: P12: measure whether a skill or prompt change did anything, blinded.
4
+ ---
5
+
6
+ # goblin-eval (P12)
7
+
8
+ Use when a skill or prompt changed and you want to know if it did anything — including when the
9
+ thing that changed is a **generated verification skill** (P6's output). This is the only mechanism
10
+ goblin-stack has for that question, and it has never been run — treat its output as a first
11
+ measurement, not a verdict. The runner is **not shipped**: every step below is a procedure an
12
+ operator executes, and no row in `manifest/enforcement.tsv` reads a lane yet.
13
+
14
+ 1. **Candidate and control run in sanitized directories.** Same task, same starting tree, one
15
+ difference. Sanitized means a **clean workspace with nothing borrowed**: no prior chats, no
16
+ prior artifacts, no prior executions. An agent that can reach a previous run can produce the
17
+ answer without using the skill, and then the eval measures nothing but the leak.
18
+ 2. **No condition token anywhere the candidate can see.** The ban is eleven tokens — `eval`,
19
+ `test`, `judge`, `experiment`, `rubric`, `score`, `compare`, `benchmark`, `candidate`,
20
+ `arena`, `control` — not only the four this skill used to name, and it covers the directory
21
+ names, the file names and the prompt. A candidate that knows it is being graded changes
22
+ behaviour.
23
+ 3. **Grade the chain from the transcript** — which files it actually opened, which commands it
24
+ ran — never the self-report.
25
+ 4. **The judge runs as `role-judge`, on a lane disjoint from the candidate's.** A model grading its
26
+ own family's output is grading itself; the lane that decides must not be the lane that wrote.
27
+ `JG-02` checks the declared lanes are disjoint, and reports an unresolved judge lane as `ADV`
28
+ rather than failing a repo for the fleet's routing. What the judge receives, what it must
29
+ refuse and how its verdict is recorded is `docs/LOOP.md` — the short form is: **a predicate and
30
+ a handle the repo can resolve, never a self-report.**
31
+ 5. **Write the record where the next reader looks.** One directory per eval, `evals/<slug>/`:
32
+ `prompt.md` (the one organic prompt, identical across lanes), `rubric.md` (3–6 criteria, for
33
+ the judge only), `manifest.tsv` (`lane · condition · label · profile · provider · model ·
34
+ family · workspace · card_id · session_id · transcript`), `transcripts/<label>.jsonl` (the
35
+ exported session, one file per lane) and `verdict.md` (the judge's per-criterion scores, each
36
+ with an `evidence:` line naming the lane and the location). The **family** is recorded, never
37
+ only the model slug — and **no model slug belongs under `skills/ manifest/ bin/ templates/
38
+ presets/ .goblin/ .hermes/`**; `evals/` sits outside that scope on purpose, because a slug is
39
+ run data, not content for a rule.
40
+ 6. **A change must improve the evaluated cases or add new evaluations.** That is the merge rule,
41
+ and it is this repo's "a check must be able to go RED" applied to skill prose: changing
42
+ instructions requires evidence about the resulting behaviour, not a judgement that the new
43
+ wording sounds better. **When the subject is a generated verification skill, the pass condition
44
+ is a number**: every seeded defect must make the harness RED, the control (the same defects
45
+ against a harness carrying no assertions) must detect **zero**, and every correction must be
46
+ RED before it is GREEN. A skill that misses a seeded defect has a hole; a sensitivity below
47
+ 1.0 is reported with the number beside it, never rounded up to a pass.
48
+ 7. **Cheap checks first; the judge last.** Regex and script assertions for what is concrete (the
49
+ right SDK, the right method, the absence of an obsolete pattern); an LLM judge with an explicit
50
+ rubric only for what needs interpretation. **Triggering is a diagnostic, not the objective** —
51
+ loading the skill on the fifth turn can be fine and completing the task without loading it can
52
+ be fine too, so the trigger result stays visible without defining pass/fail.
53
+
54
+ ## Verification
55
+
56
+ - The judge's verdict is reproducible from the transcripts alone.
57
+ - Candidates never learn that other candidates exist.
58
+ - Every lane named in `manifest.tsv` has a transcript file that exists.
59
+ - A generated verification skill is not "verified" until a record exists for it: its declared
60
+ doctor exits 0 (`VA-01` is the executable half), and the record carries the sensitivity.
61
+
62
+ ## What this cannot see
63
+
64
+ Small effects. With a handful of runs, a difference smaller than the run-to-run variance is
65
+ noise, and no amount of prose makes it a signal. It also cannot see **its own blinding**: on a
66
+ machine where a candidate holds terminal access, sibling workspaces, sibling cards and other
67
+ sessions are readable, so blinding is a discipline with a detector rather than a property. And
68
+ the runner is not shipped — with no record, this skill is a procedure, not a measurement.
@@ -0,0 +1,26 @@
1
+ ---
2
+ name: goblin-feature
3
+ description: P3: SPEC first, then land the feature in units that each end checkable.
4
+ ---
5
+
6
+ # goblin-feature (P3)
7
+
8
+ 1. **SPEC first.** A `*-SPEC.md` with a measured root cause and an `AC:` list, committed with
9
+ or before the code. No code moves until it exists.
10
+ 2. **Name the data shape before the code.** The shape decided late is the shape that gets
11
+ rewritten.
12
+ 3. **Land it in units that each end checkable.** Each unit ends with a measurement, not with
13
+ "and now the next part".
14
+ 4. **Write the SHA it landed at into the review note** (`goblin-pr-gate`, P7). "The reviewer
15
+ approved this" is a claim about a moment until it names a commit.
16
+
17
+ ## Verification
18
+
19
+ - Every `AC:` item has a checkable assertion: a backticked command, or a comparison operator.
20
+ - The round reports one line of measured numbers.
21
+ - The commit SHA is named.
22
+
23
+ ## What this cannot see
24
+
25
+ Whether the feature was the right thing to build. That is the SPEC's job, and it is a human's
26
+ decision.
@@ -0,0 +1,140 @@
1
+ ---
2
+ name: goblin-feature-map
3
+ description: The feature map - a per-project inventory of every user-facing feature, its entry points and the exact command that drives each one. Use when a project needs a map an agent can read cold, when a map's index or entry paths have drifted from the app, or when a verification skill must state which features it covers (P6 authors the map; FM-01/FM-02 keep it honest).
4
+ ---
5
+
6
+ # goblin-feature-map
7
+
8
+ A **feature map** is a directory of small files, one per user-facing feature, written for the
9
+ next agent rather than for a human: it is read cold, mid-task, by an agent that has never seen
10
+ the app. It answers "how do I reach this behaviour, with what command, and what should I see" —
11
+ without re-deriving entry points from source every session.
12
+
13
+ **It is an inventory, not a proof.** A complete map with every feature driven proves the features
14
+ behave *as described*; it proves nothing about whether the app is correct.
15
+
16
+ ## Where it lives, and who declares it
17
+
18
+ The map's location is **declared, never inferred**: `.goblin/goblin.yaml` holds
19
+
20
+ ```
21
+ feature_map: <path to the map README, relative to the repo> # empty -> FM-01/FM-02 SKIP
22
+ source_root: . # entry paths resolve against this
23
+ verify_doctor: "" # empty -> VA-01 SKIP
24
+ ```
25
+
26
+ An **empty `feature_map:` is the honest state of a project with no map yet** — `FM-01` and
27
+ `FM-02` then SKIP with that reason instead of failing a fresh install. Declaring a path and not
28
+ writing the map is a hole, and is reported as one.
29
+
30
+ ## The README — the index
31
+
32
+ `features/README.md`, four H2 sections, in this order, and no more:
33
+
34
+ ```
35
+ ## Baseline preconditions how to launch, isolate, seed, health-check
36
+ ## Driving conventions stable-handle rule, literal-command rule, restore rule
37
+ ## Proof and skip reporting what counts as proof; how a skip is reported
38
+ ## Features one line per feature file: - [Name](./<slug>.md) covers ...
39
+ ```
40
+
41
+ The `## Features` list is the index. **Every `features/*.md` file must appear in it, in the
42
+ `](./<slug>.md)` form, and every relative link must resolve** (`FM-01`). The README's own
43
+ sections are prose: no check reads them, so nothing here pretends they are enforced.
44
+
45
+ ## Each feature file — the entry contract
46
+
47
+ `features/<slug>.md`: frontmatter, then **exactly four H2s, in this order**.
48
+
49
+ ```
50
+ ---
51
+ feature: <slug> # must equal the filename stem
52
+ entry_paths: # >=1 single-token strings; each must still occur under source_root
53
+ - <route-or-path-token>
54
+ verified: <YYYY-MM-DD> # the date a human or an audit last drove this feature
55
+ ---
56
+ # <Feature name>
57
+ <one paragraph, user-visible behaviour, no implementation detail>
58
+
59
+ ## Sub-features
60
+ - `<id>` <one line>
61
+
62
+ ## How to get to it (user POV)
63
+ - one bullet per entry point
64
+
65
+ ## Driving it with <harness>
66
+ Preconditions: <what must be true first>
67
+ **<action>.** Run `<exact command>`. <observable result>
68
+
69
+ ## Gotchas
70
+ - traps that waste or invalidate a run
71
+ ```
72
+
73
+ `<harness>` is this project's own — a browser driver, a CLI, a test runner. Keep an entry path a
74
+ **single token** (a route, a file path, a CLI flag) that occurs verbatim under `source_root`,
75
+ because that is what `FM-02` greps for. Pair every user action with an **exact command** and an
76
+ **observable result**, or the entry is not drivable and belongs in `## Gotchas`.
77
+
78
+ ## What makes it rot, and what keeps it honest
79
+
80
+ | rot | what it looks like | what catches it |
81
+ |---|---|---|
82
+ | index rot | a feature file added or deleted, README not updated | `FM-01`: every feature file linked, every link resolves |
83
+ | contract rot | a section renamed, a fifth H2 added, the four reordered | `FM-01`: the four H2s, in order, exactly |
84
+ | entry-point rot | a route renamed, a flag changed — the recipe drives a dead path | `FM-02`: every `entry_paths:` token still occurs under `source_root:` |
85
+ | staleness | the app changed after the map was written; the recipe "works" by luck | `FM-02`: no entry path's file may have a commit newer than the feature's `verified:` date |
86
+ | coverage rot | a proof drives one convenient entry point while the map lists three | **not mechanically checkable** — see "what this cannot see" |
87
+ | implementation creep | the map freezes internals meant to be runtime discoveries | contract only; resist it in review |
88
+
89
+ **`FM-02` is a tripwire, not a proof.** `git log` sees a *file* change, not a *behaviour* change:
90
+ a refactor touching the file without changing the route reports stale-and-wrong, and a behaviour
91
+ change in a file the token does not appear in is missed. It can be RED-when-stale; it can never
92
+ mean "verified fresh". A `verified:` date is a **claim**, and nothing distinguishes "driven
93
+ yesterday" from "dated yesterday".
94
+
95
+ ## Upkeep — the pass that keeps a map alive
96
+
97
+ A map rots the moment the app changes, so it needs a periodic pass, not a one-time write. One
98
+ person (or one agent) per pass, in this order, with an edit scope limited to the map's own
99
+ directory:
100
+
101
+ 1. **Index hygiene.** Run `goblin-verify --only FM-01` and fix the index and the four-H2
102
+ contract.
103
+ 2. **Source wave.** One read-only subagent (or one focused read) per feature file, launched
104
+ concurrently: it reads the app and reports what the file now gets wrong. **Children never
105
+ drive and never edit.**
106
+ 3. **Reconcile.** Apply only the corrections the source wave proved; a correction the subagent
107
+ could not evidence is left alone and named.
108
+ 4. **Live pass.** Drive every feature at least once, with the command in the file. Run the
109
+ project's declared doctor (`verify_doctor:`, enforced by `VA-01`) before the first drive and
110
+ after any failed one. A feature that cannot be reached is *verified-unreachable* only with the
111
+ concrete prerequisite and the route attempted — never reported as verified through another
112
+ path.
113
+ 5. **Triage.** Sort findings into **doc drift** (the map was wrong), **harness gap** (the map was
114
+ right and the check did not exist) and **product gap** (the app is wrong). Doc drift and
115
+ harness gaps ship; a product gap is reported and kept out of the map's change.
116
+ 6. **One pull request of proven corrections**, re-reading every changed file first, and advance
117
+ `verified:` only for features actually driven. The outcome is one of `clean` / `changed` /
118
+ `blocked`, stated in writing.
119
+
120
+ ## The commands
121
+
122
+ ```
123
+ goblin-verify --only FM-01 # the index and the four-H2 entry contract
124
+ goblin-verify --only FM-02 # every entry path still resolves; no entry path changed after verified:
125
+ goblin-verify --only VA-01 # the declared verify_doctor: exits 0
126
+ ```
127
+
128
+ ## What this cannot see
129
+
130
+ - **A feature nobody listed.** The map is a list; whether the list is complete needs semantic
131
+ judgement over the app, and no command in this repo answers it. That is why it is a note here
132
+ and not a row.
133
+ - **Whether a recipe drives the *right* path rather than a convenient one.** The map lists entry
134
+ points; the driver still chooses.
135
+ - **Whether a `verified:` date means anything.** The row proves the entry path still exists; it
136
+ cannot tell a feature driven yesterday from a date typed yesterday.
137
+ - **A defect the project's own assertions do not cover.** That is the mutation pass's job, not the
138
+ map's.
139
+ - **Behaviour changes that leave the file alone.** `FM-02` watches files, and a behaviour can
140
+ change without one.
@@ -0,0 +1,28 @@
1
+ ---
2
+ name: goblin-handoff
3
+ description: P9: write or pick up a HANDOFF - state, gates, next steps, and what is NOT verified.
4
+ ---
5
+
6
+ # goblin-handoff (P9)
7
+
8
+ Use when ending a session, or picking up another's.
9
+
10
+ 1. **Commit uncommitted edits** as one internally consistent `wip:` commit. A run that dies
11
+ must lose nothing.
12
+ 2. **Write intent / verified state / next steps / what is NOT verified.** The last one is the
13
+ most important and the most often skipped.
14
+ 3. **Stale sentences get a dated parenthetical, never deletion.** Deleting erases the fact
15
+ that it was once believed true; leaving it undated re-arms the trap.
16
+ 4. **On pickup, verify inherited claims against the artifact.** A passing prior self-report is
17
+ not the proof. Re-run the gate, check the file exists, grep the vault.
18
+
19
+ ## Verification
20
+
21
+ - Every gate number in the HANDOFF carries `measured <date>`.
22
+ - The named HEAD matches `git rev-parse --short HEAD`.
23
+ - The NOT-verified section is non-empty, or explicitly says "nothing outstanding".
24
+
25
+ ## What this cannot see
26
+
27
+ Whether the prose is true. A HANDOFF's own top blocks drift out of sync with its body — grep
28
+ the whole file for a status before planning from the first hit.
@@ -0,0 +1,26 @@
1
+ ---
2
+ name: goblin-investigation
3
+ description: P1: answer a read-only question with cited evidence, or mark it unverified.
4
+ ---
5
+
6
+ # goblin-investigation (P1)
7
+
8
+ For a read-only question, or "why is this happening". **No code changes.**
9
+
10
+ 1. **Name the question as a falsifiable claim.** "The board writes to the wrong path" is a
11
+ claim; "the board is broken" is not.
12
+ 2. **Read the code paths and cite `file:line`.** A claim with no anchor is a rumour.
13
+ 3. **If the answer is observable by running something, run it instead of asking.** A question
14
+ whose answer is a fact a command can print is not the human's to answer. This is the whole
15
+ classifier: *observable by running* -> run it; *a preference or a decision* -> ask.
16
+ 4. **Write the answer with its evidence.** Command and real output, or `file:line`.
17
+
18
+ ## Verification
19
+
20
+ - Every claim carries a `file:line` or a command plus its output.
21
+ - No file is modified. Prove it: `git status --short` is unchanged from before the pass.
22
+ - Anything you could not source is marked `[unverified]`. There is no third state.
23
+
24
+ ## What this cannot see
25
+
26
+ Intent, and anything outside the repo — a service that is down, a decision not written down.
@@ -0,0 +1,74 @@
1
+ ---
2
+ name: goblin-judge
3
+ description: role-judge: decide whether a process met its own predicate, from a command's output and a pointer it can resolve - never from a report.
4
+ ---
5
+
6
+ # goblin-judge (role-judge)
7
+
8
+ A **reviewer inspects an artifact**. A **judge decides whether a process met its own predicate**.
9
+ Different input, different output, different failure mode, which is why the judge is a role of its
10
+ own and not a sentence inside another playbook.
11
+
12
+ | | reviewer (`role-review-panel`) | judge (`role-judge`) |
13
+ |---|---|---|
14
+ | input | an artifact: a diff at a SHA, a file, a spec | a **predicate** plus the evidence the predicate produced |
15
+ | question | "is this artifact right?" | "did the process meet the condition it declared?" |
16
+ | output | findings, categorised (act / consider / noted / dismissed) | a verdict from a **closed enum**, plus the handle it rests on |
17
+ | may accept | the artifact itself, at a named revision | only what the verdict can be re-derived from |
18
+ | failure mode | misses a defect, or bikesheds | **says yes** - a judge that has never returned "not done" |
19
+ | retry | re-review, possibly by another lane | re-run the *same* predicate on the recorded evidence |
20
+ | relation to the author | may be a sibling lane | **not the author, and not the author's profile** |
21
+
22
+ A reviewer's verdict is an opinion attached to a SHA, and a new head voids it. A judge's verdict
23
+ is a decision about a **process**, so it is honest only when it is reproducible from a record and
24
+ useful only when it can say **no**.
25
+
26
+ ## What it receives
27
+
28
+ - **A predicate** - one shell command, the loop's declared exit condition. Never a rubric.
29
+ - **Evidence** - a pointer the repo can resolve: `sha:<hex>` (a commit), `file:<path>`, or
30
+ `sha256:<hex>` (a file's digest). **Never the author's prose.** The single most common failure
31
+ of a judging loop is that its only input is the worker's own summary; the prompt can demand
32
+ evidence and still be handed a self-report.
33
+ - The **budget** and the iteration rows, so it can see what the loop has already tried.
34
+
35
+ ## What it outputs
36
+
37
+ One verdict, from a closed enum: **`done` | `blocked` | `continue` | wait** - plus, in this repo's
38
+ record, the **handle** the verdict rests on. `done` may only cite a handle the repo resolves.
39
+ Anything else is a claim.
40
+
41
+ ## What it must refuse
42
+
43
+ - **A self-report as evidence.** A summary is not a handle. If the only thing offered is prose,
44
+ the verdict is `continue`, not `done`.
45
+ - **Grading its own family.** The judge's profile set must be disjoint from the author's
46
+ (`JG-02`). A judge drawn from the same lane as the author is the author.
47
+ - **Ruling a goal done to unblock itself.** `blocked` blocks; it is never a completion. A
48
+ judged-done-but-unfinalised loop gets one nudge, then a block - never a completion.
49
+
50
+ ## Where the judge is called
51
+
52
+ 1. **P10 `goblin-overnight`** - the exit predicate.
53
+ 2. **P12 `goblin-eval`** - the rubric, graded from the transcript chain.
54
+ 3. **P7 `goblin-pr-gate` at S3+** - the panel's *foreman*: N lane verdicts are opinions, one
55
+ judge turns them into a decision. A panel is N opinions; a judge is one decision.
56
+ 4. **The terminal handoff gate** - the card is judged *before* it is marked done.
57
+ 5. **An automation's "did the fix land"** - the repro fails before, passes after, twice.
58
+
59
+ ## How a verdict is recorded
60
+
61
+ One row per iteration in `.goblin/loop/decisions.tsv` (the decision-log shape pstack already
62
+ uses): `ts · phase · decision · why · evidence · result`, where `decision` is `verdict:<one of the
63
+ enum>` and `evidence` is a handle. `JG-01` reads it; `JG-03` counts it. On the board, the same
64
+ turn also leaves one line in the worker log, which is a UI surface - the committed row is what a
65
+ later reader can re-derive from.
66
+
67
+ ## What this cannot see
68
+
69
+ - Whether a **resolvable handle supports the verdict** it is attached to. `JG-01` proves a commit
70
+ or a file exists; a judge may cite a real commit that has nothing to do with the claim.
71
+ - **Which lane actually produced a verdict.** No repo-local file observes which profile ran;
72
+ `JG-02` checks that the *declared* lanes are disjoint, not that a judge ran on one.
73
+ - Whether a lane that has always said yes is **bad**, from one repo's rows (`JG-03`, counted).
74
+ - Whether the predicate was the **right** one. The judge measures what it was told to measure.
@@ -0,0 +1,88 @@
1
+ ---
2
+ name: goblin-loop
3
+ description: Arm, pin, run and close one unattended loop whose exit condition is a command - and stop when it stops making progress.
4
+ ---
5
+
6
+ # goblin-loop (the loop contract)
7
+
8
+ A **loop** is an unattended run whose exit condition is a **command**. One run, one record:
9
+ `.goblin/loop/`, committed, so a reader tomorrow can re-derive every decision. An unrecorded loop
10
+ is a story.
11
+
12
+ ## 1. Arm it: the predicate comes first
13
+
14
+ Write the exit condition as **one** shell command in `.goblin/loop/predicate`, and **run it once
15
+ before iteration 1**:
16
+
17
+ .goblin/loop/predicate # one command; exits 0 == the loop is finished
18
+ .goblin/loop/predicate.sha256 # its digest, recorded at loop start
19
+ .goblin/loop/first-run # exit=<n> ts=<ISO8601>, from that first run
20
+ .goblin/loop/budget # the turn budget this run declares
21
+
22
+ A **duration is not a finish condition.** Neither is a description of one. The run-once rule is
23
+ not ceremony: it is the only way to know the command is runnable *and currently red*, which is
24
+ what makes a later green mean anything.
25
+
26
+ `.goblin/loop/budget` is capped by `loop_max_turns_ceiling` in `.goblin/goblin.yaml` (default
27
+ **20**). The number is not invented: 20 is the engine's own default turn budget for a kanban goal
28
+ loop, and a ceiling that disagreed with the engine would be a number someone made up.
29
+
30
+ ## 2. Pin it: never relax
31
+
32
+ The predicate's digest is recorded at loop start (`LP-02`). Changing it mid-loop is **not an
33
+ edit** - it is closing this loop and opening another:
34
+
35
+ .goblin/loop/closed-<date>/ # the relaxed predicate AND its pin go here, committed
36
+ # then re-record predicate.sha256 with a `previous: <archived digest>` line beside the new one
37
+
38
+ The chain is what the row checks: a silent relaxation fails, a recorded re-scope costs one line.
39
+ Whether the new predicate is *weaker* is not decidable from a digest and is recorded in
40
+ `docs/LIMITS.md` #39 rather than claimed.
41
+
42
+ The pin never updates itself, for the same reason `practice_sha256:` never re-pins itself: a
43
+ self-updating pin is the silent edit it exists to catch. The remedy is printed when the row fails
44
+ and is never automatic.
45
+
46
+ ## 3. Log every iteration
47
+
48
+ One row per iteration, appended to `.goblin/loop/decisions.tsv` - the decision-log shape, unchanged
49
+ from the template:
50
+
51
+ ts · phase · decision · why · evidence · result
52
+ decision = "verdict:<done|blocked|continue|wait>"
53
+ evidence = a pointer that resolves: sha:<hex> | file:<path> | sha256:<hex>
54
+ result = "predicate:red" | "predicate:green" | "parked" | "stuck"
55
+
56
+ `evidence` is the thing a judge's verdict rests on. A `verdict:done` row that cites nothing
57
+ resolvable is refused by `JG-01`: the verdict may only rest on a handle the repo can re-derive.
58
+
59
+ ## 4. Guard rails - what stops a night being burned
60
+
61
+ - **No progress stops the loop.** Three consecutive verdict rows with an **unchanged evidence
62
+ pointer** and a result that is not `predicate:green` is a failure (`LP-04`). What it measures
63
+ is a *changed pointer*, a proxy for progress, not progress itself - a loop that edits a file
64
+ each turn to keep the pointer moving is not caught.
65
+ - **The check cannot be weakened.** The predicate is pinned by digest (`LP-02`); the budget is
66
+ declared and capped (`LP-03`); and a card's own predicate cannot be edited by its worker at all,
67
+ which is why the file predicate needs the pin and the card predicate does not.
68
+ - **A stuck loop writes up and stops.** A run whose last row is not `predicate:green` carries
69
+ `.goblin/loop/stuck.md`: at least three non-blank lines naming the predicate it failed
70
+ (`LP-05`). The exit a worker can actually reach is `kanban_block` - naming the predicate - plus
71
+ the write-up. Thrashing is not an option the record allows.
72
+
73
+ ## 5. The morning audit
74
+
75
+ Read the rows whose `result` is not `predicate:green`, **in order**, then the last five rows for
76
+ context. That is the same discipline as P10's "read the Attention section first": the interesting
77
+ rows are the ones that did not reach the goal, and a green summary line is not one of them.
78
+
79
+ ## What this cannot see
80
+
81
+ - **Whether the predicate was the right one.** Every `LP-` row checks the process; a loop can run
82
+ perfectly against a predicate that measures the wrong thing - and a predicate that was
83
+ **vacuously true from the start** (a `grep -c` against a renamed directory exits 1, passes
84
+ `LP-01` and `LP-02`, and ends green on nothing) is the failure mode the run-once rule exists to
85
+ make visible to a human, not one a row can catch.
86
+ - **Whether the recorded `exit=` was measured or typed**, and whether the loop really ran - the
87
+ record proves a date and a code exist, not that a command produced them.
88
+ - **Cost.** A turn budget bounds turns, not tokens, and the record prices nothing.