@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/CHANGELOG.md +351 -0
  2. package/LICENSE +21 -0
  3. package/README.md +163 -2
  4. package/VERSION +1 -0
  5. package/adapters/_template/adapter.tsv +16 -0
  6. package/adapters/_template/detect.sh +10 -0
  7. package/adapters/_template/emit.sh +5 -0
  8. package/adapters/_template/verify.sh +4 -0
  9. package/adapters/claude/adapter.tsv +8 -0
  10. package/adapters/claude/detect.sh +8 -0
  11. package/adapters/claude/verify.sh +47 -0
  12. package/adapters/codex/adapter.tsv +12 -0
  13. package/adapters/codex/detect.sh +9 -0
  14. package/adapters/codex/verify.sh +45 -0
  15. package/adapters/copilot/adapter.tsv +10 -0
  16. package/adapters/copilot/detect.sh +8 -0
  17. package/adapters/copilot/verify.sh +45 -0
  18. package/adapters/cursor/adapter.tsv +11 -0
  19. package/adapters/cursor/detect.sh +10 -0
  20. package/adapters/cursor/verify.sh +45 -0
  21. package/adapters/gemini/adapter.tsv +15 -0
  22. package/adapters/gemini/detect.sh +11 -0
  23. package/adapters/gemini/verify.sh +49 -0
  24. package/adapters/hermes/adapter.tsv +9 -0
  25. package/adapters/hermes/detect.sh +8 -0
  26. package/adapters/hermes/verify.sh +27 -0
  27. package/adapters/opencode/adapter.tsv +14 -0
  28. package/adapters/opencode/detect.sh +9 -0
  29. package/adapters/opencode/verify.sh +45 -0
  30. package/automations/README.md +53 -0
  31. package/automations/bugreporter-intake.sh +145 -0
  32. package/automations/drift-audit.sh +139 -0
  33. package/automations/report.schema.tsv +10 -0
  34. package/bans/README.md +82 -0
  35. package/bans/grep-ban.sh +84 -0
  36. package/bans/layer-check.sh +57 -0
  37. package/bin/goblin +119 -0
  38. package/bin/goblin-audit +145 -0
  39. package/bin/goblin-bans +178 -0
  40. package/bin/goblin-doctor +233 -0
  41. package/bin/goblin-emit +482 -0
  42. package/bin/goblin-init +519 -0
  43. package/bin/goblin-install +720 -0
  44. package/bin/goblin-lib.sh +289 -0
  45. package/bin/goblin-model +105 -0
  46. package/bin/goblin-upgrade +572 -0
  47. package/bin/goblin-verify +2798 -0
  48. package/bin/goblin.js +48 -0
  49. package/docs/ADOPTION.md +168 -0
  50. package/docs/CI.md +187 -0
  51. package/docs/CONTRACTS.md +197 -0
  52. package/docs/DESIGN.md +92 -0
  53. package/docs/ENFORCEMENT.md +225 -0
  54. package/docs/FLOWS.md +164 -0
  55. package/docs/GUARDRAILS.md +126 -0
  56. package/docs/GUIDE.md +610 -0
  57. package/docs/INTEGRATION.md +92 -0
  58. package/docs/LIMITS.md +591 -0
  59. package/docs/LOOP.md +165 -0
  60. package/docs/RE-PLAYBOOK.md +183 -0
  61. package/docs/RISKS.md +70 -0
  62. package/docs/ROLES.md +105 -0
  63. package/manifest/bans.tsv +9 -0
  64. package/manifest/classes.tsv +61 -0
  65. package/manifest/enforcement.tsv +88 -0
  66. package/manifest/glossary.tsv +25 -0
  67. package/manifest/playbooks.tsv +16 -0
  68. package/package.json +36 -4
  69. package/presets/A-shipped-software.yaml +48 -0
  70. package/presets/B-service-config.yaml +40 -0
  71. package/presets/C-game.yaml +38 -0
  72. package/presets/D-knowledge.yaml +41 -0
  73. package/presets/E-fleet-config.yaml +42 -0
  74. package/presets/F-electron.yaml +67 -0
  75. package/roles.yaml +54 -0
  76. package/skills/goblin-bootstrap/SKILL.md +51 -0
  77. package/skills/goblin-bugfix/SKILL.md +26 -0
  78. package/skills/goblin-bugreporter/SKILL.md +52 -0
  79. package/skills/goblin-drift-audit/SKILL.md +43 -0
  80. package/skills/goblin-eval/SKILL.md +68 -0
  81. package/skills/goblin-feature/SKILL.md +26 -0
  82. package/skills/goblin-feature-map/SKILL.md +140 -0
  83. package/skills/goblin-handoff/SKILL.md +28 -0
  84. package/skills/goblin-investigation/SKILL.md +26 -0
  85. package/skills/goblin-judge/SKILL.md +74 -0
  86. package/skills/goblin-loop/SKILL.md +88 -0
  87. package/skills/goblin-mode/SKILL.md +70 -0
  88. package/skills/goblin-overnight/SKILL.md +42 -0
  89. package/skills/goblin-pr-gate/SKILL.md +42 -0
  90. package/skills/goblin-re-mobile/SKILL.md +51 -0
  91. package/skills/goblin-refactor/SKILL.md +23 -0
  92. package/skills/goblin-sweep/SKILL.md +23 -0
  93. package/skills/goblin-tdd-repro/SKILL.md +27 -0
  94. package/skills/goblin-verify-author/SKILL.md +50 -0
  95. package/skills/practice/SKILL.md +37 -0
  96. package/templates/AGENTS.md.tmpl +23 -0
  97. package/templates/HANDOFF.md.tmpl +43 -0
  98. package/templates/SPEC.md.tmpl +34 -0
  99. package/templates/audit-waiver.tsv.tmpl +10 -0
  100. package/templates/boundary-waivers.tmpl +8 -0
  101. package/templates/checks/assert.mjs.tmpl +60 -0
  102. package/templates/checks/gate.sh.tmpl +29 -0
  103. package/templates/ci/goblin-gate.yml.tmpl +46 -0
  104. package/templates/goblin.yaml.tmpl +138 -0
  105. package/templates/install-hooks.allowlist.tmpl +9 -0
  106. package/templates/loop/decisions.tsv.tmpl +1 -0
  107. package/templates/loop/predicate.tmpl +16 -0
  108. package/templates/report.yaml.tmpl +16 -0
package/docs/LOOP.md ADDED
@@ -0,0 +1,165 @@
1
+ # The loop contract — the judge, and the loop that has one
2
+
3
+ `P10 goblin-overnight` and kanban `goal_mode` both describe an unattended run. One gap joins them:
4
+ **the loop has a judge, and the judge has no contract.** This file is the contract, and it is the
5
+ prose half of the eight rows `JG-01`..`JG-03` and `LP-01`..`LP-05` mechanise.
6
+
7
+ ## 1. A reviewer inspects an artifact; a judge decides whether a *process* met its own predicate
8
+
9
+ | | reviewer (`role-review-panel`) | judge (`role-judge`) |
10
+ |---|---|---|
11
+ | input | an artifact — a diff at a SHA, a file, a spec | a **predicate** + the evidence it produced |
12
+ | question | "is this artifact right?" | "did the process meet the condition it declared?" |
13
+ | output | findings, categorised | a verdict from a **closed enum**, plus the handle it rests on |
14
+ | may accept | the artifact itself, at a named revision | only what the verdict can be re-derived from |
15
+ | failure mode | misses a defect, or bikesheds | **says yes** — a judge that has never returned "not done" |
16
+ | retry | re-review, possibly by another lane | re-run the *same* predicate on the recorded evidence |
17
+ | relation to the author | may be a sibling lane | **not the author, and not the author's profile** |
18
+
19
+ A panel is **N opinions**; a judge is **one decision**. That is why the judge is a role
20
+ (`role-judge` in `roles.yaml`) and not a sentence inside P7 or P12.
21
+
22
+ ## 2. What Hermes actually does today — read, not assumed
23
+
24
+ Every claim below was read in this box's Hermes tree (`~/.hermes/hermes-agent/`) on
25
+ 2026-09-25. Anything not read there is marked `[inferred]`.
26
+
27
+ | mechanism | measured behaviour | where it was read |
28
+ |---|---|---|
29
+ | the judge call | an **auxiliary LLM call**, `temperature=0`, max_tokens 4096, timeout 30 s, routed by `auxiliary.goal_judge.*`. The function also accepts `contract=`, `subgoals=` and `background_processes=` | `hermes_cli/goals.py:31-38`, `:858-863`, `:870-935` |
30
+ | **what the judge sees** | the goal (truncated to 2000 chars) and **the agent's own most recent response** (truncated to 4000) | `hermes_cli/goals.py:39`, `:904-905` |
31
+ | the verdict enum | `done \| blocked \| continue \| wait`, plus `skipped` for an empty goal; the legacy `{"done": bool}` shape is still accepted | `hermes_cli/goals.py:773-800`, `:888` |
32
+ | quality gates | deterministic shell commands that must pass before the judge may say `done`; a failed gate **short-circuits judging** and its bounded output (3000 chars) becomes the continuation prompt. 300 s timeout, 3 retries | `hermes_cli/goals.py:50-54`, `:60`, `:1279-1324` |
33
+ | fail-open | an unparseable or unreachable judge → `continue`. The constants for an auto-pause exist (3 consecutive parse failures, 5 consecutive transport failures) but belong to the interactive goal loop: **the kanban loop binds `_parse_failed` and `_transport_failed` and discards both** | `hermes_cli/goals.py:45`, `:48` (constants), `:923-930` (fall-through), `:1662` (discarded) |
34
+ | budget | `DEFAULT_MAX_TURNS = 20`; the kanban loop **starts `turns_used` at 1** and checks the budget **before** spending another turn; exhaustion blocks with outcome `blocked_budget` | `hermes_cli/goals.py:31`, `:1632-1634` (the default and the clamp), `:1637` (`turns_used = 1`), `:1689-1696` (the check and the block) |
35
+ | terminators | the worker's own terminal calls stop the loop (`kanban_complete`, `kanban_block`, review hand-off) | `hermes_cli/goals.py:1586-1592` |
36
+ | `wait` in a kanban loop | **downgraded to `continue`** — a worker finishes with kanban tools, not by parking | `hermes_cli/goals.py:1666-1667` |
37
+ | **what the kanban loop passes the judge** | `judge_goal(goal_text, last_response)` — **two positional arguments only: no contract, no subgoals, no gates** (the function can take all three; this call site passes none) | `hermes_cli/goals.py:1662` |
38
+ | a judged-done worker that never finalises | one finalize nudge, then a **block**, never a completion | `hermes_cli/goals.py:1676-1685` |
39
+ | the terminal handoff gate | `kanban_complete` and `kanban_request_review` are judged on the **supplied summary text**; a broken judge **allows** the handoff; the gate is skipped entirely when no auxiliary client resolves | `tools/kanban_tools.py:414-424`, `:447-477`, `:654-680`, `:783-808` |
40
+ | **progress detection** | **none.** The loop's whole state is `last_response`, `turns_used`, `nudged_to_finalize` | `hermes_cli/goals.py:1636-1638` |
41
+ | predicate pinning | **none** — `goal_text` is read once from the card, and a worker cannot edit its own card: `grep -rn kanban_edit tools/` finds **one** hit, the gate's own message text, and no tool definition | `hermes_cli/goals.py:1595-1601`, `tools/kanban_tools.py:432` |
42
+
43
+ **The three sentences that matter.**
44
+
45
+ 1. **The stop condition is a self-report.** The judge is handed the agent's own prose and asked
46
+ whether it is enough. The prompt demands concrete evidence; the judge cannot go and get it. So
47
+ *"a judge given prose instead of evidence"* is the default on every turn, not an edge case.
48
+ 2. **The one mechanical path exists and is not connected.** Quality gates would make this a real
49
+ loop, and the kanban path passes neither gates nor contract (`hermes_cli/goals.py:1662`).
50
+ 3. **No progress detector, no predicate pin — the budget is the only backstop.**
51
+
52
+ `[inferred]` the goal-judge verdicts on this box are any *good*: this card ran no `goal_mode`
53
+ card, and nothing in the repo can observe which lane graded what.
54
+
55
+ ## 3. The record
56
+
57
+ One unattended run, one committed record. `templates/loop/` ships a copy-ready `predicate` and
58
+ `decisions.tsv` header; nothing installs them, because a fresh install must not be born with a
59
+ loop record — every `LP-` row skips until a loop has actually run here.
60
+
61
+ .goblin/loop/predicate one command; exits 0 == the loop is finished
62
+ .goblin/loop/predicate.sha256 its digest, recorded at loop start (the pin); after a close,
63
+ the chain follows on `previous: <digest>` lines
64
+ .goblin/loop/first-run exit=<n> ts=<ISO8601>, run BEFORE iteration 1
65
+ .goblin/loop/budget the turn budget this run declares
66
+ .goblin/loop/decisions.tsv ts · phase · decision · why · evidence · result
67
+ .goblin/loop/stuck.md the write-up when the run ended without predicate:green
68
+ .goblin/loop/closed-<date>/ a relaxed predicate AND the pin it was closed under, archived
69
+
70
+ An iteration row's `evidence` is a **pointer that resolves**: `sha:<hex>` (a commit in
71
+ `git rev-list --all`), `file:<path>` (a path under the root), `sha256:<hex>` (the digest of a file
72
+ under `.goblin/loop/`). A `cmd:` token resolves **nothing** on purpose: the command ran, its
73
+ output is not in the record, and a verdict resting on it is the self-report `JG-01` refuses.
74
+
75
+ ## 4. The five rules
76
+
77
+ 1. **The exit predicate is a command, written before iteration 1, and run at least once before
78
+ it.** Not a duration, not a description. The *run once before* half is not ceremony: it is the
79
+ only way to know the command is runnable **and currently red**, which is what makes a later
80
+ green mean anything (`LP-01`).
81
+ 2. **Never relax.** The predicate's digest is recorded at loop start (`LP-02`). Changing it
82
+ mid-loop is not an edit; it is closing this loop and opening another, with the old predicate
83
+ **and the pin it was closed under** archived under `.goblin/loop/closed-<date>/` and committed,
84
+ and the new `predicate.sha256` naming the archived digest on a `previous: <digest>` line. That
85
+ chain is what `LP-02` checks: not that the new bar is as strong — no digest can say that
86
+ (`docs/LIMITS.md` #39) — but that the supersession is **recorded**, so a silent relaxation is a
87
+ FAIL rather than an indistinguishable re-scope. The precedent is already in the repo:
88
+ `practice_sha256:` + `IN-02` and the deliberate `--re-pin`, which never re-pins automatically
89
+ "because a self-updating pin would be the silent edit it exists to catch"
90
+ (`docs/CONTRACTS.md`).
91
+ 3. **The escape hatch is a write-up, not a silence.** A run whose last row's `result` is not
92
+ `predicate:green` carries `.goblin/loop/stuck.md` — at least three non-blank lines naming the
93
+ predicate (`LP-05`). Wording matters: the exits a worker can actually reach are `kanban_block`
94
+ and `kanban_complete`; **there is no `kanban_edit` tool** in the worker toolset, so the remedy
95
+ the rejection message names is not one a worker holds. `kanban_block` names the predicate, and
96
+ the write-up makes the stop visible.
97
+ 4. **The budget.** `goal_max_turns` on the card, mirrored into the record, capped by
98
+ `loop_max_turns_ceiling` in `.goblin/goblin.yaml`, default **20** — the value measured at
99
+ `hermes_cli/goals.py:31`. A ceiling that does not match the engine's own default is a number
100
+ someone made up (`LP-03`).
101
+ 5. **The morning audit reads the rows that did not reach the goal, in order**, then the last five
102
+ for context — `P10` step 4's "read the `Attention` section first" applied to the record.
103
+
104
+ ## 5. The guard rails — what stops a night being burned
105
+
106
+ **What stops a loop making no progress?** *Nothing in Hermes*: `run_kanban_goal_loop` has no
107
+ notion of progress (`hermes_cli/goals.py:1636-1638`), so a loop returning `continue` with the same
108
+ reason nineteen times spends nineteen turns and then blocks. The mechanism here is `LP-04`: three
109
+ consecutive verdict rows with an **unchanged evidence pointer** and a result that is not
110
+ `predicate:green` is a FAIL naming the row numbers. **What it measures is a changed pointer, not
111
+ progress** — a loop that edits a file each turn to keep the pointer moving is not caught, which is
112
+ why `LP-05` and the budget sit behind it.
113
+
114
+ **What stops a loop "succeeding" by weakening its own check?** *Partly, and partly by accident.*
115
+ A **card** predicate is already protected: a worker cannot mutate its own card
116
+ (`tools/kanban_tools.py:211-229`) and no `kanban_edit` **tool** ships - the CLI verb
117
+ `hermes kanban edit` exists, and a headless worker cannot call it. A predicate in a **file** has no
118
+ such protection, so `LP-02`'s pin is the mechanism. The wider hazard is unaddressed and stated:
119
+ the judge is a language model grading prose (`hermes_cli/goals.py:904-905`), and *"the agent optimises
120
+ exactly the gate signal, including by faking it"* is measured rather than hypothetical. **The
121
+ counter-measure is not a better prompt.** It is that a `done` verdict may only cite a handle the
122
+ repo can resolve and re-check tomorrow (`JG-01`).
123
+
124
+ **What does a genuinely stuck loop do?** Today: it burns the budget and **blocks for review**,
125
+ carrying only the last judge reason, which came from a self-report (`hermes_cli/goals.py:1692-1695`).
126
+ Approved shape: write `.goblin/loop/stuck.md`, commit it, `kanban_block` naming the predicate.
127
+ `LP-05` makes the write-up mandatory and therefore visible; **nothing can make it true.**
128
+
129
+ ## 6. The judge's failure modes, and what is mechanical
130
+
131
+ | failure mode | counter-measure | mechanically enforceable? |
132
+ |---|---|---|
133
+ | **a judge that always says yes** | count the lane's verdicts; escalate a lane with no non-`done` verdict over N; **baseline it against one known-red control verdict per wave** | **No.** A repo sees only its own rows, and a lane that has judged twice cannot be called always-yes. → `JG-03`, `advisory`, counted |
134
+ | **a judge grading its own profile** | the judge's profile set must be **disjoint** from the author's (`JG-02`, FAIL) | **Partly.** Disjoint *profiles* is fully checkable; disjoint *families* needs the mapping file and is a report (`MD-02`), and *which lane ran* is unobservable from a repo (`MD-03`) |
135
+ | **a judge given prose instead of evidence** | `done` may only cite a **typed handle** the repo can resolve (`JG-01`) | **Yes**, on the record — it still cannot see whether the handle *supports* the verdict |
136
+ | **a loop making no progress** | three consecutive identical pointers with a non-terminal result → FAIL (`LP-04`) | **Yes** on the record; a changed pointer is a proxy, not progress |
137
+ | **a loop weakening its check** | the predicate is pinned by digest at loop start (`LP-02`); relaxing it is a committed archive-and-restart, and the new pin must name the archived digest on a `previous:` line, so a SILENT relaxation is a FAIL | **Partly**: the chain is checked; whether the new predicate is weaker is not (docs/LIMITS.md #39). A card predicate is protected by the board itself |
138
+ | **a loop that thrashes** | budget ceiling + `LP-04` + the mandatory write-up (`LP-03`, `LP-04`, `LP-05`) | **Yes** |
139
+
140
+ **The JG-03 policy, stated so it is not mistaken for a check.** `JG-03` is a **counted** row, never
141
+ a gate: once per wave, run the judge against a **known-red control** — a predicate the loop
142
+ deliberately cannot satisfy — and record the verdict in this file. A judge that returns `done` on a
143
+ known-red control is a judge that always says yes, and that fact is invisible to every row here.
144
+
145
+ ## 7. What this contract cannot see
146
+
147
+ 1. **Whether a resolvable handle supports the verdict it is attached to.** `JG-01` proves a commit
148
+ or a file exists; it cannot read it. A judge may cite a real commit that has nothing to do with
149
+ the claim.
150
+ 2. **Which lane actually produced a verdict.** No repo file observes which profile ran. `JG-02`
151
+ checks that the *declared* lanes are disjoint, not that a judge ran on one — the blindness
152
+ `MD-03` records.
153
+ 3. **Whether the predicate was the right one.** P10's own skill says it first: *"It measures only
154
+ what it was told to measure."* Every `LP-` row checks the process.
155
+ 4. **Whether `first-run`'s `exit=` was measured or typed** — `HP-03`'s defect, one artifact over:
156
+ the record proves a date and a code exist, not that a command produced them.
157
+ 5. **A predicate that was vacuously true from the start.** A `grep -c` against a renamed directory
158
+ exits 1, passes `LP-01` ("it ran") and `LP-02` (nothing edited), and the loop ends green on
159
+ nothing. The only defences are the run-once rule and a human reading the predicate.
160
+ 6. **Cost.** No row prices a loop. `DEFAULT_MAX_TURNS = 20` bounds turns, not tokens, and the
161
+ auxiliary judge call per turn is a cost the record does not contain.
162
+ 7. **A worker-created loop.** The board's worker toolset exposes `goal_mode` / `goal_max_turns` on
163
+ `kanban_create`, so loops can spawn loops; `loop_max_turns_ceiling` bounds one loop's turns and
164
+ nothing bounds their number. An aggregate board budget is fleet work, not repo work.
165
+ 8. **Whether a lane that never says no is a bad lane** — see the `JG-03` policy in §6.
@@ -0,0 +1,183 @@
1
+ # RE-PLAYBOOK - P15 `goblin-re-mobile`
2
+
3
+ The P15 procedure: one shipped Android build, understood as **facts for study**, with a
4
+ reproducible, hash-manifested corpus. This document is the procedure; `docs/FLOWS.md` carries
5
+ the summary, and the four `RC-` rows in `manifest/enforcement.tsv` are what make its
6
+ verification column measurable rather than aspirational. Two houses: the **lab repo** holds
7
+ scripts, notes and hash manifests only; the **quarantine** (inside the sandbox) holds every
8
+ extracted byte and never enters a repo.
9
+
10
+ ## The fences, and what enforces each
11
+
12
+ | fence | what bites |
13
+ |---|---|
14
+ | An owned or free build only - never a cracked or pirated APK | stated here; S1 is the hard fence, and `RC-04` records what was acquired and from where |
15
+ | One dedicated sandbox, never the host | S0 preflight; nothing in this row can see the host's filesystem, so the preflight is the fence |
16
+ | No extracted byte is ever tracked in a repo | `RC-03` - the lab repo's tracked tree, allowlist plus payload-hash clause |
17
+ | Nothing expression-shaped (art, text) leaves the quarantine | `RC-01` - the exact-hash gate over `security: build_output:` at release time |
18
+ | A manifest that cannot pass vacuously | `RC-02` - the `reference-manifest/1` schema |
19
+
20
+ ## The steps
21
+
22
+ ### S0 - preflight
23
+
24
+ The sandbox exists and is the one the fences describe: a disposable LXC or VM, never the
25
+ host. Confirm it before touching a binary. **Produces:** a live sandbox with a quarantine
26
+ root. The **lab repo itself lives outside the sandbox** - it is a host repo; what sits inside
27
+ the sandbox is the un-versioned working copy and the quarantine (no `.git` there). **Enforced
28
+ by:** S0 itself - no `RC-` row inspects the host.
29
+
30
+ ### S1 - acquire
31
+
32
+ An owned or free build only: a store listing with a **published hash** (md5 or sha256).
33
+ Cracked or pirated APKs are a hard fence - not a judgment call, a stop. **Produces:** the
34
+ APK file inside the sandbox. **Enforced by:** `RC-04`, the acquisition record.
35
+
36
+ ### S2 - provenance
37
+
38
+ Verify the download against the store's published md5/sha256 **before anything else touches
39
+ it**. **Produces:** the measured digest, written into the acquisition record with the
40
+ target slug + version and the source store. **Enforced by:** `RC-04` - a header naming
41
+ target slug + version + store + checksum, and a 64-hex token which, **if present**, must
42
+ equal the apk row's sha256 (a header carrying only the md5 passes - the comparison is
43
+ conditional in the engine).
44
+
45
+ ### S3 - triage
46
+
47
+ The engine verdict, cheapest test first. Marker rules over the archive listing: `.so`
48
+ libraries - `lib/*/libunity.so` = Unity IL2CPP (with `libil2cpp.so`), `libue4.so` = Unreal,
49
+ `libcocos2d*.so` = Cocos, `libgodot*` = Godot; `assets/bin/Data/Managed` + `libmono` =
50
+ Unity Mono; `global-metadata.dat` = Unity IL2CPP; no `.so` and packed assets in `assets/`
51
+ = custom Java/Dalvik engine. **Produces:** the verdict line in the triage note.
52
+ **Enforced by:** nothing mechanised - the verdict is a note, and `RC-03` sees only that
53
+ the note lives under `notes/`.
54
+
55
+ ### S4 - static decompile
56
+
57
+ apktool/jadx into the **quarantine**, never into any repo. **Produces:** the decompiled
58
+ tree under the sandbox quarantine root. **Enforced by:** `RC-03` - the lab allowlist has
59
+ no `build/`, `dist/`, or decompiled tree.
60
+
61
+ ### S5 - carve the containers
62
+
63
+ A carver written **from decompiled evidence, not guessed**: read the loader in the
64
+ decompiled tree, recover the transform (key, offset, container layout), reimplement, and
65
+ only then run it. The precedent: Kairosoft's whole-file-XOR `.dat` containers, whose
66
+ loader was found at `kairo/android/util/g.java` and whose 16-byte XOR key came from
67
+ `b.a.a()` in the decompiled tree - the carver was written from that evidence, and its
68
+ parse was cross-checked against size tables *inside* each container (16/16 consistent).
69
+ **Produces:** extracted payloads under the quarantine root, and the carver script in the
70
+ lab repo (`scripts/`). **Enforced by:** `RC-03` - the script is tracked, the payloads are
71
+ not.
72
+
73
+ ### S6 - manifest the corpus
74
+
75
+ sha256 of **every** extracted payload into `manifests/*.sha256`, with a header naming
76
+ target + version + source store + anchor, and the **apk row itself**. This `.sha256`
77
+ manifest is the sha256sum-format artifact: it is what `sha256sum -c` verifies and what
78
+ `RC-03`/`RC-04` read. It is **not** what `RC-01`/`RC-02` consume - those rows parse a
79
+ separate `reference-manifest/1` **JSON** (`generated_from`, `reference_app`, `entries[]`,
80
+ `entry_count`), declared via `reference_manifest:` in `.goblin/goblin.yaml`. That JSON is an
81
+ **optional** artifact, produced only if this lab ever ships a build output that could carry
82
+ corpus bytes; a study-only lab never ships, so none is produced here - a deliberate
83
+ omission, not a gap. Near-duplicate manifest rows are not linted - a known limitation.
84
+ **Produces:** the manifest. **Enforced by:** `RC-04` (the header + the apk row) and `RC-03`
85
+ (the payload-hash clause); `RC-02` exists for the *optional* JSON (the `reference-manifest/1`
86
+ schema so `RC-01` cannot pass vacuously - `entries` non-empty, `entry_count` honest, 64-hex
87
+ digests, byte counts).
88
+
89
+ ### S7 - dossier
90
+
91
+ Facts and numbers, never expression: measurements, counts, class/method citations
92
+ (`file:line`), and the digests already in the manifest. **No extracted art, no copied
93
+ text** - an asset's dimensions and its hash are facts; the asset itself is expression and
94
+ stays in the quarantine. **Produces:** `notes/<date>-<target>.md`. **Enforced by:**
95
+ `RC-03` (the note is tracked; nothing it quotes may be a tracked payload byte) and, at
96
+ release time, `RC-01` (any corpus hash that shows up in a shipping build fails it).
97
+
98
+ ### S8 - runtime analysis, deferred by design
99
+
100
+ Frida hooking needs a device or emulator and an owned build, and static facts come first:
101
+ a hash-manifested corpus and a written dossier are reproducible; a hook session is not.
102
+ S8 is **not** cut - it is deferred, and a future pass may pick it up with its own fence
103
+ (the device, the owned-build check, its own quarantine).
104
+
105
+ ### S9 - retention / teardown
106
+
107
+ The corpus lives only in the sandbox quarantine; the lab repo keeps scripts, notes and
108
+ hashes only. **Delete or keep is a recorded decision** - recorded in the **lab repo** (a
109
+ note under `notes/`, or the HANDOFF NEXT section), with the date. **Enforced by:** `RC-03`,
110
+ re-run after any teardown: the tracked tree must
111
+ still be scripts/notes/manifests/docs only, and no tracked file may hash to a manifest row.
112
+
113
+ ## Verifying the lab repo - the harness must be installed there first
114
+
115
+ `goblin-verify` requires an **installed harness** in the repo it points at: against a bare
116
+ lab repo it exits 2 with `not installed: <root>/.goblin/goblin.yaml is absent`. `--source`
117
+ relocates the enforcement manifest, not the target's config requirement, so the bridge is a
118
+ one-time install **into the lab repo itself**:
119
+
120
+ goblin-install --target <lab-repo> --class A
121
+
122
+ A class-A install is sufficient and violates nothing - the lab's own files are not touched;
123
+ it only adds the harness scaffolding: `.goblin/` (config + the verifier), `checks/`,
124
+ `reviews/`, `.github/workflows/`, and the `HANDOFF.md` scaffolding. Note: `HANDOFF.md`'s
125
+ placeholder commit gives `HP-05`'s one expected day-one red until a real commit is named.
126
+ After that one-time install, `goblin-verify --only RC-03` / `RC-04` run against the lab repo
127
+ (`RC-03`/`RC-04` need the installed harness; `RC-01`/`RC-02` additionally need a declared
128
+ `reference_manifest:`, which a study-only lab deliberately leaves empty).
129
+
130
+ ## What the `RC-` rows bite
131
+
132
+ | row | gate | bites at |
133
+ |---|---|---|
134
+ | `RC-01` | build output | S7 / release: no file in the shipping tree matches a corpus hash (four clauses: skip-when-empty, fail-when-absent, the hash clause, and the manifest itself must not ship) |
135
+ | `RC-02` | the manifest | S6: the manifest's shape - so `RC-01` cannot pass vacuously |
136
+ | `RC-03` | the lab tree | S4-S9: every tracked path is on the allowlist and no tracked byte equals a payload |
137
+ | `RC-04` | the manifest header | S1-S2: the acquisition record exists and names target, version, store, checksum |
138
+
139
+ ## First real run - 2026-09-27, Oh!Edo Towns Lite
140
+
141
+ - **Target:** `edotownsL-1.0.9.apk`, package `net.kairosoft.android.edotownsL`, version
142
+ 1.0.9 (versionCode 10), 6,119,826 bytes, md5 `23974f8582359aa32eac30ad42a7744e`,
143
+ acquired from Aptoide (`pool.apk.aptoide.com`) - a free listing with a published md5.
144
+ - **Sandbox:** LXC 217 `re-lab`; quarantine root `/opt/re-lab/refs/edotownsL-1.0.9/`.
145
+ - **Verdict (S3):** Java/Dalvik, **custom Kairosoft `.dat` containers** - no engine
146
+ `.so` matched; 16 `.dat` files in `assets/`, with `xls.dat` the design-data suspect
147
+ (it turned out to be 3 localisation text entries, not balance data - the dossier
148
+ records the measurement, not the hope).
149
+ - **S5 precedent:** the carver (`scripts/karve-dat.py`) written from the decompiled
150
+ loader, not guessed; 312 payloads extracted, 16/16 containers parse-consistent.
151
+ - **Lab repo:** `~/projects/re-lab` - `scripts/triage-apk.sh`,
152
+ `scripts/karve-dat.py`, `manifests/edotownsL-1.0.9.sha256`,
153
+ `notes/2026-09-27-edotowns-lite-triage.md`.
154
+
155
+ ## Verification
156
+
157
+ - The corpus manifest verifies `sha256sum -c` **where the corpus lives** (in the sandbox,
158
+ against the anchor the manifest header names).
159
+ - `goblin-verify` against the **lab repo needs the harness installed there first** (see
160
+ "Verifying the lab repo" above): one `goblin-install --target <lab-repo> --class A`, after
161
+ which `--only RC-03` / `RC-04` are runnable. `goblin-verify --only RC-01` / `RC-02` /
162
+ `RC-03` / `RC-04` return the exits their rows
163
+ define - an exact hash inside the build output, a weak manifest and a tracked payload
164
+ each fail the build.
165
+ - Every negative control NC-1..NC-6 was shown RED and then restored: NC-1/NC-2 under
166
+ `RC-01` (a payload, then the manifest itself, inside the build output), NC-3 under
167
+ `RC-02` (`entries` truncated to `[]`, with the pair-proving `RC-01`-stays-GREEN half),
168
+ NC-4 under `RC-01` (the declared manifest file deleted), NC-5/NC-6 under `RC-03`
169
+ (a payload git-added under an allowed path, then a tracked path outside the
170
+ allowlist). `tests/t-verify-red.sh` is the file that runs them.
171
+
172
+ ## What this cannot see
173
+
174
+ - **A re-encoded, resized or recoloured asset** passes `RC-01`'s exact-hash gate - the
175
+ gate is sha256 equality, and a byte that differs is a different byte.
176
+ - **Copied text inside a shipped string** is level 3 and invisible to the hash gate.
177
+ - **A weak manifest authored by hand** is `RC-02`'s problem only if it is *malformed* -
178
+ a well-formed manifest of the wrong bytes is exactly what the schema cannot judge
179
+ (`LIMITS.md` #28, the weakened-input defect; `installed.json` is unsigned, #18).
180
+ - **Nothing here proves a fact is CORRECT** - only that it is traceable to the corpus.
181
+ The dossier's measurements are as good as the commands that produced them, and a
182
+ misread loader produces a confidently wrong carver. `RC-04`, the weakest of the four,
183
+ proves the acquisition record *exists*, never that the number came from the store.
package/docs/RISKS.md ADDED
@@ -0,0 +1,70 @@
1
+ # Risk register and non-goals
2
+
3
+ ## Risks
4
+
5
+ | # | Risk | Counter-measure | Status |
6
+ |---|---|---|---|
7
+ | K1 | **Model promotion churn** breaks a role binding | Roles are capabilities, never slugs; one mapping file; `MD-01` lints every artifact for a hardcoded model name; `MD-02` reports family equality. A campaign that changes nine profiles changes nothing here. | Handled by design |
8
+ | K2 | **Profile drift across the fleet** | The attack surface is removed, not managed: project procedures live in the repo's `.hermes/skills/`, so there is no second copy to drift. `SK-02` hashes what is installed. The existing divergence outside goblin-stack's scope is a separate, escalated cleanup. | Handled in scope; fleet-side cleanup escalated |
9
+ | K3 | **Skills duplicating between profiles** | Same as K2. The precedence order is stated in `docs/INTEGRATION.md` so a future agent knows which copy wins instead of guessing. | Handled by design |
10
+ | K4 | **The vendored copy rots** (a repo sits at an old version) | `installed.json` records version and hashes; `goblin-verify` reports the installed version; `--upgrade` prints created/updated/unchanged. Accepted for repos that stop being worked on — the alternative (a network call at verify time) violates the offline dependency contract. | Accepted, with detection |
11
+ | K5 | **The harness decays into prose** | `IN-03` plus `SK-03`: a row without a command must say `advisory`, and the advisory count is capped (`advisory_ceiling`, default 10). The cap is the ratchet. | Handled by design |
12
+ | K6 | **A HANDOFF carries a stale number** | `HP-03` requires `measured <date>`; `HP-05` requires a commit that exists. | Partial: proves a date exists, not that the number is fresh. V1 anchored the row on the gate names the config declares (and skips the template's example sentence), so the real gate line can no longer lose its date in silence - but nothing re-measures the number, and a gate-bearing line that names no declared gate and carries no gate-shaped keyword is still unseen |
13
+ | K7 | **The gate is theatre** (self-skipping checks, admin bypass) | `PG-05` FAILs a conditional **job**, a conditional **step**, a job with no step and a workflow with no `jobs:` - re-declared at W4, because the old predicate counted "every step is guarded" and therefore passed the measured real shape (one unguarded step deciding whether the guarded gate step runs) while GitHub reported Success. `PG-06` then requires the gate CI runs to be the gate the project declares. `PG-04` documents the bypass. *"A required check that self-skips reports success… and a protected branch whose only admin is the person pushing protects nothing."* | Mechanised where a file can be read (`docs/CI.md` §1); the required-check list, the bypass switch and the push identity are forge state and stay documented, not solvable from the repo (`docs/LIMITS.md` #13, #34) |
14
+ | K8 | **Cost** | Roles plus effort tokens; the panel only at S3+; sweeps are cron-and-one-line; no playbook fans out without a named predicate. | Handled by policy |
15
+ | K9 | **The builder may only install into its own repo plus a scratch copy** | The install proof target is a throwaway copy under the scratch directory; the recipe is in `docs/ADOPTION.md`. | Constraint, satisfied |
16
+ | K10 | **A fleet-config repo is the least-governed artifact in an estate** | The E-class preset plus an artifact-scoped gate (a commit exists). The underlying staleness bug in a backup job is named and escalated — goblin-stack can detect staleness but cannot fix another repository. | Escalated |
17
+ | K11 | **The pre-change tree is unknown in a repo with no git history** | Repos with no `.git` are ordered *after* `git init`. `HS-02` is **skipped with a reason** rather than faked while no pinned commit exists. | Handled by ordering |
18
+ | K12 | **A check green on both trees** (the failure mode the REPLAY exists for) | `HS-02` runs the harness set against the pinned pre-change commit and requires **every** harness to be RED there. The shipped scaffold harness is deliberately such a check and is therefore reported as unproven until it is replaced. | Handled by design; see `docs/LIMITS.md` |
19
+ | K13 | **A nightly automation files the same defect twice, or files one that is not there** | A content-only dedup key (`--idempotency-key`, checked by `AU-02`) plus the board's own `recent_success` and `active_pr` guards; a report whose `revision` does not resolve is a refusal, not a card; `--max-runtime`, `--max-retries 1` and the failure limit auto-block a looping card. The producer's own ceiling bounds filings per day. | Handled by design; the key is a dedup, not a mutex (`docs/LIMITS.md` #20) |
20
+ | K14 | **A waiver becomes a permanent blind spot** - a dated exception nobody re-decides, or an audit record nobody re-takes, so `SC-07` passes against facts that are no longer true | The record's date is checked against `security.audit_max_age_days` (90) and every waiver's date against `security.waiver_max_age_days` (180); the waiver count is printed on the gate line so the debt is loud even when the row passes; `goblin-audit` writes a refusal (exit 5) instead of an empty record, so "no advisories" can never mean "the parse found nothing". | Handled by design; the policy is unmeasured against a real registry (`docs/LIMITS.md` #22) |
21
+ | K15 | **goblin-stack installs agent-authored skills, and the one controlled study of that practice puts it BELOW the no-skill baseline.** SkillsBench 1.1's self-generated condition (the agent authors its own Skills before solving) reports all three tested configurations below their no-Skills baseline (-8.1, -11.3, -11.5 points), while curated Skills rose +16.6 points across 18 configurations. A generated skill accepted after a skim is a different proposition from a written one. | `P6` hands every generated verification skill to `P12`, and `verified:` does not advance until an eval record exists; the record's shape, the eleven-token ban, the cheap-checks-first ladder and the pass condition (every seeded defect detected, the control at zero, every correction RED before GREEN) are specified in `skills/goblin-eval/SKILL.md`. The evidence is cited with its pin in `docs/LIMITS.md` #15. | **Stated requirement, not an enforced one**: the runner is not shipped and no row reads a lane (`docs/LIMITS.md` #31) |
22
+ | K16 | **The feature map rots, or claims coverage it does not have** - a route renamed under a recipe that still "works", a feature file nobody indexed, a `verified:` date nobody drove | `FM-01` (every feature file indexed, the four-H2 entry contract, the slug matches the filename), `FM-02` (every declared entry path still resolves under `source_root:`, no entry path changed after its `verified:` date), `VA-01` (the declared `verify_doctor:` exits 0); the upkeep pass and the rot table live in `skills/goblin-feature-map/SKILL.md`. | Handled by design for everything the map LISTED; completeness is not checkable (`docs/LIMITS.md` #30) |
23
+ | K17 | **An unattended loop grades itself** - the judge is a language model handed the worker's own prose, the lane that judges can be the lane that wrote, and Hermes has **no progress detector**: `run_kanban_goal_loop` carries no progress state, so a loop returning `continue` for the same reason nineteen times spends nineteen turns and then blocks | `JG-01` (a `done` verdict may only cite a handle the repo can resolve - a commit in `git rev-list --all`, a path under the root, a `sha256:` of a file under `.goblin/loop/`), `JG-02` (the declared judge lane must be **disjoint** from the author's - a FAIL, not a report), `JG-03` (counted: a lane with no non-`done` verdict is escalated), and `LP-01`..`LP-05` (one predicate command, run and recorded before iteration 1, pinned by digest, budgeted under a ceiling, no three identical pointers without a green, and a write-up when it ends red). The contract is `docs/LOOP.md`. | **Partial, and stated**: disjoint *profiles* is gated; disjoint *families* needs the mapping file (`MD-02`, ADV) and *which lane ran* is unobservable from a repo (`MD-03`); the judge lane resolves to no profile on this box, so `JG-02` reports ADV with its one-line remedy rather than failing a repo for the fleet's routing. The judge's own failure modes (prose-not-evidence, a lane that always says yes, no progress detector, cost) are `docs/LIMITS.md` #32 and #33 |
24
+
25
+ | K18 | **The shipped workflow is mistaken for a gate** — a file in `.github/workflows/` with no required-check entry, with the admin-bypass switch on, or pushed under the sole admin's own identity, is decoration that reads as enforcement | The template's header states the four settings that make a workflow a gate, `docs/CI.md` §1 argues them with the vendor's own words, `PG-05` refuses a conditional job or step and `PG-06` refuses a workflow that runs some other truth, `CL-01` makes `ci-gate` a class contract (absent where the class forbids it), and the verifier's "cannot see" footer names the lane on every run | The file half is mechanised; the forge half is not observable from a repo, and `PG-04` is the row that says so (`docs/LIMITS.md` #34) |
26
+
27
+ ## The advisory rows, named
28
+
29
+ Ten rows are labelled `advisory` (this sentence said nine until 2026-09-25: the tenth, `JG-03`,
30
+ landed with G2's judge lane), and the count is capped by `SK-03` (default ceiling 10 — the cap is
31
+ now **full**, `advisory 10 of ceiling 10`).
32
+ Nine carry **no executable check at all**:
33
+
34
+ - **HP-04** — a stale sentence is corrected in place with a dated parenthetical, never deleted.
35
+ - **HS-03** — source probes read text with comments blanked first.
36
+ - **CM-02** — the commit message was written to a file, not passed inline.
37
+ - **MD-03** — role-pinned fan-out goes through the kanban, not a model-less subagent spawn.
38
+ - **PG-04** — never bypass what the forge enforces.
39
+ - **DOC-01** — a significant change updates the docs that teach it.
40
+ - **DOC-02** — system-level changes are recorded wherever the project's standard says they live.
41
+ - **SC-09** — auth is applied consistently across sibling routes.
42
+ - **JG-03** — a judge lane that has never returned a non-`done` verdict is escalated. The history
43
+ that would show a bad lane lives across cards and repos, so the counter-measure is policy (one
44
+ known-red control verdict per wave, `docs/LOOP.md`), not a command.
45
+
46
+ One is advisory-labelled but still **reports its state** as `ADV`:
47
+
48
+ - **MD-02** — the review lane is a different model family from the code lane.
49
+
50
+ `HP-04`, `CM-02` and `MD-02` additionally print a heuristic when run with `--only`. A heuristic
51
+ is not a check: it never fails a run. A counted rule is still not an enforced one, and the cap
52
+ is a policy, not a proof.
53
+
54
+ ## Non-goals
55
+
56
+ Stated so no future reader infers them:
57
+
58
+ - not a plugin or a marketplace package; no slash commands;
59
+ - **no auto-merge** — no reviewed source ships it unconditionally, and merging is a different
60
+ decision from a green gate;
61
+ - not a fleet orchestrator — the board owns that;
62
+ - does not provision, migrate or verify models;
63
+ - does not write the vault;
64
+ - does not replace any project's existing gate (adopt, don't replace);
65
+ - is not a monorepo tool and not a CI tool. **Amended at W4:** it now places **at most one**
66
+ workflow, into its own target, for a class that requires or permits `ci-gate` — and it runs
67
+ nothing for that repo, holds no credentials, and cannot arm a check. See `docs/CI.md`;
68
+ - does not manage profile skill libraries;
69
+ - writes nothing outside the target repo;
70
+ - does not attempt to make the *prose* rules enforceable — it counts them instead.
package/docs/ROLES.md ADDED
@@ -0,0 +1,105 @@
1
+ # Roles and models
2
+
3
+ A **profile** is a worker identity. A **role** is what the work needs. Conflating them is how
4
+ five documents once gave five different model answers — a role name used as a model name.
5
+
6
+ goblin-stack keeps them separate and **never names a model**. `MD-01` lints every reusable rule
7
+ for a hardcoded model name; `MD-02` reports, but cannot change, whether the review lane and the
8
+ code lane resolve to the same family.
9
+
10
+ ## The role vocabulary
11
+
12
+ `roles.yaml` declares each role's capability and the default **profile** that carries it:
13
+
14
+ | role | what it needs | default profile | used by |
15
+ |---|---|---|---|
16
+ | `code` | fast, cheap, mechanical correctness | `coder` | P2, P4, P5, P10 |
17
+ | `judgment` | strongest available reasoning and prose | `architect` | P1, P3 (spec half), P6, P8, P9 |
18
+ | `review-panel` | **a list** — N independent verdict lanes, each its own lane | `[reviewer, architect]` | P7 at stakes S3+ |
19
+ | `synthesis` | merges many outputs into one artifact | `architect` | P11, P12 |
20
+ | `investigate` | read-only exploration returning a distilled summary | `researcher` | P1 (read-only half) |
21
+ | `judge` | decides whether a process met its own predicate, from a command's output and a pointer it can resolve, never from a report | `judge` (a **single** lane) | P10, P12, P7 at S3+, the terminal handoff gate |
22
+
23
+ The list-valued panel is the sharpest configuration idea retained here: **one lane runs per list
24
+ entry, so the list length sets the lane count.** Lane count is configuration, not code.
25
+
26
+ A **panel is N opinions; a judge is one decision** — which is why `role-judge` is a role and not a
27
+ sentence inside P7 or P12. The judge's contract (what it receives, what it must refuse, how its
28
+ verdict is recorded) is `docs/LOOP.md`; the pair of rules that make it real are that a `done`
29
+ verdict may only cite a handle the repo can resolve (`JG-01`) and that the judge's lane must be
30
+ disjoint from the author's (`JG-02`).
31
+
32
+ ## The model-mapping contract
33
+
34
+ The machine-specific mapping (`profile → provider/model/effort`) is read at run time from the
35
+ file declared as `models_file:` in `.goblin/goblin.yaml`. That is **the single documented
36
+ install-time machine input**, and goblin-stack **reads it and never writes it**.
37
+
38
+ Why not ship a copy of the mapping: the mapping file is a live routing source owned and
39
+ rewritten by another tool. Adding a foreign key to a file another tool owns is the precedent
40
+ for what happens when two tools share a file with unstated ownership. One place to change a
41
+ model, and the model layer stays where it belongs.
42
+
43
+ Resolution:
44
+
45
+ bin/goblin-model <role> # one line per lane: profile provider model effort
46
+ bin/goblin-model --list # the roles and their capabilities
47
+ bin/goblin-model review-panel # the panel: one line per lane
48
+
49
+ `bin/goblin-model` is **checkout-only**. `goblin-install` copies four scripts into a target's
50
+ `.goblin/bin/` — `goblin-verify`, `goblin-lib.sh`, `goblin-audit` and `goblin-bans` — so an
51
+ adopted repo has no `goblin-model` command (`ls .goblin/bin/` →
52
+ `goblin-audit goblin-bans goblin-lib.sh goblin-verify`). The installed path
53
+ for the same resolution is the `resolve_role_models` helper inside `.goblin/bin/goblin-verify`,
54
+ which is what `MD-02` calls; `bin/goblin-model` exists for a human at a checkout, is covered only
55
+ by `bash -n` in `tests/run-tests.sh`, and has no `enforcement.tsv` row because it enforces
56
+ nothing — it prints. Naming it here is the alternative R6 §2.2 allows to shipping a row for it.
57
+
58
+ Absent on this machine, every role resolves to `unknown` and the model-dependent checks report
59
+ advisory. Absent is not an error — goblin-stack is portable, and another machine has no such
60
+ file. That is the one documented exception to the portability rule, and `PT-01` enforces it:
61
+ the path is a config *value* in `.goblin/goblin.yaml`, never a literal inside a rule.
62
+
63
+ ## The fan-out rule: role-pinned work goes through the kanban
64
+
65
+ Measured, from the live tool schemas:
66
+
67
+ - a bare subagent spawn takes **no** model or provider parameter, so it cannot honour a role —
68
+ it would silently run every lane on one model, which is exactly the review failure the panel
69
+ exists to prevent;
70
+ - a board card **does** take a model and a provider, and a review request takes a reviewer
71
+ profile.
72
+
73
+ Therefore: a panel is N board cards parented to the change, each carrying the resolved
74
+ `(provider, model)`, the pinned SHA, the diff, and one focus. `goblin-mode` forbids inventing a
75
+ fan-out that a subagent spawn cannot honour, and `MD-03` records that no repo-local file can
76
+ observe which tool created a worker — so this rule is stated here and enforced at board level.
77
+
78
+ ## The measured caveat
79
+
80
+ Today's mapping puts every switchable profile on the **same** model, so `review-panel` resolves
81
+ to the same family as `code`, and **`judge` resolves to no lane at all**. `MD-02` therefore reports
82
+ the state and is labelled advisory: goblin-stack cannot choose the fleet's models, and a harness
83
+ must not fail a repo for a fleet-wide campaign. This is a reported number, not a silent assumption.
84
+
85
+ The judge lane, measured on this box on 2026-09-25:
86
+
87
+ bash bin/goblin-model judge -> judge unknown unknown unknown (rc 0)
88
+ bash bin/goblin-model code -> coder <provider> <model> <effort>
89
+ bash bin/goblin-model review-panel -> reviewer <...> · architect <...>
90
+
91
+ `~/projects/fleet-model.yaml` has no `judge:` entry, and `hermes profile list` reports no `judge`
92
+ profile. So **`MD-02`'s rule is not satisfiable on this box today through the mapping file**: the
93
+ judge and the author would run on one family. A different family exists only as a per-profile
94
+ alias held in another profile's own config (a comment in the mapping file names it), and **the
95
+ mapping file cannot express an alias** — it maps a profile to a provider/model pair. Making the
96
+ constraint real is a **fleet-side** change (add the `judge` profile, add the entry, re-apply),
97
+ never a goblin-stack one; `JG-02` therefore gates the *declared lanes* being disjoint and reports
98
+ an unresolved judge lane as `ADV` with the one-line remedy rather than failing a repo for the
99
+ fleet's routing.
100
+
101
+ ## Budget
102
+
103
+ `effort` is one word per role, read from the same mapping file. The panel is bounded — lanes
104
+ only at S3+ — because parallel lanes cost about N times the tokens, and token usage is the
105
+ dominant term in the variance of agent outcomes. No playbook fans out without a named predicate.
@@ -0,0 +1,9 @@
1
+ id ban globs detect replacement escape reviewer source
2
+ BN-01 No `any` in application TypeScript. src bash .goblin/bans/grep-ban.sh -e ':[[:space:]]*any\b' -e '<any>' -e '<any,' -e ',any>' -e 'any\[\]' src Use the real type, or narrow an `unknown` at the boundary `// BAN-OK(BN-01): <reason>` on the offending line (a non-empty reason is required), or list the path under bans_exempt: architecture R2 section A12 (Dune rule 2); R1 section 3
3
+ BN-02 No `@ts-ignore` / `@ts-expect-error` suppressions. src bash .goblin/bans/grep-ban.sh -e '@ts-(ignore|expect-error|nocheck)' src Fix the type, or narrow with `unknown` at the boundary `// BAN-OK(BN-02): <reason>` architecture R1 section 6 (suppressions are stink)
4
+ BN-03 No direct network call from a component. src/components app bash .goblin/bans/grep-ban.sh -e '\bfetch[[:space:]]*\(' src/components app Put the call in the data layer (src/data/*) `// BAN-OK(BN-03): <the endpoint, and when it migrates>` architecture R2 section A12; E1 (overreach)
5
+ BN-05 No import across a declared layer boundary. src bash .goblin/bans/layer-check.sh Move the code, or import through the declared accessor declare the pair under layers:, or move the file architecture R2 section A12 (Dune's dependency-graph check)
6
+ BN-06 No renderer with Node access (`nodeIntegration: true`). src app electron bash .goblin/bans/grep-ban.sh -e 'nodeIntegration[[:space:]]*:[[:space:]]*true' src app electron Pass data over IPC through a preload script (contextBridge); leave nodeIntegration false `// BAN-OK(BN-06): <why the renderer needs a Node primitive>` architecture G6 section B.2 item 1 (Electron security S13/S14)
7
+ BN-07 No renderer with context isolation or the process sandbox turned off. src app electron bash .goblin/bans/grep-ban.sh -e 'contextIsolation[[:space:]]*:[[:space:]]*false' -e 'sandbox[[:space:]]*:[[:space:]]*false' src app electron Keep `contextIsolation: true` and `sandbox: true`; expose a narrow API from the preload script `// BAN-OK(BN-07): <why isolation is off, and what replaces it>` architecture G6 section B.2 item 2 (Electron security S14)
8
+ BN-08 No dangerous webPreferences. src app electron bash .goblin/bans/grep-ban.sh -e 'webSecurity[[:space:]]*:[[:space:]]*false' -e 'allowRunningInsecureContent[[:space:]]*:[[:space:]]*true' -e 'enableBlinkFeatures' -e 'allowpopups' src app electron Fix the origin or the CSP; a disabled web security model is not a workaround `// BAN-OK(BN-08): <the check, and when it goes>` architecture G6 section B.2 item 9 (Electron security S14 checklist items 6, 8, 10, 11)
9
+ BN-09 No synchronous IPC and no `@electron/remote`. src app electron bash .goblin/bans/grep-ban.sh -e 'sendSync[[:space:]]*\(' -e '@electron/remote' src app electron Use `ipcRenderer.invoke` / `ipcMain.handle` (async), and a preload-exposed API instead of @electron/remote `// BAN-OK(BN-09): <why a blocking call is the only option here>` architecture G6 section B.2 item 11 (Electron performance S13)
@@ -0,0 +1,61 @@
1
+ class part need
2
+ A handoff R
3
+ B handoff R
4
+ C handoff R
5
+ D handoff R
6
+ E handoff R
7
+ F handoff R
8
+ A spec R
9
+ B spec R
10
+ C spec R
11
+ D spec -
12
+ E spec R
13
+ F spec R
14
+ A gate R
15
+ B gate R
16
+ C gate R
17
+ D gate O
18
+ E gate R
19
+ F gate R
20
+ A replay R
21
+ B replay -
22
+ C replay R
23
+ D replay -
24
+ E replay O
25
+ F replay R
26
+ A ratchet R
27
+ B ratchet O
28
+ C ratchet O
29
+ D ratchet -
30
+ E ratchet O
31
+ F ratchet R
32
+ A pr-gate O
33
+ B pr-gate -
34
+ C pr-gate O
35
+ D pr-gate -
36
+ E pr-gate O
37
+ F pr-gate O
38
+ A review-panel O
39
+ B review-panel -
40
+ C review-panel R
41
+ D review-panel -
42
+ E review-panel O
43
+ F review-panel O
44
+ A playbooks R
45
+ B playbooks R
46
+ C playbooks R
47
+ D playbooks R
48
+ E playbooks R
49
+ F playbooks R
50
+ A tokens O
51
+ B tokens -
52
+ C tokens -
53
+ D tokens -
54
+ E tokens -
55
+ F tokens O
56
+ A ci-gate R
57
+ B ci-gate -
58
+ C ci-gate O
59
+ D ci-gate -
60
+ E ci-gate O
61
+ F ci-gate R