@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +351 -0
- package/LICENSE +21 -0
- package/README.md +163 -2
- package/VERSION +1 -0
- package/adapters/_template/adapter.tsv +16 -0
- package/adapters/_template/detect.sh +10 -0
- package/adapters/_template/emit.sh +5 -0
- package/adapters/_template/verify.sh +4 -0
- package/adapters/claude/adapter.tsv +8 -0
- package/adapters/claude/detect.sh +8 -0
- package/adapters/claude/verify.sh +47 -0
- package/adapters/codex/adapter.tsv +12 -0
- package/adapters/codex/detect.sh +9 -0
- package/adapters/codex/verify.sh +45 -0
- package/adapters/copilot/adapter.tsv +10 -0
- package/adapters/copilot/detect.sh +8 -0
- package/adapters/copilot/verify.sh +45 -0
- package/adapters/cursor/adapter.tsv +11 -0
- package/adapters/cursor/detect.sh +10 -0
- package/adapters/cursor/verify.sh +45 -0
- package/adapters/gemini/adapter.tsv +15 -0
- package/adapters/gemini/detect.sh +11 -0
- package/adapters/gemini/verify.sh +49 -0
- package/adapters/hermes/adapter.tsv +9 -0
- package/adapters/hermes/detect.sh +8 -0
- package/adapters/hermes/verify.sh +27 -0
- package/adapters/opencode/adapter.tsv +14 -0
- package/adapters/opencode/detect.sh +9 -0
- package/adapters/opencode/verify.sh +45 -0
- package/automations/README.md +53 -0
- package/automations/bugreporter-intake.sh +145 -0
- package/automations/drift-audit.sh +139 -0
- package/automations/report.schema.tsv +10 -0
- package/bans/README.md +82 -0
- package/bans/grep-ban.sh +84 -0
- package/bans/layer-check.sh +57 -0
- package/bin/goblin +119 -0
- package/bin/goblin-audit +145 -0
- package/bin/goblin-bans +178 -0
- package/bin/goblin-doctor +233 -0
- package/bin/goblin-emit +482 -0
- package/bin/goblin-init +519 -0
- package/bin/goblin-install +720 -0
- package/bin/goblin-lib.sh +289 -0
- package/bin/goblin-model +105 -0
- package/bin/goblin-upgrade +572 -0
- package/bin/goblin-verify +2798 -0
- package/bin/goblin.js +48 -0
- package/docs/ADOPTION.md +168 -0
- package/docs/CI.md +187 -0
- package/docs/CONTRACTS.md +197 -0
- package/docs/DESIGN.md +92 -0
- package/docs/ENFORCEMENT.md +225 -0
- package/docs/FLOWS.md +164 -0
- package/docs/GUARDRAILS.md +126 -0
- package/docs/GUIDE.md +610 -0
- package/docs/INTEGRATION.md +92 -0
- package/docs/LIMITS.md +591 -0
- package/docs/LOOP.md +165 -0
- package/docs/RE-PLAYBOOK.md +183 -0
- package/docs/RISKS.md +70 -0
- package/docs/ROLES.md +105 -0
- package/manifest/bans.tsv +9 -0
- package/manifest/classes.tsv +61 -0
- package/manifest/enforcement.tsv +88 -0
- package/manifest/glossary.tsv +25 -0
- package/manifest/playbooks.tsv +16 -0
- package/package.json +36 -4
- package/presets/A-shipped-software.yaml +48 -0
- package/presets/B-service-config.yaml +40 -0
- package/presets/C-game.yaml +38 -0
- package/presets/D-knowledge.yaml +41 -0
- package/presets/E-fleet-config.yaml +42 -0
- package/presets/F-electron.yaml +67 -0
- package/roles.yaml +54 -0
- package/skills/goblin-bootstrap/SKILL.md +51 -0
- package/skills/goblin-bugfix/SKILL.md +26 -0
- package/skills/goblin-bugreporter/SKILL.md +52 -0
- package/skills/goblin-drift-audit/SKILL.md +43 -0
- package/skills/goblin-eval/SKILL.md +68 -0
- package/skills/goblin-feature/SKILL.md +26 -0
- package/skills/goblin-feature-map/SKILL.md +140 -0
- package/skills/goblin-handoff/SKILL.md +28 -0
- package/skills/goblin-investigation/SKILL.md +26 -0
- package/skills/goblin-judge/SKILL.md +74 -0
- package/skills/goblin-loop/SKILL.md +88 -0
- package/skills/goblin-mode/SKILL.md +70 -0
- package/skills/goblin-overnight/SKILL.md +42 -0
- package/skills/goblin-pr-gate/SKILL.md +42 -0
- package/skills/goblin-re-mobile/SKILL.md +51 -0
- package/skills/goblin-refactor/SKILL.md +23 -0
- package/skills/goblin-sweep/SKILL.md +23 -0
- package/skills/goblin-tdd-repro/SKILL.md +27 -0
- package/skills/goblin-verify-author/SKILL.md +50 -0
- package/skills/practice/SKILL.md +37 -0
- package/templates/AGENTS.md.tmpl +23 -0
- package/templates/HANDOFF.md.tmpl +43 -0
- package/templates/SPEC.md.tmpl +34 -0
- package/templates/audit-waiver.tsv.tmpl +10 -0
- package/templates/boundary-waivers.tmpl +8 -0
- package/templates/checks/assert.mjs.tmpl +60 -0
- package/templates/checks/gate.sh.tmpl +29 -0
- package/templates/ci/goblin-gate.yml.tmpl +46 -0
- package/templates/goblin.yaml.tmpl +138 -0
- package/templates/install-hooks.allowlist.tmpl +9 -0
- package/templates/loop/decisions.tsv.tmpl +1 -0
- package/templates/loop/predicate.tmpl +16 -0
- package/templates/report.yaml.tmpl +16 -0
package/docs/LOOP.md
ADDED
|
@@ -0,0 +1,165 @@
|
|
|
1
|
+
# The loop contract — the judge, and the loop that has one
|
|
2
|
+
|
|
3
|
+
`P10 goblin-overnight` and kanban `goal_mode` both describe an unattended run. One gap joins them:
|
|
4
|
+
**the loop has a judge, and the judge has no contract.** This file is the contract, and it is the
|
|
5
|
+
prose half of the eight rows `JG-01`..`JG-03` and `LP-01`..`LP-05` mechanise.
|
|
6
|
+
|
|
7
|
+
## 1. A reviewer inspects an artifact; a judge decides whether a *process* met its own predicate
|
|
8
|
+
|
|
9
|
+
| | reviewer (`role-review-panel`) | judge (`role-judge`) |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
| input | an artifact — a diff at a SHA, a file, a spec | a **predicate** + the evidence it produced |
|
|
12
|
+
| question | "is this artifact right?" | "did the process meet the condition it declared?" |
|
|
13
|
+
| output | findings, categorised | a verdict from a **closed enum**, plus the handle it rests on |
|
|
14
|
+
| may accept | the artifact itself, at a named revision | only what the verdict can be re-derived from |
|
|
15
|
+
| failure mode | misses a defect, or bikesheds | **says yes** — a judge that has never returned "not done" |
|
|
16
|
+
| retry | re-review, possibly by another lane | re-run the *same* predicate on the recorded evidence |
|
|
17
|
+
| relation to the author | may be a sibling lane | **not the author, and not the author's profile** |
|
|
18
|
+
|
|
19
|
+
A panel is **N opinions**; a judge is **one decision**. That is why the judge is a role
|
|
20
|
+
(`role-judge` in `roles.yaml`) and not a sentence inside P7 or P12.
|
|
21
|
+
|
|
22
|
+
## 2. What Hermes actually does today — read, not assumed
|
|
23
|
+
|
|
24
|
+
Every claim below was read in this box's Hermes tree (`~/.hermes/hermes-agent/`) on
|
|
25
|
+
2026-09-25. Anything not read there is marked `[inferred]`.
|
|
26
|
+
|
|
27
|
+
| mechanism | measured behaviour | where it was read |
|
|
28
|
+
|---|---|---|
|
|
29
|
+
| the judge call | an **auxiliary LLM call**, `temperature=0`, max_tokens 4096, timeout 30 s, routed by `auxiliary.goal_judge.*`. The function also accepts `contract=`, `subgoals=` and `background_processes=` | `hermes_cli/goals.py:31-38`, `:858-863`, `:870-935` |
|
|
30
|
+
| **what the judge sees** | the goal (truncated to 2000 chars) and **the agent's own most recent response** (truncated to 4000) | `hermes_cli/goals.py:39`, `:904-905` |
|
|
31
|
+
| the verdict enum | `done \| blocked \| continue \| wait`, plus `skipped` for an empty goal; the legacy `{"done": bool}` shape is still accepted | `hermes_cli/goals.py:773-800`, `:888` |
|
|
32
|
+
| quality gates | deterministic shell commands that must pass before the judge may say `done`; a failed gate **short-circuits judging** and its bounded output (3000 chars) becomes the continuation prompt. 300 s timeout, 3 retries | `hermes_cli/goals.py:50-54`, `:60`, `:1279-1324` |
|
|
33
|
+
| fail-open | an unparseable or unreachable judge → `continue`. The constants for an auto-pause exist (3 consecutive parse failures, 5 consecutive transport failures) but belong to the interactive goal loop: **the kanban loop binds `_parse_failed` and `_transport_failed` and discards both** | `hermes_cli/goals.py:45`, `:48` (constants), `:923-930` (fall-through), `:1662` (discarded) |
|
|
34
|
+
| budget | `DEFAULT_MAX_TURNS = 20`; the kanban loop **starts `turns_used` at 1** and checks the budget **before** spending another turn; exhaustion blocks with outcome `blocked_budget` | `hermes_cli/goals.py:31`, `:1632-1634` (the default and the clamp), `:1637` (`turns_used = 1`), `:1689-1696` (the check and the block) |
|
|
35
|
+
| terminators | the worker's own terminal calls stop the loop (`kanban_complete`, `kanban_block`, review hand-off) | `hermes_cli/goals.py:1586-1592` |
|
|
36
|
+
| `wait` in a kanban loop | **downgraded to `continue`** — a worker finishes with kanban tools, not by parking | `hermes_cli/goals.py:1666-1667` |
|
|
37
|
+
| **what the kanban loop passes the judge** | `judge_goal(goal_text, last_response)` — **two positional arguments only: no contract, no subgoals, no gates** (the function can take all three; this call site passes none) | `hermes_cli/goals.py:1662` |
|
|
38
|
+
| a judged-done worker that never finalises | one finalize nudge, then a **block**, never a completion | `hermes_cli/goals.py:1676-1685` |
|
|
39
|
+
| the terminal handoff gate | `kanban_complete` and `kanban_request_review` are judged on the **supplied summary text**; a broken judge **allows** the handoff; the gate is skipped entirely when no auxiliary client resolves | `tools/kanban_tools.py:414-424`, `:447-477`, `:654-680`, `:783-808` |
|
|
40
|
+
| **progress detection** | **none.** The loop's whole state is `last_response`, `turns_used`, `nudged_to_finalize` | `hermes_cli/goals.py:1636-1638` |
|
|
41
|
+
| predicate pinning | **none** — `goal_text` is read once from the card, and a worker cannot edit its own card: `grep -rn kanban_edit tools/` finds **one** hit, the gate's own message text, and no tool definition | `hermes_cli/goals.py:1595-1601`, `tools/kanban_tools.py:432` |
|
|
42
|
+
|
|
43
|
+
**The three sentences that matter.**
|
|
44
|
+
|
|
45
|
+
1. **The stop condition is a self-report.** The judge is handed the agent's own prose and asked
|
|
46
|
+
whether it is enough. The prompt demands concrete evidence; the judge cannot go and get it. So
|
|
47
|
+
*"a judge given prose instead of evidence"* is the default on every turn, not an edge case.
|
|
48
|
+
2. **The one mechanical path exists and is not connected.** Quality gates would make this a real
|
|
49
|
+
loop, and the kanban path passes neither gates nor contract (`hermes_cli/goals.py:1662`).
|
|
50
|
+
3. **No progress detector, no predicate pin — the budget is the only backstop.**
|
|
51
|
+
|
|
52
|
+
`[inferred]` the goal-judge verdicts on this box are any *good*: this card ran no `goal_mode`
|
|
53
|
+
card, and nothing in the repo can observe which lane graded what.
|
|
54
|
+
|
|
55
|
+
## 3. The record
|
|
56
|
+
|
|
57
|
+
One unattended run, one committed record. `templates/loop/` ships a copy-ready `predicate` and
|
|
58
|
+
`decisions.tsv` header; nothing installs them, because a fresh install must not be born with a
|
|
59
|
+
loop record — every `LP-` row skips until a loop has actually run here.
|
|
60
|
+
|
|
61
|
+
.goblin/loop/predicate one command; exits 0 == the loop is finished
|
|
62
|
+
.goblin/loop/predicate.sha256 its digest, recorded at loop start (the pin); after a close,
|
|
63
|
+
the chain follows on `previous: <digest>` lines
|
|
64
|
+
.goblin/loop/first-run exit=<n> ts=<ISO8601>, run BEFORE iteration 1
|
|
65
|
+
.goblin/loop/budget the turn budget this run declares
|
|
66
|
+
.goblin/loop/decisions.tsv ts · phase · decision · why · evidence · result
|
|
67
|
+
.goblin/loop/stuck.md the write-up when the run ended without predicate:green
|
|
68
|
+
.goblin/loop/closed-<date>/ a relaxed predicate AND the pin it was closed under, archived
|
|
69
|
+
|
|
70
|
+
An iteration row's `evidence` is a **pointer that resolves**: `sha:<hex>` (a commit in
|
|
71
|
+
`git rev-list --all`), `file:<path>` (a path under the root), `sha256:<hex>` (the digest of a file
|
|
72
|
+
under `.goblin/loop/`). A `cmd:` token resolves **nothing** on purpose: the command ran, its
|
|
73
|
+
output is not in the record, and a verdict resting on it is the self-report `JG-01` refuses.
|
|
74
|
+
|
|
75
|
+
## 4. The five rules
|
|
76
|
+
|
|
77
|
+
1. **The exit predicate is a command, written before iteration 1, and run at least once before
|
|
78
|
+
it.** Not a duration, not a description. The *run once before* half is not ceremony: it is the
|
|
79
|
+
only way to know the command is runnable **and currently red**, which is what makes a later
|
|
80
|
+
green mean anything (`LP-01`).
|
|
81
|
+
2. **Never relax.** The predicate's digest is recorded at loop start (`LP-02`). Changing it
|
|
82
|
+
mid-loop is not an edit; it is closing this loop and opening another, with the old predicate
|
|
83
|
+
**and the pin it was closed under** archived under `.goblin/loop/closed-<date>/` and committed,
|
|
84
|
+
and the new `predicate.sha256` naming the archived digest on a `previous: <digest>` line. That
|
|
85
|
+
chain is what `LP-02` checks: not that the new bar is as strong — no digest can say that
|
|
86
|
+
(`docs/LIMITS.md` #39) — but that the supersession is **recorded**, so a silent relaxation is a
|
|
87
|
+
FAIL rather than an indistinguishable re-scope. The precedent is already in the repo:
|
|
88
|
+
`practice_sha256:` + `IN-02` and the deliberate `--re-pin`, which never re-pins automatically
|
|
89
|
+
"because a self-updating pin would be the silent edit it exists to catch"
|
|
90
|
+
(`docs/CONTRACTS.md`).
|
|
91
|
+
3. **The escape hatch is a write-up, not a silence.** A run whose last row's `result` is not
|
|
92
|
+
`predicate:green` carries `.goblin/loop/stuck.md` — at least three non-blank lines naming the
|
|
93
|
+
predicate (`LP-05`). Wording matters: the exits a worker can actually reach are `kanban_block`
|
|
94
|
+
and `kanban_complete`; **there is no `kanban_edit` tool** in the worker toolset, so the remedy
|
|
95
|
+
the rejection message names is not one a worker holds. `kanban_block` names the predicate, and
|
|
96
|
+
the write-up makes the stop visible.
|
|
97
|
+
4. **The budget.** `goal_max_turns` on the card, mirrored into the record, capped by
|
|
98
|
+
`loop_max_turns_ceiling` in `.goblin/goblin.yaml`, default **20** — the value measured at
|
|
99
|
+
`hermes_cli/goals.py:31`. A ceiling that does not match the engine's own default is a number
|
|
100
|
+
someone made up (`LP-03`).
|
|
101
|
+
5. **The morning audit reads the rows that did not reach the goal, in order**, then the last five
|
|
102
|
+
for context — `P10` step 4's "read the `Attention` section first" applied to the record.
|
|
103
|
+
|
|
104
|
+
## 5. The guard rails — what stops a night being burned
|
|
105
|
+
|
|
106
|
+
**What stops a loop making no progress?** *Nothing in Hermes*: `run_kanban_goal_loop` has no
|
|
107
|
+
notion of progress (`hermes_cli/goals.py:1636-1638`), so a loop returning `continue` with the same
|
|
108
|
+
reason nineteen times spends nineteen turns and then blocks. The mechanism here is `LP-04`: three
|
|
109
|
+
consecutive verdict rows with an **unchanged evidence pointer** and a result that is not
|
|
110
|
+
`predicate:green` is a FAIL naming the row numbers. **What it measures is a changed pointer, not
|
|
111
|
+
progress** — a loop that edits a file each turn to keep the pointer moving is not caught, which is
|
|
112
|
+
why `LP-05` and the budget sit behind it.
|
|
113
|
+
|
|
114
|
+
**What stops a loop "succeeding" by weakening its own check?** *Partly, and partly by accident.*
|
|
115
|
+
A **card** predicate is already protected: a worker cannot mutate its own card
|
|
116
|
+
(`tools/kanban_tools.py:211-229`) and no `kanban_edit` **tool** ships - the CLI verb
|
|
117
|
+
`hermes kanban edit` exists, and a headless worker cannot call it. A predicate in a **file** has no
|
|
118
|
+
such protection, so `LP-02`'s pin is the mechanism. The wider hazard is unaddressed and stated:
|
|
119
|
+
the judge is a language model grading prose (`hermes_cli/goals.py:904-905`), and *"the agent optimises
|
|
120
|
+
exactly the gate signal, including by faking it"* is measured rather than hypothetical. **The
|
|
121
|
+
counter-measure is not a better prompt.** It is that a `done` verdict may only cite a handle the
|
|
122
|
+
repo can resolve and re-check tomorrow (`JG-01`).
|
|
123
|
+
|
|
124
|
+
**What does a genuinely stuck loop do?** Today: it burns the budget and **blocks for review**,
|
|
125
|
+
carrying only the last judge reason, which came from a self-report (`hermes_cli/goals.py:1692-1695`).
|
|
126
|
+
Approved shape: write `.goblin/loop/stuck.md`, commit it, `kanban_block` naming the predicate.
|
|
127
|
+
`LP-05` makes the write-up mandatory and therefore visible; **nothing can make it true.**
|
|
128
|
+
|
|
129
|
+
## 6. The judge's failure modes, and what is mechanical
|
|
130
|
+
|
|
131
|
+
| failure mode | counter-measure | mechanically enforceable? |
|
|
132
|
+
|---|---|---|
|
|
133
|
+
| **a judge that always says yes** | count the lane's verdicts; escalate a lane with no non-`done` verdict over N; **baseline it against one known-red control verdict per wave** | **No.** A repo sees only its own rows, and a lane that has judged twice cannot be called always-yes. → `JG-03`, `advisory`, counted |
|
|
134
|
+
| **a judge grading its own profile** | the judge's profile set must be **disjoint** from the author's (`JG-02`, FAIL) | **Partly.** Disjoint *profiles* is fully checkable; disjoint *families* needs the mapping file and is a report (`MD-02`), and *which lane ran* is unobservable from a repo (`MD-03`) |
|
|
135
|
+
| **a judge given prose instead of evidence** | `done` may only cite a **typed handle** the repo can resolve (`JG-01`) | **Yes**, on the record — it still cannot see whether the handle *supports* the verdict |
|
|
136
|
+
| **a loop making no progress** | three consecutive identical pointers with a non-terminal result → FAIL (`LP-04`) | **Yes** on the record; a changed pointer is a proxy, not progress |
|
|
137
|
+
| **a loop weakening its check** | the predicate is pinned by digest at loop start (`LP-02`); relaxing it is a committed archive-and-restart, and the new pin must name the archived digest on a `previous:` line, so a SILENT relaxation is a FAIL | **Partly**: the chain is checked; whether the new predicate is weaker is not (docs/LIMITS.md #39). A card predicate is protected by the board itself |
|
|
138
|
+
| **a loop that thrashes** | budget ceiling + `LP-04` + the mandatory write-up (`LP-03`, `LP-04`, `LP-05`) | **Yes** |
|
|
139
|
+
|
|
140
|
+
**The JG-03 policy, stated so it is not mistaken for a check.** `JG-03` is a **counted** row, never
|
|
141
|
+
a gate: once per wave, run the judge against a **known-red control** — a predicate the loop
|
|
142
|
+
deliberately cannot satisfy — and record the verdict in this file. A judge that returns `done` on a
|
|
143
|
+
known-red control is a judge that always says yes, and that fact is invisible to every row here.
|
|
144
|
+
|
|
145
|
+
## 7. What this contract cannot see
|
|
146
|
+
|
|
147
|
+
1. **Whether a resolvable handle supports the verdict it is attached to.** `JG-01` proves a commit
|
|
148
|
+
or a file exists; it cannot read it. A judge may cite a real commit that has nothing to do with
|
|
149
|
+
the claim.
|
|
150
|
+
2. **Which lane actually produced a verdict.** No repo file observes which profile ran. `JG-02`
|
|
151
|
+
checks that the *declared* lanes are disjoint, not that a judge ran on one — the blindness
|
|
152
|
+
`MD-03` records.
|
|
153
|
+
3. **Whether the predicate was the right one.** P10's own skill says it first: *"It measures only
|
|
154
|
+
what it was told to measure."* Every `LP-` row checks the process.
|
|
155
|
+
4. **Whether `first-run`'s `exit=` was measured or typed** — `HP-03`'s defect, one artifact over:
|
|
156
|
+
the record proves a date and a code exist, not that a command produced them.
|
|
157
|
+
5. **A predicate that was vacuously true from the start.** A `grep -c` against a renamed directory
|
|
158
|
+
exits 1, passes `LP-01` ("it ran") and `LP-02` (nothing edited), and the loop ends green on
|
|
159
|
+
nothing. The only defences are the run-once rule and a human reading the predicate.
|
|
160
|
+
6. **Cost.** No row prices a loop. `DEFAULT_MAX_TURNS = 20` bounds turns, not tokens, and the
|
|
161
|
+
auxiliary judge call per turn is a cost the record does not contain.
|
|
162
|
+
7. **A worker-created loop.** The board's worker toolset exposes `goal_mode` / `goal_max_turns` on
|
|
163
|
+
`kanban_create`, so loops can spawn loops; `loop_max_turns_ceiling` bounds one loop's turns and
|
|
164
|
+
nothing bounds their number. An aggregate board budget is fleet work, not repo work.
|
|
165
|
+
8. **Whether a lane that never says no is a bad lane** — see the `JG-03` policy in §6.
|
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
# RE-PLAYBOOK - P15 `goblin-re-mobile`
|
|
2
|
+
|
|
3
|
+
The P15 procedure: one shipped Android build, understood as **facts for study**, with a
|
|
4
|
+
reproducible, hash-manifested corpus. This document is the procedure; `docs/FLOWS.md` carries
|
|
5
|
+
the summary, and the four `RC-` rows in `manifest/enforcement.tsv` are what make its
|
|
6
|
+
verification column measurable rather than aspirational. Two houses: the **lab repo** holds
|
|
7
|
+
scripts, notes and hash manifests only; the **quarantine** (inside the sandbox) holds every
|
|
8
|
+
extracted byte and never enters a repo.
|
|
9
|
+
|
|
10
|
+
## The fences, and what enforces each
|
|
11
|
+
|
|
12
|
+
| fence | what bites |
|
|
13
|
+
|---|---|
|
|
14
|
+
| An owned or free build only - never a cracked or pirated APK | stated here; S1 is the hard fence, and `RC-04` records what was acquired and from where |
|
|
15
|
+
| One dedicated sandbox, never the host | S0 preflight; nothing in this row can see the host's filesystem, so the preflight is the fence |
|
|
16
|
+
| No extracted byte is ever tracked in a repo | `RC-03` - the lab repo's tracked tree, allowlist plus payload-hash clause |
|
|
17
|
+
| Nothing expression-shaped (art, text) leaves the quarantine | `RC-01` - the exact-hash gate over `security: build_output:` at release time |
|
|
18
|
+
| A manifest that cannot pass vacuously | `RC-02` - the `reference-manifest/1` schema |
|
|
19
|
+
|
|
20
|
+
## The steps
|
|
21
|
+
|
|
22
|
+
### S0 - preflight
|
|
23
|
+
|
|
24
|
+
The sandbox exists and is the one the fences describe: a disposable LXC or VM, never the
|
|
25
|
+
host. Confirm it before touching a binary. **Produces:** a live sandbox with a quarantine
|
|
26
|
+
root. The **lab repo itself lives outside the sandbox** - it is a host repo; what sits inside
|
|
27
|
+
the sandbox is the un-versioned working copy and the quarantine (no `.git` there). **Enforced
|
|
28
|
+
by:** S0 itself - no `RC-` row inspects the host.
|
|
29
|
+
|
|
30
|
+
### S1 - acquire
|
|
31
|
+
|
|
32
|
+
An owned or free build only: a store listing with a **published hash** (md5 or sha256).
|
|
33
|
+
Cracked or pirated APKs are a hard fence - not a judgment call, a stop. **Produces:** the
|
|
34
|
+
APK file inside the sandbox. **Enforced by:** `RC-04`, the acquisition record.
|
|
35
|
+
|
|
36
|
+
### S2 - provenance
|
|
37
|
+
|
|
38
|
+
Verify the download against the store's published md5/sha256 **before anything else touches
|
|
39
|
+
it**. **Produces:** the measured digest, written into the acquisition record with the
|
|
40
|
+
target slug + version and the source store. **Enforced by:** `RC-04` - a header naming
|
|
41
|
+
target slug + version + store + checksum, and a 64-hex token which, **if present**, must
|
|
42
|
+
equal the apk row's sha256 (a header carrying only the md5 passes - the comparison is
|
|
43
|
+
conditional in the engine).
|
|
44
|
+
|
|
45
|
+
### S3 - triage
|
|
46
|
+
|
|
47
|
+
The engine verdict, cheapest test first. Marker rules over the archive listing: `.so`
|
|
48
|
+
libraries - `lib/*/libunity.so` = Unity IL2CPP (with `libil2cpp.so`), `libue4.so` = Unreal,
|
|
49
|
+
`libcocos2d*.so` = Cocos, `libgodot*` = Godot; `assets/bin/Data/Managed` + `libmono` =
|
|
50
|
+
Unity Mono; `global-metadata.dat` = Unity IL2CPP; no `.so` and packed assets in `assets/`
|
|
51
|
+
= custom Java/Dalvik engine. **Produces:** the verdict line in the triage note.
|
|
52
|
+
**Enforced by:** nothing mechanised - the verdict is a note, and `RC-03` sees only that
|
|
53
|
+
the note lives under `notes/`.
|
|
54
|
+
|
|
55
|
+
### S4 - static decompile
|
|
56
|
+
|
|
57
|
+
apktool/jadx into the **quarantine**, never into any repo. **Produces:** the decompiled
|
|
58
|
+
tree under the sandbox quarantine root. **Enforced by:** `RC-03` - the lab allowlist has
|
|
59
|
+
no `build/`, `dist/`, or decompiled tree.
|
|
60
|
+
|
|
61
|
+
### S5 - carve the containers
|
|
62
|
+
|
|
63
|
+
A carver written **from decompiled evidence, not guessed**: read the loader in the
|
|
64
|
+
decompiled tree, recover the transform (key, offset, container layout), reimplement, and
|
|
65
|
+
only then run it. The precedent: Kairosoft's whole-file-XOR `.dat` containers, whose
|
|
66
|
+
loader was found at `kairo/android/util/g.java` and whose 16-byte XOR key came from
|
|
67
|
+
`b.a.a()` in the decompiled tree - the carver was written from that evidence, and its
|
|
68
|
+
parse was cross-checked against size tables *inside* each container (16/16 consistent).
|
|
69
|
+
**Produces:** extracted payloads under the quarantine root, and the carver script in the
|
|
70
|
+
lab repo (`scripts/`). **Enforced by:** `RC-03` - the script is tracked, the payloads are
|
|
71
|
+
not.
|
|
72
|
+
|
|
73
|
+
### S6 - manifest the corpus
|
|
74
|
+
|
|
75
|
+
sha256 of **every** extracted payload into `manifests/*.sha256`, with a header naming
|
|
76
|
+
target + version + source store + anchor, and the **apk row itself**. This `.sha256`
|
|
77
|
+
manifest is the sha256sum-format artifact: it is what `sha256sum -c` verifies and what
|
|
78
|
+
`RC-03`/`RC-04` read. It is **not** what `RC-01`/`RC-02` consume - those rows parse a
|
|
79
|
+
separate `reference-manifest/1` **JSON** (`generated_from`, `reference_app`, `entries[]`,
|
|
80
|
+
`entry_count`), declared via `reference_manifest:` in `.goblin/goblin.yaml`. That JSON is an
|
|
81
|
+
**optional** artifact, produced only if this lab ever ships a build output that could carry
|
|
82
|
+
corpus bytes; a study-only lab never ships, so none is produced here - a deliberate
|
|
83
|
+
omission, not a gap. Near-duplicate manifest rows are not linted - a known limitation.
|
|
84
|
+
**Produces:** the manifest. **Enforced by:** `RC-04` (the header + the apk row) and `RC-03`
|
|
85
|
+
(the payload-hash clause); `RC-02` exists for the *optional* JSON (the `reference-manifest/1`
|
|
86
|
+
schema so `RC-01` cannot pass vacuously - `entries` non-empty, `entry_count` honest, 64-hex
|
|
87
|
+
digests, byte counts).
|
|
88
|
+
|
|
89
|
+
### S7 - dossier
|
|
90
|
+
|
|
91
|
+
Facts and numbers, never expression: measurements, counts, class/method citations
|
|
92
|
+
(`file:line`), and the digests already in the manifest. **No extracted art, no copied
|
|
93
|
+
text** - an asset's dimensions and its hash are facts; the asset itself is expression and
|
|
94
|
+
stays in the quarantine. **Produces:** `notes/<date>-<target>.md`. **Enforced by:**
|
|
95
|
+
`RC-03` (the note is tracked; nothing it quotes may be a tracked payload byte) and, at
|
|
96
|
+
release time, `RC-01` (any corpus hash that shows up in a shipping build fails it).
|
|
97
|
+
|
|
98
|
+
### S8 - runtime analysis, deferred by design
|
|
99
|
+
|
|
100
|
+
Frida hooking needs a device or emulator and an owned build, and static facts come first:
|
|
101
|
+
a hash-manifested corpus and a written dossier are reproducible; a hook session is not.
|
|
102
|
+
S8 is **not** cut - it is deferred, and a future pass may pick it up with its own fence
|
|
103
|
+
(the device, the owned-build check, its own quarantine).
|
|
104
|
+
|
|
105
|
+
### S9 - retention / teardown
|
|
106
|
+
|
|
107
|
+
The corpus lives only in the sandbox quarantine; the lab repo keeps scripts, notes and
|
|
108
|
+
hashes only. **Delete or keep is a recorded decision** - recorded in the **lab repo** (a
|
|
109
|
+
note under `notes/`, or the HANDOFF NEXT section), with the date. **Enforced by:** `RC-03`,
|
|
110
|
+
re-run after any teardown: the tracked tree must
|
|
111
|
+
still be scripts/notes/manifests/docs only, and no tracked file may hash to a manifest row.
|
|
112
|
+
|
|
113
|
+
## Verifying the lab repo - the harness must be installed there first
|
|
114
|
+
|
|
115
|
+
`goblin-verify` requires an **installed harness** in the repo it points at: against a bare
|
|
116
|
+
lab repo it exits 2 with `not installed: <root>/.goblin/goblin.yaml is absent`. `--source`
|
|
117
|
+
relocates the enforcement manifest, not the target's config requirement, so the bridge is a
|
|
118
|
+
one-time install **into the lab repo itself**:
|
|
119
|
+
|
|
120
|
+
goblin-install --target <lab-repo> --class A
|
|
121
|
+
|
|
122
|
+
A class-A install is sufficient and violates nothing - the lab's own files are not touched;
|
|
123
|
+
it only adds the harness scaffolding: `.goblin/` (config + the verifier), `checks/`,
|
|
124
|
+
`reviews/`, `.github/workflows/`, and the `HANDOFF.md` scaffolding. Note: `HANDOFF.md`'s
|
|
125
|
+
placeholder commit gives `HP-05`'s one expected day-one red until a real commit is named.
|
|
126
|
+
After that one-time install, `goblin-verify --only RC-03` / `RC-04` run against the lab repo
|
|
127
|
+
(`RC-03`/`RC-04` need the installed harness; `RC-01`/`RC-02` additionally need a declared
|
|
128
|
+
`reference_manifest:`, which a study-only lab deliberately leaves empty).
|
|
129
|
+
|
|
130
|
+
## What the `RC-` rows bite
|
|
131
|
+
|
|
132
|
+
| row | gate | bites at |
|
|
133
|
+
|---|---|---|
|
|
134
|
+
| `RC-01` | build output | S7 / release: no file in the shipping tree matches a corpus hash (four clauses: skip-when-empty, fail-when-absent, the hash clause, and the manifest itself must not ship) |
|
|
135
|
+
| `RC-02` | the manifest | S6: the manifest's shape - so `RC-01` cannot pass vacuously |
|
|
136
|
+
| `RC-03` | the lab tree | S4-S9: every tracked path is on the allowlist and no tracked byte equals a payload |
|
|
137
|
+
| `RC-04` | the manifest header | S1-S2: the acquisition record exists and names target, version, store, checksum |
|
|
138
|
+
|
|
139
|
+
## First real run - 2026-09-27, Oh!Edo Towns Lite
|
|
140
|
+
|
|
141
|
+
- **Target:** `edotownsL-1.0.9.apk`, package `net.kairosoft.android.edotownsL`, version
|
|
142
|
+
1.0.9 (versionCode 10), 6,119,826 bytes, md5 `23974f8582359aa32eac30ad42a7744e`,
|
|
143
|
+
acquired from Aptoide (`pool.apk.aptoide.com`) - a free listing with a published md5.
|
|
144
|
+
- **Sandbox:** LXC 217 `re-lab`; quarantine root `/opt/re-lab/refs/edotownsL-1.0.9/`.
|
|
145
|
+
- **Verdict (S3):** Java/Dalvik, **custom Kairosoft `.dat` containers** - no engine
|
|
146
|
+
`.so` matched; 16 `.dat` files in `assets/`, with `xls.dat` the design-data suspect
|
|
147
|
+
(it turned out to be 3 localisation text entries, not balance data - the dossier
|
|
148
|
+
records the measurement, not the hope).
|
|
149
|
+
- **S5 precedent:** the carver (`scripts/karve-dat.py`) written from the decompiled
|
|
150
|
+
loader, not guessed; 312 payloads extracted, 16/16 containers parse-consistent.
|
|
151
|
+
- **Lab repo:** `~/projects/re-lab` - `scripts/triage-apk.sh`,
|
|
152
|
+
`scripts/karve-dat.py`, `manifests/edotownsL-1.0.9.sha256`,
|
|
153
|
+
`notes/2026-09-27-edotowns-lite-triage.md`.
|
|
154
|
+
|
|
155
|
+
## Verification
|
|
156
|
+
|
|
157
|
+
- The corpus manifest verifies `sha256sum -c` **where the corpus lives** (in the sandbox,
|
|
158
|
+
against the anchor the manifest header names).
|
|
159
|
+
- `goblin-verify` against the **lab repo needs the harness installed there first** (see
|
|
160
|
+
"Verifying the lab repo" above): one `goblin-install --target <lab-repo> --class A`, after
|
|
161
|
+
which `--only RC-03` / `RC-04` are runnable. `goblin-verify --only RC-01` / `RC-02` /
|
|
162
|
+
`RC-03` / `RC-04` return the exits their rows
|
|
163
|
+
define - an exact hash inside the build output, a weak manifest and a tracked payload
|
|
164
|
+
each fail the build.
|
|
165
|
+
- Every negative control NC-1..NC-6 was shown RED and then restored: NC-1/NC-2 under
|
|
166
|
+
`RC-01` (a payload, then the manifest itself, inside the build output), NC-3 under
|
|
167
|
+
`RC-02` (`entries` truncated to `[]`, with the pair-proving `RC-01`-stays-GREEN half),
|
|
168
|
+
NC-4 under `RC-01` (the declared manifest file deleted), NC-5/NC-6 under `RC-03`
|
|
169
|
+
(a payload git-added under an allowed path, then a tracked path outside the
|
|
170
|
+
allowlist). `tests/t-verify-red.sh` is the file that runs them.
|
|
171
|
+
|
|
172
|
+
## What this cannot see
|
|
173
|
+
|
|
174
|
+
- **A re-encoded, resized or recoloured asset** passes `RC-01`'s exact-hash gate - the
|
|
175
|
+
gate is sha256 equality, and a byte that differs is a different byte.
|
|
176
|
+
- **Copied text inside a shipped string** is level 3 and invisible to the hash gate.
|
|
177
|
+
- **A weak manifest authored by hand** is `RC-02`'s problem only if it is *malformed* -
|
|
178
|
+
a well-formed manifest of the wrong bytes is exactly what the schema cannot judge
|
|
179
|
+
(`LIMITS.md` #28, the weakened-input defect; `installed.json` is unsigned, #18).
|
|
180
|
+
- **Nothing here proves a fact is CORRECT** - only that it is traceable to the corpus.
|
|
181
|
+
The dossier's measurements are as good as the commands that produced them, and a
|
|
182
|
+
misread loader produces a confidently wrong carver. `RC-04`, the weakest of the four,
|
|
183
|
+
proves the acquisition record *exists*, never that the number came from the store.
|
package/docs/RISKS.md
ADDED
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# Risk register and non-goals
|
|
2
|
+
|
|
3
|
+
## Risks
|
|
4
|
+
|
|
5
|
+
| # | Risk | Counter-measure | Status |
|
|
6
|
+
|---|---|---|---|
|
|
7
|
+
| K1 | **Model promotion churn** breaks a role binding | Roles are capabilities, never slugs; one mapping file; `MD-01` lints every artifact for a hardcoded model name; `MD-02` reports family equality. A campaign that changes nine profiles changes nothing here. | Handled by design |
|
|
8
|
+
| K2 | **Profile drift across the fleet** | The attack surface is removed, not managed: project procedures live in the repo's `.hermes/skills/`, so there is no second copy to drift. `SK-02` hashes what is installed. The existing divergence outside goblin-stack's scope is a separate, escalated cleanup. | Handled in scope; fleet-side cleanup escalated |
|
|
9
|
+
| K3 | **Skills duplicating between profiles** | Same as K2. The precedence order is stated in `docs/INTEGRATION.md` so a future agent knows which copy wins instead of guessing. | Handled by design |
|
|
10
|
+
| K4 | **The vendored copy rots** (a repo sits at an old version) | `installed.json` records version and hashes; `goblin-verify` reports the installed version; `--upgrade` prints created/updated/unchanged. Accepted for repos that stop being worked on — the alternative (a network call at verify time) violates the offline dependency contract. | Accepted, with detection |
|
|
11
|
+
| K5 | **The harness decays into prose** | `IN-03` plus `SK-03`: a row without a command must say `advisory`, and the advisory count is capped (`advisory_ceiling`, default 10). The cap is the ratchet. | Handled by design |
|
|
12
|
+
| K6 | **A HANDOFF carries a stale number** | `HP-03` requires `measured <date>`; `HP-05` requires a commit that exists. | Partial: proves a date exists, not that the number is fresh. V1 anchored the row on the gate names the config declares (and skips the template's example sentence), so the real gate line can no longer lose its date in silence - but nothing re-measures the number, and a gate-bearing line that names no declared gate and carries no gate-shaped keyword is still unseen |
|
|
13
|
+
| K7 | **The gate is theatre** (self-skipping checks, admin bypass) | `PG-05` FAILs a conditional **job**, a conditional **step**, a job with no step and a workflow with no `jobs:` - re-declared at W4, because the old predicate counted "every step is guarded" and therefore passed the measured real shape (one unguarded step deciding whether the guarded gate step runs) while GitHub reported Success. `PG-06` then requires the gate CI runs to be the gate the project declares. `PG-04` documents the bypass. *"A required check that self-skips reports success… and a protected branch whose only admin is the person pushing protects nothing."* | Mechanised where a file can be read (`docs/CI.md` §1); the required-check list, the bypass switch and the push identity are forge state and stay documented, not solvable from the repo (`docs/LIMITS.md` #13, #34) |
|
|
14
|
+
| K8 | **Cost** | Roles plus effort tokens; the panel only at S3+; sweeps are cron-and-one-line; no playbook fans out without a named predicate. | Handled by policy |
|
|
15
|
+
| K9 | **The builder may only install into its own repo plus a scratch copy** | The install proof target is a throwaway copy under the scratch directory; the recipe is in `docs/ADOPTION.md`. | Constraint, satisfied |
|
|
16
|
+
| K10 | **A fleet-config repo is the least-governed artifact in an estate** | The E-class preset plus an artifact-scoped gate (a commit exists). The underlying staleness bug in a backup job is named and escalated — goblin-stack can detect staleness but cannot fix another repository. | Escalated |
|
|
17
|
+
| K11 | **The pre-change tree is unknown in a repo with no git history** | Repos with no `.git` are ordered *after* `git init`. `HS-02` is **skipped with a reason** rather than faked while no pinned commit exists. | Handled by ordering |
|
|
18
|
+
| K12 | **A check green on both trees** (the failure mode the REPLAY exists for) | `HS-02` runs the harness set against the pinned pre-change commit and requires **every** harness to be RED there. The shipped scaffold harness is deliberately such a check and is therefore reported as unproven until it is replaced. | Handled by design; see `docs/LIMITS.md` |
|
|
19
|
+
| K13 | **A nightly automation files the same defect twice, or files one that is not there** | A content-only dedup key (`--idempotency-key`, checked by `AU-02`) plus the board's own `recent_success` and `active_pr` guards; a report whose `revision` does not resolve is a refusal, not a card; `--max-runtime`, `--max-retries 1` and the failure limit auto-block a looping card. The producer's own ceiling bounds filings per day. | Handled by design; the key is a dedup, not a mutex (`docs/LIMITS.md` #20) |
|
|
20
|
+
| K14 | **A waiver becomes a permanent blind spot** - a dated exception nobody re-decides, or an audit record nobody re-takes, so `SC-07` passes against facts that are no longer true | The record's date is checked against `security.audit_max_age_days` (90) and every waiver's date against `security.waiver_max_age_days` (180); the waiver count is printed on the gate line so the debt is loud even when the row passes; `goblin-audit` writes a refusal (exit 5) instead of an empty record, so "no advisories" can never mean "the parse found nothing". | Handled by design; the policy is unmeasured against a real registry (`docs/LIMITS.md` #22) |
|
|
21
|
+
| K15 | **goblin-stack installs agent-authored skills, and the one controlled study of that practice puts it BELOW the no-skill baseline.** SkillsBench 1.1's self-generated condition (the agent authors its own Skills before solving) reports all three tested configurations below their no-Skills baseline (-8.1, -11.3, -11.5 points), while curated Skills rose +16.6 points across 18 configurations. A generated skill accepted after a skim is a different proposition from a written one. | `P6` hands every generated verification skill to `P12`, and `verified:` does not advance until an eval record exists; the record's shape, the eleven-token ban, the cheap-checks-first ladder and the pass condition (every seeded defect detected, the control at zero, every correction RED before GREEN) are specified in `skills/goblin-eval/SKILL.md`. The evidence is cited with its pin in `docs/LIMITS.md` #15. | **Stated requirement, not an enforced one**: the runner is not shipped and no row reads a lane (`docs/LIMITS.md` #31) |
|
|
22
|
+
| K16 | **The feature map rots, or claims coverage it does not have** - a route renamed under a recipe that still "works", a feature file nobody indexed, a `verified:` date nobody drove | `FM-01` (every feature file indexed, the four-H2 entry contract, the slug matches the filename), `FM-02` (every declared entry path still resolves under `source_root:`, no entry path changed after its `verified:` date), `VA-01` (the declared `verify_doctor:` exits 0); the upkeep pass and the rot table live in `skills/goblin-feature-map/SKILL.md`. | Handled by design for everything the map LISTED; completeness is not checkable (`docs/LIMITS.md` #30) |
|
|
23
|
+
| K17 | **An unattended loop grades itself** - the judge is a language model handed the worker's own prose, the lane that judges can be the lane that wrote, and Hermes has **no progress detector**: `run_kanban_goal_loop` carries no progress state, so a loop returning `continue` for the same reason nineteen times spends nineteen turns and then blocks | `JG-01` (a `done` verdict may only cite a handle the repo can resolve - a commit in `git rev-list --all`, a path under the root, a `sha256:` of a file under `.goblin/loop/`), `JG-02` (the declared judge lane must be **disjoint** from the author's - a FAIL, not a report), `JG-03` (counted: a lane with no non-`done` verdict is escalated), and `LP-01`..`LP-05` (one predicate command, run and recorded before iteration 1, pinned by digest, budgeted under a ceiling, no three identical pointers without a green, and a write-up when it ends red). The contract is `docs/LOOP.md`. | **Partial, and stated**: disjoint *profiles* is gated; disjoint *families* needs the mapping file (`MD-02`, ADV) and *which lane ran* is unobservable from a repo (`MD-03`); the judge lane resolves to no profile on this box, so `JG-02` reports ADV with its one-line remedy rather than failing a repo for the fleet's routing. The judge's own failure modes (prose-not-evidence, a lane that always says yes, no progress detector, cost) are `docs/LIMITS.md` #32 and #33 |
|
|
24
|
+
|
|
25
|
+
| K18 | **The shipped workflow is mistaken for a gate** — a file in `.github/workflows/` with no required-check entry, with the admin-bypass switch on, or pushed under the sole admin's own identity, is decoration that reads as enforcement | The template's header states the four settings that make a workflow a gate, `docs/CI.md` §1 argues them with the vendor's own words, `PG-05` refuses a conditional job or step and `PG-06` refuses a workflow that runs some other truth, `CL-01` makes `ci-gate` a class contract (absent where the class forbids it), and the verifier's "cannot see" footer names the lane on every run | The file half is mechanised; the forge half is not observable from a repo, and `PG-04` is the row that says so (`docs/LIMITS.md` #34) |
|
|
26
|
+
|
|
27
|
+
## The advisory rows, named
|
|
28
|
+
|
|
29
|
+
Ten rows are labelled `advisory` (this sentence said nine until 2026-09-25: the tenth, `JG-03`,
|
|
30
|
+
landed with G2's judge lane), and the count is capped by `SK-03` (default ceiling 10 — the cap is
|
|
31
|
+
now **full**, `advisory 10 of ceiling 10`).
|
|
32
|
+
Nine carry **no executable check at all**:
|
|
33
|
+
|
|
34
|
+
- **HP-04** — a stale sentence is corrected in place with a dated parenthetical, never deleted.
|
|
35
|
+
- **HS-03** — source probes read text with comments blanked first.
|
|
36
|
+
- **CM-02** — the commit message was written to a file, not passed inline.
|
|
37
|
+
- **MD-03** — role-pinned fan-out goes through the kanban, not a model-less subagent spawn.
|
|
38
|
+
- **PG-04** — never bypass what the forge enforces.
|
|
39
|
+
- **DOC-01** — a significant change updates the docs that teach it.
|
|
40
|
+
- **DOC-02** — system-level changes are recorded wherever the project's standard says they live.
|
|
41
|
+
- **SC-09** — auth is applied consistently across sibling routes.
|
|
42
|
+
- **JG-03** — a judge lane that has never returned a non-`done` verdict is escalated. The history
|
|
43
|
+
that would show a bad lane lives across cards and repos, so the counter-measure is policy (one
|
|
44
|
+
known-red control verdict per wave, `docs/LOOP.md`), not a command.
|
|
45
|
+
|
|
46
|
+
One is advisory-labelled but still **reports its state** as `ADV`:
|
|
47
|
+
|
|
48
|
+
- **MD-02** — the review lane is a different model family from the code lane.
|
|
49
|
+
|
|
50
|
+
`HP-04`, `CM-02` and `MD-02` additionally print a heuristic when run with `--only`. A heuristic
|
|
51
|
+
is not a check: it never fails a run. A counted rule is still not an enforced one, and the cap
|
|
52
|
+
is a policy, not a proof.
|
|
53
|
+
|
|
54
|
+
## Non-goals
|
|
55
|
+
|
|
56
|
+
Stated so no future reader infers them:
|
|
57
|
+
|
|
58
|
+
- not a plugin or a marketplace package; no slash commands;
|
|
59
|
+
- **no auto-merge** — no reviewed source ships it unconditionally, and merging is a different
|
|
60
|
+
decision from a green gate;
|
|
61
|
+
- not a fleet orchestrator — the board owns that;
|
|
62
|
+
- does not provision, migrate or verify models;
|
|
63
|
+
- does not write the vault;
|
|
64
|
+
- does not replace any project's existing gate (adopt, don't replace);
|
|
65
|
+
- is not a monorepo tool and not a CI tool. **Amended at W4:** it now places **at most one**
|
|
66
|
+
workflow, into its own target, for a class that requires or permits `ci-gate` — and it runs
|
|
67
|
+
nothing for that repo, holds no credentials, and cannot arm a check. See `docs/CI.md`;
|
|
68
|
+
- does not manage profile skill libraries;
|
|
69
|
+
- writes nothing outside the target repo;
|
|
70
|
+
- does not attempt to make the *prose* rules enforceable — it counts them instead.
|
package/docs/ROLES.md
ADDED
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
# Roles and models
|
|
2
|
+
|
|
3
|
+
A **profile** is a worker identity. A **role** is what the work needs. Conflating them is how
|
|
4
|
+
five documents once gave five different model answers — a role name used as a model name.
|
|
5
|
+
|
|
6
|
+
goblin-stack keeps them separate and **never names a model**. `MD-01` lints every reusable rule
|
|
7
|
+
for a hardcoded model name; `MD-02` reports, but cannot change, whether the review lane and the
|
|
8
|
+
code lane resolve to the same family.
|
|
9
|
+
|
|
10
|
+
## The role vocabulary
|
|
11
|
+
|
|
12
|
+
`roles.yaml` declares each role's capability and the default **profile** that carries it:
|
|
13
|
+
|
|
14
|
+
| role | what it needs | default profile | used by |
|
|
15
|
+
|---|---|---|---|
|
|
16
|
+
| `code` | fast, cheap, mechanical correctness | `coder` | P2, P4, P5, P10 |
|
|
17
|
+
| `judgment` | strongest available reasoning and prose | `architect` | P1, P3 (spec half), P6, P8, P9 |
|
|
18
|
+
| `review-panel` | **a list** — N independent verdict lanes, each its own lane | `[reviewer, architect]` | P7 at stakes S3+ |
|
|
19
|
+
| `synthesis` | merges many outputs into one artifact | `architect` | P11, P12 |
|
|
20
|
+
| `investigate` | read-only exploration returning a distilled summary | `researcher` | P1 (read-only half) |
|
|
21
|
+
| `judge` | decides whether a process met its own predicate, from a command's output and a pointer it can resolve, never from a report | `judge` (a **single** lane) | P10, P12, P7 at S3+, the terminal handoff gate |
|
|
22
|
+
|
|
23
|
+
The list-valued panel is the sharpest configuration idea retained here: **one lane runs per list
|
|
24
|
+
entry, so the list length sets the lane count.** Lane count is configuration, not code.
|
|
25
|
+
|
|
26
|
+
A **panel is N opinions; a judge is one decision** — which is why `role-judge` is a role and not a
|
|
27
|
+
sentence inside P7 or P12. The judge's contract (what it receives, what it must refuse, how its
|
|
28
|
+
verdict is recorded) is `docs/LOOP.md`; the pair of rules that make it real are that a `done`
|
|
29
|
+
verdict may only cite a handle the repo can resolve (`JG-01`) and that the judge's lane must be
|
|
30
|
+
disjoint from the author's (`JG-02`).
|
|
31
|
+
|
|
32
|
+
## The model-mapping contract
|
|
33
|
+
|
|
34
|
+
The machine-specific mapping (`profile → provider/model/effort`) is read at run time from the
|
|
35
|
+
file declared as `models_file:` in `.goblin/goblin.yaml`. That is **the single documented
|
|
36
|
+
install-time machine input**, and goblin-stack **reads it and never writes it**.
|
|
37
|
+
|
|
38
|
+
Why not ship a copy of the mapping: the mapping file is a live routing source owned and
|
|
39
|
+
rewritten by another tool. Adding a foreign key to a file another tool owns is the precedent
|
|
40
|
+
for what happens when two tools share a file with unstated ownership. One place to change a
|
|
41
|
+
model, and the model layer stays where it belongs.
|
|
42
|
+
|
|
43
|
+
Resolution:
|
|
44
|
+
|
|
45
|
+
bin/goblin-model <role> # one line per lane: profile provider model effort
|
|
46
|
+
bin/goblin-model --list # the roles and their capabilities
|
|
47
|
+
bin/goblin-model review-panel # the panel: one line per lane
|
|
48
|
+
|
|
49
|
+
`bin/goblin-model` is **checkout-only**. `goblin-install` copies four scripts into a target's
|
|
50
|
+
`.goblin/bin/` — `goblin-verify`, `goblin-lib.sh`, `goblin-audit` and `goblin-bans` — so an
|
|
51
|
+
adopted repo has no `goblin-model` command (`ls .goblin/bin/` →
|
|
52
|
+
`goblin-audit goblin-bans goblin-lib.sh goblin-verify`). The installed path
|
|
53
|
+
for the same resolution is the `resolve_role_models` helper inside `.goblin/bin/goblin-verify`,
|
|
54
|
+
which is what `MD-02` calls; `bin/goblin-model` exists for a human at a checkout, is covered only
|
|
55
|
+
by `bash -n` in `tests/run-tests.sh`, and has no `enforcement.tsv` row because it enforces
|
|
56
|
+
nothing — it prints. Naming it here is the alternative R6 §2.2 allows to shipping a row for it.
|
|
57
|
+
|
|
58
|
+
Absent on this machine, every role resolves to `unknown` and the model-dependent checks report
|
|
59
|
+
advisory. Absent is not an error — goblin-stack is portable, and another machine has no such
|
|
60
|
+
file. That is the one documented exception to the portability rule, and `PT-01` enforces it:
|
|
61
|
+
the path is a config *value* in `.goblin/goblin.yaml`, never a literal inside a rule.
|
|
62
|
+
|
|
63
|
+
## The fan-out rule: role-pinned work goes through the kanban
|
|
64
|
+
|
|
65
|
+
Measured, from the live tool schemas:
|
|
66
|
+
|
|
67
|
+
- a bare subagent spawn takes **no** model or provider parameter, so it cannot honour a role —
|
|
68
|
+
it would silently run every lane on one model, which is exactly the review failure the panel
|
|
69
|
+
exists to prevent;
|
|
70
|
+
- a board card **does** take a model and a provider, and a review request takes a reviewer
|
|
71
|
+
profile.
|
|
72
|
+
|
|
73
|
+
Therefore: a panel is N board cards parented to the change, each carrying the resolved
|
|
74
|
+
`(provider, model)`, the pinned SHA, the diff, and one focus. `goblin-mode` forbids inventing a
|
|
75
|
+
fan-out that a subagent spawn cannot honour, and `MD-03` records that no repo-local file can
|
|
76
|
+
observe which tool created a worker — so this rule is stated here and enforced at board level.
|
|
77
|
+
|
|
78
|
+
## The measured caveat
|
|
79
|
+
|
|
80
|
+
Today's mapping puts every switchable profile on the **same** model, so `review-panel` resolves
|
|
81
|
+
to the same family as `code`, and **`judge` resolves to no lane at all**. `MD-02` therefore reports
|
|
82
|
+
the state and is labelled advisory: goblin-stack cannot choose the fleet's models, and a harness
|
|
83
|
+
must not fail a repo for a fleet-wide campaign. This is a reported number, not a silent assumption.
|
|
84
|
+
|
|
85
|
+
The judge lane, measured on this box on 2026-09-25:
|
|
86
|
+
|
|
87
|
+
bash bin/goblin-model judge -> judge unknown unknown unknown (rc 0)
|
|
88
|
+
bash bin/goblin-model code -> coder <provider> <model> <effort>
|
|
89
|
+
bash bin/goblin-model review-panel -> reviewer <...> · architect <...>
|
|
90
|
+
|
|
91
|
+
`~/projects/fleet-model.yaml` has no `judge:` entry, and `hermes profile list` reports no `judge`
|
|
92
|
+
profile. So **`MD-02`'s rule is not satisfiable on this box today through the mapping file**: the
|
|
93
|
+
judge and the author would run on one family. A different family exists only as a per-profile
|
|
94
|
+
alias held in another profile's own config (a comment in the mapping file names it), and **the
|
|
95
|
+
mapping file cannot express an alias** — it maps a profile to a provider/model pair. Making the
|
|
96
|
+
constraint real is a **fleet-side** change (add the `judge` profile, add the entry, re-apply),
|
|
97
|
+
never a goblin-stack one; `JG-02` therefore gates the *declared lanes* being disjoint and reports
|
|
98
|
+
an unresolved judge lane as `ADV` with the one-line remedy rather than failing a repo for the
|
|
99
|
+
fleet's routing.
|
|
100
|
+
|
|
101
|
+
## Budget
|
|
102
|
+
|
|
103
|
+
`effort` is one word per role, read from the same mapping file. The panel is bounded — lanes
|
|
104
|
+
only at S3+ — because parallel lanes cost about N times the tokens, and token usage is the
|
|
105
|
+
dominant term in the variance of agent outcomes. No playbook fans out without a named predicate.
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
id ban globs detect replacement escape reviewer source
|
|
2
|
+
BN-01 No `any` in application TypeScript. src bash .goblin/bans/grep-ban.sh -e ':[[:space:]]*any\b' -e '<any>' -e '<any,' -e ',any>' -e 'any\[\]' src Use the real type, or narrow an `unknown` at the boundary `// BAN-OK(BN-01): <reason>` on the offending line (a non-empty reason is required), or list the path under bans_exempt: architecture R2 section A12 (Dune rule 2); R1 section 3
|
|
3
|
+
BN-02 No `@ts-ignore` / `@ts-expect-error` suppressions. src bash .goblin/bans/grep-ban.sh -e '@ts-(ignore|expect-error|nocheck)' src Fix the type, or narrow with `unknown` at the boundary `// BAN-OK(BN-02): <reason>` architecture R1 section 6 (suppressions are stink)
|
|
4
|
+
BN-03 No direct network call from a component. src/components app bash .goblin/bans/grep-ban.sh -e '\bfetch[[:space:]]*\(' src/components app Put the call in the data layer (src/data/*) `// BAN-OK(BN-03): <the endpoint, and when it migrates>` architecture R2 section A12; E1 (overreach)
|
|
5
|
+
BN-05 No import across a declared layer boundary. src bash .goblin/bans/layer-check.sh Move the code, or import through the declared accessor declare the pair under layers:, or move the file architecture R2 section A12 (Dune's dependency-graph check)
|
|
6
|
+
BN-06 No renderer with Node access (`nodeIntegration: true`). src app electron bash .goblin/bans/grep-ban.sh -e 'nodeIntegration[[:space:]]*:[[:space:]]*true' src app electron Pass data over IPC through a preload script (contextBridge); leave nodeIntegration false `// BAN-OK(BN-06): <why the renderer needs a Node primitive>` architecture G6 section B.2 item 1 (Electron security S13/S14)
|
|
7
|
+
BN-07 No renderer with context isolation or the process sandbox turned off. src app electron bash .goblin/bans/grep-ban.sh -e 'contextIsolation[[:space:]]*:[[:space:]]*false' -e 'sandbox[[:space:]]*:[[:space:]]*false' src app electron Keep `contextIsolation: true` and `sandbox: true`; expose a narrow API from the preload script `// BAN-OK(BN-07): <why isolation is off, and what replaces it>` architecture G6 section B.2 item 2 (Electron security S14)
|
|
8
|
+
BN-08 No dangerous webPreferences. src app electron bash .goblin/bans/grep-ban.sh -e 'webSecurity[[:space:]]*:[[:space:]]*false' -e 'allowRunningInsecureContent[[:space:]]*:[[:space:]]*true' -e 'enableBlinkFeatures' -e 'allowpopups' src app electron Fix the origin or the CSP; a disabled web security model is not a workaround `// BAN-OK(BN-08): <the check, and when it goes>` architecture G6 section B.2 item 9 (Electron security S14 checklist items 6, 8, 10, 11)
|
|
9
|
+
BN-09 No synchronous IPC and no `@electron/remote`. src app electron bash .goblin/bans/grep-ban.sh -e 'sendSync[[:space:]]*\(' -e '@electron/remote' src app electron Use `ipcRenderer.invoke` / `ipcMain.handle` (async), and a preload-exposed API instead of @electron/remote `// BAN-OK(BN-09): <why a blocking call is the only option here>` architecture G6 section B.2 item 11 (Electron performance S13)
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
class part need
|
|
2
|
+
A handoff R
|
|
3
|
+
B handoff R
|
|
4
|
+
C handoff R
|
|
5
|
+
D handoff R
|
|
6
|
+
E handoff R
|
|
7
|
+
F handoff R
|
|
8
|
+
A spec R
|
|
9
|
+
B spec R
|
|
10
|
+
C spec R
|
|
11
|
+
D spec -
|
|
12
|
+
E spec R
|
|
13
|
+
F spec R
|
|
14
|
+
A gate R
|
|
15
|
+
B gate R
|
|
16
|
+
C gate R
|
|
17
|
+
D gate O
|
|
18
|
+
E gate R
|
|
19
|
+
F gate R
|
|
20
|
+
A replay R
|
|
21
|
+
B replay -
|
|
22
|
+
C replay R
|
|
23
|
+
D replay -
|
|
24
|
+
E replay O
|
|
25
|
+
F replay R
|
|
26
|
+
A ratchet R
|
|
27
|
+
B ratchet O
|
|
28
|
+
C ratchet O
|
|
29
|
+
D ratchet -
|
|
30
|
+
E ratchet O
|
|
31
|
+
F ratchet R
|
|
32
|
+
A pr-gate O
|
|
33
|
+
B pr-gate -
|
|
34
|
+
C pr-gate O
|
|
35
|
+
D pr-gate -
|
|
36
|
+
E pr-gate O
|
|
37
|
+
F pr-gate O
|
|
38
|
+
A review-panel O
|
|
39
|
+
B review-panel -
|
|
40
|
+
C review-panel R
|
|
41
|
+
D review-panel -
|
|
42
|
+
E review-panel O
|
|
43
|
+
F review-panel O
|
|
44
|
+
A playbooks R
|
|
45
|
+
B playbooks R
|
|
46
|
+
C playbooks R
|
|
47
|
+
D playbooks R
|
|
48
|
+
E playbooks R
|
|
49
|
+
F playbooks R
|
|
50
|
+
A tokens O
|
|
51
|
+
B tokens -
|
|
52
|
+
C tokens -
|
|
53
|
+
D tokens -
|
|
54
|
+
E tokens -
|
|
55
|
+
F tokens O
|
|
56
|
+
A ci-gate R
|
|
57
|
+
B ci-gate -
|
|
58
|
+
C ci-gate O
|
|
59
|
+
D ci-gate -
|
|
60
|
+
E ci-gate O
|
|
61
|
+
F ci-gate R
|