@techgoblin/gobstack 0.0.0-stage → 0.4.4-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +351 -0
- package/LICENSE +21 -0
- package/README.md +163 -2
- package/VERSION +1 -0
- package/adapters/_template/adapter.tsv +16 -0
- package/adapters/_template/detect.sh +10 -0
- package/adapters/_template/emit.sh +5 -0
- package/adapters/_template/verify.sh +4 -0
- package/adapters/claude/adapter.tsv +8 -0
- package/adapters/claude/detect.sh +8 -0
- package/adapters/claude/verify.sh +47 -0
- package/adapters/codex/adapter.tsv +12 -0
- package/adapters/codex/detect.sh +9 -0
- package/adapters/codex/verify.sh +45 -0
- package/adapters/copilot/adapter.tsv +10 -0
- package/adapters/copilot/detect.sh +8 -0
- package/adapters/copilot/verify.sh +45 -0
- package/adapters/cursor/adapter.tsv +11 -0
- package/adapters/cursor/detect.sh +10 -0
- package/adapters/cursor/verify.sh +45 -0
- package/adapters/gemini/adapter.tsv +15 -0
- package/adapters/gemini/detect.sh +11 -0
- package/adapters/gemini/verify.sh +49 -0
- package/adapters/hermes/adapter.tsv +9 -0
- package/adapters/hermes/detect.sh +8 -0
- package/adapters/hermes/verify.sh +27 -0
- package/adapters/opencode/adapter.tsv +14 -0
- package/adapters/opencode/detect.sh +9 -0
- package/adapters/opencode/verify.sh +45 -0
- package/automations/README.md +53 -0
- package/automations/bugreporter-intake.sh +145 -0
- package/automations/drift-audit.sh +139 -0
- package/automations/report.schema.tsv +10 -0
- package/bans/README.md +82 -0
- package/bans/grep-ban.sh +84 -0
- package/bans/layer-check.sh +57 -0
- package/bin/goblin +119 -0
- package/bin/goblin-audit +145 -0
- package/bin/goblin-bans +178 -0
- package/bin/goblin-doctor +233 -0
- package/bin/goblin-emit +482 -0
- package/bin/goblin-init +519 -0
- package/bin/goblin-install +720 -0
- package/bin/goblin-lib.sh +289 -0
- package/bin/goblin-model +105 -0
- package/bin/goblin-upgrade +572 -0
- package/bin/goblin-verify +2798 -0
- package/bin/goblin.js +48 -0
- package/docs/ADOPTION.md +168 -0
- package/docs/CI.md +187 -0
- package/docs/CONTRACTS.md +197 -0
- package/docs/DESIGN.md +92 -0
- package/docs/ENFORCEMENT.md +225 -0
- package/docs/FLOWS.md +164 -0
- package/docs/GUARDRAILS.md +126 -0
- package/docs/GUIDE.md +610 -0
- package/docs/INTEGRATION.md +92 -0
- package/docs/LIMITS.md +591 -0
- package/docs/LOOP.md +165 -0
- package/docs/RE-PLAYBOOK.md +183 -0
- package/docs/RISKS.md +70 -0
- package/docs/ROLES.md +105 -0
- package/manifest/bans.tsv +9 -0
- package/manifest/classes.tsv +61 -0
- package/manifest/enforcement.tsv +88 -0
- package/manifest/glossary.tsv +25 -0
- package/manifest/playbooks.tsv +16 -0
- package/package.json +36 -4
- package/presets/A-shipped-software.yaml +48 -0
- package/presets/B-service-config.yaml +40 -0
- package/presets/C-game.yaml +38 -0
- package/presets/D-knowledge.yaml +41 -0
- package/presets/E-fleet-config.yaml +42 -0
- package/presets/F-electron.yaml +67 -0
- package/roles.yaml +54 -0
- package/skills/goblin-bootstrap/SKILL.md +51 -0
- package/skills/goblin-bugfix/SKILL.md +26 -0
- package/skills/goblin-bugreporter/SKILL.md +52 -0
- package/skills/goblin-drift-audit/SKILL.md +43 -0
- package/skills/goblin-eval/SKILL.md +68 -0
- package/skills/goblin-feature/SKILL.md +26 -0
- package/skills/goblin-feature-map/SKILL.md +140 -0
- package/skills/goblin-handoff/SKILL.md +28 -0
- package/skills/goblin-investigation/SKILL.md +26 -0
- package/skills/goblin-judge/SKILL.md +74 -0
- package/skills/goblin-loop/SKILL.md +88 -0
- package/skills/goblin-mode/SKILL.md +70 -0
- package/skills/goblin-overnight/SKILL.md +42 -0
- package/skills/goblin-pr-gate/SKILL.md +42 -0
- package/skills/goblin-re-mobile/SKILL.md +51 -0
- package/skills/goblin-refactor/SKILL.md +23 -0
- package/skills/goblin-sweep/SKILL.md +23 -0
- package/skills/goblin-tdd-repro/SKILL.md +27 -0
- package/skills/goblin-verify-author/SKILL.md +50 -0
- package/skills/practice/SKILL.md +37 -0
- package/templates/AGENTS.md.tmpl +23 -0
- package/templates/HANDOFF.md.tmpl +43 -0
- package/templates/SPEC.md.tmpl +34 -0
- package/templates/audit-waiver.tsv.tmpl +10 -0
- package/templates/boundary-waivers.tmpl +8 -0
- package/templates/checks/assert.mjs.tmpl +60 -0
- package/templates/checks/gate.sh.tmpl +29 -0
- package/templates/ci/goblin-gate.yml.tmpl +46 -0
- package/templates/goblin.yaml.tmpl +138 -0
- package/templates/install-hooks.allowlist.tmpl +9 -0
- package/templates/loop/decisions.tsv.tmpl +1 -0
- package/templates/loop/predicate.tmpl +16 -0
- package/templates/report.yaml.tmpl +16 -0
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
id scope rule enforced_by artifact check if_not_why
|
|
2
|
+
IN-01 target The install exists and records its version + every file's hash. script .goblin/installed.json goblin-verify --only IN-01 — (W1: the check is a builtin so the engine.mode=global clause can run — a global, declaration-only repo has no install record and SKIPs with `global engine mode — no per-repo install record`; the vendored clauses are the old one-liner: the record exists and names its version)
|
|
3
|
+
IN-02 target Every installed file still matches its recorded hash. script .goblin/installed.json goblin-verify --only IN-02 —
|
|
4
|
+
IN-03 target The verifier's own manifest is complete: every rule has a check or is advisory. script manifest/enforcement.tsv goblin-verify --only IN-03 — (this row is the reason the matrix cannot rot; the second clause is D6's shape in general: a row that carries no check must be labelled advisory, or it claims verification it does not perform. The third clause is Z1-5: `enforced_by` is documented as a closed enum in docs/ENFORCEMENT.md and was read by NOTHING, so a typo in that cell changed nothing - `script|lint|gate|advisory` plus `test`, the source-scope value whose check is tests/run-tests.sh. W1: the check runs against whichever manifest the engine actually resolved (the chain in bin/goblin-verify), no longer the hardcoded per-repo path — the rule's meaning is untouched, only the path input follows the engine)
|
|
5
|
+
IN-04 target No file goblin-stack did not create has been overwritten. script .goblin/installed.json goblin-verify --only IN-04 Detects a file the installer recorded as pre-existing (a `refused` entry) that has since vanished, or that is listed as installed anyway. The second clause is an internal-consistency guard: with correct code a refused path is never written, so it fires only if the installer regresses. The negative control exercises the vanished branch.
|
|
6
|
+
HP-01 target HANDOFF.md exists at the root. gate HANDOFF.md test -f HANDOFF.md —
|
|
7
|
+
HP-02 target The HANDOFF carries its five required sections. lint HANDOFF.md for h in 'START HERE' 'STATE|STATUS' 'GATES?' 'NEXT STEPS|NEXT' 'NOT VERIFIED|UNVERIFIED|NOT PROVEN|UNPROVEN|PENDING[^.]*DEVICE TEST'; do grep -qiE "^#{2,3}[[:space:]]+[^[:alpha:]]*($h)\b" HANDOFF.md || { echo "missing section: $h"; exit 1; }; done Partial: proves a heading exists whose first word names the slot, not that the section's content is complete. The slot may use the repo's own vocabulary (the model repo's `### NEXT` passes; `Status`, `Gate`, `Unverified`, `Pending <x> device test` are accepted); a heading that merely contains the word (`## BOARD STATE`) does not satisfy `State`.
|
|
8
|
+
HP-03 target Every gate number is a measurement with a date, never a copy. lint HANDOFF.md goblin-verify --only HP-03 Builtin, anchored on the gate NAMES `.goblin/goblin.yaml` declares (GT-01's source of truth) rather than on a hardcoded keyword list: the shipped gates are named `commit` and `todo_ceiling`, neither of which the old list matched, so on a fresh install the only line it could see was the template's own example sentence, and a real gate line could lose its `date` and the row stayed GREEN (G8-2). A line whose text says "example of the required form" is template prose, not a gate number, and is skipped, so the template cannot satisfy the row. Partial: it proves a dated line inside the `Gates` section exists and that every gate-bearing line there carries a date - not that the number was re-measured that day, and not a gate-bearing line that names no declared gate and carries no gate-shaped keyword (so a line the config does not declare and the keyword list does not recognise is unseen). Historical gate lines outside that section are exempt by design - PROJECT-PRACTICE section 1's stale-sentence rule requires them to be kept.
|
|
9
|
+
HP-04 target A stale sentence is corrected in place with a dated parenthetical, never deleted. advisory HANDOFF.md advisory Detecting a silent deletion needs semantic judgement; a diff heuristic (>=5 removed non-empty lines with no 'corrected' addition) is too noisy to gate on. goblin-verify --only HP-04 prints the heuristic as a warning only.
|
|
10
|
+
HP-05 target The HANDOFF names the HEAD it describes. gate HANDOFF.md goblin-verify --only HP-05 Deviation from the design spec, with the reason: the spec's literal check is `grep -q "$(git rev-parse --short HEAD)" HANDOFF.md`, which can never pass - committing the HANDOFF moves HEAD, so the file can only name a commit that is now an ancestor. The mechanised form is therefore 'the HANDOFF names a commit that exists in this repo AND is an ancestor of HEAD', which still catches the defect it exists for (a review/handoff artifact that names no commit at all).
|
|
11
|
+
SP-01 target A *-SPEC.md file exists at the repo root (any round, not the current one - round-scoping arrives with the W6 staged chain). gate *-SPEC.md ls ./*-SPEC.md >/dev/null 2>&1 Skipped when class: D and disabled: [spec].
|
|
12
|
+
SP-02 target The SPEC is committed, not left untracked. script git ls-files test -z "$(git ls-files --others --exclude-standard -- '*-SPEC.md')" —
|
|
13
|
+
SP-03 target Every AC: item is checkable without a human. lint *-SPEC.md awk '/^[[:space:]]*[-*][[:space:]]/ && /AC[0-9]*:/ && !/`|==|===|exit|<|>/ {print; bad=1} END{exit bad}' ./*-SPEC.md Partial: structural only - a checkable-looking bullet can still be unfalsifiable. The matcher is `AC[0-9]*:` because the shipped template writes `- AC1: ...`; the literal `/AC:/` only saw a bullet that spelled the label without a number.
|
|
14
|
+
GT-01 target The gate set is declared, never inferred from the stack. gate .goblin/goblin.yaml goblin-verify --only GT-01 Builtin: every `- name:` under `gates:` is a DECLARED gate and each one must carry a `cmd:` - a gate whose `cmd:` was deleted, blanked or re-indented FAILs here instead of vanishing from the count (G8-3, the condition of G8's own 9/10 sentence). The count is of declarations, not of runnable pairs. It cannot see whether a declared command is the RIGHT gate for the project: it proves a command exists, not that it is meaningful.
|
|
15
|
+
GT-02 target Every declared gate runs and exits 0. gate goblin.yaml gates: goblin-verify --only GT-02 —
|
|
16
|
+
GT-03 target The round reports one line of measured numbers. script .goblin/last-gate-line test -f .goblin/last-gate-line && [ .goblin/last-gate-line -nt "$(git rev-parse --git-dir)/logs/HEAD" ] The freshness reference is HEAD's reflog (`.git/logs/HEAD`), which every HEAD movement rewrites - a commit in an attached or a detached worktree included, and independently of whether the refs are packed. The clause it replaced read `.git/HEAD`, a file a commit never rewrites (only branch operations do), so a commit landing after the measured line left the row GREEN while the round had moved on (D2, measured: `.git/HEAD`'s mtime unchanged across a real commit, `--only GT-03` exit 0). Two limits remain, recorded rather than hidden: a repo with the reflog disabled (`core.logAllRefUpdates=false`) has no reference to compare against, and `-nt` against a missing path is true, so the clause passes vacuously and the row then proves only that a measured line EXISTS (a repo with no commit yet is the same case); and the clause reads ANY HEAD movement as staleness, so a checkout, a branch rename or a reset FAILs it until the next gate run rewrites the line. That second behaviour is what makes a full run self-freshening: `GT-02` writes the line earlier in the same pass, so the row asserts that the line in front of you came from THIS run. `tests/t-gt03-freshness.sh` is the control, in both directions.
|
|
17
|
+
GT-04 target A ratchet is declared with a ceiling. gate goblin.yaml ratchet: goblin-verify --only GT-04 —
|
|
18
|
+
GT-05 target The ratchet has not risen. gate ratchet.cmd goblin-verify --only GT-05 —
|
|
19
|
+
HS-01 target Asserting harnesses follow the house shape. lint harness_dir/* goblin-verify --only HS-01 A declared harness_dir that is absent FAILS when the class scaffolds one (config `scaffold_checks: yes`, classes A and C); a class that ships no harness dir (B/D/E) SKIPs with a reason. Keying the skip off the path alone let one config line switch this row and HS-02 off. A report utility in the same dir is counted and reported separately rather than failing the run.
|
|
20
|
+
HS-02 target A check green on both trees proves nothing - the REPLAY must show RED pre-change. gate harness_dir/*, replay.env, replay.cmd goblin-verify --only HS-02 The declared `replay.cmd` is EXECUTED with `{name}` replaced by each harness's name, in the pre-change worktree, with `replay.env=<commit>` set (Z1-4). Two clauses that used to be unheard: a command that interpolates no `{name}` FAILs (it cannot be running the harness it names, so nothing was replayed), and a command that cannot be executed at all - exit 126 or 127 - FAILs rather than counting as a RED harness. The harness's name is substituted SHELL-QUOTED (`printf %q`), because the name is part of the command text the shell parses; unquoted, a name carrying `;` or `#` reached the shell as syntax (Z2-3, the G8-1 surface class), and the control for it carries a metacharacter-bearing file name. What it cannot see: a command that runs *something else* under the harness's name and exits non-zero, and whether the harness tests the right path rather than merely failing on this tree.
|
|
21
|
+
HS-03 target Source probes read text with comments blanked first. advisory harness_dir/* advisory Recognising 'this probe reads source text' is semantic; a grep for the blanking helper produces false FAILs on harnesses that do not probe source.
|
|
22
|
+
CM-01 target Commits carry the owner identity, not an ambient one. Current scope: this gates the identity of HEAD (the last commit) at the moment of the run - earlier commits by other authors are not scanned. gate git log test "$(git log -1 --format='%ae')" = "$(grep '^owner_email:' .goblin/goblin.yaml | cut -d' ' -f2)" —
|
|
23
|
+
CM-02 target The commit message was written to a file, not passed with -m. advisory git log advisory A backtick lost to command substitution leaves no trace a later check can read. Reported as a heuristic (unbalanced backticks in a body) only.
|
|
24
|
+
CM-03 target Commit-as-you-go: the working tree is not carrying a dead run's work. script git status --porcelain goblin-verify --only CM-03 —
|
|
25
|
+
MD-01 target No model name is hardcoded in any reusable rule. lint skills/ manifest/ bin/ templates/ presets/ .goblin/ .hermes/ for d in skills manifest bin templates presets .goblin .hermes; do [ -d "$d" ] || continue; grep -rniE '(d[e]epseek|cl[a]ude|g[p]t-[0-9]|gr[o]k|g[e]mini|g[l]m-[0-9]|k[i]mi)[a-z0-9.:_-]*' "$d" && exit 1; done; exit 0 — (the pattern is written with character classes so this row cannot match itself; tests/t-verify-red.sh proves it still catches a real model name. SCOPE, stated so a repo-wide grep is not re-reported as a hole (G8-10): this row reads the seven directories an install writes rules into - skills/ manifest/ bin/ templates/ presets/ .goblin/ .hermes/. It does NOT read tests/, docs/, or the source automations/ directory; tests/ is where a control must be free to WRITE the thing it guards, and the source directories are covered by MD-01's own body in tests/run-tests.sh, which scans skills manifest bin templates presets automations. Both control strings are assembled at run time, so a repo-wide grep over the whole tree is 0 and the row's own scope is 0 by construction.)
|
|
26
|
+
MD-02 target The review lane is a different model family from the code lane. advisory roles.yaml + mapping file goblin-verify --only MD-02 goblin-stack cannot choose the fleet's models; today's map resolves both roles to the same family. Reported as ADV, never gated. W5-6: the JUDGE lane's resolved model is compared with the code lane's too, because `JG-02` proves only that the declared profile NAMES are disjoint; it is still ADV, never gated, and the measured state of this box is recorded in `docs/LIMITS.md` #38.
|
|
27
|
+
MD-03 target Role-pinned fan-out goes through kanban, not a model-less subagent spawn. advisory skills/* advisory It is a fleet-runtime property: no repo-local file can observe which tool created a worker. Enforced at board level, described in docs/INTEGRATION.md.
|
|
28
|
+
PG-01 target The reviewed artifact is named by SHA, and that SHA exists. gate reviews/*.md goblin-verify --only PG-01 —
|
|
29
|
+
PG-02 target The gate is chosen by the change, not the repo, and the tier's evidence exists. script reviews/*.md goblin-verify --only PG-02 —
|
|
30
|
+
PG-03 target A new head voids the verdict. gate reviews/*.md + git goblin-verify --only PG-03 —
|
|
31
|
+
PG-04 target Never bypass what the forge enforces. advisory — advisory Needs the forge: GitHub's restrictions do not apply to admins, and a sole-admin repo has nobody the gate binds. Not observable from the repo.
|
|
32
|
+
PG-05 target No required check that self-skips. lint .github/workflows/* goblin-verify --only PG-05 Text, not a YAML parser. Three clauses per job: the job declares at least one `run:`/`uses:` step (else the check runs nothing); the JOB carries no `if:` (else the whole required check self-skips); and no STEP carries an `if:` (else that step - possibly the gate step - self-skips). The old body flagged only 'every step guarded', which PASSED the real shape: in the estate's one existing workflow a deliberately UNGUARDED credential step decides whether the guarded compile step runs, so `guarded < steps` and the row reported PASS on the workflow it exists to catch (G8 section 3, re-measured V3-4). Deliberately strict, and the strictness is the point: GitHub reports a SKIPPED job as Success even when it is a required check (docs S1/S2), so a conditional step is a step that can green-light a commit whose gate never ran. False positives it cannot avoid: a `#` inside a quoted string is read as a comment, a flow-style (`jobs: {...}`) mapping is refused rather than parsed, and a conditional step that is genuinely safe is indistinguishable from the trap - the remedy is to move the condition into the declared command, or into a second job that is not the required check. It cannot see branch-protection state, the required-check list, or whether the workflow ever ran (`PG-04`), and a repo with no workflow passes with the count printed on the line.
|
|
33
|
+
PG-06 target The gate CI runs is the gate the project declares. gate .github/workflows/* + goblin.yaml gates: goblin-verify --only PG-06 The declared set comes from `g_yaml_gates`, the reader `GT-01` uses, so a gate cannot vanish from the comparison in silence (G8-3). A workflow passes when the whole declared set is RUN: the verifier with no `--only` (a full run executes every declared gate through `GT-02`), or each declared gate command verbatim. It reads TEXT with comments blanked first, and only `run:` payloads and block-scalar bodies count - a command sitting in a `name:`, `env:` or `with:` value is dropped, which is what stops a comment or a label from satisfying the row. What it cannot see: that the forge marks that job a REQUIRED check, that the job is the one the forge waits on, or that the workflow can fail at all (`PG-05`). A repo with no workflow reports SKIP with that reason instead of a vacuous pass. The lane's own enforcement is the TARGET repo's, not goblin-stack's (`docs/LIMITS.md` #34): goblin-stack can prove the file invokes the gate and cannot make the forge run it, so this row is a text reading of the target's CI and never a claim that CI gated the SHA.
|
|
34
|
+
DS-01 target Runtime data is not test fixture: a gate run must not write it. script goblin.yaml runtime_data: goblin-verify --only DS-01 —
|
|
35
|
+
DS-02 target Snapshot before, verify after. gate same as DS-01 goblin-verify --only DS-02 —
|
|
36
|
+
DOC-01 target A significant change updates the docs that teach it. advisory README.md, docs/ advisory 'Significant' is a judgement; a diff-size heuristic fails on the cases that matter.
|
|
37
|
+
DOC-02 target System-level changes are recorded wherever the project's standard says they live. advisory — advisory Where the recording lives may be owned by a stricter external rule than goblin-stack may add; the repo can only state it.
|
|
38
|
+
SK-01 target Every shipped skill has name + description frontmatter. lint .hermes/skills/*/SKILL.md goblin-verify --only SK-01 W1: the check is a builtin so the engine.mode=global clause can run - in global mode the procedure tier is emitted per platform (not carried in this repo) and the row SKIPs with that reason instead of passing vacuously on a repo with no skills (§2.5 names SK-01 alongside SK-02/SK-04). In vendored mode it is exactly the old loop: every SKILL.md must open with frontmatter carrying both name and description.
|
|
39
|
+
SK-02 target The installed skills match their recorded hashes (no drift). script .hermes/skills/, installed.json goblin-verify --only SK-02 —
|
|
40
|
+
SK-03 target A rule with no mechanism is labelled advisory, and the advisory count is reported. script manifest/enforcement.tsv goblin-verify --only SK-03 —
|
|
41
|
+
SK-04 target Every shipped skill says what it cannot see. lint .hermes/skills/*/SKILL.md goblin-verify --only SK-04 Partial: proves the section exists, not that what it says is complete or true - the limit every prose rule carries. Every shipped skill already carries it, so the row is GREEN on a fresh install and RED only under a real violation. W1: the check is a builtin so the engine.mode=global clause can run - in global mode the procedure tier is emitted per platform (not carried in this repo) and the row SKIPs with that reason.
|
|
42
|
+
PT-01 target No tenant-specific string inside a reusable rule. lint skills/ manifest/ bin/ templates/ presets/ .goblin/ .hermes/ for d in skills manifest bin templates presets .goblin .hermes; do [ -d "$d" ] || continue; grep -rniE --exclude=goblin.yaml --exclude=installed.json '(h[a]rvey|tech-g[o]blin|/h[o]me/[a-z]+|g[o]blin-ui|op[e]n-door|sup[r]eme|bb[t]ech|c[l]v)' "$d" && exit 1; done; exit 0 — (the same directory list MD-01 uses: the rules an install actually writes live in `.goblin/` and `.hermes/`, not in the source layout. Two documented exceptions: `.goblin/goblin.yaml`, which holds `models_file:`/`practice:` - per-machine config, not a rule - and `.goblin/installed.json`, which since W1 records the machine's absolute `engine_dir` in its `engine:` block - both per-machine facts, not rules - each excluded by name. The pattern is written with character classes so this row cannot match itself.)
|
|
43
|
+
PT-02 target The default branch is declared, not assumed. gate goblin.yaml branch: goblin-verify --only PT-02 —
|
|
44
|
+
CL-01 target Every part the class requires is present, and every part it forbids is absent. script goblin.yaml class: + manifest/classes.tsv goblin-verify --only CL-01 —
|
|
45
|
+
CL-02 target An archive: true project verifies GREEN without a HANDOFF or gates. script goblin.yaml archive: goblin-verify --only CL-02 Falsifiable: FAILs when `archive:` is not `true`/`false`, and when the config's value disagrees with the one the install recorded in `.goblin/installed.json` (so the waiver cannot be flipped on by hand). It cannot observe the *effect* of the waiver on the other rows without re-entering the runner.
|
|
46
|
+
SC-01 target No secret file is tracked. lint repo n=$(git ls-files | grep -iE '(^|/)\.env|\.pem$|\.key$' | grep -vcE '\.(example|sample|template)$'); printf '%s tracked secret file(s)\n' "$n"; [ "$n" = 0 ] Partial: it sees tracked PATHS, never contents - a secret pasted into a tracked file is invisible here, and the pattern is a name family, so a credential inside `config.ts` is missed by construction.
|
|
47
|
+
SC-02 target The ignore rules cover the whole secret family. gate .gitignore + git check-ignore goblin-verify --only SC-02 Builtin, and behavioural: clause 1 reads `.gitignore`; clause 2 asks git's own matcher (`git check-ignore`) for `.env`, `.env.local` and `.env.production` one path at a time, so a rule that looks right but does not match still fails. It cannot see a secret already in git history, or one committed under a name the family does not cover. SKIPs with a reason when there is no `.gitignore` and no `package.json`.
|
|
48
|
+
SC-03 target No client-visible name is secret-shaped, and no build output carries a secret literal. lint repo + security's build_output goblin-verify --only SC-03 Builtin, two clauses: the `NEXT_PUBLIC_*_(SECRET|TOKEN|KEY|PASSWORD|PRIVATE)` name pattern over source, and known secret prefixes over the declared build output. Partial: a prefix pattern cannot see a secret that does not look like one, and clause 2 says "no build output to scan" on the line rather than skipping silently. The source scan excludes `.goblin/` and `.hermes/`, so it cannot match the row text that describes it.
|
|
49
|
+
SC-04 target Every cookie write carries its flags. lint repo goblin-verify --only SC-04 Builtin, same-statement only: a write spread over three lines is not seen, and the row says so. It REPORTS that a JS-written cookie is readable by any script rather than failing on it - that is a design fact, not a bug - so the pass is about the flags, never about the choice.
|
|
50
|
+
SC-05 target Every write route validates its input, or is waived. gate security's write_routes goblin-verify --only SC-05 Builtin: it proves a validator is CALLED (`safeParse|zod|valibot|yup|ajv|superstruct|validate(`), never that the schema is right - a schema that accepts everything passes. `.goblin/boundary-waivers` is the escape hatch, and the waived count is printed, so a silent pile-up is visible.
|
|
51
|
+
SC-06 target A lockfile exists, and the repo tracks it. gate package.json + lockfile goblin-verify --only SC-06 Builtin: presence, then `git ls-files --error-unmatch`. It cannot see that the lockfile is STALE relative to `package.json` - resolving that needs the package manager, which is a deliberate network-shaped step, not a check. SKIPs with a reason when there is no `package.json`.
|
|
52
|
+
SC-07 target The dependency audit record is present, dated, fresh, and clean-or-waived. gate .goblin/audit.tsv, .goblin/audit-waiver.tsv goblin-verify --only SC-07 Builtin, OFFLINE by construction: it reads the record and never the network. The record comes from `.goblin/bin/goblin-audit`, run once, deliberately; the waiver count is printed on the gate line so the debt is loud even when the row passes. It cannot see an advisory the registry did not know on the day the record was taken. SKIPs with a reason when no record exists yet.
|
|
53
|
+
SC-08 target No dependency runs an install-time script that is not on the allowlist. gate package-lock.json + .goblin/install-hooks.allowlist goblin-verify --only SC-08 Builtin over `package-lock.json`'s `hasInstallScript`: a pnpm/yarn lockfile has no such field, so those repos get a SKIP with that reason rather than a vacuous pass, and a MINIFIED (one-line) lockfile is read like a pretty-printed one - the text is normalised into the pretty shape before the line-anchored reader sees it, so a hook inside a one-line lock FAILs instead of being reported as `0 install hook(s)` (Z2-2: before that, a one-line lock PASSED vacuously). A `lockfileVersion` 1 lock records its hooks under `dependencies`, not under `node_modules/` keys, so the line-anchored reader enumerates nothing there either and prints the same `0 install hook(s), 0 allowlisted` PASS over a surface it did not read (measured at 0.4.2; npm <= 6 records no `hasInstallScript` at all, so for a genuine v1 lock this is blindness rather than a wrong answer). Lowest-value row of the ten - regression detection - and the first to cut if the matrix gets heavy.
|
|
54
|
+
SC-09 target Auth is applied consistently across sibling routes. advisory — advisory Prose on purpose: "consistently" is a semantic judgement about a private surface no repo here has yet. Counted (9 of ceiling 10) so the matrix cannot quietly grow prose.
|
|
55
|
+
PF-01 target The perf baseline names the commit it measured. lint .goblin/goblin.yaml perf: + ratchet: goblin-verify --only PF-01 Builtin: the metric must equal `ratchet.name` so the budget and the measurement cannot silently disagree, the value must be numeric, the date must exist, the baseline commit must exist AND be an ancestor of HEAD (`HP-05`'s mechanic, reused rather than re-derived), and `ratchet.ceiling` must equal `perf.baseline_value` - otherwise a one-line ceiling raise passes while the row prints the contradiction, which is `I raised the budget and never measured again` (G8-6b). It cannot see whether the metric is the right one for the product, and it never re-measures: re-anchoring is a deliberate operator action. SKIPs with a reason when the class declares no metric, or none has been recorded yet.
|
|
56
|
+
AU-01 target An automation's producer is deterministic and network-free. lint automations/*.sh goblin-verify --only AU-01 Partial: proves that no line of a producer begins with a network or forge verb, not that the script is otherwise deterministic. W1: in global mode the first producer glob resolves under the engine dir (the resolution chain); the repo-local globs stay. New clause: a repo with NO producer anywhere (repo or engine) SKIPs with `no automation producer found (repo or engine)` instead of FAILing — the old FAIL was a born-RED artifact of the glob; a repo that declares its own automations but has no producer still FAILs.
|
|
57
|
+
AU-02 target A report's dedup key is a function of content only - no date, no run id. script reports/*/report.yaml goblin-verify --only AU-02 Builtin: it recomputes the key from the report's own `repo` and `symptom` and requires the recorded `dedup_key` to equal it, then refuses a key carrying a date. Skipped with a reason when the repo holds no report - nothing to dedup. It cannot see whether two reports should have been one: a normalisation that merges two genuinely different symptoms is a duplicate card, not a lost report.
|
|
58
|
+
AU-03 target A reporter run leaves the tree and the harness untouched. gate git status, harness_dir goblin-verify --only AU-03 Builtin: asserts a clean working tree, and - when HEAD is a reporter commit - that no path under the declared harness_dir appears in it. Skipped with a reason when the repo holds no reports/ - no reporter has run here. It cannot see a reporter that edited the tree and committed the edit as part of the report.
|
|
59
|
+
AU-04 target An automation's skill declares its own write surface. lint .hermes/skills/goblin-bugreporter, .hermes/skills/goblin-drift-audit goblin-verify --only AU-04 Partial: proves every installed automation skill carries a `## Write surface` section. W1: the check is a builtin so the engine.mode=global clause can run - in global mode the procedure tier is emitted per platform (not carried in this repo) and the row SKIPs with that reason.
|
|
60
|
+
BN-00 target Every ban has an enforcement row, every ban row names a replacement, and the ban table is not empty. script .goblin/manifest/bans.tsv + .goblin/manifest/enforcement.tsv goblin-verify --only BN-00 — (this row is the reason the ban list cannot decay into prose: IN-03's shape applied to bans.tsv, and it agrees in both directions)
|
|
61
|
+
BN-01 target No `any` in application TypeScript. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-01 Text probe, not an AST: a `: any` inside a string or a comment is reported, and `Record<string, any>` (no leading colon) is missed. The AST form needs a parser the no-npm contract (docs/CONTRACTS.md) forbids (docs/LIMITS.md #27). SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
62
|
+
BN-02 target No `@ts-ignore` / `@ts-expect-error` suppressions. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-02 Text probe: it sees the directive wherever it appears, including inside a string, and cannot tell a suppression hiding a real error from one on a line that would compile anyway. SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
63
|
+
BN-03 target No direct network call from a component. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-03 Text probe over the declared component globs: it stops the call and cannot tell whether a data layer was written or the call merely moved into a helper. SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
64
|
+
BN-05 target No import across a declared layer boundary. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-05 Reads the `layers:` list; an empty list SKIPs with a reason, never a vacuous pass. It matches an import path naming the target directory's last segment - module aliases and dynamic imports are not seen. SKIPs when no file matches its globs.
|
|
65
|
+
BN-06 target No renderer with Node access (`nodeIntegration: true`). lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-06 Text probe over the ban table's globs, the same mechanism as BN-01..BN-05: it sees `nodeIntegration: true` wherever it appears, including inside a string, and cannot see a webPreferences object built at run time or spread in from another module. The STRONGER form is a runtime measurement - the renderer prints `process.contextIsolated` and `process.sandboxed` and the check requires true/true - and that needs a real Electron process, which the dependency contract (docs/CONTRACTS.md) does not allow a shipped rule to launch: it is the project's host gate (docs/LIMITS.md #34). SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
66
|
+
BN-07 target No renderer with context isolation or the process sandbox turned off. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-07 One probe for two properties because Electron's own documentation makes them one: disabling `contextIsolation` "also disables process sandboxing", so a repo that has turned either off has lost both. Text probe, with the same false-positive set as BN-06. SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
67
|
+
BN-08 target No dangerous webPreferences. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-08 Four one-line patterns from Electron's own security checklist (`webSecurity: false`, `allowRunningInsecureContent: true`, `enableBlinkFeatures`, `<webview allowpopups>`). Text probe: `enableBlinkFeatures` is banned by name rather than by value, so the string is reported even in a comment. SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
68
|
+
BN-09 target No synchronous IPC and no `@electron/remote`. lint .goblin/manifest/bans.tsv + .goblin/bans/ goblin-verify --only BN-09 The banned-list shape the wave's note 9 asks for, applied to Electron: `sendSync(` and `@electron/remote` block the renderer's own thread, which is the freeze the class exists to prevent. Text probe - it sees the call site, not the call graph, so a wrapper around `sendSync` in a file the globs do not match is missed. SKIPs when the ban is not in `bans:` or its globs match no file.
|
|
69
|
+
FM-01 target Every feature file is indexed from the map README, declares its slug and at least one entry path, and carries the four-H2 entry contract. lint <feature_map dir> goblin-verify --only FM-01 SKIPs (exit 3) when feature_map: is empty - a fresh install has no map and must not be born RED (the D8 shape). When a map IS declared: the README must exist, every features/*.md must be linked from it in the (./<slug>.md) form and every relative .md link must resolve, each feature file's `feature:` must equal its filename stem, it must declare >=1 `entry_paths:`, and its H2s must be exactly Sub-features / How to get to it (user POV) / Driving it with <harness> / Gotchas, in that order. Partial: the README's own H2s are prose this row does not read, and 'the map lists every user-facing feature' is not mechanically checkable - that is docs/LIMITS.md #30, not a row.
|
|
70
|
+
FM-02 target Every entry point a feature declares still resolves in source, and no entry path changed after the map was verified. lint <feature_map dir> + git goblin-verify --only FM-02 SKIPs (exit 3) when feature_map: is empty. A tripwire, not a proof. The token is searched under source_root with occurrences under the map's own directory excluded - without that exclusion the map's own entry-path list satisfies the search and the row could never go RED. Freshness is `git log -1 --format=%cs` on the resolved file against the feature's `verified:` date, and git sees a FILE change, not a behaviour change: the row can be RED-when-stale and never GREEN-means-fresh. A token that also occurs in a vendored copy or a build artifact is read as resolved, and a `verified:` date is itself a claim the row cannot test (docs/LIMITS.md #30). W5-4: the search skips the harness's own directories (`.goblin/`, `.hermes/`, the declared `harness_dir`) and the map's own directory, so a stub map whose token occurs only in the install no longer resolves; a token that occurs only in the target's own `docs/`, `tests/` or build output still does (docs/LIMITS.md #37). Z1-6: the resolved file must be TRACKED (`git ls-files --error-unmatch`) before its date is compared - an untracked file used to make the freshness clause skip in silence, so a map could claim `verified: 2020-01-01` over source that was never committed.
|
|
71
|
+
VA-01 target The generated verification skill's doctor command runs and exits 0. gate .goblin/goblin.yaml verify_doctor: goblin-verify --only VA-01 SKIPs (exit 3) when verify_doctor: is empty (the replay.commit: "" shape). Runs the DECLARED command exactly as GT-02 runs a declared gate, and never a string read out of file content (the v0.2 blocker). Closes P6's stated-but-unenforced clause 'a generated skill that was never executed is a draft': the doctor is the smallest executable proof that the skill's own instructions still run - and it proves only that, never that the doctor tests the right path.
|
|
72
|
+
JG-01 target A judge verdict is recorded as a row whose evidence resolves. script .goblin/loop/decisions.tsv goblin-verify --only JG-01 Builtin: it proves the named handle EXISTS - a commit `git rev-list --all` knows, a path under the root, or a `sha256:` matching a file under `.goblin/loop/` - never that the handle SUPPORTS the verdict beside it: a judge may cite a real commit that has nothing to do with the claim. It also cannot see whether the verdict was reached by the declared judge lane. A row with fewer columns than the header is a FAIL, not a silent skip. SKIPs (exit 3) when there is no loop record: nothing to check is reported as a reason, never as a pass.
|
|
73
|
+
JG-02 target The judge lane is disjoint from the author lane. gate roles.yaml + the mapping file goblin-verify --only JG-02 Builtin: it proves the DECLARED lane sets are disjoint (profile names), not that a judge ran on one - no repo-local file observes which profile ran (the blindness `MD-03` records), and family equality is not this row's business (`MD-02` owns it). An unresolved lane is an ADV carrying the one-line remedy, never a FAIL: a box with no mapping file must not be failed for the fleet's routing. Both sentinels count as unresolved - `?` (the installed path) and `unknown` (the checkout path) - because a check that greps for one passes on the other.
|
|
74
|
+
JG-03 target A judge lane that has never returned a non-`done` verdict is escalated. advisory .goblin/loop/decisions.tsv advisory A lane that has returned two verdicts cannot be called always-yes, and a repo sees only its own rows; the history that would show a bad lane lives across cards and repos. The counter-measure is policy - one known-red control verdict per wave, recorded in docs/LOOP.md - and it is counted, not enforced.
|
|
75
|
+
LP-01 target The exit predicate is a command, and it ran before iteration 1. script .goblin/loop/predicate, .goblin/loop/first-run goblin-verify --only LP-01 Builtin: it proves the predicate is exactly ONE command and that a recorded first run exists whose timestamp is at or before the first log row - not that the command was actually run, and not that the `exit=` value was measured rather than typed (`HP-03`'s defect, one artifact over). It deliberately does NOT re-run the predicate: a stop condition is legitimately red before the loop finishes, so re-running it would fail a repo whose loop is working, and PROJECT-PRACTICE section 7 forbids a probe that writes. Timestamps compare on their `YYYY-MM-DDThh:mm:ss` prefix, so records written in different UTC offsets compare as text.
|
|
76
|
+
LP-02 target The predicate is pinned at loop start and is never relaxed. script .goblin/loop/predicate.sha256 goblin-verify --only LP-02 Builtin: it proves the predicate file on disk still hashes to the recorded pin - never that the predicate is the RIGHT one, and never who edited it. The digest covers the predicate file alone: a loop that quietly relaxes a predicate living somewhere else is not seen. The remedy is printed with the failure and is never automatic, the `IN-02` `--re-pin` shape. W5-7: a close-and-reopen is auditable - every `closed-<date>/` archive must hold its predicate AND the pin it was closed under, and the live pin must name the archived digest on a `previous:` line - which catches a SILENT relaxation, never a WEAKER one (nothing in bash judges that: `docs/LIMITS.md` #39).
|
|
77
|
+
LP-03 target The loop declares a budget, and the record never exceeds it. gate .goblin/loop/budget, .goblin/loop/decisions.tsv goblin-verify --only LP-03 Builtin: it proves the declared budget is a positive integer at or under the configured `loop_max_turns_ceiling`, and that the record holds no more verdict rows than the budget - not that the budget is affordable. Cost is not a field the record carries: a turn budget bounds turns, and the auxiliary judge call per turn is unpriced (`docs/LIMITS.md` #32). A worker cannot edit its own card body, so a CARD predicate is already protected; this row is the file-predicate half.
|
|
78
|
+
LP-04 target A loop making no progress stops instead of thrashing. gate .goblin/loop/decisions.tsv goblin-verify --only LP-04 Builtin: it measures a CHANGED EVIDENCE POINTER, which is a proxy for progress, not progress. Three consecutive verdict rows with an identical non-empty pointer and a non-`predicate:green` result is a FAIL that names the row numbers. A loop that edits a file each turn to keep the pointer moving is not caught, which is why LP-05 and the budget sit beside it. Nothing in the Hermes kanban goal loop detects a lack of progress at all - `run_kanban_goal_loop` carries no progress state.
|
|
79
|
+
LP-05 target A loop that ended without its predicate green carries a written-up reason. script .goblin/loop/stuck.md, .goblin/loop/decisions.tsv goblin-verify --only LP-05 Builtin: the three-non-blank-line threshold is arbitrary and stated as such - it proves a write-up EXISTS, the way `SP-03` proves a bullet LOOKS checkable. Nothing can make the write-up true, and the row cannot see whether it names the right dead end or whether the loop stopped too early. A loop whose last row IS `predicate:green` needs no write-up and the row passes without reading one.
|
|
80
|
+
RC-01 target No file in the shipping tree matches a reference-corpus hash. gate reference-manifest.json + security: build_output: goblin-verify --only RC-01 Four clauses: (1) `reference_manifest:` empty -> SKIP with that reason (the `FM-01`/`VA-01` shape - not born RED); (2) declared but missing or unparseable -> FAIL, never a SKIP; (3) any file under the declared `security: build_output:` whose sha256 appears in an entry's sha256 -> FAIL, naming path + entry; (4) the manifest itself must not sit inside the build output, at the declared path or as a byte-identical copy - it is a listing of every corpus hash, so shipping it ships the corpus's shape. Cannot see: a re-encoded/resized/recoloured asset (level 2), copied text in a shipped string (level 3), or a manifest authored weak - that is `RC-02`.
|
|
81
|
+
RC-02 target The reference manifest is shaped so `RC-01` cannot pass vacuously. lint reference-manifest.json goblin-verify --only RC-02 Schema `reference-manifest/1`; `generated_from` non-empty; `reference_app.package`/`.version`/`apk_sha256` (64 hex); `entries` non-empty; every entry carries `path`, 64-hex `sha256`, numeric `bytes`; `entry_count` equals `len(entries)`. SKIPs with `RC-01`'s reason when the key is empty. Exists because `RC-01` alone carries the `bans.tsv` weakness: a weakened input passes the check it feeds (`LIMITS.md` #28). It cannot see whether the entries are the *right* hashes - that is as strong as tamper hashing, and `installed.json` is unsigned (`LIMITS.md` #18).
|
|
82
|
+
RC-03 target A lab repo tracks no extracted byte: scripts, notes, manifests, docs only. script manifests/*.sha256 + git ls-files goblin-verify --only RC-03 Two clauses: every tracked path falls under a declared allowlist (`scripts/`, `notes/`, `manifests/`, root docs - the harness's own `.goblin/` and `.hermes/` files and the `.gitignore` block it appends are the install, not the lab's content, and are excluded the way `FM-02` excludes them); and no tracked file's sha256 equals any hash in a `manifests/*.sha256`. The manifest itself lists paths and hashes, which the lab repo's own README explicitly permits - the check is about *bytes*, and it must not read the manifest as a violation. SKIPs with a reason when no `manifests/` exists. Cannot see a payload renamed and re-encoded - which is why the allowlist is a second, independent trap.
|
|
83
|
+
RC-04 target The acquisition record exists, and the manifest names the target and its source. lint manifests/*.sha256 header goblin-verify --only RC-04 The header block must name a target slug + version, a source store, and a checksum (md5 or sha256 hex); the manifest must hold a `*.apk` row, the header's named apk must be that row, and a 64-hex token in the header, if present, must equal the row's sha256 (a header carrying only the md5 passes - the comparison is conditional in the engine). Weakest of the four: it proves a record EXISTS, not that the number came from the store (`HP-03`'s defect; the `JG-01` shape). It is the first to cut if the matrix gets heavy.
|
|
84
|
+
PR-01 source The installer never writes outside its target. test tests/t-install-* tests/run-tests.sh —
|
|
85
|
+
PR-02 source A second install is a no-op, and an upgrade reports created/updated/unchanged. test tests/t-install-idempotent.sh tests/run-tests.sh —
|
|
86
|
+
PR-03 source Every target-scope check goes RED under its own violation. test tests/t-verify-red.sh tests/run-tests.sh — (the negative control the verifier re-runs)
|
|
87
|
+
PR-04 source The repo is portable: no personal path in any reusable rule. lint repo tests/run-tests.sh (the PT-01 body over the source tree, plus tests/) —
|
|
88
|
+
PR-05 source The automation producer is silent when there is nothing to report. test tests/t-automation-silent.sh tests/run-tests.sh — (the mutation is the control: the same producer, on the same fixture, with one installed file edited, must go from an empty stdout to a record and exit 1. A producer that stays quiet after the mutation is not silent, it is broken.)
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
term definition source
|
|
2
|
+
playbook A named, ordered procedure with a measurable verification step. goblin-stack ships 15. R6 sec.3
|
|
3
|
+
skill A Hermes SKILL.md directory. goblin-stack installs its skills into the target repo at .hermes/skills/ (project tier), never into a profile. R6 sec.0.1 / sec.7.3
|
|
4
|
+
profile A worker identity in the Hermes fleet (architect, coder, reviewer, ...). Owns memory and skills. R4
|
|
5
|
+
role What the work needs: code | judgment | review-panel | synthesis | investigate | judge. A role maps to a profile; it never names a model. R6 sec.4.1
|
|
6
|
+
class A project category A-E that selects which parts are required, optional or off. See manifest/classes.tsv. R7 sec.4
|
|
7
|
+
part One installable unit a class requires or forbids: handoff, spec, gate, replay, ratchet, pr-gate, review-panel, playbooks, tokens. R7 sec.5
|
|
8
|
+
gate A declared command that must exit 0. Declared per project, never inferred from the stack. PP sec.4
|
|
9
|
+
ratchet A count that must not rise, with a declared ceiling; a rise is allowed only when re-anchored with the arithmetic (old + N new = new). PP sec.4
|
|
10
|
+
HANDOFF The session-boundary contract at the repo root: START HERE / State / Gates / Next steps / NOT verified. PP sec.1
|
|
11
|
+
SPEC A per-round document written before implementation: measured root cause plus an AC: list checkable without a human. PP sec.2
|
|
12
|
+
REPLAY Re-running each assertion against the pinned pre-change commit and requiring it to be RED there. A check green on both trees proves nothing. PP sec.3
|
|
13
|
+
harness An asserting check file: a failures counter, an assert(name, ok, detail) printer, and a non-zero exit on failure. PP sec.3
|
|
14
|
+
stakes The S0-S4 ladder that chooses the review gate strength by the change, not by the repo. R5 sec.3.2
|
|
15
|
+
patch-id git patch-id --stable of base..head. A new head voids a verdict; a matching commit message does not restore it. R1 sec.10
|
|
16
|
+
advisory A rule with no executable check. Counted in the verify summary and capped by advisory_ceiling; never a silent pass. R6 sec.5
|
|
17
|
+
opt-out A part recorded in disabled: so its required checks report SKIP (opt-out) instead of failing. R6 sec.6.3
|
|
18
|
+
archive A class-D flag: verify requires no HANDOFF and no gates, and the summary says so. R7 sec.2.4
|
|
19
|
+
sweep P11: the same change or question across many projects, enumerated by glob, one card each. R6 sec.3
|
|
20
|
+
overnight P10: an unattended run over a checkable predicate written before iteration 1. R6 sec.3
|
|
21
|
+
review-panel N independent verdict lanes, each its own card with its own resolved model. Only at stakes S3+. R6 sec.4.3
|
|
22
|
+
judge A role that decides whether a PROCESS met its own predicate, from a command's output and a pointer it can resolve - never from a report. One lane per verdict; never the author's lane. docs/LOOP.md
|
|
23
|
+
predicate One shell command that exits 0 when the loop is finished. Written before iteration 1 and run once before it, so its first state is known to be red. A duration is not a predicate. docs/LOOP.md + pstack guide/07
|
|
24
|
+
loop One unattended run over a predicate, recorded in the committed .goblin/loop/ (predicate, pin, first-run, budget, one decisions.tsv row per iteration). docs/LOOP.md
|
|
25
|
+
never-relax The predicate's digest is recorded at loop start and never updates itself; relaxing it is closing this loop and opening another, with the old predicate archived under .goblin/loop/closed-<date>/. docs/LOOP.md
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
id playbook when steps verification profiles role
|
|
2
|
+
P1 goblin-investigation a read-only question, or "why is this happening" 1 name the question as a falsifiable claim | 2 read the code paths, cite file:line | 3 if the answer is observable by running something, run it instead of asking | 4 write the answer with its evidence every claim carries a file:line or a command+output; no code changes; a claim you cannot source is marked [unverified] any investigate
|
|
3
|
+
P2 goblin-bugfix a reported defect 1 reproduce it yourself | 2 state the root cause with a measurement, never 'should be' | 3 fix the root, not the symptom | 4 prove absence on the same surface, with a negative control the repro fails before and passes after, on the same command; unit tests show branch behaviour, not bug absence coder code
|
|
4
|
+
P3 goblin-feature new behaviour 1 SPEC first (measured root cause + AC: list) | 2 name the data shape before the code | 3 land it in units that each end checkable | 4 write the SHA it landed at into the review note every AC: item has a checkable assertion; the gate line is measured; the SHA is named architect -> coder judgment -> code
|
|
5
|
+
P4 goblin-refactor a behaviour-preserving reshape 1 pin the contract first (characterization test / snapshot / equivalence harness) | 2 shape only, no behaviour | 3 delete the legacy path in the same change the pin is a real assertion run before and after; a type check and lint are not a pin coder code
|
|
6
|
+
P5 goblin-tdd-repro a defect where a regression test is cheap 1 write the failing test | 2 confirm it fails for the intended reason | 3 smallest production fix | 4 revert the fix -> the test MUST fail -> restore the RED-again step is captured; prefer no new test over a bad test; the skip path is explicit, never silent coder code
|
|
7
|
+
P6 goblin-verify-author a project has no live lane, or its gates drift 1 read the repo, not the user, for entry points | 2 write the gate/check set into the harness dir with the house harness shape | 3 execute it once end to end | 4 add the REPLAY block | 5 seed the feature map and declare feature_map:/source_root: (goblin-feature-map) | 6 hand the generated skill to P12: verified: does not advance until the eval record exists a generated skill that was never executed is a draft, not a deliverable; the harness prints PASS/FAIL and exits non-zero on failure; the REPLAY shows RED pre-change; the map's index and four-H2 contract hold (FM-01), every declared entry path still resolves (FM-02), and the declared verify_doctor: exits 0 (VA-01) architect judgment
|
|
8
|
+
P7 goblin-pr-gate anything that should be reviewed before it lands 1 classify stakes S0-S4 | 2 run the gate set at the candidate SHA and record the numbers | 3 write reviews/<slug>-<head7>.md with head/base/patch-id/stakes/checks-run | 4 evaluate the panel rule for S3+ | 5 at S3+ the foreman is role-judge: one decision from N lane verdicts, never a vote count | 6 re-check the patch-id before landing the patch-id of base..head still matches the recorded one; the review note names a SHA that exists in git rev-list; for S2+ a check ran on that SHA; at S3+ the judge is a lane disjoint from the author's (JG-02) reviewer (+ architect for S3) review-panel
|
|
9
|
+
P8 goblin-bootstrap adopting goblin-stack in a repo, or starting one 1 classify the project (A-E) | 2 goblin-install --class <x> | 3 goblin-verify GREEN | 4 fix .gitignore BEFORE any git init | 5 first HANDOFF, first SPEC, first check script goblin-verify exits 0 and the created-file list matches installed.json; a repo with no gate declares one and records its first measured numbers architect judgment
|
|
10
|
+
P9 goblin-handoff ending a session, or picking up another's 1 commit uncommitted edits as one internally-consistent wip: commit | 2 write intent / verified state / next steps / what is NOT verified | 3 stale sentences get a dated parenthetical, never deletion | 4 on pickup: verify inherited claims against the artifact every gate number in the HANDOFF carries 'measured <date>'; the named HEAD matches git rev-parse --short HEAD; the NOT-verified section is non-empty or says 'nothing outstanding' any judgment
|
|
11
|
+
P10 goblin-overnight an unattended run over a predicate 1 the exit condition is a checkable predicate written before iteration 1 | 2 it never gets relaxed | 3 an escape hatch: a genuine dead end writes up why and stops | 4 the morning audit reads the Attention section first the predicate is a command in the card body and was run at least once (its first run recorded before iteration 1); goal_max_turns is set and at or under loop_max_turns_ceiling; every landed change has a P7 verdict row; any verdict that says done cites a handle the repo resolves (JG-01); a run that ends without its predicate green carries .goblin/loop/stuck.md (LP-05) default + coder code + judge
|
|
12
|
+
P11 goblin-sweep the same change or question across projects 1 enumerate targets with a shell glob, not a memory | 2 classify each (A-E); an archive project is skipped, not processed | 3 one card per project, parents=[sweep] | 4 collect one line per project: what changed / what was refused / what is unfindable the per-project line carries the command it ran; the sweep report states its own coverage (n of m projects, and names the skipped ones) default synthesis
|
|
13
|
+
P12 goblin-eval a skill or prompt changed, and you want to know if it did anything 1 candidate and control run in sanitized directories | 2 no eval/test/judge/rubric token anywhere the candidate sees | 3 grade the chain from the transcript (which files it actually opened), never self-report | 4 the judge runs on a different model family | 5 write the record to evals/<slug>/ (prompt, rubric, manifest.tsv, transcripts/, verdict.md) | 6 a change must improve the evaluated cases or add new evaluations; a generated verification skill passes only on a measured sensitivity the judge's verdict is reproducible from the transcripts; candidates never learn other candidates exist; the record exists and every lane names a transcript file that exists; a generated verification skill is verified only by a record, with its sensitivity (detected of seeded) printed beside the control's number researcher synthesis
|
|
14
|
+
P13 goblin-bugreporter an event delivered a report 1 validate intake | 2 freeze repo+revision | 3 reproduce (R1 then R2 then R3) | 4 write repro.md | 5 create the fix card only on `reproduced` | 6 complete with the verdict repro.md carries a command, a revision that exists in git rev-list, and a RED `pre` row; the fix card exists iff the verdict is `reproduced`; the tree is clean after the run researcher investigate
|
|
15
|
+
P14 goblin-drift-audit a recorded claim disagrees with the artifact 1 enumerate targets by glob | 2 compute the drift record | 3 print nothing when clean | 4 one card per drifting repo | 5 report n of m the summary names each repo and the command it ran; a clean run prints nothing; skipped targets are named architect judgment
|
|
16
|
+
P15 goblin-re-mobile one shipped Android build must be understood as facts for study, with a reproducible, hash-manifested corpus 1 S0 preflight: the sandbox exists and is the one the fences describe | 2 S1 acquire, S2 verify provenance against the published hash | 3 S3 triage: the engine verdict, cheapest test first | 4 S4 static decompile, S5 carve the containers, S6 manifest the corpus | 5 S7 dossier: facts and numbers, never expression | 6 S8 deferred by design; S9 retention/teardown the corpus manifest verifies `sha256sum -c` where the corpus lives; `RC-01`, `RC-02` and `RC-03` return the exits their rows define (an exact hash inside the build output, a weak manifest and a tracked payload each fail the build); every negative control NC-1..NC-6 was shown RED and then restored coder code
|
package/package.json
CHANGED
|
@@ -1,6 +1,38 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@techgoblin/gobstack",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"
|
|
5
|
-
"
|
|
6
|
-
|
|
3
|
+
"version": "0.4.4-beta.1",
|
|
4
|
+
"description": "Agent-discipline toolkit: one verify command, an enforcement matrix, and LIMITS. bash engine, npm shim.",
|
|
5
|
+
"bin": {
|
|
6
|
+
"goblin": "bin/goblin.js"
|
|
7
|
+
},
|
|
8
|
+
"publishConfig": {
|
|
9
|
+
"access": "public"
|
|
10
|
+
},
|
|
11
|
+
"engines": {
|
|
12
|
+
"node": ">=18"
|
|
13
|
+
},
|
|
14
|
+
"license": "MIT",
|
|
15
|
+
"files": [
|
|
16
|
+
"bin/",
|
|
17
|
+
"adapters/",
|
|
18
|
+
"manifest/",
|
|
19
|
+
"presets/",
|
|
20
|
+
"templates/",
|
|
21
|
+
"skills/",
|
|
22
|
+
"automations/",
|
|
23
|
+
"bans/",
|
|
24
|
+
"docs/",
|
|
25
|
+
"roles.yaml",
|
|
26
|
+
"VERSION",
|
|
27
|
+
"CHANGELOG.md"
|
|
28
|
+
],
|
|
29
|
+
"repository": {
|
|
30
|
+
"type": "git",
|
|
31
|
+
"url": "git+https://github.com/harveynguyen33/goblin-stack.git"
|
|
32
|
+
},
|
|
33
|
+
"keywords": [
|
|
34
|
+
"agent-discipline",
|
|
35
|
+
"enforcement",
|
|
36
|
+
"verification"
|
|
37
|
+
]
|
|
38
|
+
}
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# presets/A-shipped-software.yaml — class A: shipped software.
|
|
2
|
+
#
|
|
3
|
+
# A class is not a stringency level. It selects WHICH PARTS are required, optional or off
|
|
4
|
+
# (manifest/classes.tsv), and it supplies the default gate/ratchet shape below. The gate
|
|
5
|
+
# command is a floor, not the project's gate set: replace it with the real commands
|
|
6
|
+
# (playbook P8 step 3). goblin-stack never infers a gate from the stack it finds.
|
|
7
|
+
label: Shipped software
|
|
8
|
+
done_means: a gate set reports measured numbers, a round lands, the artifact deploys or publishes
|
|
9
|
+
harness_dir: checks
|
|
10
|
+
scaffold_checks: yes
|
|
11
|
+
gate_name: commit
|
|
12
|
+
gate_cmd: git rev-parse --verify --quiet HEAD
|
|
13
|
+
# The TODO count is a GATE here, not the ratchet (G4 D2): the ratchet carries the performance
|
|
14
|
+
# budget, which is the number a shipped-software round is supposed to move. 160 is deliberately
|
|
15
|
+
# generous - a floor against decay, not the project's real budget. Replace it with yours.
|
|
16
|
+
gate2_name: todo_ceiling
|
|
17
|
+
gate2_cmd: test "$(grep -rniE '\b(TODO|FIXME)\b' --include='*.ts' --include='*.tsx' --include='*.js' --include='*.jsx' --include='*.mjs' . 2>/dev/null | wc -l)" -le 160
|
|
18
|
+
# The perf lane's number lives in the ratchet, measured: ratchet.name names the metric,
|
|
19
|
+
# ratchet.cmd is the ONE command that produces it, ratchet.ceiling is the measured baseline.
|
|
20
|
+
# `find ... -exec cat {} + | wc -c` is used rather than `du`, because du reports block sizes and
|
|
21
|
+
# is not deterministic across filesystems (G4 C1).
|
|
22
|
+
ratchet_name: client_js_bytes
|
|
23
|
+
ratchet_cmd: find .next/static dist/assets dist -type f -name '*.js' -exec cat {} + 2>/dev/null | wc -c
|
|
24
|
+
ratchet_ceiling: measure
|
|
25
|
+
replay_env: GOBLIN_PRE_COMMIT
|
|
26
|
+
replay_cmd: node checks/{name}.mjs
|
|
27
|
+
runtime_data: .goblin/state.json
|
|
28
|
+
# SC-02 requires the whole secret family to be ignored; SC-03 clause 2 scans `sec_build_output`
|
|
29
|
+
# for secret-shaped literals; SC-07 runs `sec_audit_cmd` once, deliberately, via goblin-audit.
|
|
30
|
+
sec_gitignore_family: yes
|
|
31
|
+
sec_build_output: .next/static dist/assets dist
|
|
32
|
+
sec_audit_cmd: npm audit --json
|
|
33
|
+
sec_audit_max_age_days: 90
|
|
34
|
+
sec_waiver_max_age_days: 180
|
|
35
|
+
sec_write_routes: app src/app
|
|
36
|
+
perf_metric: client_js_bytes
|
|
37
|
+
perf_cmd: find .next/static dist/assets dist -type f -name '*.js' -exec cat {} + 2>/dev/null | wc -c
|
|
38
|
+
perf_baseline_commit: ""
|
|
39
|
+
perf_baseline_value: 0
|
|
40
|
+
perf_measured: ""
|
|
41
|
+
perf_host_gate: ""
|
|
42
|
+
notes: class A also requires a SPEC before every change, and the REPLAY on every new assertion.
|
|
43
|
+
# The ban list this class turns on (G5). An unlisted ban SKIPs with a reason.
|
|
44
|
+
bans: BN-01, BN-02, BN-05
|
|
45
|
+
|
|
46
|
+
# The loop ceiling (G2, LP-03): the most turns ONE loop record may declare in its budget.
|
|
47
|
+
# The same number for every class - a ceiling, not a per-class policy. 20 is the engine default.
|
|
48
|
+
loop_max_turns_ceiling: 20
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# presets/B-service-config.yaml — class B: service / configuration.
|
|
2
|
+
#
|
|
3
|
+
# "Done" means a contract (schema, route, API) is unchanged, or the change is intentional
|
|
4
|
+
# and migrated. The class turns off the REPLAY and the PR gate, because a config repo has no
|
|
5
|
+
# pre-change tree to pin and no remote to gate against.
|
|
6
|
+
label: Service / configuration
|
|
7
|
+
done_means: a contract (schema, route, API) is unchanged, or the change is intentional and migrated
|
|
8
|
+
harness_dir: checks
|
|
9
|
+
scaffold_checks: no
|
|
10
|
+
gate_name: commit
|
|
11
|
+
gate_cmd: git rev-parse --verify --quiet HEAD
|
|
12
|
+
ratchet_name: todo_markers
|
|
13
|
+
ratchet_cmd: grep -rniE '\b(TODO|FIXME)\b' --include='*.ts' --include='*.tsx' --include='*.js' --include='*.jsx' --include='*.mjs' --include='*.py' . 2>/dev/null | wc -l
|
|
14
|
+
ratchet_ceiling: measure
|
|
15
|
+
replay_env: GOBLIN_PRE_COMMIT
|
|
16
|
+
replay_cmd: node checks/{name}.mjs
|
|
17
|
+
runtime_data: .goblin/state.json
|
|
18
|
+
# A service has real inputs and real dependencies, so the security rows bite here.
|
|
19
|
+
sec_gitignore_family: yes
|
|
20
|
+
sec_build_output: dist
|
|
21
|
+
sec_audit_cmd: npm audit --json
|
|
22
|
+
sec_audit_max_age_days: 90
|
|
23
|
+
sec_waiver_max_age_days: 180
|
|
24
|
+
sec_write_routes: src app
|
|
25
|
+
# No perf metric is declared for this class by default (G4 D1): build bytes are measurable, but
|
|
26
|
+
# latency - the number a service is judged on - is a HOST gate (`curl -w '%{time_total}'`), and a
|
|
27
|
+
# host gate cannot be a hermetic ratchet. Declare metric+cmd here the day you pick a build metric.
|
|
28
|
+
perf_metric: ""
|
|
29
|
+
perf_cmd: ""
|
|
30
|
+
perf_baseline_commit: ""
|
|
31
|
+
perf_baseline_value: 0
|
|
32
|
+
perf_measured: ""
|
|
33
|
+
perf_host_gate: ""
|
|
34
|
+
notes: declare the schema-shape snapshot as the gate; content is optional, schema is required.
|
|
35
|
+
# The ban list this class turns on (G5). An unlisted ban SKIPs with a reason.
|
|
36
|
+
bans: ""
|
|
37
|
+
|
|
38
|
+
# The loop ceiling (G2, LP-03): the most turns ONE loop record may declare in its budget.
|
|
39
|
+
# The same number for every class - a ceiling, not a per-class policy. 20 is the engine default.
|
|
40
|
+
loop_max_turns_ceiling: 20
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
# presets/C-game.yaml — class C: game.
|
|
2
|
+
#
|
|
3
|
+
# "Done" means a suite green in the Editor AND a human feel verdict. The verdict is a
|
|
4
|
+
# first-class deliverable, which is why this class requires the review-panel part: the panel
|
|
5
|
+
# is the review, and it is the one part that cannot be automated away.
|
|
6
|
+
label: Game
|
|
7
|
+
done_means: a suite green in the Editor and a human feel verdict, the verdict being a first-class deliverable
|
|
8
|
+
harness_dir: checks
|
|
9
|
+
scaffold_checks: yes
|
|
10
|
+
gate_name: commit
|
|
11
|
+
gate_cmd: git rev-parse --verify --quiet HEAD
|
|
12
|
+
ratchet_name: todo_markers
|
|
13
|
+
ratchet_cmd: grep -rniE '\b(TODO|FIXME)\b' --include='*.cs' --include='*.ts' --include='*.js' . 2>/dev/null | wc -l
|
|
14
|
+
ratchet_ceiling: measure
|
|
15
|
+
replay_env: GOBLIN_PRE_COMMIT
|
|
16
|
+
replay_cmd: node checks/{name}.mjs
|
|
17
|
+
runtime_data: .goblin/state.json
|
|
18
|
+
sec_gitignore_family: yes
|
|
19
|
+
sec_build_output: Builds Library
|
|
20
|
+
sec_audit_cmd: ""
|
|
21
|
+
sec_audit_max_age_days: 90
|
|
22
|
+
sec_waiver_max_age_days: 180
|
|
23
|
+
sec_write_routes: ""
|
|
24
|
+
# The game's perf number is frame time on the target device - a HOST gate, declared and named,
|
|
25
|
+
# never a hermetic ratchet (G4 C3.4). The engine's Editor run is already a host gate.
|
|
26
|
+
perf_metric: ""
|
|
27
|
+
perf_cmd: ""
|
|
28
|
+
perf_baseline_commit: ""
|
|
29
|
+
perf_baseline_value: 0
|
|
30
|
+
perf_measured: ""
|
|
31
|
+
perf_host_gate: "Editor run - frame time on the target device"
|
|
32
|
+
notes: the Editor suite is a host gate, not a hermetic one - declare it as such and do not expect it to run in CI.
|
|
33
|
+
# The ban list this class turns on (G5). An unlisted ban SKIPs with a reason.
|
|
34
|
+
bans: BN-01, BN-02, BN-05
|
|
35
|
+
|
|
36
|
+
# The loop ceiling (G2, LP-03): the most turns ONE loop record may declare in its budget.
|
|
37
|
+
# The same number for every class - a ceiling, not a per-class policy. 20 is the engine default.
|
|
38
|
+
loop_max_turns_ceiling: 20
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# presets/D-knowledge.yaml — class D: knowledge / research.
|
|
2
|
+
#
|
|
3
|
+
# "Done" means a question is answered with sources and the answer is findable. A research
|
|
4
|
+
# note is not a spec, so the SPEC part is OFF, not optional. So are the REPLAY and the PR
|
|
5
|
+
# gate: there is no pre-change tree to pin and no compiled artifact to gate.
|
|
6
|
+
# Use --archive for an input directory that is not a build target at all.
|
|
7
|
+
label: Knowledge / research
|
|
8
|
+
done_means: a question is answered with sources and the answer is findable
|
|
9
|
+
harness_dir: checks
|
|
10
|
+
scaffold_checks: no
|
|
11
|
+
gate_name: commit
|
|
12
|
+
gate_cmd: git rev-parse --verify --quiet HEAD
|
|
13
|
+
ratchet_name: ""
|
|
14
|
+
ratchet_cmd: ""
|
|
15
|
+
ratchet_ceiling: 0
|
|
16
|
+
replay_env: GOBLIN_PRE_COMMIT
|
|
17
|
+
replay_cmd: node checks/{name}.mjs
|
|
18
|
+
runtime_data: .goblin/state.json
|
|
19
|
+
# A knowledge repo has no client surface and no dependency manifest, so SC-03's build-output
|
|
20
|
+
# clause, SC-05..SC-08 skip with a reason. The SECRET FAMILY is not one of those: a note file is
|
|
21
|
+
# one `git add` away from a pasted token, so SC-02 is required here too and the installer writes
|
|
22
|
+
# the family into `.gitignore` for every class.
|
|
23
|
+
sec_gitignore_family: yes
|
|
24
|
+
sec_build_output: ""
|
|
25
|
+
sec_audit_cmd: ""
|
|
26
|
+
sec_audit_max_age_days: 90
|
|
27
|
+
sec_waiver_max_age_days: 180
|
|
28
|
+
sec_write_routes: ""
|
|
29
|
+
perf_metric: ""
|
|
30
|
+
perf_cmd: ""
|
|
31
|
+
perf_baseline_commit: ""
|
|
32
|
+
perf_baseline_value: 0
|
|
33
|
+
perf_measured: ""
|
|
34
|
+
perf_host_gate: ""
|
|
35
|
+
notes: mark every unsourced claim [unverified]; the freshness/lint gate over notes is optional.
|
|
36
|
+
# The ban list this class turns on (G5). An unlisted ban SKIPs with a reason.
|
|
37
|
+
bans: ""
|
|
38
|
+
|
|
39
|
+
# The loop ceiling (G2, LP-03): the most turns ONE loop record may declare in its budget.
|
|
40
|
+
# The same number for every class - a ceiling, not a per-class policy. 20 is the engine default.
|
|
41
|
+
loop_max_turns_ceiling: 20
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# presets/E-fleet-config.yaml — class E: agent-fleet configuration.
|
|
2
|
+
#
|
|
3
|
+
# "Done" means a config change is applied, verified against the ARTIFACT, and versioned.
|
|
4
|
+
# The gate therefore checks the artifact (a commit exists, dated), never the intention.
|
|
5
|
+
label: Agent-fleet config
|
|
6
|
+
done_means: a config change is applied, verified against the artifact, and versioned
|
|
7
|
+
harness_dir: checks
|
|
8
|
+
scaffold_checks: no
|
|
9
|
+
gate_name: commit
|
|
10
|
+
gate_cmd: git rev-parse --verify --quiet HEAD
|
|
11
|
+
ratchet_name: todo_markers
|
|
12
|
+
# The harness's own vendored directories are excluded: without it, a fresh class-E install
|
|
13
|
+
# counts the ratchet command line itself (`.goblin/goblin.yaml` holds the literal TODO|FIXME)
|
|
14
|
+
# and goblin-stack's own skill prose (`.hermes/skills/*/SKILL.md`), so its gate was born RED
|
|
15
|
+
# against a ceiling measured before those files were written (D8). The fleet's real config —
|
|
16
|
+
# `*.yaml`/`*.yml`/`*.sh` at the root and under `profiles/` — is still counted.
|
|
17
|
+
ratchet_cmd: grep -rniE '\b(TODO|FIXME)\b' --include='*.sh' --include='*.yaml' --include='*.yml' --include='*.md' --exclude-dir=.goblin --exclude-dir=.hermes . 2>/dev/null | wc -l
|
|
18
|
+
ratchet_ceiling: measure
|
|
19
|
+
replay_env: GOBLIN_PRE_COMMIT
|
|
20
|
+
replay_cmd: node checks/{name}.mjs
|
|
21
|
+
runtime_data: .goblin/state.json
|
|
22
|
+
# A fleet config has no client surface; the secret family still matters (a profile file is one
|
|
23
|
+
# `git add` away from a live token), so SC-01/SC-02 are the rows that bite here.
|
|
24
|
+
sec_gitignore_family: yes
|
|
25
|
+
sec_build_output: ""
|
|
26
|
+
sec_audit_cmd: ""
|
|
27
|
+
sec_audit_max_age_days: 90
|
|
28
|
+
sec_waiver_max_age_days: 180
|
|
29
|
+
sec_write_routes: ""
|
|
30
|
+
perf_metric: ""
|
|
31
|
+
perf_cmd: ""
|
|
32
|
+
perf_baseline_commit: ""
|
|
33
|
+
perf_baseline_value: 0
|
|
34
|
+
perf_measured: ""
|
|
35
|
+
perf_host_gate: ""
|
|
36
|
+
notes: the order is dry-run, then one profile, then verify, then apply-all.
|
|
37
|
+
# The ban list this class turns on (G5). An unlisted ban SKIPs with a reason.
|
|
38
|
+
bans: BN-02
|
|
39
|
+
|
|
40
|
+
# The loop ceiling (G2, LP-03): the most turns ONE loop record may declare in its budget.
|
|
41
|
+
# The same number for every class - a ceiling, not a per-class policy. 20 is the engine default.
|
|
42
|
+
loop_max_turns_ceiling: 20
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# presets/F-electron.yaml — class F: desktop shell (Electron).
|
|
2
|
+
#
|
|
3
|
+
# A class is not a stringency level: it selects WHICH PARTS are required, optional or off
|
|
4
|
+
# (manifest/classes.tsv) and supplies the default gate/ratchet shape. A desktop shell needs a
|
|
5
|
+
# part no other class has — a HOST gate, a number measured on a machine with a display — and it
|
|
6
|
+
# forbids a thing the others allow: a renderer that reaches the filesystem or Node directly.
|
|
7
|
+
# That is a new row-set, not a flag on class A (G6 section B.2, Part B).
|
|
8
|
+
#
|
|
9
|
+
# THE PERF LANE IS THE SHIPPED RATCHET, NOT A SECOND ONE (G6 section B.3). There is one
|
|
10
|
+
# mechanism — `ratchet: {name, cmd, ceiling}`, enforced by GT-04/GT-05 and pinned to a commit by
|
|
11
|
+
# PF-01 — and this class reuses it verbatim. What the ratchet measures here is the one hermetic
|
|
12
|
+
# perf number a desktop shell has: the packaged bundle's byte count. A fat bundle is a slow
|
|
13
|
+
# cold start on every machine, and the number needs no browser, no display and no dependency.
|
|
14
|
+
#
|
|
15
|
+
# THE FPS NUMBER IS A HOST GATE, and this is a DEVIATION from G6 section B.3, recorded with its
|
|
16
|
+
# measured reason. G6 wants `main_thread_busy_pct` in the ratchet. It cannot be: the instrument
|
|
17
|
+
# that produces it (CDP `Performance.getMetrics` over a real Chromium, or
|
|
18
|
+
# `app.getAppMetrics()[i].cpu.percentCPUUsage` inside a real Electron) needs Playwright or
|
|
19
|
+
# Electron plus a GUI, and the dependency contract (docs/CONTRACTS.md) allows a shipped rule
|
|
20
|
+
# nothing but bash/git/awk/sed/grep/python3. A `ratchet.cmd` that cannot run on a fresh install
|
|
21
|
+
# makes a fresh install BORN RED, which is the one thing every class must not be. So the probe
|
|
22
|
+
# belongs to the project, next to the code it measures, and the number it produces is declared
|
|
23
|
+
# here as a host gate and carried in the HANDOFF with its date (the class C pattern).
|
|
24
|
+
#
|
|
25
|
+
# Why frame time is the WRONG number, measured in G6 on a real Chromium (CDP, 0 to 32 ms of
|
|
26
|
+
# work per frame): p50 frame time stayed FLAT at 16.70 ms while the main thread went from 1.8%
|
|
27
|
+
# to 54.5% busy, and the dropped-frame count was non-monotone (0,0,0,1,3,11,0 — the worst
|
|
28
|
+
# workload read 0). A frame-time gate at 16.7 ms is green on the idle tree AND on the loaded
|
|
29
|
+
# tree: PROJECT-PRACTICE section 3, reproduced live in the exact metric the note proposed.
|
|
30
|
+
# `main_thread_busy_pct` is the only monotone instrument in that sweep (1.8 -> 13.5 -> 26.0 ->
|
|
31
|
+
# 54.5 -> 76.9 -> 99.5%), which is why it is named here as the host gate's metric.
|
|
32
|
+
# docs/CI.md carries the sweep; docs/LIMITS.md #34 carries the gap.
|
|
33
|
+
label: Desktop shell
|
|
34
|
+
done_means: the renderer is isolated from Node, the main process is not busy, and the packaged bundle ships no dev dependency
|
|
35
|
+
harness_dir: checks
|
|
36
|
+
scaffold_checks: yes
|
|
37
|
+
gate_name: commit
|
|
38
|
+
gate_cmd: git rev-parse --verify --quiet HEAD
|
|
39
|
+
ratchet_name: app_bundle_bytes
|
|
40
|
+
ratchet_cmd: find dist out release -type f -exec cat {} + 2>/dev/null | wc -c
|
|
41
|
+
ratchet_ceiling: measure
|
|
42
|
+
replay_env: GOBLIN_PRE_COMMIT
|
|
43
|
+
replay_cmd: node checks/{name}.mjs
|
|
44
|
+
runtime_data: .goblin/state.json
|
|
45
|
+
sec_gitignore_family: yes
|
|
46
|
+
sec_build_output: dist out release
|
|
47
|
+
sec_audit_cmd: npm audit --json
|
|
48
|
+
sec_audit_max_age_days: 90
|
|
49
|
+
sec_waiver_max_age_days: 180
|
|
50
|
+
sec_write_routes: ""
|
|
51
|
+
# perf_metric must EQUAL ratchet_name (PF-01 fails when the budget and the measurement disagree),
|
|
52
|
+
# so the hermetic metric is named here too, and the host gate carries the FPS number beside it.
|
|
53
|
+
perf_metric: app_bundle_bytes
|
|
54
|
+
perf_cmd: find dist out release -type f -exec cat {} + 2>/dev/null | wc -c
|
|
55
|
+
perf_baseline_commit: ""
|
|
56
|
+
perf_baseline_value: 0
|
|
57
|
+
perf_measured: ""
|
|
58
|
+
perf_host_gate: "Electron run - main_thread_busy_pct with a window open (app.getAppMetrics()[i].cpu.percentCPUUsage, or CDP Performance.getMetrics), on a machine with a display"
|
|
59
|
+
# The nine electron bans (G5's mechanism, G6's failure surface): nodeIntegration, context
|
|
60
|
+
# isolation / sandbox, the dangerous webPreferences, and synchronous IPC / @electron/remote.
|
|
61
|
+
# An unlisted ban SKIPs with a reason, so the other classes are unaffected by this list.
|
|
62
|
+
bans: BN-01, BN-02, BN-05, BN-06, BN-07, BN-08, BN-09
|
|
63
|
+
notes: a perf number measured on one box is not a user's experience - the renderer half can be a CI job, the Electron half is a host gate and cannot be.
|
|
64
|
+
|
|
65
|
+
# The loop ceiling (G2, LP-03): the most turns ONE loop record may declare in its budget.
|
|
66
|
+
# The same number for every class - a ceiling, not a per-class policy. 20 is the engine default.
|
|
67
|
+
loop_max_turns_ceiling: 20
|
package/roles.yaml
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# roles.yaml — the role catalogue. A ROLE is what the work needs; a PROFILE is a worker
|
|
2
|
+
# identity. They are kept separate on purpose: five documents once gave five different model
|
|
3
|
+
# answers because a role name was used as a model name.
|
|
4
|
+
#
|
|
5
|
+
# THIS FILE NAMES NO MODEL. It names a capability and the default profile that carries it.
|
|
6
|
+
# The machine-specific mapping (profile -> provider/model/effort) is read at run time from
|
|
7
|
+
# the mapping file declared as models_file: in .goblin/goblin.yaml — the single documented
|
|
8
|
+
# install-time machine input. bin/goblin-model resolves a role through it.
|
|
9
|
+
#
|
|
10
|
+
# profiles: a list. One lane runs per entry, so the LIST LENGTH is the lane count — that is
|
|
11
|
+
# how the review panel is bounded by configuration rather than by code.
|
|
12
|
+
|
|
13
|
+
role-code:
|
|
14
|
+
capability: fast, cheap, mechanical correctness
|
|
15
|
+
profiles: [coder]
|
|
16
|
+
used_by: P2, P4, P5, P10
|
|
17
|
+
budget: one effort token per role, read from the mapping file
|
|
18
|
+
|
|
19
|
+
role-judgment:
|
|
20
|
+
capability: strongest available reasoning and prose
|
|
21
|
+
profiles: [architect]
|
|
22
|
+
used_by: P1, P3 (spec half), P6, P8, P9
|
|
23
|
+
|
|
24
|
+
role-review-panel:
|
|
25
|
+
capability: N independent verdict lanes, each its own lane
|
|
26
|
+
profiles: [reviewer, architect]
|
|
27
|
+
used_by: P7 at stakes S3 and above
|
|
28
|
+
budget: bounded — lanes only at S3+, because five parallel lanes cost about five times the tokens
|
|
29
|
+
|
|
30
|
+
role-synthesis:
|
|
31
|
+
capability: merges many outputs into one artifact
|
|
32
|
+
profiles: [architect]
|
|
33
|
+
used_by: P11, P12
|
|
34
|
+
|
|
35
|
+
role-investigate:
|
|
36
|
+
capability: read-only exploration returning a distilled summary
|
|
37
|
+
profiles: [researcher]
|
|
38
|
+
used_by: P1 (read-only half)
|
|
39
|
+
|
|
40
|
+
role-judge:
|
|
41
|
+
capability: decides whether a process met its own predicate, from a command's output and a
|
|
42
|
+
pointer it can resolve, never from a report
|
|
43
|
+
profiles: [judge]
|
|
44
|
+
used_by: P10 (the exit predicate), P12 (grading), P7 at S3+ (the panel's foreman), the terminal
|
|
45
|
+
handoff gate, an automation agent's "did the fix land"
|
|
46
|
+
budget: one lane per verdict - a verdict is a decision, so this role never fans out
|
|
47
|
+
|
|
48
|
+
# Role-pinned fan-out goes through the kanban, never through a bare subagent spawn:
|
|
49
|
+
# delegate_task takes no model or provider parameter, so it cannot honour a role.
|
|
50
|
+
# kanban_create takes both. docs/ROLES.md states this as a rule.
|
|
51
|
+
fanout:
|
|
52
|
+
carrier: kanban
|
|
53
|
+
rule: one card per lane, each carrying the resolved provider/model, the pinned SHA, the diff and one focus
|
|
54
|
+
judge_is_not_a_panel: role-judge is a single lane; N independent verdicts is role-review-panel
|