@bongos/core 1.20.42 → 1.20.44

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (68) hide show
  1. package/.bongos-core.json +99 -74
  2. package/.claude/skills/backlog-review/SKILL.md +1 -0
  3. package/.claude/skills/blocker-review/SKILL.md +1 -1
  4. package/.claude/skills/bug-triage/SKILL.md +1 -0
  5. package/.claude/skills/builder-backup/SKILL.md +1 -0
  6. package/.claude/skills/builder-claim/SKILL.md +1 -1
  7. package/.claude/skills/builder-cost/SKILL.md +1 -0
  8. package/.claude/skills/builder-exit/SKILL.md +1 -0
  9. package/.claude/skills/builder-key/SKILL.md +1 -1
  10. package/.claude/skills/builder-reauth/SKILL.md +1 -1
  11. package/.claude/skills/builder-redteam/SKILL.md +1 -1
  12. package/.claude/skills/builder-sequence/SKILL.md +1 -1
  13. package/.claude/skills/builder-setup/SKILL.md +1 -1
  14. package/.claude/skills/collab-review/SKILL.md +1 -1
  15. package/.claude/skills/design/SKILL.md +1 -1
  16. package/.claude/skills/design-sync/SKILL.md +1 -0
  17. package/.claude/skills/feedback/SKILL.md +1 -1
  18. package/.claude/skills/figma-design-sync/SKILL.md +1 -0
  19. package/.claude/skills/goal-close/SKILL.md +184 -0
  20. package/.claude/skills/goal-create/SKILL.md +1 -1
  21. package/.claude/skills/goal-review/SKILL.md +5 -102
  22. package/.claude/skills/goal-uat/SKILL.md +5 -81
  23. package/.claude/skills/grade-audit/SKILL.md +6 -85
  24. package/.claude/skills/grade-recover/SKILL.md +1 -1
  25. package/.claude/skills/grade-sweep/SKILL.md +174 -0
  26. package/.claude/skills/grader-health/SKILL.md +5 -72
  27. package/.claude/skills/idea-triage/SKILL.md +1 -0
  28. package/.claude/skills/merge-mode/SKILL.md +1 -1
  29. package/.claude/skills/new-project/SKILL.md +1 -0
  30. package/.claude/skills/owner-review/SKILL.md +40 -0
  31. package/.claude/skills/planning-session/SKILL.md +1 -0
  32. package/.claude/skills/priority-session/SKILL.md +1 -0
  33. package/.claude/skills/read-session-export/SKILL.md +1 -1
  34. package/.claude/skills/recall/SKILL.md +1 -1
  35. package/.claude/skills/scan-before-install/SKILL.md +1 -1
  36. package/.claude/skills/session-handoff/SKILL.md +1 -1
  37. package/.claude/skills/worktree-clean/SKILL.md +1 -1
  38. package/docs/api/openapi.json +1 -1
  39. package/docs/file-map.md +12 -10
  40. package/docs/module-api-changelog.md +4 -0
  41. package/docs/modules-contract.md +3 -0
  42. package/docs/onboarding/slash-commands.md +17 -11
  43. package/docs/packs/artist.md +1 -1
  44. package/docs/packs/engineer.md +4 -5
  45. package/docs/page-readings.json +21 -21
  46. package/modules/copy-desk/module.json +1 -1
  47. package/{.claude → modules/copy-desk}/skills/tweak/SKILL.md +1 -0
  48. package/modules/hall-ui/public/collab.css +8 -1
  49. package/modules/hall-ui/public/oversight.css +0 -3
  50. package/modules/lifecycle/dependency-advisory.js +2 -2
  51. package/package-lock.json +2 -2
  52. package/package.json +1 -1
  53. package/release-notes.json +16 -0
  54. package/scripts/gds/fitness-ratchets.js +2 -0
  55. package/scripts/gds/module-artifact.js +52 -2
  56. package/scripts/gds/module.js +57 -4
  57. package/scripts/gds/skill-lint.js +32 -9
  58. package/src/bongos/routes/modules.js +5 -4
  59. package/src/module-api.js +1 -1
  60. package/src/module-loader/manifest-schema.js +28 -0
  61. package/tests/hall_lead_row_home.mjs +21 -0
  62. package/tests/module_manifest.mjs +22 -0
  63. package/tests/module_store_publish.mjs +88 -4
  64. package/tests/module_store_publish_route.mjs +20 -1
  65. package/tests/skill_grade_audit.mjs +34 -17
  66. package/tests/skill_grader_health.mjs +28 -10
  67. package/tests/skill_lint.mjs +37 -0
  68. package/tests/skill_menu_ruling.mjs +99 -0
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: collab-review
3
3
  description: >-
4
- Walk the help requests and task recommendations addressed to YOU, verifying each against the live record, one decision at a time. The answering half of /ask-for-help. Triggers: "/collab-review", "what have people asked me", "answer my asks", "work my collab queue", or a pasted "Copy for Session Start" prompt.
4
+ Walk the help requests and task recommendations addressed to YOU, one decision at a time. Triggers: "/collab-review", "what have people asked me", "answer my asks", "work my collab queue", a pasted "Copy for Session Start" prompt.
5
5
  plain: >-
6
6
  Goes through the requests other people have sent you, such as asks for help or suggested tasks, and brings you one decision at a time.
7
7
  reach-for: >-
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: design
3
3
  description: >-
4
- The UI design session — the playbook for ui-discipline work: layouts, pages, components, design systems, the hall, status and landing surfaces. World-first: the branding pack and DESIGN.md come before any taste rule. Triggers: "/design", "design this page", "let's work on the UI", or a claimed ui task.
4
+ The UI design playbook: layouts, pages, components, design systems, the hall, status and landing. World-first: the branding pack and DESIGN.md before any taste rule. Triggers: "/design", "design this page", "let's work on the UI", a claimed ui task.
5
5
  plain: >-
6
6
  The guided way to design or change how a page or screen looks, starting from the project's own style.
7
7
  reach-for: >-
@@ -2,6 +2,7 @@
2
2
  name: design-sync
3
3
  description: >-
4
4
  Export the repo design system into Claude Design, then land its generated UI code back as a task. Triggers: "/design-sync", "sync design", "export tokens to Claude Design", "import to Claude Design", "push design system".
5
+ disable-model-invocation: true
5
6
  plain: >-
6
7
  Sends the project's design style to Claude Design, then brings the screens designed there back into the project as a task.
7
8
  reach-for: >-
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: feedback
3
3
  description: >-
4
- Pull the latest Bongos feedback bundle into this session — the walkthrough transcript and screenshots, by absolute path. Local and read-only. Triggers: "/feedback", "pull in my latest feedback", "load my bongos feedback", "I just recorded a walkthrough", "grab the latest feedback bundle", "act on my screen recording".
4
+ Pull the latest Bongos feedback bundle (walkthrough transcript, screenshots) into this session; read-only. Triggers: "/feedback", "pull in my latest feedback", "load my bongos feedback", "I just recorded a walkthrough", "act on my screen recording".
5
5
  plain: >-
6
6
  Loads your most recent recorded walkthrough, what you said and the screenshots, into this conversation.
7
7
  reach-for: >-
@@ -2,6 +2,7 @@
2
2
  name: figma-design-sync
3
3
  description: >-
4
4
  Round-trip the UI design system with Figma — push tokens and surfaces, then land a designer's edit back as code. Not otb-figma-sync, which is pixel-art tiles only. Triggers: "sync design to Figma", "push tokens to Figma", "pull my Figma edit into code", "figma round-trip".
5
+ disable-model-invocation: true
5
6
  plain: >-
6
7
  Keeps the project's design style and a Figma file in step: sends the style to Figma, and turns a designer's edit there back into the project.
7
8
  reach-for: >-
@@ -0,0 +1,184 @@
1
+ ---
2
+ name: goal-close
3
+ description: >-
4
+ Close a goal's criteria: test the ones awaiting UAT on the live site and sign off, then decide the ones a UAT can't close (work abandoned or unlinked). Metic+. Triggers: "/goal-close", "/goal-uat", "/goal-review", "what's awaiting UAT".
5
+ disable-model-invocation: true
6
+ plain: >-
7
+ Walks you through trying finished goal work on the live site so a person can sign it off, then helps decide the checks nobody can test any more.
8
+ reach-for: >-
9
+ When work is waiting for someone to try it for real, or a goal is stuck on checks nobody can finish.
10
+ cost: >-
11
+ Free. Your sign-off or decision is recorded; nothing else changes.
12
+ ---
13
+
14
+ You are running the **goal-close** session: the closing end of a goal. It is one skill with two parts that used to be two commands (task 1004471):
15
+
16
+ - **Part 1 — UAT** (was `/goal-uat`): criteria whose work shipped, tested on the live site and signed off.
17
+ - **Part 2 — residue** (was `/goal-review`): criteria a UAT cannot close, because all their work was abandoned or none was ever linked.
18
+
19
+ **Which part to run.** `/goal-close` runs Part 1 then Part 2. `/goal-close uat` (or the old `/goal-uat`) runs Part 1 only; `/goal-close residue` (or the old `/goal-review`) runs Part 2 only. If the builder named a goal, keep only that goal's rows in both parts.
20
+
21
+ > **The other end of this lifecycle is [`/goal-create`](../goal-create/SKILL.md)** — it opens a goal (scope wall, criteria, seed tasks).
22
+
23
+ ## Rank gate — Metic+ only (both parts)
24
+
25
+ Signing off rides `criterion.review`, and an override rides `POST /api/bongos/done-when/:criterionId/satisfy` (`requireRank('metic','archon')`, pinned in `route-rank-check.MODULE_EXPECTED_RANKS.lifecycle`). The server enforces both; this check just stops a Xenos early instead of after a wall of 403s. Markdown never grants authority ([`docs/canonical-permissions.md`](../../../docs/canonical-permissions.md), [ADR 0016](../../../docs/adr/0016-trust-boundary-server-enforced-permissions.md)).
26
+
27
+ **Enforce the gate at the very top of the session, before reading anything else:**
28
+
29
+ 1. Call `node scripts/gds/api.js GET /api/bongos/me` and read `builder.rank`. Compare case-insensitively (lowercase it before testing).
30
+ 2. **If `rank` lowercases to `xenos` or `thetes`** → stop immediately. Tell the builder: *"Closing a goal's criteria is a Metic+ step — signing one off or overriding it closes out a slice of a version and can auto-achieve a goal. Ask an Archon to promote you, or ask a Metic to run it."* Do not list either queue.
31
+ 3. **If `rank` lowercases to `metic` or `archon`** → proceed.
32
+ 4. **If `rank` is absent** (a deployment where the three-rank model hasn't shipped) → treat the session as available, exactly as `/idea-triage` does in its pre-rank passthrough mode, and note that in the session log.
33
+
34
+ ---
35
+
36
+ ## Part 1 — UAT: test on the live site, then sign off
37
+
38
+ A criterion no longer closes because its linked tasks shipped ([ADR 0351](../../../docs/adr/0351-a-criterion-closes-on-a-uat.md), task 1004392, superseding ADR 0183's "shipped is enough"). Once its work has shipped it reads **Awaiting UAT**, and it closes only when a person who did **not** ship that work does what the criterion describes **on the live site** and signs it off.
39
+
40
+ This is the owner's rule because criteria were being closed by any tasks that happened to be linked to them, whether or not the thing the criterion describes was real. A UAT is the check against the criterion's **words**, not against the task list. Hold it to that.
41
+
42
+ ### The three checks (what you are verifying)
43
+
44
+ | Check | Passes when | Who decides |
45
+ |---|---|---|
46
+ | **Code** | every linked task is done and at least one shipped | automatic |
47
+ | **Live** | the shipped work is running on the deployed site | read automatically where the site can tell; otherwise the signer confirms it |
48
+ | **UAT** | a signed-in person performed the criterion on the live site; a recording is stored | the signer, who must not have shipped any linked task (the project owner excepted) |
49
+
50
+ A criterion marked **backend-only** (set when it was written, visible on the criterion) has no screen to record, so it takes a recording-free **backend sign-off** instead. It still needs a non-shipper to sign it.
51
+
52
+ ### Step 1: the queue
53
+
54
+ ```
55
+ node scripts/gds/uat.js --queue
56
+ ```
57
+
58
+ (Add `--version <id>` to scope it.) If nothing is awaiting UAT, say so and move on to Part 2 (or stop, if only Part 1 was asked for).
59
+
60
+ ### Step 2: one criterion at a time
61
+
62
+ For each, run `node scripts/gds/uat.js <criterion-id>` and present, in plain words:
63
+
64
+ - the criterion's text, in full;
65
+ - its goal, and whether it is the goal's **last** open criterion (then signing it off achieves the goal);
66
+ - the Live line (running, not yet, or "this site cannot tell");
67
+ - whether it is backend-only.
68
+
69
+ Then tell the tester exactly what to do **on the live site**, derived from the criterion's words: which page, what to click, what they should see. Write it as a short numbered script. If the criterion's words describe something the tester cannot find on the live site, that is the answer: **it is not met.** Say so and go to Step 4 (reject) rather than looking for a reading under which it passes.
70
+
71
+ ### Step 3: record and sign off
72
+
73
+ The tester records their screen doing the script (any screen recorder that saves **mp4 or webm**, up to **100 MB**; Windows: Win+Alt+R with Xbox Game Bar, or the Snipping Tool's record button; macOS: Cmd+Shift+5). Then:
74
+
75
+ ```
76
+ node scripts/gds/uat.js <criterion-id> --recording <path-to-file> --note "<one line: what was checked>"
77
+ ```
78
+
79
+ For a backend-only criterion, after the tester has checked it by whatever means proves it (a log line, an API read, a test run against live):
80
+
81
+ ```
82
+ node scripts/gds/uat.js <criterion-id> --backend-signoff --note "<how it was checked>"
83
+ ```
84
+
85
+ Add `--live-attested` only when the server says this site cannot tell whether the work is live, and only if the tester really did it on the live site.
86
+
87
+ **Refusals, in the words to give the builder:**
88
+
89
+ | Refusal | Meaning |
90
+ |---|---|
91
+ | `signer_shipped_this_work` | they shipped linked work; someone else signs this one |
92
+ | `work_not_live` | shipped but not deployed yet; sign off after the deploy |
93
+ | `criterion_not_awaiting_uat` | linked work still open (or none linked, or all abandoned: that is Part 2) |
94
+ | `criterion_needs_recording` / `criterion_is_backend_only` | wrong kind of sign-off for this criterion |
95
+ | `live_attestation_required` | confirm it was tested on the live site (`--live-attested`) |
96
+
97
+ A sign-off on a goal's last open criterion reports the goal achieved; say so.
98
+
99
+ ### Step 4: when it is NOT met
100
+
101
+ The shipped work does not do what the criterion says. Do not sign off. File the missing work as a task linked to the criterion (`POST /api/bongos/tasks`, then `POST /api/bongos/tasks/<id>/criteria` with `{"criterion_id": <criterion-id>}`) so the criterion drops back to Open until that ships. That is what happened to `wa7-government` (task 1004400).
102
+
103
+ ### What Part 1 does NOT do
104
+
105
+ - It never uses `POST /done-when/:id/satisfy`. That is the **override** (Part 2): it requires a written `override_reason` and the criterion reads "Satisfied (override)" for good. Use it only when the owner asks, and never to get past a refusal above.
106
+ - It never marks a criterion backend-only to avoid recording one. Backend-only is set while a criterion is still open (`uat.js <id> --backend-only on`), and the server refuses it once the criterion is awaiting UAT.
107
+
108
+ ---
109
+
110
+ ## Part 2 — residue: the criteria a UAT cannot close
111
+
112
+ This part mirrors `/idea-triage`, for the *closing* end of the work hierarchy (ADR 0086 §6). The review queue (`GET /done-when/pending-review`) also holds criteria nothing was delivered for, and a UAT cannot test those (every row carries `uat_state`; skip the `awaiting_uat` ones, they are Part 1's):
113
+
114
+ 1. **A criterion whose linked tasks were ALL ABANDONED.** Nothing shipped, so there is nothing to test. *Somebody must decide whether the criterion still means anything now that its work is gone*: reject it (file the work that will really deliver it, linked to it), or, if it was genuinely met another way, override it with the reason.
115
+ 2. **A criterion with NO linked tasks.** It cannot even reach Awaiting UAT. Link the tasks that deliver it (better — it makes the board true); it then goes through Part 1 like any other.
116
+
117
+ **Rejecting is the valuable verdict here**; an override is the exception and says so forever ("Satisfied (override)").
118
+
119
+ ### Step 1 — fetch the review queue
120
+
121
+ ```
122
+ node scripts/gds/api.js GET /api/bongos/done-when/pending-review
123
+ ```
124
+
125
+ (Add `?version=BONGOS-V1` to scope to one version. `api.js` is the cross-platform Bongos API helper — it signs the request with your session token.)
126
+
127
+ If no row is left once the `awaiting_uat` ones are skipped, print "review queue empty — no criteria pending review today" and finish.
128
+
129
+ Each row carries: `id` (the criterion's numeric id — what you POST to), `criterion_id` (the slug) + `criterion_md` (the prose), `version_id`, `goal_id` + `goal_title` + `goal_status`, `task_total` (how many shipped tasks fulfilled it), and **`is_last_in_goal`** — `true` when this is the goal's only remaining unsatisfied criterion, so **closing it will auto-achieve the goal**.
130
+
131
+ ### Step 2 — walk the list, criterion by criterion
132
+
133
+ Present each criterion with: its `criterion_md`, its goal (`goal_title`), how many tasks shipped under it (`task_total`), and — when `is_last_in_goal` is true — a clear flag:
134
+
135
+ > ⚠️ This is the LAST open criterion in **{goal_title}** — closing it will mark the whole goal **achieved**.
136
+
137
+ Then ask the user (concise; frame in product/outcome terms):
138
+
139
+ > Criterion C{n} — "{criterion_md}". None of its linked work was delivered. **Reject and file the real work, override (it was met another way), or defer?**
140
+
141
+ ### Step 3 — apply the verdict
142
+
143
+ **Override** (only for a criterion met some other way, never for one awaiting UAT) → it closes WITHOUT a UAT, so the server requires the reason and records it; the criterion reads "Satisfied (override)" from then on:
144
+
145
+ ```
146
+ node scripts/gds/api.js POST /api/bongos/done-when/<criterion_id>/satisfy --body '{"override_reason":"<why this is met without a UAT>"}'
147
+ ```
148
+
149
+ The response is `{ criterion, goal_achieved, goal }`. **If `goal_achieved` is true**, announce it: *"Goal '{goal.title}' is now achieved — all its criteria are closed."* (That goal now awaits disposition — see Step 4.)
150
+
151
+ **Reject** → on review it is **not actually met** (the work missed something, or "done" was mis-scoped). There is no "unsatisfy the flag" — the durable fix is to **add the missing work**: capture what's still needed (a follow-up task linked to this criterion, or an idea via `node scripts/gds/capture.js "..."`), and note the rejection in the session log. Once an unshipped task is linked to the criterion, it drops out of `pending_review` on its own. If no follow-up is filed, the criterion simply stays flagged and resurfaces next session — an honest signal that a human said "not done" but nothing was queued to close the gap. **Do not close a criterion you would not stake the version on.**
152
+
153
+ **Note / defer** is implicit — leave the criterion open by not POSTing. It stays in the queue and surfaces again next session (same as `/idea-triage`'s defer).
154
+
155
+ ### Step 4 — achieved goals (disposition)
156
+
157
+ Any goal that flipped to `achieved` this session (or is already `achieved`) awaits a disposition: **archive** it (done, put it to rest) or **carry it forward** into the next version's scope. Surface each newly-achieved goal to the user and record the intended disposition in the session log.
158
+
159
+ > The archive / reopen / carry-forward **actions** (the routes that mutate goal lifecycle) land in **BV1.R64** ([task 1518](https://example.com/builders#/task/1518)) — they are not wired here. For this session, *report* achieved goals and capture the intent; once R64 ships, this step gains the verbs to execute it.
160
+
161
+ ### Step 5 — dependency check (task 1002827, idea 1000276)
162
+
163
+ A cheap housekeeping pass over the same corpus this session already has open: run the likely-missing-dependency detector and surface any `high`-confidence candidate for a human to triage.
164
+
165
+ ```
166
+ node scripts/gds/audit-deps.js --min high
167
+ ```
168
+
169
+ Each line names an open task and an unshipped task whose files it overlaps with no declared dependency between them. **Read both task descriptions before backfilling an edge** — most overlap is parallel work on a shared file, not a build order; a wrong edge gates the dependent (`POST /claims` → 409 `DEPS_NOT_SHIPPED`) until the named blocker ships. Report the count in the summary; zero is the common case and needs no further comment. This is advisory only — it never blocks the review, and skipping it on a rushed pass is fine (it resurfaces next session, same as a deferred criterion).
170
+
171
+ ---
172
+
173
+ ## Summary to give at the end
174
+
175
+ - **Part 1:** criteria walked · signed off (ids) · goals achieved · rejected, with the follow-up task filed for each · left waiting (not live yet, or no tester available).
176
+ - **Part 2:** criteria walked · overridden (count + ids) · goals auto-achieved (count + titles) · rejected (count + ids + what follow-up was filed) · deferred (count + ids — these surface again next session) · the new queue count · the dependency-check count.
177
+
178
+ ## Why we run this
179
+
180
+ Without a closing cadence, criteria pile up "done but unconfirmed" and goals never flip to achieved — the version looks perpetually in-progress even when the work has shipped. A fast pass keeps the rollup honest. As with `/idea-triage`, the cadence is the discipline; a quick walk with mostly defers beats skipping it.
181
+
182
+ ## Tone for the user
183
+
184
+ Frame each decision in terms of outcome — "this criterion was 'players see each other'; both tasks shipped — is that actually true in the live game?" not "flip satisfied=true on criterion 45." Decide the mechanics yourself; ask the owner only about whether the thing is *really* done.
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: goal-create
3
3
  description: >-
4
- Plan and create ONE goal — the module-scoped workspace between a version and its criteria — drafting its done-when criteria and seed tasks. Metic+ only. Triggers: "/goal-create", "create a goal", "what goal should this live in", or a task created with no goal.
4
+ Plan and create ONE goal (the module-scoped workspace between a version and its criteria) with its done-when criteria and seed tasks. Metic+. Triggers: "/goal-create", "create a goal", "what goal should this live in", a task created with no goal.
5
5
  plain: >-
6
6
  Helps you set up a new goal: what it is for, how you will know it is done, and the first tasks to get there.
7
7
  reach-for: >-
@@ -1,111 +1,14 @@
1
1
  ---
2
2
  name: goal-review
3
3
  description: >-
4
- Decide criteria a UAT can't close: work abandoned or unlinked. Metic+. Triggers: "/goal-review".
4
+ Old name for /goal-close residue: decide criteria a UAT can't close (work abandoned or unlinked). Metic+. Triggers: "/goal-review".
5
+ disable-model-invocation: true
5
6
  plain: >-
6
- Helps you decide what to do with a goal's done-when checks that can no longer be tested, because their work was dropped or never linked.
7
+ The older name for deciding what to do with a goal's checks that can no longer be tested, because their work was dropped or never linked.
7
8
  reach-for: >-
8
- When a goal is stuck on checks nobody can finish.
9
+ When you remember the old name; it opens the same walk as closing a goal.
9
10
  cost: >-
10
11
  Free. It only closes or changes a check when you decide to.
11
12
  ---
12
13
 
13
- You are running the **goal-review** session: the residue of the criterion close. It mirrors `/idea-triage`, for the *closing* end of the work hierarchy (ADR 0086 §6).
14
-
15
- > **Criteria close on a UAT now, and that is [`/goal-uat`](../goal-uat/SKILL.md), not this skill** ([ADR 0351](../../../docs/adr/0351-a-criterion-closes-on-a-uat.md), task 1004392, superseding ADR 0183's "shipped is enough"). A criterion whose linked work has shipped reads **Awaiting UAT** and closes when someone who did not ship it tests it on the live site and signs it off. If the builder wants to close finished work, send them to `/goal-uat`.
16
- >
17
- > **The other end of this lifecycle is [`/goal-create`](../goal-create/SKILL.md)** — it opens a goal (scope wall, criteria, seed tasks).
18
-
19
- **What is left for this skill.** The review queue (`GET /done-when/pending-review`) also holds criteria nothing was delivered for, and a UAT cannot test those (every row carries `uat_state`; skip the `awaiting_uat` ones, they are `/goal-uat`'s):
20
-
21
- 1. **A criterion whose linked tasks were ALL ABANDONED.** Nothing shipped, so there is nothing to test. *Somebody must decide whether the criterion still means anything now that its work is gone*: reject it (file the work that will really deliver it, linked to it), or, if it was genuinely met another way, override it with the reason.
22
- 2. **A criterion with NO linked tasks.** It cannot even reach Awaiting UAT. Link the tasks that deliver it (better — it makes the board true); it then goes through `/goal-uat` like any other.
23
-
24
- **Rejecting is the valuable verdict here**; an override is the exception and says so forever ("Satisfied (override)").
25
-
26
- ## Rank gate — Metic+ only
27
-
28
- **Gated to Metic+ (Metic, Archon); not Xenos.** Confirming a criterion sets `satisfied=true` and can auto-achieve a goal — a trusted-builder operation that closes out a slice of a version (the "who confirms" resolution in ADR 0086 §6: the planner, not only the Archon). The server enforces this — `POST /api/bongos/done-when/:criterionId/satisfy` is `requireRank('metic','archon')` and pinned in `route-rank-check.MODULE_EXPECTED_RANKS.lifecycle` — so this check just stops a Xenos early instead of after a wall of 403s. Markdown never grants authority ([`docs/canonical-permissions.md`](../../../docs/canonical-permissions.md), [ADR 0016](../../../docs/adr/0016-trust-boundary-server-enforced-permissions.md)).
29
-
30
- **Enforce the gate at the very top of the session, before reading anything else:**
31
-
32
- 1. Call `node scripts/gds/api.js GET /api/bongos/me` and read `builder.rank`. Compare case-insensitively (lowercase it before testing).
33
- 2. **If `rank` lowercases to `xenos`** → stop immediately. Tell the builder: *"Goal review is a Metic+ operation — confirming a criterion closes out a slice of a version and can auto-achieve a goal. You're currently Xenos. Ask an Archon to promote you if you need to confirm criteria."* Do not list the queue, do not propose verdicts.
34
- 3. **If `rank` lowercases to `metic` or `archon`** → proceed with the steps below.
35
- 4. **If `rank` is absent** (a deployment where the three-rank model hasn't shipped) → treat the session as available, exactly as `/idea-triage` does in its pre-rank passthrough mode, and note that in the session log.
36
-
37
- ## What this skill does
38
-
39
- 1. Lists flagged criteria via `GET /api/bongos/done-when/pending-review` (optionally `?version=<id>`).
40
- 2. For each row that is NOT `awaiting_uat`, prompts the user with: **reject** (leave it open and file the work that delivers it), **override** (it was met another way: close it with a written reason), or **note / defer** (it resurfaces next session).
41
- 3. Applies an override via `POST /api/bongos/done-when/:criterionId/satisfy` with `override_reason` — and reports any goal that auto-achieved as a result.
42
-
43
- ## How to use
44
-
45
- **Step 1 — fetch the review queue.**
46
-
47
- ```
48
- node scripts/gds/api.js GET /api/bongos/done-when/pending-review
49
- ```
50
-
51
- (Add `?version=BONGOS-V1` to scope to one version. `api.js` is the cross-platform Bongos API helper — it signs the request with your session token.)
52
-
53
- If `count` is 0, print "review queue empty — no criteria pending review today" and exit.
54
-
55
- Each row carries: `id` (the criterion's numeric id — what you POST to), `criterion_id` (the slug) + `criterion_md` (the prose), `version_id`, `goal_id` + `goal_title` + `goal_status`, `task_total` (how many shipped tasks fulfilled it), and **`is_last_in_goal`** — `true` when this is the goal's only remaining unsatisfied criterion, so **confirming it will auto-achieve the goal**.
56
-
57
- **Step 2 — walk the list, criterion by criterion.**
58
-
59
- Present each criterion with: its `criterion_md`, its goal (`goal_title`), how many tasks shipped under it (`task_total`), and — when `is_last_in_goal` is true — a clear flag:
60
-
61
- > ⚠️ This is the LAST open criterion in **{goal_title}** — confirming it will mark the whole goal **achieved**.
62
-
63
- Then ask the user (concise; frame in product/outcome terms):
64
-
65
- > Criterion C{n} — "{criterion_md}". None of its linked work was delivered. **Reject and file the real work, override (it was met another way), or defer?**
66
-
67
- **Step 3 — apply the verdict.**
68
-
69
- **Override** (only for a criterion met some other way, never for one awaiting UAT) → it closes WITHOUT a UAT, so the server requires the reason and records it; the criterion reads "Satisfied (override)" from then on:
70
-
71
- ```
72
- node scripts/gds/api.js POST /api/bongos/done-when/<criterion_id>/satisfy --body '{"override_reason":"<why this is met without a UAT>"}'
73
- ```
74
-
75
- The response is `{ criterion, goal_achieved, goal }`. **If `goal_achieved` is true**, announce it: *"Goal '{goal.title}' is now achieved — all its criteria are closed."* (That goal now awaits disposition — see Step 4.)
76
-
77
- **Reject** → the criterion's tasks all shipped, but on review it is **not actually met** (the work missed something, or "done" was mis-scoped). There is no "unsatisfy the flag" — the flag is *derived* from shipped tasks, so the durable fix is to **add the missing work**: capture what's still needed (a follow-up task linked to this criterion, or an idea via `node scripts/gds/capture.js "..."`), and note the rejection in the session log. Once an unshipped task is linked to the criterion, it drops out of `pending_review` on its own. If no follow-up is filed, the criterion simply stays flagged and resurfaces next session — an honest signal that a human said "not done" but nothing was queued to close the gap. **Do not confirm a criterion you would not stake the version on.**
78
-
79
- **Note / defer** is implicit — leave the criterion un-confirmed by not POSTing. It stays in the queue and surfaces again next session (same as `/idea-triage`'s defer).
80
-
81
- **Step 4 — achieved goals (disposition).**
82
-
83
- Any goal that flipped to `achieved` this session (or is already `achieved`) awaits a disposition: **archive** it (done, put it to rest) or **carry it forward** into the next version's scope. Surface each newly-achieved goal to the user and record the intended disposition in the session log.
84
-
85
- > The archive / reopen / carry-forward **actions** (the routes that mutate goal lifecycle) land in **BV1.R64** ([task 1518](https://example.com/builders#/task/1518)) — they are not wired here. For this session, *report* achieved goals and capture the intent; once R64 ships, this step gains the verbs to execute it.
86
-
87
- **Step 5 — surface a summary.**
88
-
89
- Report to the user:
90
- - Total criteria walked
91
- - Confirmed (count + ids)
92
- - Goals auto-achieved (count + titles)
93
- - Rejected (count + ids + what follow-up was filed)
94
- - Deferred (count + ids — these surface again next session)
95
- - New queue `count`
96
-
97
- **Step 6 — dependency check (task 1002827, idea 1000276).** A cheap housekeeping pass over the same corpus this session already has open: run the likely-missing-dependency detector and surface any `high`-confidence candidate for a human to triage.
98
-
99
- ```
100
- node scripts/gds/audit-deps.js --min high
101
- ```
102
-
103
- Each line names an open task and an unshipped task whose files it overlaps with no declared dependency between them. **Read both task descriptions before backfilling an edge** — most overlap is parallel work on a shared file, not a build order; a wrong edge gates the dependent (`POST /claims` → 409 `DEPS_NOT_SHIPPED`) until the named blocker ships. Report the count in the same summary as Step 5; zero is the common case and needs no further comment. This is advisory only — it never blocks the review, and skipping it on a rushed pass is fine (it resurfaces next session, same as a deferred criterion).
104
-
105
- ## Why we run this
106
-
107
- Without a closing cadence, criteria pile up "done but unconfirmed" and goals never flip to achieved — the version looks perpetually in-progress even when the work has shipped. A fast pass keeps the rollup honest. As with `/idea-triage`, the cadence is the discipline; a quick walk with mostly defers beats skipping it.
108
-
109
- ## Tone for the user
110
-
111
- Lars is the prompter (see CLAUDE.md §2). Frame each decision in terms of outcome — "this criterion was 'players see each other'; both tasks shipped — is that actually true in the live game?" not "flip satisfied=true on criterion 45." Decide the mechanics yourself; ask him only about whether the thing is *really* done.
14
+ **This command is an alias (task 1004471).** `/goal-review` and `/goal-uat` were merged into one skill, `/goal-close`. Read [`.claude/skills/goal-close/SKILL.md`](../goal-close/SKILL.md) and run its rank gate, then **Part 2 — residue** only (the same as `/goal-close residue`), with every rule that file states: rejecting is the valuable verdict, and an override always carries its written reason.
@@ -1,90 +1,14 @@
1
1
  ---
2
2
  name: goal-uat
3
3
  description: >-
4
- Test criteria awaiting UAT on the live site, then sign off. Metic+. Triggers: "/goal-uat", "what's awaiting UAT".
4
+ Old name for /goal-close uat: test criteria awaiting UAT on the live site, then sign off. Metic+. Triggers: "/goal-uat".
5
+ disable-model-invocation: true
5
6
  plain: >-
6
- Walks you through trying a finished goal on the live site, so a person can confirm it really works before it is signed off.
7
+ The older name for trying a finished goal on the live site, so a person can confirm it really works before it is signed off.
7
8
  reach-for: >-
8
- When work is waiting for someone to try it for real.
9
+ When you remember the old name; it opens the same walk as closing a goal.
9
10
  cost: >-
10
11
  Free. Your sign-off is recorded; nothing else changes.
11
12
  ---
12
13
 
13
- You are running the **goal-uat** session: the closing step of a goal. A criterion no longer closes because its linked tasks shipped ([ADR 0351](../../../docs/adr/0351-a-criterion-closes-on-a-uat.md), task 1004392). Once its work has shipped it reads **Awaiting UAT**, and it closes only when a person who did **not** ship that work does what the criterion describes **on the live site** and signs it off.
14
-
15
- This is the owner's rule because criteria were being closed by any tasks that happened to be linked to them, whether or not the thing the criterion describes was real. A UAT is the check against the criterion's **words**, not against the task list. Hold it to that.
16
-
17
- ## Rank gate: Metic+ only
18
-
19
- Signing off rides `criterion.review` (Metic+); the server enforces it. Check first so a Xenos is not walked into a wall of 403s: `node scripts/gds/api.js GET /api/bongos/me` and read `builder.rank`. If it lowercases to `xenos` or `thetes`, stop and say: *"Signing off a UAT is a Metic+ step. Ask an Archon to promote you, or ask a Metic to run it."*
20
-
21
- ## The three checks (what you are verifying)
22
-
23
- | Check | Passes when | Who decides |
24
- |---|---|---|
25
- | **Code** | every linked task is done and at least one shipped | automatic |
26
- | **Live** | the shipped work is running on the deployed site | read automatically where the site can tell; otherwise the signer confirms it |
27
- | **UAT** | a signed-in person performed the criterion on the live site; a recording is stored | the signer, who must not have shipped any linked task (the project owner excepted) |
28
-
29
- A criterion marked **backend-only** (set when it was written, visible on the criterion) has no screen to record, so it takes a recording-free **backend sign-off** instead. It still needs a non-shipper to sign it.
30
-
31
- ## Step 1: the queue
32
-
33
- ```
34
- node scripts/gds/uat.js --queue
35
- ```
36
-
37
- (Add `--version <id>` to scope it.) If nothing is awaiting UAT, say so and stop. If the builder named a goal, keep only that goal's rows.
38
-
39
- ## Step 2: one criterion at a time
40
-
41
- For each, run `node scripts/gds/uat.js <criterion-id>` and present, in plain words:
42
-
43
- - the criterion's text, in full;
44
- - its goal, and whether it is the goal's **last** open criterion (then signing it off achieves the goal);
45
- - the Live line (running, not yet, or "this site cannot tell");
46
- - whether it is backend-only.
47
-
48
- Then tell the tester exactly what to do **on the live site**, derived from the criterion's words: which page, what to click, what they should see. Write it as a short numbered script. If the criterion's words describe something the tester cannot find on the live site, that is the answer: **it is not met.** Say so and go to Step 4 (reject) rather than looking for a reading under which it passes.
49
-
50
- ## Step 3: record and sign off
51
-
52
- The tester records their screen doing the script (any screen recorder that saves **mp4 or webm**, up to **100 MB**; Windows: Win+Alt+R with Xbox Game Bar, or the Snipping Tool's record button; macOS: Cmd+Shift+5). Then:
53
-
54
- ```
55
- node scripts/gds/uat.js <criterion-id> --recording <path-to-file> --note "<one line: what was checked>"
56
- ```
57
-
58
- For a backend-only criterion, after the tester has checked it by whatever means proves it (a log line, an API read, a test run against live):
59
-
60
- ```
61
- node scripts/gds/uat.js <criterion-id> --backend-signoff --note "<how it was checked>"
62
- ```
63
-
64
- Add `--live-attested` only when the server says this site cannot tell whether the work is live, and only if the tester really did it on the live site.
65
-
66
- **Refusals, in the words to give the builder:**
67
-
68
- | Refusal | Meaning |
69
- |---|---|
70
- | `signer_shipped_this_work` | they shipped linked work; someone else signs this one |
71
- | `work_not_live` | shipped but not deployed yet; sign off after the deploy |
72
- | `criterion_not_awaiting_uat` | linked work still open (or none linked, or all abandoned: that is `/goal-review`) |
73
- | `criterion_needs_recording` / `criterion_is_backend_only` | wrong kind of sign-off for this criterion |
74
- | `live_attestation_required` | confirm it was tested on the live site (`--live-attested`) |
75
-
76
- A sign-off on a goal's last open criterion reports the goal achieved; say so.
77
-
78
- ## Step 4: when it is NOT met
79
-
80
- The shipped work does not do what the criterion says. Do not sign off. File the missing work as a task linked to the criterion (`POST /api/bongos/tasks`, then `POST /api/bongos/tasks/<id>/criteria` with `{"criterion_id": <criterion-id>}`) so the criterion drops back to Open until that ships. That is what happened to `wa7-government` (task 1004400).
81
-
82
- ## What this skill does NOT do
83
-
84
- - It never uses `POST /done-when/:id/satisfy`. That is the **override** now: it requires a written `override_reason` and the criterion reads "Satisfied (override)" for good. Use it only when the owner asks, and never to get past a refusal above.
85
- - It never marks a criterion backend-only to avoid recording one. Backend-only is set while a criterion is still open (`uat.js <id> --backend-only on`), and the server refuses it once the criterion is awaiting UAT.
86
- - The all-abandoned and no-linked-task residue is `/goal-review`'s, not this skill's.
87
-
88
- ## Summary to give at the end
89
-
90
- Criteria walked · signed off (ids) · goals achieved · rejected, with the follow-up task filed for each · left waiting (not live yet, or no tester available).
14
+ **This command is an alias (task 1004471).** `/goal-uat` and `/goal-review` were merged into one skill, `/goal-close`. Read [`.claude/skills/goal-close/SKILL.md`](../goal-close/SKILL.md) and run its rank gate, then **Part 1 — UAT** only (the same as `/goal-close uat`), with every rule that file states: hold each criterion to its words on the live site, and never use the override to get past a refusal.
@@ -1,93 +1,14 @@
1
1
  ---
2
2
  name: grade-audit
3
3
  description: >-
4
- Accountability sweep over recent grades: override ledger, dropped findings on passing grades, false-pass spot-check, outage roll-up. Queues follow-ups only — never confirms, flips a status, or re-grades. Triggers: "/grade-audit", "audit recent grades", "what shipped past the grader".
4
+ Old name for /grade-sweep audit: the accountability half of the grade sweep (override ledger, dropped findings, false-pass spot-check). Queues follow-ups only. Triggers: "/grade-audit".
5
+ disable-model-invocation: true
5
6
  plain: >-
6
- Looks back over recent quality reviews for anything that slipped through: overrides, ignored findings, or passes that look wrong.
7
+ The older name for the look back over recent quality reviews, for anything that slipped through.
7
8
  reach-for: >-
8
- Now and then, to check that the quality reviews are being trusted properly.
9
+ When you remember the old name; it opens the same check as the grade sweep.
9
10
  cost: >-
10
- Free. It only files follow-up tasks; it never changes a review or a task's status.
11
+ Free. It only files follow-up notes; it never changes a review or a task.
11
12
  ---
12
13
 
13
- You are auditing the last week of grades for accountability gaps: work that shipped past a failing grade unaccounted for, real findings that rode passing grades into production and evaporated, and passes the panel structurally could not have judged. You **describe and queue**. You never confirm, flip, re-grade, or add a gate — every follow-up lands as a queued idea or a handoff to another skill.
14
-
15
- ## The sweep window
16
-
17
- ```bash
18
- node scripts/gds/api.js GET "/api/bongos/public/recent-shipped?days=7"
19
- ```
20
-
21
- For each shipped task in the window, read its persisted grade:
22
-
23
- ```bash
24
- node scripts/gds/api.js GET "/api/bongos/tasks/<id>?include=grade"
25
- ```
26
-
27
- Bucket by grade shape: `passed`, failed, `signals.panel_outcome='unavailable'` (outage — no quality signal existed), or no grade row at all.
28
-
29
- **Then split off the no-artifact species before any leg reads the bucket** (task 1004011, ADR 0320). A grade with `signals.no_artifact_grade === true` judged the builder's **claim that nothing was needed** — the handoff notes and value summary against the task description — not a diff. There is no diff, no panel, and one generalist shot (`signals.lenses_dropped` names the four lenses that did not run). It is a real grade and a real gate, but it measures a different thing, so:
30
-
31
- - **Do not average or count it alongside diff grades.** A 7.5 over three paragraphs of prose and a 7.5 over a 400-line diff are not the same number.
32
- - **Leg 2 still applies, and matters more here.** A `blocker`/`major` on a no-artifact pass usually means the *evidence* was thin — queue it the same way.
33
- - **Leg 3 does not apply** — there is no merged diff to re-read for runtime behaviour. The equivalent question, worth asking on any no-artifact pass that shipped on assertion alone, is: *could a reader today re-run or re-read what the notes claim?* If not, queue it.
34
- - **A no-artifact FAIL is not an override candidate for Leg 1** unless it also landed. The builder's remedy is better evidence, not a permission.
35
-
36
- Note `task_grades.grader_kind` is **not** the discriminator: the server records `'subagent'` for everything arriving through `POST /tasks/:id/grade`. `signals.no_artifact_grade` is the durable marker.
37
-
38
- ## Leg 1 — the override ledger
39
-
40
- For every shipped task whose grade **failed** (or is absent), the land needed an override. Cross-reference the ledger (Archon — `override_request.decide`):
41
-
42
- ```bash
43
- node scripts/gds/api.js GET "/api/bongos/override-requests?status=approved"
44
- ```
45
-
46
- - Split **infra vs real** first: a fail that is actually `panel_outcome='unavailable'` (or the errored-worker zero-issue shape) is an outage artifact, not an overridden quality verdict — route those to the Leg 4 roll-up.
47
- - A shipped task with a real failing grade and **no approved override-request row** landed via the raw confirm — the un-audited path. Name it in the report and print its unaddressed `issues[]` verbatim. That set is the ledger's debt.
48
- - For each such task, queue the debt (see the queueing rule below) — do not chase the lander, do not un-ship anything.
49
-
50
- ## Leg 2 — dropped findings on passing grades
51
-
52
- A pass with `blocker`/`major` entries in `issues[]` shipped real findings that no process ever picks up (the evaluation counted 274 findings on passing ships with zero conversion paths). For each:
53
-
54
- - **Worker-attribution filter:** findings attributed to the advisory **Narc** are the known noise class — triage them by hand (read the finding, decide), never auto-trust them into the queue. Findings from the gating workers (Quality, Hacker, Efficiency) or the deterministic pre-passes are higher-confidence — queue them unless plainly stale.
55
- - **Queueing rule (idempotent):** first read the open inbox (`node scripts/gds/api.js GET /api/bongos/inbox`) and skip any finding already queued — the marker is the deterministic title prefix. Then:
56
-
57
- ```bash
58
- node scripts/gds/capture.js "[grade-audit] task <id>: <finding summary>" --kind bug
59
- ```
60
-
61
- One idea per surviving finding, titled exactly `[grade-audit] task <id>: …` so a re-run of this audit files nothing twice.
62
-
63
- ## Leg 3 — false-pass spot-check (runtime-behavior blindness)
64
-
65
- The re-review found 2/8 material misses, both the same root cause: **the panel grades diff text and never reasons about runtime** — a CI workflow whose default shallow checkout broke on first run scored functional_fidelity 10/10, and an SSRF guard bypassable via HTTP redirect passed at 9.5. For the riskiest ships in the window — security surfaces, CI/workflow files, the largest diffs, and 9.5s with suspiciously few findings — re-read the merged diff asking one question: *what does this code do at runtime that the diff text does not show?* (defaults it inherits, redirects it follows, first-run behavior, tool/CI defaults). A suspicion is queued via the Leg 2 rule, prefixed the same way — never a re-grade, never a flipped verdict.
66
-
67
- ## Leg 4 — outage roll-up
68
-
69
- Count the window's `panel_outcome='unavailable'` rows. One-off outages just get the `--regrade` pointer in the report; a streak (2+ consecutive, or clustered on one builder) is a health problem — hand off to **/grader-health**, which owns outage diagnosis. Do not diagnose here.
70
-
71
- ## Report shape
72
-
73
- End with the ledger, most severe first:
74
-
75
- 1. **Un-accounted overrides** — task, findings printed verbatim, queued-idea ref. A task that landed with no approved override-request row is the headline.
76
- 2. **Queued findings** — what was filed this run, what was skipped as already-queued, what was held back by the Narc filter (and why).
77
- 3. **Spot-check suspicions** — task, the runtime question the panel couldn't answer, queued-idea ref.
78
- 4. **Outage roll-up** — count + the /grader-health handoff if a streak.
79
-
80
- A clean week is four one-liners.
81
-
82
- ## Constraints
83
-
84
- - **Never confirm, never flip, never re-grade.** A verdict is advisory and never self-executing (ADR 0158 §2) — this skill's entire output is a report plus queued `idea_inbox` rows, and it adds no approval step to anyone's ship (ADR 0162, which retired review gates). If a leg tempts you to "just fix it", the fix is a queued idea or a claimed task, not an action inside the audit.
85
- - **Manual cadence.** Invoke by hand (weekly is the intended rhythm). Do not wire it to cron or depend on `.claude/scheduled-tasks/`.
86
- - **Idempotent re-runs.** The `[grade-audit] task <id>:` title marker + the inbox pre-check are what make running it twice harmless — keep both.
87
- - **Archon surfaces degrade gracefully.** Without `override_request.decide`, Leg 1 can only report "ledger unreadable at this rank" — say so rather than skipping silently.
88
-
89
- ## Files this skill touches
90
-
91
- - Reads: `GET /api/bongos/public/recent-shipped`, `GET /api/bongos/tasks/:id?include=grade`, `GET /api/bongos/override-requests?status=approved` (Archon), `GET /api/bongos/inbox`.
92
- - Writes: `idea_inbox` rows via `node scripts/gds/capture.js` (queued follow-ups only).
93
- - Never calls: any confirm, status, grade, or regrade endpoint.
14
+ **This command is an alias (task 1004471).** `/grade-audit` and `/grader-health` were merged into one skill, `/grade-sweep`. Read [`.claude/skills/grade-sweep/SKILL.md`](../grade-sweep/SKILL.md) and run its **Part B — grade audit** only (the same as `/grade-sweep audit`), with every rule that file states: describe and queue, never confirm, flip a status or re-grade.
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: grade-recover
3
3
  description: >-
4
- Diagnose why a ship parked at completed — grader outage, fabricated fail, or genuine quality fail — and run the right recovery. Adds no gate to the ship path. Triggers: "/grade-recover N", "my grade failed", "grader unavailable", or a ship that landed at completed.
4
+ Diagnose why a ship parked at completed (grader outage, fabricated fail, or a real quality fail) and run the right recovery. Triggers: "/grade-recover N", "my grade failed", "grader unavailable", or a ship that landed at completed.
5
5
  plain: >-
6
6
  Works out why a quality review held your work back, whether the reviewer broke, got it wrong, or found a real problem, and takes the right next step.
7
7
  reach-for: >-