@bongos/core 1.20.42 → 1.20.44
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.bongos-core.json +99 -74
- package/.claude/skills/backlog-review/SKILL.md +1 -0
- package/.claude/skills/blocker-review/SKILL.md +1 -1
- package/.claude/skills/bug-triage/SKILL.md +1 -0
- package/.claude/skills/builder-backup/SKILL.md +1 -0
- package/.claude/skills/builder-claim/SKILL.md +1 -1
- package/.claude/skills/builder-cost/SKILL.md +1 -0
- package/.claude/skills/builder-exit/SKILL.md +1 -0
- package/.claude/skills/builder-key/SKILL.md +1 -1
- package/.claude/skills/builder-reauth/SKILL.md +1 -1
- package/.claude/skills/builder-redteam/SKILL.md +1 -1
- package/.claude/skills/builder-sequence/SKILL.md +1 -1
- package/.claude/skills/builder-setup/SKILL.md +1 -1
- package/.claude/skills/collab-review/SKILL.md +1 -1
- package/.claude/skills/design/SKILL.md +1 -1
- package/.claude/skills/design-sync/SKILL.md +1 -0
- package/.claude/skills/feedback/SKILL.md +1 -1
- package/.claude/skills/figma-design-sync/SKILL.md +1 -0
- package/.claude/skills/goal-close/SKILL.md +184 -0
- package/.claude/skills/goal-create/SKILL.md +1 -1
- package/.claude/skills/goal-review/SKILL.md +5 -102
- package/.claude/skills/goal-uat/SKILL.md +5 -81
- package/.claude/skills/grade-audit/SKILL.md +6 -85
- package/.claude/skills/grade-recover/SKILL.md +1 -1
- package/.claude/skills/grade-sweep/SKILL.md +174 -0
- package/.claude/skills/grader-health/SKILL.md +5 -72
- package/.claude/skills/idea-triage/SKILL.md +1 -0
- package/.claude/skills/merge-mode/SKILL.md +1 -1
- package/.claude/skills/new-project/SKILL.md +1 -0
- package/.claude/skills/owner-review/SKILL.md +40 -0
- package/.claude/skills/planning-session/SKILL.md +1 -0
- package/.claude/skills/priority-session/SKILL.md +1 -0
- package/.claude/skills/read-session-export/SKILL.md +1 -1
- package/.claude/skills/recall/SKILL.md +1 -1
- package/.claude/skills/scan-before-install/SKILL.md +1 -1
- package/.claude/skills/session-handoff/SKILL.md +1 -1
- package/.claude/skills/worktree-clean/SKILL.md +1 -1
- package/docs/api/openapi.json +1 -1
- package/docs/file-map.md +12 -10
- package/docs/module-api-changelog.md +4 -0
- package/docs/modules-contract.md +3 -0
- package/docs/onboarding/slash-commands.md +17 -11
- package/docs/packs/artist.md +1 -1
- package/docs/packs/engineer.md +4 -5
- package/docs/page-readings.json +21 -21
- package/modules/copy-desk/module.json +1 -1
- package/{.claude → modules/copy-desk}/skills/tweak/SKILL.md +1 -0
- package/modules/hall-ui/public/collab.css +8 -1
- package/modules/hall-ui/public/oversight.css +0 -3
- package/modules/lifecycle/dependency-advisory.js +2 -2
- package/package-lock.json +2 -2
- package/package.json +1 -1
- package/release-notes.json +16 -0
- package/scripts/gds/fitness-ratchets.js +2 -0
- package/scripts/gds/module-artifact.js +52 -2
- package/scripts/gds/module.js +57 -4
- package/scripts/gds/skill-lint.js +32 -9
- package/src/bongos/routes/modules.js +5 -4
- package/src/module-api.js +1 -1
- package/src/module-loader/manifest-schema.js +28 -0
- package/tests/hall_lead_row_home.mjs +21 -0
- package/tests/module_manifest.mjs +22 -0
- package/tests/module_store_publish.mjs +88 -4
- package/tests/module_store_publish_route.mjs +20 -1
- package/tests/skill_grade_audit.mjs +34 -17
- package/tests/skill_grader_health.mjs +28 -10
- package/tests/skill_lint.mjs +37 -0
- package/tests/skill_menu_ruling.mjs +99 -0
|
@@ -0,0 +1,174 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: grade-sweep
|
|
3
|
+
description: >-
|
|
4
|
+
Read-only sweep over recent grades: grader health (outage runs, pass-rate drift, cost, stale model pins) then the accountability audit (override ledger, dropped findings, false passes). Queues follow-ups only. Triggers: "/grade-sweep", "/grader-health", "/grade-audit".
|
|
5
|
+
disable-model-invocation: true
|
|
6
|
+
plain: >-
|
|
7
|
+
Checks that the automatic quality reviewer is working well, then looks back over recent reviews for anything that slipped through.
|
|
8
|
+
reach-for: >-
|
|
9
|
+
Now and then, or when reviews seem to be failing strangely.
|
|
10
|
+
cost: >-
|
|
11
|
+
Free. It only files follow-up notes and never changes a review or a task. An optional live test of the reviewer costs a small amount.
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
You are sweeping the last weeks of grades. It is one skill with two parts that used to be two commands (task 1004471):
|
|
15
|
+
|
|
16
|
+
- **Part A — health** (was `/grader-health`): is the grading panel *running* well?
|
|
17
|
+
- **Part B — audit** (was `/grade-audit`): did anything *ship past* it unaccounted for?
|
|
18
|
+
|
|
19
|
+
**Which part to run.** `/grade-sweep` runs A then B — A comes first because B's outage roll-up hands streaks to A, and an outage must be ruled out before a fail is read as a quality verdict. `/grade-sweep health` (or the old `/grader-health`) runs Part A only; `/grade-sweep audit` (or the old `/grade-audit`) runs Part B only. `--probe` applies to Part A only.
|
|
20
|
+
|
|
21
|
+
Everything here **describes and queues**. You never confirm, flip a status, re-grade, or add a gate — the builder (or `/grade-recover`) acts on what you report, and Part B's only writes are queued `idea_inbox` rows.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## Part A — grader health (read-only)
|
|
26
|
+
|
|
27
|
+
You read live surfaces, compare against known baselines, and emit findings + exact remediation commands. Part A writes nothing.
|
|
28
|
+
|
|
29
|
+
### The four axes
|
|
30
|
+
|
|
31
|
+
#### 1 — Outage runs (the pattern this part exists to catch)
|
|
32
|
+
|
|
33
|
+
A grader outage used to be recorded as a 0.0 "quality failure": one builder accumulated **21 zero-score rows** over a month before anyone noticed, and the public tile showed a phantom 5-day 0%-pass crater. Since the outage-as-zero fix, aggregates exclude outage rounds and bucket them as `n_unavailable` — so the streak is now directly readable.
|
|
34
|
+
|
|
35
|
+
1. Read the public trend: `node scripts/gds/api.js GET "/api/bongos/public/grades?days=30"` — scan `trend[]` for days where the average is 0 (or null) with passes 0 and n ≥ 2, and for any nonzero `n_unavailable` streak.
|
|
36
|
+
2. (Archon) Read the per-builder split: `node scripts/gds/api.js GET "/api/bongos/grades/by-builder?days=30"` — a single builder with pass_rate 0.000 or a large `n_unavailable` bucket is the 21×0.0 shape localized to one machine or one environment.
|
|
37
|
+
3. **Confirm outage vs rejection before alarming anyone.** For each task in a suspect window: `node scripts/gds/api.js GET "/api/bongos/tasks/<id>?include=grade"` — `signals.panel_outcome='unavailable'` (with `grader_unavailable_reason` and per-worker errors in `signals.workers`) is an outage, **not a quality rejection**; a real FAIL carries findings in `issues[]`.
|
|
38
|
+
4. For every task parked at `completed` by a confirmed outage, emit the exact remediation line, one per task:
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
node scripts/gds/ship.js <id> --regrade
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
5. Check the local spawn leg: the panel spawns `claude -p` subprocesses, so verify the CLI is present and authenticated, and run `node scripts/gds/doctor.js` for the hooks/worktree leg.
|
|
45
|
+
|
|
46
|
+
#### 2 — Pass-rate drift
|
|
47
|
+
|
|
48
|
+
Baselines: the 14-agent evaluation window measured a **90.6%** pass rate, already above the 87% that ADR 0024 called "suspiciously high" and built the adversarial panel to fix. From `/api/bongos/public/grades` compute the rolling pass rate and compare:
|
|
49
|
+
|
|
50
|
+
- **Creeping up** past the baseline → grade inflation; the panel is rubber-stamping (report; candidate causes: score compression, reused grades, trivial-path overuse).
|
|
51
|
+
- **Collapsing** far below it → either a real quality regression or an unconfirmed outage window — axis 1 disambiguates.
|
|
52
|
+
|
|
53
|
+
#### 3 — Cost anomalies
|
|
54
|
+
|
|
55
|
+
Benchmark: mean panel cost **$2.14** per grade, and ~90% of it is **fixed** re-exploration cost, independent of diff size. (Measured with Quality on the haiku tier; since task 1003318 Quality runs on claude-sonnet-5, so a mean sitting somewhat above $2.14 is the new normal until re-baselined — a rise alone is not an anomaly.) Read grader spend and divide by the grade count from axis 2's window:
|
|
56
|
+
|
|
57
|
+
```
|
|
58
|
+
node scripts/gds/api.js GET "/api/bongos/public/cost-summary?source=grader&days=30" --quiet \
|
|
59
|
+
| node -e "let s='';process.stdin.on('data',d=>s+=d).on('end',()=>{const j=JSON.parse(s);
|
|
60
|
+
if(j.applied_filters?.source!=='grader'){console.error('FILTER DROPPED — do not divide');process.exit(1)}
|
|
61
|
+
console.log('grader 30d:', j.filtered_total_usd)})"
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
**Read `filtered_total_usd`, never `total_usd`.** The filters drill into the breakdown; they deliberately do *not* rescope the headline, so `total_usd` stays instance-wide (the treasury chart depends on that) and is off by more than an order of magnitude here. `by_category` / `by_builder` / `by_source` *are* filtered — under `?source=grader`, `by_builder` attributes each panel to the builder whose ship triggered it, which is why human logins appear. The `applied_filters` guard above makes a silently-dropped filter fail loudly instead of yielding a wrong mean.
|
|
65
|
+
|
|
66
|
+
- Mean drifting well above the benchmark → panel re-rolls/retries or timeout-and-retry loops (not "big diffs" — the fixed-cost share means diff size barely moves the mean).
|
|
67
|
+
- Near-zero spend with grades still recording → the trivial path or grade reuse dominating; spot-check that gating lenses are actually running.
|
|
68
|
+
|
|
69
|
+
#### 4 — Model pins
|
|
70
|
+
|
|
71
|
+
The evaluation found every grade ran on a generation-stale pin. Read the live pin from the repo: `DEFAULT_GRADER_MODEL` and `selectGraderModel` in `modules/grading/grader.js`, plus the rubric version in `modules/grading/grader-rubric.json`. Report when the pinned family is a generation behind the current Claude family, or when `selectGraderModel` would throw on a current builder model id (it throws on ids lacking opus/sonnet/haiku — a stranded-ship risk).
|
|
72
|
+
|
|
73
|
+
### Optional: `--probe` (live calibration, spends real money)
|
|
74
|
+
|
|
75
|
+
With `--probe`, run the live panel calibration: `node scripts/gds/grade-smoke.js`. It grades known-good and known-bad fixtures against the real panel and reports miscalibration. **Costs ~$0.30–0.90 per run; Metic+ only.** Never run it implicitly — only when the user passed `--probe` or asked for a live calibration.
|
|
76
|
+
|
|
77
|
+
### Related: do the workers fail differently?
|
|
78
|
+
|
|
79
|
+
Part A checks whether the grader is RUNNING well. Whether its four workers carry independent signal — or are four copies of one opinion — is a different question with its own read-only report: `node scripts/gds/grade-correlation-audit.js` (pairwise agreement + Cohen's kappa between workers, each worker's kappa against the owner's override and manual-confirm decisions, refused below a minimum sample; no model calls, Metic+ for the override read). Point the owner at it when the question is panel composition rather than health.
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
## Part B — grade audit (describe and queue)
|
|
84
|
+
|
|
85
|
+
You are auditing the last week of grades for accountability gaps: work that shipped past a failing grade unaccounted for, real findings that rode passing grades into production and evaporated, and passes the panel structurally could not have judged. Every follow-up lands as a queued idea or a handoff to another skill.
|
|
86
|
+
|
|
87
|
+
### The sweep window
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
node scripts/gds/api.js GET "/api/bongos/public/recent-shipped?days=7"
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
For each shipped task in the window, read its persisted grade:
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
node scripts/gds/api.js GET "/api/bongos/tasks/<id>?include=grade"
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Bucket by grade shape: `passed`, failed, `signals.panel_outcome='unavailable'` (outage — no quality signal existed), or no grade row at all.
|
|
100
|
+
|
|
101
|
+
**Then split off the no-artifact species before any leg reads the bucket** (task 1004011, ADR 0320). A grade with `signals.no_artifact_grade === true` judged the builder's **claim that nothing was needed** — the handoff notes and value summary against the task description — not a diff. There is no diff, no panel, and one generalist shot (`signals.lenses_dropped` names the four lenses that did not run). It is a real grade and a real gate, but it measures a different thing, so:
|
|
102
|
+
|
|
103
|
+
- **Do not average or count it alongside diff grades.** A 7.5 over three paragraphs of prose and a 7.5 over a 400-line diff are not the same number.
|
|
104
|
+
- **Leg 2 still applies, and matters more here.** A `blocker`/`major` on a no-artifact pass usually means the *evidence* was thin — queue it the same way.
|
|
105
|
+
- **Leg 3 does not apply** — there is no merged diff to re-read for runtime behaviour. The equivalent question, worth asking on any no-artifact pass that shipped on assertion alone, is: *could a reader today re-run or re-read what the notes claim?* If not, queue it.
|
|
106
|
+
- **A no-artifact FAIL is not an override candidate for Leg 1** unless it also landed. The builder's remedy is better evidence, not a permission.
|
|
107
|
+
|
|
108
|
+
Note `task_grades.grader_kind` is **not** the discriminator: the server records `'subagent'` for everything arriving through `POST /tasks/:id/grade`. `signals.no_artifact_grade` is the durable marker.
|
|
109
|
+
|
|
110
|
+
### Leg 1 — the override ledger
|
|
111
|
+
|
|
112
|
+
For every shipped task whose grade **failed** (or is absent), the land needed an override. Cross-reference the ledger (Archon — `override_request.decide`):
|
|
113
|
+
|
|
114
|
+
```bash
|
|
115
|
+
node scripts/gds/api.js GET "/api/bongos/override-requests?status=approved"
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
- Split **infra vs real** first: a fail that is actually `panel_outcome='unavailable'` (or the errored-worker zero-issue shape) is an outage artifact, not an overridden quality verdict — route those to the Leg 4 roll-up.
|
|
119
|
+
- A shipped task with a real failing grade and **no approved override-request row** landed via the raw confirm — the un-audited path. Name it in the report and print its unaddressed `issues[]` verbatim. That set is the ledger's debt.
|
|
120
|
+
- For each such task, queue the debt (see the queueing rule below) — do not chase the lander, do not un-ship anything.
|
|
121
|
+
|
|
122
|
+
### Leg 2 — dropped findings on passing grades
|
|
123
|
+
|
|
124
|
+
A pass with `blocker`/`major` entries in `issues[]` shipped real findings that no process ever picks up (the evaluation counted 274 findings on passing ships with zero conversion paths). For each:
|
|
125
|
+
|
|
126
|
+
- **Worker-attribution filter:** findings attributed to the advisory **Narc** are the known noise class — triage them by hand (read the finding, decide), never auto-trust them into the queue. Findings from the gating workers (Quality, Hacker, Efficiency) or the deterministic pre-passes are higher-confidence — queue them unless plainly stale.
|
|
127
|
+
- **Queueing rule (idempotent):** first read the open inbox (`node scripts/gds/api.js GET /api/bongos/inbox`) and skip any finding already queued — the marker is the deterministic title prefix. Then:
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
node scripts/gds/capture.js "[grade-audit] task <id>: <finding summary>" --kind bug
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
One idea per surviving finding, titled exactly `[grade-audit] task <id>: …` so a re-run of this audit files nothing twice. (The marker keeps the old command's name on purpose: rows already in the inbox carry it.)
|
|
134
|
+
|
|
135
|
+
### Leg 3 — false-pass spot-check (runtime-behavior blindness)
|
|
136
|
+
|
|
137
|
+
The re-review found 2/8 material misses, both the same root cause: **the panel grades diff text and never reasons about runtime** — a CI workflow whose default shallow checkout broke on first run scored functional_fidelity 10/10, and an SSRF guard bypassable via HTTP redirect passed at 9.5. For the riskiest ships in the window — security surfaces, CI/workflow files, the largest diffs, and 9.5s with suspiciously few findings — re-read the merged diff asking one question: *what does this code do at runtime that the diff text does not show?* (defaults it inherits, redirects it follows, first-run behavior, tool/CI defaults). A suspicion is queued via the Leg 2 rule, prefixed the same way — never a re-grade, never a flipped verdict.
|
|
138
|
+
|
|
139
|
+
### Leg 4 — outage roll-up
|
|
140
|
+
|
|
141
|
+
Count the window's `panel_outcome='unavailable'` rows. One-off outages just get the `--regrade` pointer in the report; a streak (2+ consecutive, or clustered on one builder) is a health problem — that is **Part A's axis 1** (`/grade-sweep health`), which owns outage diagnosis. Do not diagnose it inside Part B.
|
|
142
|
+
|
|
143
|
+
---
|
|
144
|
+
|
|
145
|
+
## Report shape
|
|
146
|
+
|
|
147
|
+
End with one compact list, most severe first.
|
|
148
|
+
|
|
149
|
+
**Part A:** axis, evidence (the numbers you read), verdict (outage / drift / anomaly / stale-pin / healthy), and the exact remediation command or the skill to hand off to (`/grade-recover` for a single stuck task; a Bongos task via `capture.js` for systemic fixes). A clean bill of health is one line per axis.
|
|
150
|
+
|
|
151
|
+
**Part B:**
|
|
152
|
+
|
|
153
|
+
1. **Un-accounted overrides** — task, findings printed verbatim, queued-idea ref. A task that landed with no approved override-request row is the headline.
|
|
154
|
+
2. **Queued findings** — what was filed this run, what was skipped as already-queued, what was held back by the Narc filter (and why).
|
|
155
|
+
3. **Spot-check suspicions** — task, the runtime question the panel couldn't answer, queued-idea ref.
|
|
156
|
+
4. **Outage roll-up** — count, and whether Part A found a streak.
|
|
157
|
+
|
|
158
|
+
A clean week is four one-liners.
|
|
159
|
+
|
|
160
|
+
## Constraints
|
|
161
|
+
|
|
162
|
+
- **Never confirm, never flip, never re-grade.** A verdict is advisory and never self-executing (ADR 0158 §2) — this skill's entire output is a report plus queued `idea_inbox` rows, and it adds no approval step to anyone's ship (ADR 0162, which retired review gates). If a leg tempts you to "just fix it", the fix is a queued idea or a claimed task, not an action inside the sweep. The remediation lines are FOR the builder to run.
|
|
163
|
+
- **Confirm before alarming.** Axis 1's outage-vs-rejection check comes before any "the grader is down" conclusion — a real FAIL streak on one surface is signal, not outage.
|
|
164
|
+
- **`--probe` spends money** — explicit opt-in only, Metic+.
|
|
165
|
+
- **Manual cadence.** Invoke by hand (weekly is the intended rhythm). Do not wire it to cron or depend on `.claude/scheduled-tasks/`. It is hidden from the model's menu (`disable-model-invocation`), so it runs only when someone types it.
|
|
166
|
+
- **Idempotent re-runs.** The `[grade-audit] task <id>:` title marker + the inbox pre-check are what make running it twice harmless — keep both.
|
|
167
|
+
- **Archon surfaces degrade gracefully.** Without `override_request.decide`, Leg 1 can only report "ledger unreadable at this rank"; without the by-builder read, axis 1 step 2 says so — never skip silently.
|
|
168
|
+
|
|
169
|
+
## Files this skill touches
|
|
170
|
+
|
|
171
|
+
- Reads: `GET /api/bongos/public/grades`, `GET /api/bongos/grades/by-builder` (Archon), `GET /api/bongos/public/cost-summary?source=grader` (the `filtered_total_usd` key), `GET /api/bongos/public/recent-shipped`, `GET /api/bongos/tasks/:id?include=grade`, `GET /api/bongos/override-requests?status=approved` (Archon), `GET /api/bongos/inbox`, `modules/grading/grader.js`, `modules/grading/grader-rubric.json`.
|
|
172
|
+
- Runs (read-only): `node scripts/gds/doctor.js`; with `--probe` only: `node scripts/gds/grade-smoke.js`.
|
|
173
|
+
- Writes: `idea_inbox` rows via `node scripts/gds/capture.js` (Part B's queued follow-ups only).
|
|
174
|
+
- Never calls: any confirm, status, grade, or regrade endpoint.
|
|
@@ -1,81 +1,14 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: grader-health
|
|
3
3
|
description: >-
|
|
4
|
-
|
|
4
|
+
Old name for /grade-sweep health: the read-only grader health check (outage runs, pass-rate drift, cost, stale model pins). --probe runs a paid calibration. Triggers: "/grader-health".
|
|
5
|
+
disable-model-invocation: true
|
|
5
6
|
plain: >-
|
|
6
|
-
|
|
7
|
+
The older name for the check that the automatic quality reviewer is working well.
|
|
7
8
|
reach-for: >-
|
|
8
|
-
When
|
|
9
|
+
When you remember the old name; it opens the same check as the grade sweep.
|
|
9
10
|
cost: >-
|
|
10
11
|
Free and read-only. An optional live test of the reviewer costs a small amount.
|
|
11
12
|
---
|
|
12
13
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
## The four axes
|
|
16
|
-
|
|
17
|
-
### 1 — Outage runs (the pattern this skill exists to catch)
|
|
18
|
-
|
|
19
|
-
A grader outage used to be recorded as a 0.0 "quality failure": one builder accumulated **21 zero-score rows** over a month before anyone noticed, and the public tile showed a phantom 5-day 0%-pass crater. Since the outage-as-zero fix, aggregates exclude outage rounds and bucket them as `n_unavailable` — so the streak is now directly readable.
|
|
20
|
-
|
|
21
|
-
1. Read the public trend: `node scripts/gds/api.js GET "/api/bongos/public/grades?days=30"` — scan `trend[]` for days where the average is 0 (or null) with passes 0 and n ≥ 2, and for any nonzero `n_unavailable` streak.
|
|
22
|
-
2. (Archon) Read the per-builder split: `node scripts/gds/api.js GET "/api/bongos/grades/by-builder?days=30"` — a single builder with pass_rate 0.000 or a large `n_unavailable` bucket is the 21×0.0 shape localized to one machine or one environment.
|
|
23
|
-
3. **Confirm outage vs rejection before alarming anyone.** For each task in a suspect window: `node scripts/gds/api.js GET "/api/bongos/tasks/<id>?include=grade"` — `signals.panel_outcome='unavailable'` (with `grader_unavailable_reason` and per-worker errors in `signals.workers`) is an outage, **not a quality rejection**; a real FAIL carries findings in `issues[]`.
|
|
24
|
-
4. For every task parked at `completed` by a confirmed outage, emit the exact remediation line, one per task:
|
|
25
|
-
|
|
26
|
-
```bash
|
|
27
|
-
node scripts/gds/ship.js <id> --regrade
|
|
28
|
-
```
|
|
29
|
-
|
|
30
|
-
5. Check the local spawn leg: the panel spawns `claude -p` subprocesses, so verify the CLI is present and authenticated, and run `node scripts/gds/doctor.js` for the hooks/worktree leg.
|
|
31
|
-
|
|
32
|
-
### 2 — Pass-rate drift
|
|
33
|
-
|
|
34
|
-
Baselines: the 14-agent evaluation window measured a **90.6%** pass rate, already above the 87% that ADR 0024 called "suspiciously high" and built the adversarial panel to fix. From `/api/bongos/public/grades` compute the rolling pass rate and compare:
|
|
35
|
-
|
|
36
|
-
- **Creeping up** past the baseline → grade inflation; the panel is rubber-stamping (report; candidate causes: score compression, reused grades, trivial-path overuse).
|
|
37
|
-
- **Collapsing** far below it → either a real quality regression or an unconfirmed outage window — axis 1 disambiguates.
|
|
38
|
-
|
|
39
|
-
### 3 — Cost anomalies
|
|
40
|
-
|
|
41
|
-
Benchmark: mean panel cost **$2.14** per grade, and ~90% of it is **fixed** re-exploration cost, independent of diff size. (Measured with Quality on the haiku tier; since task 1003318 Quality runs on claude-sonnet-5, so a mean sitting somewhat above $2.14 is the new normal until re-baselined — a rise alone is not an anomaly.) Read grader spend and divide by the grade count from axis 2's window:
|
|
42
|
-
|
|
43
|
-
```
|
|
44
|
-
node scripts/gds/api.js GET "/api/bongos/public/cost-summary?source=grader&days=30" --quiet \
|
|
45
|
-
| node -e "let s='';process.stdin.on('data',d=>s+=d).on('end',()=>{const j=JSON.parse(s);
|
|
46
|
-
if(j.applied_filters?.source!=='grader'){console.error('FILTER DROPPED — do not divide');process.exit(1)}
|
|
47
|
-
console.log('grader 30d:', j.filtered_total_usd)})"
|
|
48
|
-
```
|
|
49
|
-
|
|
50
|
-
**Read `filtered_total_usd`, never `total_usd`.** The filters drill into the breakdown; they deliberately do *not* rescope the headline, so `total_usd` stays instance-wide (the treasury chart depends on that) and is off by more than an order of magnitude here. `by_category` / `by_builder` / `by_source` *are* filtered — under `?source=grader`, `by_builder` attributes each panel to the builder whose ship triggered it, which is why human logins appear. The `applied_filters` guard above makes a silently-dropped filter fail loudly instead of yielding a wrong mean.
|
|
51
|
-
|
|
52
|
-
- Mean drifting well above the benchmark → panel re-rolls/retries or timeout-and-retry loops (not "big diffs" — the fixed-cost share means diff size barely moves the mean).
|
|
53
|
-
- Near-zero spend with grades still recording → the trivial path or grade reuse dominating; spot-check that gating lenses are actually running.
|
|
54
|
-
|
|
55
|
-
### 4 — Model pins
|
|
56
|
-
|
|
57
|
-
The evaluation found every grade ran on a generation-stale pin. Read the live pin from the repo: `DEFAULT_GRADER_MODEL` and `selectGraderModel` in `modules/grading/grader.js`, plus the rubric version in `modules/grading/grader-rubric.json`. Report when the pinned family is a generation behind the current Claude family, or when `selectGraderModel` would throw on a current builder model id (it throws on ids lacking opus/sonnet/haiku — a stranded-ship risk).
|
|
58
|
-
|
|
59
|
-
## Optional: `--probe` (live calibration, spends real money)
|
|
60
|
-
|
|
61
|
-
With `--probe`, run the live panel calibration: `node scripts/gds/grade-smoke.js`. It grades known-good and known-bad fixtures against the real panel and reports miscalibration. **Costs ~$0.30–0.90 per run; Metic+ only.** Never run it implicitly — only when the user passed `--probe` or asked for a live calibration.
|
|
62
|
-
|
|
63
|
-
## Related: do the workers fail differently?
|
|
64
|
-
|
|
65
|
-
This skill checks whether the grader is RUNNING well. Whether its four workers carry independent signal — or are four copies of one opinion — is a different question with its own read-only report: `node scripts/gds/grade-correlation-audit.js` (pairwise agreement + Cohen's kappa between workers, each worker's kappa against the owner's override and manual-confirm decisions, refused below a minimum sample; no model calls, Metic+ for the override read). Point the owner at it when the question is panel composition rather than health.
|
|
66
|
-
|
|
67
|
-
## Report shape
|
|
68
|
-
|
|
69
|
-
End with a compact findings list, most severe first: axis, evidence (the numbers you read), verdict (outage / drift / anomaly / stale-pin / healthy), and the exact remediation command or the skill to hand off to (`/grade-recover` for a single stuck task; a Bongos task via `capture.js` for systemic fixes). A clean bill of health is one line per axis.
|
|
70
|
-
|
|
71
|
-
## Constraints
|
|
72
|
-
|
|
73
|
-
- **Read-only, describe-only (ADR 0158 §2 — a verdict never executes itself).** No confirms, no regrades run by this skill, no status flips, and no new approval step on anyone's ship (ADR 0162). The remediation lines are FOR the builder to run.
|
|
74
|
-
- **Confirm before alarming.** Axis 1's outage-vs-rejection check comes before any "the grader is down" conclusion — a real FAIL streak on one surface is signal, not outage.
|
|
75
|
-
- **`--probe` spends money** — explicit opt-in only, Metic+.
|
|
76
|
-
|
|
77
|
-
## Files this skill touches
|
|
78
|
-
|
|
79
|
-
- Reads: `GET /api/bongos/public/grades`, `GET /api/bongos/grades/by-builder` (Archon), `GET /api/bongos/public/cost-summary?source=grader` (the `filtered_total_usd` key), `GET /api/bongos/tasks/:id?include=grade`, `modules/grading/grader.js`, `modules/grading/grader-rubric.json`.
|
|
80
|
-
- Runs (read-only): `node scripts/gds/doctor.js`; with `--probe` only: `node scripts/gds/grade-smoke.js`.
|
|
81
|
-
- Writes: nothing.
|
|
14
|
+
**This command is an alias (task 1004471).** `/grader-health` and `/grade-audit` were merged into one skill, `/grade-sweep`. Read [`.claude/skills/grade-sweep/SKILL.md`](../grade-sweep/SKILL.md) and run its **Part A — grader health** only (the same as `/grade-sweep health`; pass `--probe` through if given), with every rule that file states: read-only, confirm an outage before alarming anyone.
|
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
name: idea-triage
|
|
3
3
|
description: >-
|
|
4
4
|
Daily walk of the idea_inbox, which holds only HOMELESS work — anything naming a goal became a task at filing time. Promote, discard or merge each open idea. Metic+ only. Triggers: "/idea-triage", "triage ideas", "review the inbox", "walk the idea inbox", or a scheduled daily run.
|
|
5
|
+
disable-model-invocation: true
|
|
5
6
|
plain: >-
|
|
6
7
|
Goes through the ideas that have no home yet and decides what happens to each: turn it into work, drop it, or merge it with another.
|
|
7
8
|
reach-for: >-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: merge-mode
|
|
3
3
|
description: >-
|
|
4
|
-
Manual fallback for the
|
|
4
|
+
Manual fallback for the /builder-ship auto-merge: walks each confirmed task through merge, smoke and deploy. Triggers: "/merge-mode", "merge the queue", "land confirmed tasks", "land the queue", or a /builder-ship that reported the auto-merge bailed.
|
|
5
5
|
requires: [main-checkout, push-credential, gh, droplet-ssh]
|
|
6
6
|
plain: >-
|
|
7
7
|
The manual backup for finishing hand-ins: it takes approved work that did not merge by itself and puts it live.
|
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
name: new-project
|
|
3
3
|
description: >-
|
|
4
4
|
Runbook from zero to a live, owned STANDALONE instance: name, address, scaffold, provision, OAuth app, first sign-in, verify, brand. Triggers: "/new-project", "stand up a new instance", "start a new project".
|
|
5
|
+
disable-model-invocation: true
|
|
5
6
|
plain: >-
|
|
6
7
|
Takes you step by step from nothing to a brand-new project website of your own.
|
|
7
8
|
reach-for: >-
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: owner-review
|
|
3
|
+
description: >-
|
|
4
|
+
The owner's review in one walk: backlog (what needs a go), bugs (dedupe), ideas (promote, drop, merge), then priorities. Metic+. Triggers: "/owner-review", "review the backlog", "triage bugs", "triage ideas", "review the inbox", "reprioritize".
|
|
5
|
+
plain: >-
|
|
6
|
+
One place to go through everything waiting on you: work that needs a go-ahead, reported problems, new ideas, and what matters most next.
|
|
7
|
+
reach-for: >-
|
|
8
|
+
Once a day, or whenever you want to clear what is waiting on your decision.
|
|
9
|
+
cost: >-
|
|
10
|
+
Uses your session. Nothing changes until you decide each item.
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
You are running the **owner review**: the one front door to the four queues that wait on a person's decision (task 1004471). It walks them in this order, because each step makes the next one smaller:
|
|
14
|
+
|
|
15
|
+
| Step | Queue | The step's own playbook | Old command (still typeable) |
|
|
16
|
+
|---|---|---|---|
|
|
17
|
+
| 1. **backlog** | tasks at `status='backlog'` — waiting for a person to say go | [`.claude/skills/backlog-review/SKILL.md`](../backlog-review/SKILL.md) | `/backlog-review` |
|
|
18
|
+
| 2. **bugs** | open `kind=bug` tasks — batch the near-duplicates, then merge or won't-fix | [`.claude/skills/bug-triage/SKILL.md`](../bug-triage/SKILL.md) | `/bug-triage` |
|
|
19
|
+
| 3. **ideas** | the `idea_inbox` — promote, discard or merge each homeless idea | [`.claude/skills/idea-triage/SKILL.md`](../idea-triage/SKILL.md) | `/idea-triage` |
|
|
20
|
+
| 4. **priorities** | a few plain questions that reweight what is left and suggest what to claim next | [`.claude/skills/priority-session/SKILL.md`](../priority-session/SKILL.md) | `/priority-session` |
|
|
21
|
+
|
|
22
|
+
## Where to start
|
|
23
|
+
|
|
24
|
+
- `/owner-review` with nothing else → all four, in order.
|
|
25
|
+
- `/owner-review <step>` (`backlog`, `bugs`, `ideas` or `priorities`) → that step only.
|
|
26
|
+
- **Plain words open the matching step directly — do not walk the earlier ones first:** "review the backlog", "what needs a nod", "walk the backlog" → **backlog**; "triage bugs", "walk the bug queue", "dedupe bugs" → **bugs**; "triage ideas", "review the inbox", "walk the idea inbox" → **ideas**; "run a priority session", "reprioritize the inbox" → **priorities**. When one step finishes, offer the next in one line ("Next is bugs — go on?"); stop if they say no.
|
|
27
|
+
|
|
28
|
+
## How to run a step
|
|
29
|
+
|
|
30
|
+
The four step playbooks are **hidden** skills (`disable-model-invocation: true`): a person can still type their old `/name`, but you cannot load them through the Skill tool. So for each step, **Read that step's `SKILL.md` with the Read tool and follow it in full** — its rank gate, its queue read, its one-decision-at-a-time walk, its write rules and its summary. This file adds no rule of its own to any step and overrides none; where they differ, the step's playbook wins.
|
|
31
|
+
|
|
32
|
+
**The rank gate runs once.** All four steps are Metic+. Check it at the top — `node scripts/gds/api.js GET /api/bongos/me`, lowercase `builder.rank` — and if it is `xenos` or `thetes`, stop before reading any queue: *"The owner review is a Metic+ step — it decides what the whole team works on. Ask an Archon to promote you."* An absent `rank` follows the pre-rank passthrough each step describes.
|
|
33
|
+
|
|
34
|
+
## Between steps
|
|
35
|
+
|
|
36
|
+
Give a two-line tally of the step just finished (decided / deferred) before offering the next one. At the end, one summary per step that ran, in that step's own summary shape — nothing re-derived.
|
|
37
|
+
|
|
38
|
+
## What stays separate
|
|
39
|
+
|
|
40
|
+
Blockers have their own daily walk, `/blocker-review`, and one blocker at a time is `/blocker-solve`. Closing a goal's criteria is `/goal-close`. The unattended nightly twin of step 3 (`.claude/scheduled-tasks/idea-triage-nightly/`) reads the idea step's playbook by its path, which is why that file stays where it is.
|
|
@@ -3,6 +3,7 @@ name: planning-session
|
|
|
3
3
|
description: >-
|
|
4
4
|
Structured planning for a version — scope criteria, a written spec, an owner interview on every non-obvious decision, then the seeded task list. Metic+ only. Triggers: "/planning-session", "open a planning session", "let's plan V[N]", "plan the next version", "scope a new version".
|
|
5
5
|
requires: [droplet-ssh]
|
|
6
|
+
disable-model-invocation: true
|
|
6
7
|
plain: >-
|
|
7
8
|
A structured planning conversation for the next version: what it must achieve, a written plan, and the owner's call on every open question.
|
|
8
9
|
reach-for: >-
|
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
name: priority-session
|
|
3
3
|
description: >-
|
|
4
4
|
A trusted builder answers a few plain questions; the answers reweight the idea_inbox and suggest what to claim next. Metic+ only. Triggers: "/priority-session", "run a priority session", "what should I work on next", "reprioritize the inbox".
|
|
5
|
+
disable-model-invocation: true
|
|
5
6
|
plain: >-
|
|
6
7
|
Asks you a few simple questions about what matters most, then re-ranks the ideas and suggests what to work on next.
|
|
7
8
|
reach-for: >-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: read-session-export
|
|
3
3
|
description: >-
|
|
4
|
-
Read a Claude Code /export zip
|
|
4
|
+
Read a Claude Code /export zip to answer questions about a past session; never creates one. Triggers: a dropped session-export-*.zip, "read/summarize this session", "what happened in this session", "what was I/you thinking", "/read-session-export".
|
|
5
5
|
plain: >-
|
|
6
6
|
Reads a saved copy of a past conversation with your assistant and answers questions about what happened in it.
|
|
7
7
|
reach-for: >-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: recall
|
|
3
3
|
description: >-
|
|
4
|
-
|
|
4
|
+
Read-only, rank-scoped search over repo docs and DB prose; use it instead of grepping. Triggers: "/recall", "what do we know about X", "have we done X before", "did a past session hit X", "search the docs/ADRs for X", "is there an ADR about X".
|
|
5
5
|
plain: >-
|
|
6
6
|
Searches everything the project already knows, its documents and past decisions, in one go.
|
|
7
7
|
reach-for: >-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: scan-before-install
|
|
3
3
|
description: >-
|
|
4
|
-
Vet a third-party GitHub repo or Claude plugin BEFORE installing: quarantined fetch, deterministic floor, subagent panel
|
|
4
|
+
Vet a third-party GitHub repo or Claude plugin BEFORE installing: quarantined fetch, deterministic floor, subagent panel, verdict. Never installs. Triggers: "/scan-before-install", "is this plugin safe to install", "vet this GitHub repo".
|
|
5
5
|
plain: >-
|
|
6
6
|
Checks an outside add-on or code project for safety problems before you install it.
|
|
7
7
|
reach-for: >-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: session-handoff
|
|
3
3
|
description: >-
|
|
4
|
-
Emit a paste-ready next-steps prompt so a FRESH session starts with curated context.
|
|
4
|
+
Emit a paste-ready next-steps prompt so a FRESH session starts with curated context. Triggers: "/session-handoff", "hand off to a fresh session", "give me a handoff prompt", "wrap up context for next time", "continue this in a new session".
|
|
5
5
|
plain: >-
|
|
6
6
|
Writes a ready-to-paste note so a fresh conversation can pick up exactly where this one left off.
|
|
7
7
|
reach-for: >-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: worktree-clean
|
|
3
3
|
description: >-
|
|
4
|
-
Remove per-claim worktrees whose work landed, and sweep the husks a locked remove leaves, via junction-safe worktree.js
|
|
4
|
+
Remove per-claim worktrees whose work landed, and sweep the husks a locked remove leaves, via junction-safe worktree.js, never raw git worktree remove. Triggers: "/worktree-clean", "prune stale worktrees", "worktree husks".
|
|
5
5
|
plain: >-
|
|
6
6
|
Tidies away old workspace folders on your computer whose work has already been merged.
|
|
7
7
|
reach-for: >-
|
package/docs/api/openapi.json
CHANGED
|
@@ -19282,7 +19282,7 @@
|
|
|
19282
19282
|
"store"
|
|
19283
19283
|
],
|
|
19284
19284
|
"summary": "POST /store/modules/:key/versions",
|
|
19285
|
-
"description": "POST /api/bongos/store/modules/:key/versions — publish one module version to the store (task 1004271, ADR 0338 D1). Body: the gzip tarball `bongos module publish` builds (scripts/gds/module-artifact.js), sent as application/gzip. The route trusts nothing but the bytes: it recomputes every hash, re-runs the publish denylist
|
|
19285
|
+
"description": "POST /api/bongos/store/modules/:key/versions — publish one module version to the store (task 1004271, ADR 0338 D1). Body: the gzip tarball `bongos module publish` builds (scripts/gds/module-artifact.js), sent as application/gzip. The route trusts nothing but the bytes: it recomputes every hash, re-runs the publish denylist, validates module.json and checks that HOWTO.md has its required sections (ADR 0347 D4) itself (verifyModuleArtifact), then keeps the tarball on the control plane's disk and INSERTs the version row in one transaction (module-store.js). A key's first publish makes the caller its author; after that only the author may publish, a delisted key takes nothing new, and a version is published once and never changed. rank: metic+archon — it puts code in a shared store other instances install from. Gate: requirePermission('module.submit') — the atom that already guards filing a module into a shared queue; a dedicated `module.publish` atom is a follow-up once who-may-sell is decided.\n\n**Rank:** `metic+archon` — Metic or Archon rank (review/triage powers).\n\n**Permissions:** `module.submit` (all required).",
|
|
19286
19286
|
"x-rank": "metic+archon",
|
|
19287
19287
|
"x-source": "src/bongos/routes/modules.js",
|
|
19288
19288
|
"x-permissions": [
|
package/docs/file-map.md
CHANGED
|
@@ -113,7 +113,7 @@
|
|
|
113
113
|
│ ├── status-ui/ ← CORE-DOMAIN WEB-SURFACE carve (BV1.R84, default-on): the public status dashboard at status.<apex>, moved from public-status/. NO routes/db/port — declares contributes.webSurfaces:[{host:"status.",dir:"public"}], served by the loader web-surface seam (moduleWebSurfaces → serve-internal). CLAUDE.md, module.json, public/{index.html,status.js,style.css,cursors/,snapshots/}.
|
|
114
114
|
│ ├── hall-ui/ ← CORE-DOMAIN WEB-SURFACE carve (BV1.R85, default-on): the authenticated builders' hall at builders.<apex>, moved from public-builders/ (~13k LOC, the largest). NO routes/db/port — declares contributes.webSurfaces:[{host:"builders.",dir:"public"}] for discovery, but UNLIKE status-ui keeps its DEDICATED serving block in serve-internal (the /builders shim + canonical redirects + citizen/Archon page gates + asset-stamping + hall-widget injection), partitioned out of the generic web-surface loop and served from BUILDERS_DIR. CLAUDE.md, module.json, records/ (one file per page family — what each family's redesign settled; CLAUDE.md points at them, and a new record is a new file there so concurrent hall tasks stop colliding on one insertion point, task 1003499), public/{index.html,builders.js,style.css,work,watch,gate,settings,ranks,sessions,primer,diagrams,atlas,...}.{html,js},cursors/,chime.wav. Four ARCHETYPE sheets sit beside the per-page ones: `record.css` (the task + idea records, task 1003305), `room.css` (the READING ROOMS — the primer + the diagrams index, task 1003306: the reading measure and rhythm, the table of contents, the progress hairline and the framed figure diagram-viewer.js upgrades on both pages), `oversight.css` (the OVERSIGHT family — government, watch, gate, task 1003469 (harbor was its fourth until task 1003890 removed the dev box): the row, the figure band inside a panel, the group head, the inline form, the grant toggle and the shared confirm modal; named oversight because the vocabulary wall bans "governance" on the government surface) and `panel.css` (the PANEL family — settings, profile, task 1003308 (pair was its third until task 1003890 removed the dev box): the row, the switch, the field with its label above, the chip control, the four voices and the sign-in gate; the hall's my-projects page was DELETED in the same task — "my projects" is the hub's user-level page, and the hall keeps no link or redirect; the family's fourth page `not-ready.html` was deleted by task 1003450, along with the emptied `requireNonXenosPage` gate that was the only thing redirecting to it). A page sheet then holds only what one page has and the others do not — `primer.css` the acknowledge gate, `diagrams.css` the legend strip, `gate.css` the hold reasons and file chips, `watch.css` the severities, the report reader and the grade trend, `settings.css` the rail layout, the group head, the token ladder and the voices grid, `profile.css` the identity band and the shelf, `roster.css` the Builders roster's identity cell (was `people.css`, shared with the cross-project `/people` directory until task 1003444 deleted that page — `/people` and its legacy door `/scouting` now 302 to `/roster`), `thinking.css` THE SKY's one sheet (task 1004231 / sky part 3 — `thinking.html` at `/thinking` is the prototype the owner adopted as the design of record, `docs/design/mocks/sky/`, built over `GET /sky` + `GET /sky/mine`; its script is the seven `sky-*.js` files on one shared `window.OTBSky`; it replaced THE ROOM of task 1004098, which survives as the list drawer, and the old `/sky` page, whose URL 301s here; `idea-objects.js` beside it is the shared five-object vocabulary).
|
|
115
115
|
│ ├── lifecycle/ ← CORE-DOMAIN carve (BV1.R86, default-on): the build state machine — tasks · claims · versions · done-when · goals · dependencies · claim→ship→grade→credit · github-push + merge-lock + conflict-resolve + publish-reconciler. The LAST, most-coupled carve; registers the `lifecycle` kernel port (createTask, classifyTaskKind, tallyPeerVotes, work-tracking reads). CLAUDE.md, module.json, routes/{tasks,claims,versions,done-when,goals,dependencies,gate-approvals,analytics}.js, lifecycle.js, github-push.js, merge-lock.js, conflict-resolve.js, publish-reconciler.js, task-visuals.js + task-visual-db.js + routes/visuals.js (the optional ship-time visual, task 1003109 — the route file is NOT named task-*.js on purpose; see the module CLAUDE.md).
|
|
116
|
-
│ ├── copy-desk/ ← CORE-DOMAIN module (task 1003113 / R02 of goal 1000074, default-on): the COPY half of the design contract — ADR 0081 gave design TOKENS a repo-owned source of truth, text had none. Owns the FLAG (anyone in any role marks a string or a whole surface as needing an artist's attention, with a reason required in BOTH the route and a CHECK constraint) and the QUEUE an artist reads it from. READS R01's committed registry (docs/copy-registry.json, generated by scripts/gds/copy-inventory.js) and never writes it. NO LIVE CMS is the hard non-goal: no column here could hold replacement text, and copy_no_cms.mjs pins that. Key is `copy-desk`, not `copy` — the branding contract already owns a `copy` block, so the bare word reds the "core carries no module code" fitness check. R03 (task 1003114, ADR 0233) added the PROPOSAL: an artist writes the replacement wording and it becomes a Bongos task carrying a fenced copy-proposal patch — no row, NO second migration (a proposals table is the forbidden column with extra steps), and scripts/gds/copy-apply.js lands it as a real diff under a claim. BV2.TW05 (task 1004316, ADR 0341 D8) added Tweak Mode's three derived READS — GET /copy-desk/pages, /copy-desk/pages/:pageId, /copy-desk/tally: page status and count, changelog, drift, "N of M tweaked" per surface, and the artist's tally — over the page-tweak tasks (lifecycle port), docs/page-inventory.json and docs/page-readings.json (both now in the publish manifest), with no status table. BV2.TW06 (task 1004317, ADR 0341 D2, D5, D6) added its first two page WRITES: POST /copy-desk/pages/:pageId/claim (an artist-craft builder opens or takes the page's tweak round as a web-claimed task, through the lifecycle port's webClaimPageTweak, under an advisory lock on the page) and POST /copy-desk/pages/:pageId/asks (anyone files a page ask, a copy_desk_flags row with scope 'page', added by copy_desk_002_page_asks.sql). CLAUDE.md, module.json, registry.js, flags.js, proposals.js, pages.js (the page-tweak block format), page-status.js (the pure derivation), page-data.js (the page artifacts' read side), routes/copy-desk.js, migrations/copy_desk_001_flags.sql, migrations/copy_desk_002_page_asks.sql, tests/{copy_flags,copy_queue,copy_no_cms,copy_proposals,copy_page_status}.mjs + tests/fixtures/page-tweak.cjs. Its hall surface lives with hall-ui (public/copy-desk.{html,js,css}), per the one-key-per-web-surface precedent.
|
|
116
|
+
│ ├── copy-desk/ ← CORE-DOMAIN module (task 1003113 / R02 of goal 1000074, default-on): the COPY half of the design contract — ADR 0081 gave design TOKENS a repo-owned source of truth, text had none. Owns the FLAG (anyone in any role marks a string or a whole surface as needing an artist's attention, with a reason required in BOTH the route and a CHECK constraint) and the QUEUE an artist reads it from. READS R01's committed registry (docs/copy-registry.json, generated by scripts/gds/copy-inventory.js) and never writes it. NO LIVE CMS is the hard non-goal: no column here could hold replacement text, and copy_no_cms.mjs pins that. Key is `copy-desk`, not `copy` — the branding contract already owns a `copy` block, so the bare word reds the "core carries no module code" fitness check. R03 (task 1003114, ADR 0233) added the PROPOSAL: an artist writes the replacement wording and it becomes a Bongos task carrying a fenced copy-proposal patch — no row, NO second migration (a proposals table is the forbidden column with extra steps), and scripts/gds/copy-apply.js lands it as a real diff under a claim. BV2.TW05 (task 1004316, ADR 0341 D8) added Tweak Mode's three derived READS — GET /copy-desk/pages, /copy-desk/pages/:pageId, /copy-desk/tally: page status and count, changelog, drift, "N of M tweaked" per surface, and the artist's tally — over the page-tweak tasks (lifecycle port), docs/page-inventory.json and docs/page-readings.json (both now in the publish manifest), with no status table. BV2.TW06 (task 1004317, ADR 0341 D2, D5, D6) added its first two page WRITES: POST /copy-desk/pages/:pageId/claim (an artist-craft builder opens or takes the page's tweak round as a web-claimed task, through the lifecycle port's webClaimPageTweak, under an advisory lock on the page) and POST /copy-desk/pages/:pageId/asks (anyone files a page ask, a copy_desk_flags row with scope 'page', added by copy_desk_002_page_asks.sql). CLAUDE.md, module.json, registry.js, flags.js, proposals.js, pages.js (the page-tweak block format), page-status.js (the pure derivation), page-data.js (the page artifacts' read side), routes/copy-desk.js, migrations/copy_desk_001_flags.sql, migrations/copy_desk_002_page_asks.sql, tests/{copy_flags,copy_queue,copy_no_cms,copy_proposals,copy_page_status}.mjs + tests/fixtures/page-tweak.cjs. Its hall surface lives with hall-ui (public/copy-desk.{html,js,css}), per the one-key-per-web-surface precedent. skills/tweak/SKILL.md is the /tweak skill (contributes.skills; moved out of .claude/skills/ in task 1004471, hidden from the model listing, typed as /tweak).
|
|
117
117
|
│ ├── npm-release/ ← the deploy page for a project that publishes an npm package from GitHub (task 1002622, default-OFF — on only in the Bongos hall, which publishes @bongos/core). "Where the work is": every task finished in the last 14 days, placed in the stage it has actually reached (not merged · merged, no version yet · published, not running here · running here, not released · released · in no version), plus the version history — each version's tasks. The version↔task join reads the release-notes.json the package itself carries (scripts/gds/release-notes.js writes it at pack time), streamed out of the NEWEST version's npm tarball in-process (no child process, no sync call) when the running core's own copy stops short. module.json, work.js (the reading + the stage placement — the packument, tarball stream and tar reader it uses are the core's src/bongos/package-registry.js, reached through the doorway's readPackageRegistry since task 1004296), routes/work.js (GET /npm-release/work, gated core.pin.move like the page), public/work.{js,css} (the two sections /deploy lends it: #nr-work, #nr-versions). Task 1004301 added the same line on a task's own page: routes/task-where.js (GET /npm-release/task/:id, any signed-in builder — it shows nothing not already public) + public/task-widget.js, injected into task.html by the loader's uiSections seam (hallWidgetScripts file task-widget.js) to fill #task-contrib on otb:task-shown. Tests sit at the repo root — tests/npm_release_{work,page}.mjs — because a default-off module's own tests/ is skipped by the unit runner.
|
|
118
118
|
│ ├── render-deploy/ ← the deploy page for a project whose app deploys through Render (task 1004270, default-OFF). The service's Render deploys, each tied to its commit and the Bongos tasks in it, what is live now, and deploy + roll back behind a two-step confirm the server also holds. The owner's Render key is never stored: it rides each request in a header. Uses the one Render client, scripts/gds/render-api.js, and its ownerId guard. `deploys.js` (shaping + ledger), `render.js` (key + client), `routes/` (app, history, act), `public/deploy.js` (the page, mounted in /deploy), `migrations/` (one two-id table)
|
|
119
119
|
│ ├── ui-design/ ← CORE-DOMAIN module (task 1003324 / ADR 0197, default-on): the design capability every builder on every instance gets — ships METHOD, never a world. Absorbs ADR 0081's layer: adapters/{claude-design,figma}/index.js (moved from src/ui/adapters/), scripts/{design-sync,figma-design-sync,validate-design}.js (moved from scripts/gds/, which keeps thin shims), docs/design-contract.md (moved from docs/). Declaration-only: contributes.skills design (the /design playbook, new) + design-sync + figma-design-sync; no routes/port/migrations, so disabling it touches no served surface. CLAUDE.md separates the platform FLOORS (the Fifteen Rule, the dark twins, no page :root, the 24px floor, AA in both modes, the kill switch) from the instance's WORLD (its packs + DESIGN.md). kit/ (task 1003320) is the look-before-you-ship kit: lib.js (the floors, the states/probes contracts, the action + expectation vocabularies, the WCAG maths, the stub spawner), serve.js (any surface through the doorway's real transforms + fixtures/), render.js (states × 1440/390/320 × dark/light, audited), probe.js (PASS/FAIL contracts), check-mock.js (platform rules + the instance's DESIGN.md taste bans, then the detector), tells.js (task 1003329 / ADR 0223: the anti-pattern detector as two tiers — floors FAIL, craft tells WARN, kit-ignore waivers per file; the rendered twins live in lib.js factsInPage), contrast.js + png.js; a page's <page>.states.json / <page>.probes.json sit BESIDE the surface (modules/public-landing/public/, modules/hall-ui/public/). Recipe: docs/recipes/ui-look-before-you-ship.md. The thirteen design-style skills that lived in skills/ are the opt-in design-styles module since task 1004470 (below); design-taste-frontend-v1 was deleted. styles/ (task 1003323 / ADR 0219) is the style library: <name>/pack.json (the theme.ui overlay, exactly the fifteen) + DESIGN.md (the look's derived recipes, thirteen in lockstep with the pack) + mock.html (the specimen); chrome-world (pinned to the neutral pack), expedition, grove, blueprint (the first cool look, authored through /style at task 1003475); STYLE=<name> on the kit lays a look over the instance's pack through GDS_BRANDING_FILE; tests/ui_design_styles.mjs runs the world/cosmos/hall token suites once per alternate (UI_DESIGN_PACK). config/design-tokens.* + design-sources.json STAY in config/ as the host boundary; public-landing keeps its own key (first customer, not owner).
|
|
@@ -364,11 +364,11 @@ tests/
|
|
|
364
364
|
```
|
|
365
365
|
.claude/skills/
|
|
366
366
|
├── ask-for-help/SKILL.md ← /ask-for-help — turn 'ask <builder> to help with task N' into a filed collab help request (POST /help-requests): resolve the name to a builder id, one addressee (person XOR craft) and one context (task XOR goal), a what_is_stuck answerable without this session, then settle it; routes the near-neighbours (recommendation vs blocker) and states that the addressed read is a pull, not a push (task 1003826)
|
|
367
|
-
├── backlog-review/SKILL.md ← /backlog-review — daily walk of status=backlog, the pre-workable state a human must say go on: splits rows waiting on a PERSON (promote / kill / water) from rows waiting on a live dep TRIGGER (counted, never walked, migration 163) and surfaces rows stranded behind an abandoned dep; runs scripts/gds/backlog-review.js (task 1003746); Metic+
|
|
367
|
+
├── backlog-review/SKILL.md ← /backlog-review — HIDDEN (typed-only), step of /owner-review (task 1004471); daily walk of status=backlog, the pre-workable state a human must say go on: splits rows waiting on a PERSON (promote / kill / water) from rows waiting on a live dep TRIGGER (counted, never walked, migration 163) and surfaces rows stranded behind an abandoned dep; runs scripts/gds/backlog-review.js (task 1003746); Metic+
|
|
368
368
|
├── blocker-review/SKILL.md ← /blocker-review — daily review of open Bongos blockers (resolve / escalate / note); Metic+
|
|
369
369
|
├── blocker-solve/SKILL.md ← /blocker-solve N — drive ONE blocker to done: do the doable parts, hand back owner-only steps, verify, auto-resolve (auto-promotes waiters); Metic+
|
|
370
370
|
├── bongos-feedback/SKILL.md ← /bongos-feedback — send feedback about Bongos itself upstream to the Cloud Bongos maintainers via `bongos feedback` (a bug lands as a backlog task in the maintenance goal, an idea in the hub inbox); backed by scripts/gds/feedback-send.js (task 1004462)
|
|
371
|
-
├── bug-triage/SKILL.md ← /bug-triage — daily walk through open kind=bug tasks (batch-cluster near-dups, merge / won't-fix); the idea-triage counterpart built on the new tasks.merge primitive (task 1001478); Metic+
|
|
371
|
+
├── bug-triage/SKILL.md ← /bug-triage — HIDDEN (typed-only), step of /owner-review (task 1004471); daily walk through open kind=bug tasks (batch-cluster near-dups, merge / won't-fix); the idea-triage counterpart built on the new tasks.merge primitive (task 1001478); Metic+
|
|
372
372
|
├── builder-backup/SKILL.md ← /builder-backup — check DB backup status or trigger a fresh local dump before risky changes (migrations, destructive SQL); no SSH required
|
|
373
373
|
├── builder-claim/SKILL.md ← /builder-claim N — atomic claim; prints the claim-time context pack + discipline playbook routing
|
|
374
374
|
├── builder-cost/SKILL.md ← /builder-cost — log a cost-ledger entry (api|compute|infra|domain|art|other)
|
|
@@ -389,17 +389,20 @@ tests/
|
|
|
389
389
|
├── feedback/SKILL.md ← /feedback — pull the latest Bongos feedback bundle (prompt.md + screenshot file paths) into a Claude Code session; backed by scripts/gds/feedback-latest.js (BV1.R13, ADR 0087)
|
|
390
390
|
├── figma-design-sync/SKILL.md ← sync UI-surface tokens/surfaces with Figma (push a frame plan, land a designer's Code Connect / MCP snapshot edit back as code) — the ADR 0081 human escape hatch; distinct from otb-figma-sync's pixel-art tile review (task 2048)
|
|
391
391
|
├── fix-task/SKILL.md ← /fix-task N — claim, fix (test-first), verify, and ship a single Bongos task by id end to end; defers to /builder-claim + /builder-ship for their own mechanics
|
|
392
|
+
├── goal-close/SKILL.md ← /goal-close — HIDDEN (typed-only); the merged goal-uat + goal-review: Part 1 tests criteria awaiting UAT on the live site and signs them off, Part 2 decides the ones a UAT cannot close (abandoned or unlinked work) (task 1004471, ADR 0351); Metic+
|
|
392
393
|
├── goal-create/SKILL.md ← /goal-create — plan and create ONE goal (the module-scoped workspace between version and criteria, ADR 0086) in one session: <redacted>, scope wall, 2-4 criteria, seed tasks; also the auto-pickup when a task is created with no goal to live in (goal_advisory); Metic+ (Archon for a protected-module scope)
|
|
393
|
-
├── goal-review/SKILL.md ← /goal-review —
|
|
394
|
-
├── goal-uat/SKILL.md ← /goal-uat —
|
|
395
|
-
├── grade-audit/SKILL.md ← /grade-audit —
|
|
394
|
+
├── goal-review/SKILL.md ← /goal-review — HIDDEN alias (task 1004471): runs Part 2 (the residue a UAT cannot close) of /goal-close, so the old name still works when typed
|
|
395
|
+
├── goal-uat/SKILL.md ← /goal-uat — HIDDEN alias (task 1004471): runs Part 1 (UAT sign-off) of /goal-close, so the old name still works when typed
|
|
396
|
+
├── grade-audit/SKILL.md ← /grade-audit — HIDDEN alias (task 1004471): runs Part B (the accountability audit) of /grade-sweep, so the old name still works when typed
|
|
396
397
|
├── grade-recover/SKILL.md ← /grade-recover N — the 'my grade failed, now what' loop: diagnose outage vs fabricated fail vs genuine fail from the persisted grade, then run the right recovery (--regrade or the auditable override-request); adds no ship-path gate (ADR 0162, task 1002655)
|
|
397
|
-
├──
|
|
398
|
-
├──
|
|
398
|
+
├── grade-sweep/SKILL.md ← /grade-sweep — HIDDEN (typed-only); the merged grader-health + grade-audit: Part A reads grader health (outages, pass-rate drift, cost, model pins), Part B audits what shipped past it and queues follow-ups only; never confirms or re-grades (task 1004471)
|
|
399
|
+
├── grader-health/SKILL.md ← /grader-health — HIDDEN alias (task 1004471): runs Part A (grader health) of /grade-sweep, so the old name still works when typed
|
|
400
|
+
├── idea-triage/SKILL.md ← /idea-triage — HIDDEN (typed-only), step of /owner-review (task 1004471); daily walk through the Bongos idea_inbox (promote / discard / merge); Metic+
|
|
399
401
|
├── merge-mode/SKILL.md ← /merge-mode — manual fallback for the auto-merge: land every confirmed task (merge + smoke + deploy)
|
|
400
402
|
├── new-project/SKILL.md ← /new-project — the canonical zero-to-live onboarding runbook for a new STANDALONE instance (name+address → init → scaffold → provision → OAuth+verified-secret+sign-in → verify → brand); the CLI + hall wizard are forms over it ([#1981](https://example.com/builders#/task/1981))
|
|
403
|
+
├── owner-review/SKILL.md ← /owner-review — the LISTED front door to the four queues waiting on a person's decision: backlog → bugs → ideas → priorities, or one step by name or plain words; reads each step's hidden playbook by path (task 1004471); Metic+
|
|
401
404
|
├── planning-session/SKILL.md ← /planning-session — run a structured planning session for an upcoming version (set criteria, brainstorm, rank, seed Bongos); Metic+
|
|
402
|
-
├── priority-session/SKILL.md ← /priority-session — a trusted builder answers plain-speech questions; the answers reweight the idea_inbox + suggest what to build next
|
|
405
|
+
├── priority-session/SKILL.md ← /priority-session — HIDDEN (typed-only), step of /owner-review (task 1004471); a trusted builder answers plain-speech questions; the answers reweight the idea_inbox + suggest what to build next
|
|
403
406
|
├── read-session-export/SKILL.md ← /read-session-export — read a Claude Code /export zip fast (prompts + thinking + tool calls → stdout); reads the existing export, doesn't make one. Backed by scripts/gds/read-session-export.js
|
|
404
407
|
├── recall/SKILL.md ← /recall — one-call project-knowledge search across repo docs + DB prose (tasks, learnings, session logs)
|
|
405
408
|
├── scan-before-install/SKILL.md ← /scan-before-install — quarantine-fetch a third-party repo/plugin, deterministic floor + tiered subagent panel over the scrubbed mirror, arithmetic verdict + honest-limits report to the builder's own config dir; prints install commands (pinned SHA only), never installs (goal 1000055)
|
|
@@ -407,7 +410,6 @@ tests/
|
|
|
407
410
|
├── ship-check/SKILL.md ← /ship-check — run every freshness + fitness check ship.js will run (repo-map/session-index/file-map/api-docs/copy-registry --check, linkify, skill-lint, fitness; --tests adds the DB-free lane; an ADVISORY render-fit row renders the changed pages via scripts/gds/render-check.js, ADR 0329) in ONE command before shipping, with the healing command per stale artifact (task 1003548; script scripts/gds/ship-check.js)
|
|
408
411
|
├── status/SKILL.md ← /status — what's the status of X / what's left for criterion Cn (read-only rollup)
|
|
409
412
|
├── strand-fix/SKILL.md ← /strand-fix N — walk ONE stranded confirmed task (strand:no_branch_no_tip / land_not_proven — the shape the reconciler gave up on) to shipped: find the work (main commit / unpushed branch / conflicting PR / nowhere), do the one sanctioned move, verify the flip; also clears a 409 REBASE_REQUIRED claim gate (task 1003365)
|
|
410
|
-
├── tweak/SKILL.md ← /tweak — apply the next SUBMITTED page tweak (Tweak Mode, ADR 0341): claim the round, copy-apply.js --batch every line (refusals named, never skipped), tweak-renders.js before/after desktop+phone light+dark into named task-visual slots, ship so the passed grade waits at completed for the artist (task 1004322)
|
|
411
413
|
└── worktree-clean/SKILL.md ← /worktree-clean — list the per-claim worktrees whose branch already landed on main (scripts/gds/stale-worktrees.js) and remove each through the junction-safe worktree.js remove; never raw git worktree remove (task 1003365)
|
|
412
414
|
```
|
|
413
415
|
<!-- END GENERATED FILE-MAP .claude/skills/ -->
|
|
@@ -2741,5 +2741,9 @@ is load-bearing: the script throws rather than guess if it is missing, and
|
|
|
2741
2741
|
landed since 1.20.40 with no explicit bump. run 36798842066. (task 1002620)
|
|
2742
2742
|
1.20.42 — CI auto-patch (publish-on-merge, ADR 0161): carrier for merges
|
|
2743
2743
|
landed since 1.20.41 with no explicit bump. run 36801585300. (task 1002620)
|
|
2744
|
+
1.20.43 — CI auto-patch (publish-on-merge, ADR 0161): carrier for merges
|
|
2745
|
+
landed since 1.20.42 with no explicit bump. run 36802367948. (task 1002620)
|
|
2746
|
+
1.20.44 — CI auto-patch (publish-on-merge, ADR 0161): carrier for merges
|
|
2747
|
+
landed since 1.20.43 with no explicit bump. run 36804013644. (task 1002620)
|
|
2744
2748
|
---------------------------------------------------------------------------
|
|
2745
2749
|
```
|
package/docs/modules-contract.md
CHANGED
|
@@ -113,6 +113,7 @@ Every module must have a `module.json` at its root (`modules/<key>/module.json`)
|
|
|
113
113
|
"consumes": [], // OPTIONAL. Seam ports this module resolves (required caps).
|
|
114
114
|
"prerequisites": { "modules": [] }, // OPTIONAL. Other modules that must be enabled first.
|
|
115
115
|
"spend": { "requiresPayer": true }, // OPTIONAL. This module spends MONEY for whoever calls it.
|
|
116
|
+
"howto": { "artifactUrl": "https://claude.ai/…" }, // OPTIONAL. A Claude page beside HOWTO.md — never instead of it (ADR 0347 D3).
|
|
116
117
|
"dependencies": { "discord.js": "^14.26.4" } // OPTIONAL. External npm deps this module needs at runtime.
|
|
117
118
|
}
|
|
118
119
|
```
|
|
@@ -321,6 +322,7 @@ This creates `modules/<key>/` with:
|
|
|
321
322
|
- `migrations/` — empty directory.
|
|
322
323
|
- `ui/` — empty directory.
|
|
323
324
|
- `CLAUDE.md` — per-module reference explaining the boundary, the seam pattern, and what to keep out of the module.
|
|
325
|
+
- `HOWTO.md` — the how-to for the person who installs the module: the five required sections as headings, each with its prompt in an HTML comment (see §6).
|
|
324
326
|
|
|
325
327
|
### 2. Write the module
|
|
326
328
|
|
|
@@ -359,6 +361,7 @@ Run `bongos upgrade` to confirm coreVersion compatibility, then `node scripts/gd
|
|
|
359
361
|
- Your first publish of a key makes you its author; only the author can publish later versions, and a delisted key takes none.
|
|
360
362
|
- The publish denylist applies (ADR 0098), plus a refusal of credential-named files (`.env*`, `*.pem`, `*.key`, `id_rsa`, …) anywhere in the module. The store re-checks every hash and the denylist itself.
|
|
361
363
|
- Gate: Metic+ (`module.submit`) for now. This is not `bongos module submit`, which proposes a module *into* core.
|
|
364
|
+
- **A how-to is required** ([ADR 0347](adr/0347-every-store-module-ships-a-how-to.md)). `modules/<key>/HOWTO.md` is what a buyer reads — the store shows it as the module's page — so it is written for the person who installs the module, not for an AI working inside it (that is `CLAUDE.md`). It is plain Markdown: any AI or person can write it. Publish refuses the version unless it has these five level-2 headings, in any order and any case, each with some text under it: `## What it does`, `## Install and enable`, `## How to use it` (with one worked example), `## Configuration` ("None." is fine) and `## Limits and known issues` ("None known." is fine). HTML comments don't count as text, so the scaffold's prompts must be replaced. The refusal names each failing section. The check is one function, `checkHowto` in `scripts/gds/module-artifact.js`; the CLI runs it before uploading and the store runs it again on upload, so a hand-built upload can't skip it. `bongos module check` reports it as advice only — an upstream submit is not gated by it. Install and update don't re-check it, so a version published before the gate still installs. The file travels in the tarball, so its hash pins it to the version. An optional `howto.artifactUrl` in `module.json` (an `https://` link on `claude.ai`) is shown as an extra, never graded.
|
|
362
365
|
|
|
363
366
|
### 7. Install from the store
|
|
364
367
|
|