@jenga-ai/agent 1.2.4 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -0
- package/agents/developer.md +18 -0
- package/agents/scrum-master.md +1 -0
- package/agents/tester.md +18 -0
- package/hooks/on_session_end.sh +27 -0
- package/package.json +18 -17
- package/skills/commit/SKILL.md +11 -1
- package/skills/dev-done/SKILL.md +46 -0
- package/skills/dev-done/scripts/classify-commit-outcome.sh +114 -0
- package/skills/init/SKILL.md +7 -6
- package/skills/init/assets/scope-thresholds_template.json +7 -0
- package/skills/init/scripts/init.sh +6 -0
- package/skills/publish/SKILL.md +66 -0
- package/skills/publish/adapters/npm-ci.md +34 -0
- package/skills/publish/adapters/npm.md +18 -0
- package/skills/publish/assets/ci-contract.md +27 -0
- package/skills/publish/assets/publish.example.json +27 -0
- package/skills/publish/schemas/publish.schema.json +20 -0
- package/skills/publish/scripts/npm_ci_pipeline.sh +29 -0
- package/skills/publish/scripts/npm_stage_inspect.sh +829 -0
- package/skills/publish/scripts/npm_stage_pipeline.sh +427 -0
- package/skills/publish/scripts/publish_common.sh +16 -0
- package/skills/publish/scripts/show_history.sh +12 -5
- package/skills/publish/scripts/validate_npm_stage_env.sh +184 -0
- package/skills/publish/scripts/write_ledger_entry.sh +92 -2
- package/skills/reconcile/SKILL.md +121 -11
- package/skills/reconcile/assets/report_format.md +17 -0
- package/skills/reconcile/scripts/resolve-reconcile-scope.sh +489 -0
- package/skills/uncharted/SKILL.md +200 -21
- package/skills/uncharted/scripts/directory-triage.sh +342 -0
- package/skills/uncharted/scripts/elicitation-state.sh +457 -0
- package/templates/SCRUM_BOARD_SCHEMA.md +57 -0
- package/templates/agent-context.md.tpl +23 -11
- package/templates/copilot-instructions.md.tpl +21 -11
- package/mcp/router/README.md +0 -19
- package/mcp/router/embedder.js +0 -23
- package/mcp/router/index.js +0 -204
- package/mcp/router/matcher.js +0 -87
- package/mcp/router/package-lock.json +0 -1048
- package/mcp/router/package.json +0 -11
- package/mcp/router/skill-index.js +0 -104
- package/skills/route/SKILL.md +0 -180
|
@@ -47,9 +47,9 @@ Every mode runs the same shared investigative engine and emits the same understa
|
|
|
47
47
|
|
|
48
48
|
| Mode | Scale | What it does | Implemented by |
|
|
49
49
|
|---|---|---|---|
|
|
50
|
-
| `segment` | one file, directory, or feature |
|
|
50
|
+
| `segment` | one file, directory, or feature | Two explicit modes since E20_S08_T03: `--mode delivery` (default, unchanged) analyses a target already in the repo and proposes a standard epic/story/task; `--mode investigate` opens a conversational architecture-investigation flow instead. See **`segment`** below for the mode choice. | E40_S02 / E20_S08_T03 |
|
|
51
51
|
| `import` | an external source | Acquires a git URL, an out-of-repo path, or a pasted snippet into the repo at a user-confirmed location, then hands off to `segment`. | E40_S03 |
|
|
52
|
-
| `onboard` | the whole codebase |
|
|
52
|
+
| `onboard` | the whole codebase | Conversational by default since E20_S08_T03: discovery scripts seed a human-in-the-loop elicitation that writes `[ARCH]`-tagged board items and coarse graph nodes. `--legacy` reproduces the original fully-automated, zero-prompt, capped-**backfilled**-epic pass unchanged. Board/graph-only — never touches application code, in either mode. | E40_S04 / E20_S08_T03 |
|
|
53
53
|
|
|
54
54
|
**Dispatch rules:**
|
|
55
55
|
|
|
@@ -146,6 +146,111 @@ A document whose judgement sections are generic enough to apply to any codebase
|
|
|
146
146
|
|
|
147
147
|
---
|
|
148
148
|
|
|
149
|
+
## Conversational Elicitation
|
|
150
|
+
|
|
151
|
+
Since **E20_S08_T03**, `onboard`'s default behavior and `segment --mode investigate` both run a human-in-the-loop **conversational elicitation** instead of (or, for `onboard`, in addition to keeping available) a one-shot deterministic pass. This section defines the mechanics shared by both; each mode's own subsection below only describes what is specific to it.
|
|
152
|
+
|
|
153
|
+
**What conversational elicitation produces, and how that differs from the Understanding Document above:** the **primary** output is coarse-tier graph nodes/edges — written directly to `project/knowledge-graph/graph.json`, conforming to the stub schema at `project/knowledge-graph/STUB_SCHEMA.md` (E20_S08_T01; this is a throwaway pilot schema, swapped wholesale once E20_S01's real schema lands — do not extend it expecting stability). Every node this flow writes carries `source: "human"`, since it comes from a person confirming or correcting a proposed understanding, not from mechanical extraction. Board representation is a `[ARCH]`-tagged epic, story, or task at whichever level fits the investigated scope (see `templates/SCRUM_BOARD_SCHEMA.md`'s `[ARCH]` — Durable Architectural Inventory convention) — **not** the delivery-shaped epic/story/task proposal `segment --mode delivery` and legacy `onboard` produce, and not the fixed 7-heading Understanding Document either. A written summary is produced only when warranted, filed as the resulting board item's ordinary `-summary.md` — there is no new artifact type or separate "Understanding Document" for conversational output.
|
|
154
|
+
|
|
155
|
+
### Human-Oracle-Availability Limitation
|
|
156
|
+
|
|
157
|
+
**Read this before running a conversational elicitation on code nobody currently understands.** This is `/uncharted`'s own documented, standing, accepted limitation — not only `agents/developer.md`'s and `agents/tester.md`'s (E20_S08_T02 covers those; this is the skill's own copy of the same limitation, since the person driving `/uncharted` may never open either agent file directly).
|
|
158
|
+
|
|
159
|
+
For genuinely undocumented code, there is often no reliable human oracle to confirm or correct a proposed understanding — the person answering may not know either, or may confidently confirm a wrong answer. Per the parent story's own scrutiny (scored 3/10 on this exact point) and its solution assessment's disposition ("Accept and Descope" on Problem 1): **conversational elicitation does not solve this, and does not claim to.** It is a complementary path for codebases where *some* human context exists, not a fix for the hardest, genuinely zero-oracle case.
|
|
160
|
+
|
|
161
|
+
- **The deterministic pipeline remains the tool of record for zero-oracle codebases.** `onboard --legacy` and `segment --mode delivery` never depend on anyone confirming intent — they ground everything in mechanical evidence (file structure, dependencies, test coverage) and say so explicitly under `Open Questions` when the evidence doesn't support a conclusion. When there is no one left who understands the code, reach for one of those, not the conversational flow.
|
|
162
|
+
- **When running the conversational flow, do not manufacture confidence.** If the user's answer is uncertain, hedged, or contradicts what discovery/Investigative Mode found, write the node honestly — do not round an uncertain answer up to a confirmed one. There is no schema field yet to tag confidence (the stub schema is intentionally minimal); until one exists, say so in the node's `description` text itself (e.g. "per the user, this module retries failed charges — unconfirmed against the code, which shows only a single retry attempt") rather than silently dropping the caveat.
|
|
163
|
+
- **A confidently wrong answer is not detectable by this flow.** Corroboration against a second signal (commit history, existing docs, a second person) is the only mitigation, and it is not built here — this is the accepted residual risk, not a gap to engineer around mid-conversation.
|
|
164
|
+
|
|
165
|
+
### Directory Triage
|
|
166
|
+
|
|
167
|
+
Runs once per elicitation, between discovery and the first Investigative Mode dispatch — never skipped, and never silently absorbed into the convergence loop below, because triaging noise out is exactly what keeps that loop from wasting turns (and the user's attention) on vendor/generated directories nobody wants a graph node for.
|
|
168
|
+
|
|
169
|
+
**Step 1 — deterministic pass.** Feed the candidate directories (from `discover-subsystems.sh`'s output for `onboard`, or the target's immediate subdirectories for `segment --mode investigate`) to:
|
|
170
|
+
|
|
171
|
+
```bash
|
|
172
|
+
discover-subsystems.sh <root> | jq -r '.candidates[].path' \
|
|
173
|
+
| bash skills/uncharted/scripts/directory-triage.sh <root>
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Read `ignored` (already excluded by `.gitignore`, or matching the fixed vendor/generated pattern list — see the script's own header for the full list) and `remaining`. This is a mechanical classification; do not second-guess it or re-run judgement over an `ignored` entry.
|
|
177
|
+
|
|
178
|
+
**Step 2 — judgement pass over `remaining`.** This is the generalize-list the deterministic pattern list can never catch: content-based cases where a directory is legitimately noise or a single coarse unit for reasons a pattern match cannot see — the task's own example is "this directory is SOAP request mock-ups for a test suite." Skim `remaining` (directory names, a shallow listing, README/comment hints) and propose, for each one that looks like a generalize candidate, a single-sentence rationale and the disposition: **exclude** (like `ignored`, zero graph nodes) or **generalize** (one coarse-tier node covering the whole directory, no per-file drill-down, no Investigative Mode dispatch into it). Everything not proposed for either disposition proceeds to full investigation.
|
|
179
|
+
|
|
180
|
+
**Step 3 — one confirmation gate, both lists together.** Per the task's own acceptance criterion, both the deterministic ignore-list and the judgement-based generalize-list are confirmed with the user **before any developer/tester Investigative Mode dispatch** — not after, and not as two separate gates:
|
|
181
|
+
|
|
182
|
+
```
|
|
183
|
+
Directory triage for <target>:
|
|
184
|
+
|
|
185
|
+
Excluded (pattern/gitignore match) — N directories:
|
|
186
|
+
<path> — <reason>
|
|
187
|
+
...
|
|
188
|
+
|
|
189
|
+
Proposed for exclusion or generalization (judgement) — M directories:
|
|
190
|
+
<path> — exclude|generalize — <one-line rationale>
|
|
191
|
+
...
|
|
192
|
+
|
|
193
|
+
Everything else (K directories) proceeds to investigation.
|
|
194
|
+
|
|
195
|
+
How should this be applied?
|
|
196
|
+
1. Accept as proposed
|
|
197
|
+
2. Revise — change a disposition before proceeding
|
|
198
|
+
3. Show me the full excluded/generalize lists in detail
|
|
199
|
+
4. Other (describe below)
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
Silence, a counter-question, or an ambiguous reply is not consent — re-ask, the same convention used at every other confirmation gate in this skill. Option 3 does not count as a decision; re-present the same choice after showing the detail. Checkpoint the confirmed triage result immediately via `elicitation-state.sh checkpoint` (see Multi-Session Persistence below) so a paused-and-resumed session never re-asks a question the user already answered.
|
|
203
|
+
|
|
204
|
+
### Convergence Loop
|
|
205
|
+
|
|
206
|
+
Runs once per surviving candidate (a subsystem, a named flow, a directory) after Directory Triage. This is the "propose understanding, ask the user to confirm or correct" cycle at the center of the redesign — and the one the scrutiny flagged as having no termination bound and no defense against confirmation fatigue. Both gaps are closed mechanically, not by agent discipline alone:
|
|
207
|
+
|
|
208
|
+
1. **Dispatch Investigative Mode.** Per `agents/developer.md`'s and `agents/tester.md`'s Investigative Mode sections (E20_S08_T02), dispatch the developer to trace what the code actually does for the candidate, and the tester to trace what the test suite actually exercises and verifies for the same candidate — two distinct vantage points, not two names for the same read. Both are read-only, worktree-sandboxed, no commits, no board writes.
|
|
209
|
+
2. **Propose understanding.** From both traces, draft the candidate's coarse graph node(s)/edge(s) (per the stub schema) and a plain-language summary of what they represent.
|
|
210
|
+
3. **Risk-weighted gating — not every finding gets a prompt.** This is the fix for confirmation fatigue (solution assessment, Problem 6, Solution B — RECOMMENDED): force an explicit confirmation only for **high-uncertainty or high-impact** findings — a node whose description depends on an inference the traces don't fully support, a node with many outgoing edges (structurally central), or one the Human-Oracle-Availability Limitation above already flagged as uncertain. **Auto-accept** low-risk, high-confidence findings — the traces agree, the finding is narrow in scope, nothing about it is surprising — without a prompt, but **log every auto-accepted node** in the elicitation state's checkpoint data (see below) so the decision is auditable later, per that solution's own mitigation for "the scoring mechanism itself misjudges impact."
|
|
211
|
+
4. **Confirm/correct, one round per call to `elicitation-state.sh turn`.** For a node requiring confirmation, present the draft and ask the user to confirm or correct it (per the Interaction Pattern in `CLAUDE.md` — confirm / correct-with-detail / defer as "unconfirmed" / other). Each round, call:
|
|
212
|
+
|
|
213
|
+
```bash
|
|
214
|
+
bash skills/uncharted/scripts/elicitation-state.sh turn --id <elicitation-id> --node <node-id>
|
|
215
|
+
```
|
|
216
|
+
|
|
217
|
+
Exit `0` means keep looping (cap not yet reached) if the user corrected rather than confirmed. Exit `3` means **the hard turn cap has been reached** — the node is now marked `flagged` in the state file. **Do not loop again on that node.** Instead present it as unresolved and offer an explicit choice rather than looping indefinitely (solution assessment, Problem 10, Solution A — RECOMMENDED):
|
|
218
|
+
|
|
219
|
+
```
|
|
220
|
+
<node> has reached the confirmation round limit without converging.
|
|
221
|
+
1. Accept the current best draft as-is (flagged low-confidence)
|
|
222
|
+
2. Defer — skip this node for now, continue with the rest
|
|
223
|
+
3. Continue past the limit (explicit override)
|
|
224
|
+
4. Other (describe below)
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
Option 3 is the only way past the cap, and it is a per-node, explicit, one-time override — it does not raise the cap for the rest of the run.
|
|
228
|
+
5. **On convergence** (confirmed, corrected-and-accepted, or resolved via the cap choice above), call:
|
|
229
|
+
|
|
230
|
+
```bash
|
|
231
|
+
bash skills/uncharted/scripts/elicitation-state.sh converge --id <elicitation-id> --node <node-id> --note "<one-line summary of what was confirmed>"
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
then write the node/edge to `project/knowledge-graph/graph.json` per the stub schema, and checkpoint the elicitation state (next section) — **after every converged node**, not only at the end of the whole run.
|
|
235
|
+
|
|
236
|
+
### Multi-Session Persistence
|
|
237
|
+
|
|
238
|
+
A whole-codebase `onboard` conversation, or an investigation of a large directory, can span more sessions than fit in one sitting. State persists via the existing `SessionEnd`/queue infrastructure — no new persistence mechanism (solution assessment, Problem 11, Solution A — RECOMMENDED).
|
|
239
|
+
|
|
240
|
+
- **`init` once, at the start of an elicitation:**
|
|
241
|
+
|
|
242
|
+
```bash
|
|
243
|
+
bash skills/uncharted/scripts/elicitation-state.sh init --id <elicitation-id> --target "<path or description>" --cap 5
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
Idempotent — safe to call again on a resumed `<elicitation-id>` without resetting progress. Choose `<elicitation-id>` so it is stable and re-derivable across sessions (e.g. `onboard-<root-slug>-<date>`, or `segment-investigate-<target-slug>`), since a resuming session must be able to reconstruct it to call `init` again.
|
|
247
|
+
- **`checkpoint` after every converged node and after the Directory Triage confirmation gate** — never only at the end. This is what makes a mid-run pause lossless: `checkpoint --id <id> --json <file>` merges arbitrary progress data (triage results, draft nodes not yet converged, anything else worth surviving a pause) into the state file.
|
|
248
|
+
- **`pause` when a session must end before the elicitation has converged.** Immediately after calling `elicitation-state.sh pause --id <elicitation-id>`, write the scrum-master's own `SessionEnd` handoff (per `templates/SCRUM_BOARD_SCHEMA.md`'s `handoffs/` convention) with `status: "elicitation_paused"` and both `elicitation_id` and `state_file` set — `hooks/on_session_end.sh` routes that into an `elicitation_resume` trigger on `scrum_triggers.jsonl`, which the next scrum-master session's Drain Scrum Triggers Queue procedure picks up (`agents/scrum-master.md`).
|
|
249
|
+
- **On resume**, read `state_file` directly — every converged node, every flagged node, and the checkpoint data (including the confirmed directory-triage lists) are already there. Do not re-run Directory Triage or re-ask about an already-converged node; resume the Convergence Loop only for nodes still `pending` or explicitly deferred.
|
|
250
|
+
- **`complete` when every candidate has converged, been deferred, or been explicitly accepted past the cap.** The state file is left on disk afterward as an audit trail — nothing currently prunes a completed elicitation's state file.
|
|
251
|
+
|
|
252
|
+
---
|
|
253
|
+
|
|
149
254
|
## Modes
|
|
150
255
|
|
|
151
256
|
<!--
|
|
@@ -158,7 +263,27 @@ A document whose judgement sections are generic enough to apply to any codebase
|
|
|
158
263
|
|
|
159
264
|
### `segment`
|
|
160
265
|
|
|
161
|
-
Analyse a specific file, directory, or feature that has no board provenance
|
|
266
|
+
Analyse a specific file, directory, or feature that has no board provenance. Since **E20_S08_T03**, `segment` has two explicit modes — it no longer silently always runs one:
|
|
267
|
+
|
|
268
|
+
| Flag | Mode | What it produces |
|
|
269
|
+
|---|---|---|
|
|
270
|
+
| `--mode delivery` (default when explicit) | Delivery-shaped proposal — today's unchanged flow | A standard epic/story/task proposal, per the Understanding Document |
|
|
271
|
+
| `--mode investigate` | Conversational architecture investigation (new, E20_S08_T03) | Coarse graph nodes/edges plus a `[ARCH]`-tagged board item |
|
|
272
|
+
|
|
273
|
+
**Step 0 — choose a mode.** If `--mode` is given, skip straight to that mode's subsection below. If it is missing, do **not** default silently — present the choice, per the Interaction Pattern in `CLAUDE.md`:
|
|
274
|
+
|
|
275
|
+
```
|
|
276
|
+
/uncharted segment <target> — which mode?
|
|
277
|
+
1. Delivery-shaped proposal (produces an epic/story/task to build on this target)
|
|
278
|
+
2. Conversational architecture investigation (produces graph nodes explaining what this target does)
|
|
279
|
+
3. Other (describe below)
|
|
280
|
+
```
|
|
281
|
+
|
|
282
|
+
Silence, a counter-question, or an ambiguous reply is not consent — re-ask. This choice is the entire mechanism by which existing callers keep getting exactly today's behavior (option 1, or `--mode delivery` passed explicitly) while new callers can reach the conversational path deliberately, per the parent story's own acceptance criterion that this split must be explicit, not a silent behavior change.
|
|
283
|
+
|
|
284
|
+
#### Delivery-shaped proposal (`--mode delivery`)
|
|
285
|
+
|
|
286
|
+
Unchanged from before E20_S08_T03 — every step below is exactly what `segment` has always done.
|
|
162
287
|
|
|
163
288
|
**Step 1 — Resolve the target.** Deterministic; do not eyeball it.
|
|
164
289
|
|
|
@@ -259,6 +384,24 @@ On success, tell the user exactly which files were created, with their IDs.
|
|
|
259
384
|
|
|
260
385
|
**`/uncharted` writes board files and stops there.** It does not adapt the segment to project conventions, edit or move the code it just analysed, open a worktree, or write an execution plan. That is ordinary developer work, driven by ordinary task files, and it is the developer agent's job — the same as for a task that came from `/brainstorm` or `/pi-plan`. A segment that has reached the board is no longer a special case, and this skill growing its own integration path would be a second, divergent execution route for work the existing one already handles.
|
|
261
386
|
|
|
387
|
+
#### Conversational investigation (`--mode investigate`)
|
|
388
|
+
|
|
389
|
+
New in **E20_S08_T03**. Produces a plain-language understanding plus graph nodes for a *specific* target — the same conversational mechanics as `onboard`'s default flow, applied at single-target scale instead of whole-codebase scale. Read **Conversational Elicitation** above (Human-Oracle-Availability Limitation, Directory Triage, Convergence Loop, Multi-Session Persistence) before running this — everything below only sequences those shared mechanics for `segment`'s scope; it does not redefine them.
|
|
390
|
+
|
|
391
|
+
**Step 1 — Resolve the target.** Reuse `resolve-segment-target.sh` exactly as the delivery-shaped path's own Step 1 does above — do not reimplement resolution or the board-linkage check for this mode. A `linked` target is a normal condition here (unlike the delivery-shaped path, an existing epic/story/task referencing the target doesn't disqualify investigating it), but still tell the user before proceeding, the same as the delivery-shaped path does.
|
|
392
|
+
|
|
393
|
+
**Step 2 — Directory triage, only if the target is a directory with subdirectories.** A single-file target has nothing to triage; skip straight to Step 3. For a directory target, run the shared Directory Triage procedure above against the target's immediate subdirectories as the candidate set.
|
|
394
|
+
|
|
395
|
+
**Step 3 — Convergence loop.** Run the shared Convergence Loop procedure above. For a single-file target this is one node; for a directory target (after triage) it is one node per surviving candidate. Initialize persistence first:
|
|
396
|
+
|
|
397
|
+
```bash
|
|
398
|
+
bash skills/uncharted/scripts/elicitation-state.sh init --id "segment-investigate-$(basename "$RESOLVED_TARGET")-$(date -u +%Y%m%d)" --target "$RESOLVED_TARGET"
|
|
399
|
+
```
|
|
400
|
+
|
|
401
|
+
**Step 4 — On convergence, write the graph and the board item.** Write each converged node/edge to `project/knowledge-graph/graph.json` per the stub schema (`source: "human"`), then present a single `[ARCH]`-tagged board item proposal — epic, story, or task, whichever level fits what was actually investigated (a single flow is usually task-scale; a whole subsystem may warrant a story or, rarely, an epic) — and stop at the same kind of confirmation gate the delivery-shaped path's Step 6 uses (accept / revise / discard / other). **Nothing is written to `project/board/` before that gate is confirmed**, matching the delivery-shaped path's own "nothing written before option 1" guarantee. On acceptance, write the board item and call `elicitation-state.sh complete`.
|
|
402
|
+
|
|
403
|
+
**This mode never produces a delivery-shaped epic/story/task proposal.** If the investigation surfaces work that should actually be *built* (not just understood), say so as a follow-up recommendation in the `[ARCH]` item's own text and let the user separately invoke `--mode delivery` or `/todo` for that — conversational investigation and delivery planning stay two distinct outputs, per this mode split's own purpose.
|
|
404
|
+
|
|
262
405
|
### `import`
|
|
263
406
|
|
|
264
407
|
Pull an external source into the repo, then investigate it as a segment.
|
|
@@ -441,28 +584,33 @@ destination.
|
|
|
441
584
|
|
|
442
585
|
### `onboard`
|
|
443
586
|
|
|
444
|
-
|
|
587
|
+
Give an entire pre-existing codebase board representation at framework-adoption time. Since **E20_S08_T03**, `onboard` has two modes:
|
|
445
588
|
|
|
446
|
-
|
|
447
|
-
|
|
448
|
-
|
|
449
|
-
> -
|
|
450
|
-
|
|
451
|
-
|
|
452
|
-
|
|
589
|
+
| Invocation | Mode | Output |
|
|
590
|
+
|---|---|---|
|
|
591
|
+
| `onboard <root>` (default) | Conversational — human-in-the-loop elicitation seeded by the same discovery scripts | `[ARCH]`-tagged board items + coarse graph nodes in `project/knowledge-graph/graph.json` |
|
|
592
|
+
| `onboard <root> --legacy` | Fully-automated, zero-prompt, capped-backfilled-epic pass — E40_S04's original behavior, byte-for-byte unchanged | `provenance: backfilled` epics |
|
|
593
|
+
|
|
594
|
+
**`--legacy` is not a deprecated fallback — it is the designated tool for a codebase with no available human oracle.** Read **Human-Oracle-Availability Limitation** under Conversational Elicitation above before choosing between the two; that section states explicitly when `--legacy` is the *correct* choice, not a lesser one.
|
|
595
|
+
|
|
596
|
+
> **Build status:** the legacy pipeline (`E40_S04_T01`-`T05`: `provenance` field, subsystem discovery, cap enforcement, backfilled epic generation, `PROJECT_SUMMARY.md` population) is complete and unchanged by this rework — see **Legacy mode** below. The conversational default is new as of `E20_S08_T03` — see **Conversational default** below.
|
|
453
597
|
|
|
454
598
|
#### The hard constraint: `onboard` never touches application code
|
|
455
599
|
|
|
600
|
+
This holds in **both** modes — conversational and legacy alike.
|
|
601
|
+
|
|
456
602
|
**`onboard` mode must never modify, move, rename, delete, or restructure any file that belongs to
|
|
457
603
|
the consumer's application.** This is not a best-effort convention — it is the reason `onboard`
|
|
458
604
|
exists as a *board-only* mode rather than a generic migration tool. Its entire output surface,
|
|
459
605
|
without exception, is:
|
|
460
606
|
|
|
461
|
-
- `project/board/` (
|
|
462
|
-
- `project/rapports/analysis/` (understanding documents and the subsystem cap record)
|
|
463
|
-
- `project/PROJECT_SUMMARY.md` (populated by `E40_S04_T05`)
|
|
607
|
+
- `project/board/` (backfilled epics in `--legacy` mode; `[ARCH]`-tagged epics/stories/tasks in conversational mode)
|
|
608
|
+
- `project/rapports/analysis/` (understanding documents and the subsystem cap record — `--legacy` mode only; conversational mode's primary output is the graph, not this document type — see Conversational Elicitation above)
|
|
609
|
+
- `project/PROJECT_SUMMARY.md` (populated by `E40_S04_T05`, `--legacy` mode only)
|
|
610
|
+
- `project/knowledge-graph/graph.json` (coarse graph nodes/edges — conversational mode only, per the stub schema)
|
|
611
|
+
- `project/queue/elicitation-state/` (multi-session persistence scratch state — conversational mode only, git-ignored, not a durable artifact)
|
|
464
612
|
|
|
465
|
-
Nothing under any other path is ever created, edited, or deleted by `onboard` mode. Jenga's board
|
|
613
|
+
Nothing under any other path is ever created, edited, or deleted by `onboard` mode, in either mode. Jenga's board
|
|
466
614
|
and skills live *alongside* the application; the application never has to conform to a Jenga
|
|
467
615
|
convention to be onboarded. This guarantee is enforced in three independent places, not just
|
|
468
616
|
stated here:
|
|
@@ -476,6 +624,34 @@ stated here:
|
|
|
476
624
|
3. **A verifiable post-run check** (below): after a full `onboard` run, `git status` in the
|
|
477
625
|
analysed repository must show changes only under `project/`.
|
|
478
626
|
|
|
627
|
+
#### Conversational default
|
|
628
|
+
|
|
629
|
+
New in **E20_S08_T03**. Runs when `onboard` is invoked without `--legacy`.
|
|
630
|
+
|
|
631
|
+
**Step 1 — discovery stays scripted.** Run the exact same deterministic discovery chain the legacy pipeline uses — `discover-subsystems.sh` — unchanged. Evidence-gathering is not where this rework touches anything; only what happens with the output differs.
|
|
632
|
+
|
|
633
|
+
```bash
|
|
634
|
+
bash skills/uncharted/scripts/discover-subsystems.sh <root>
|
|
635
|
+
```
|
|
636
|
+
|
|
637
|
+
**Step 2 — directory triage.** Feed the discovery output's candidate paths into the shared Directory Triage procedure (see Conversational Elicitation above), and stop at its confirmation gate before anything else happens.
|
|
638
|
+
|
|
639
|
+
**Step 3 — initialize persistence.**
|
|
640
|
+
|
|
641
|
+
```bash
|
|
642
|
+
bash skills/uncharted/scripts/elicitation-state.sh init --id "onboard-$(basename "$(cd "$root" && pwd)")-$(date -u +%Y%m%d)" --target "<root>" --cap 5
|
|
643
|
+
```
|
|
644
|
+
|
|
645
|
+
**Step 4 — convergence loop, once per surviving candidate.** Run the shared Convergence Loop procedure (see Conversational Elicitation above) for each candidate that survived triage — this is where the legacy pipeline's `apply-subsystem-cap.sh` and `write-backfilled-epics.sh` would have silently produced a capped set of `provenance: backfilled` epics; the conversational default asks about each one instead, subject to the same risk-weighted gating and hard turn cap. There is deliberately **no subsystem cap** on the conversational path — the turn cap already bounds cost per candidate, and capping the *candidate count* the way the legacy path does would silently drop subsystems from a human-in-the-loop conversation the same way the legacy path drops them from an unattended one, which defeats the point of asking. A codebase with far more subsystems than is practical to walk through conversationally in one sitting is exactly the multi-session case Multi-Session Persistence exists for — pause, resume across sessions, rather than truncate the candidate list.
|
|
646
|
+
|
|
647
|
+
**Step 5 — on convergence, write the graph and the board item(s).** Per candidate: write the converged node(s)/edge(s) to `project/knowledge-graph/graph.json`, then present an `[ARCH]`-tagged board item proposal at whichever level fits (an individual subsystem is usually story-scale; the whole run may warrant a single `[ARCH]` epic containing one story per converged subsystem — judgement call, not a fixed rule) and stop at a confirmation gate before writing to `project/board/`, exactly as `segment --mode investigate`'s Step 4 does. Call `elicitation-state.sh complete` once every candidate has converged, been deferred, or been resolved past the turn cap.
|
|
648
|
+
|
|
649
|
+
**`PROJECT_SUMMARY.md` population still applies, unchanged in spirit.** Once the conversational pass has produced its `[ARCH]` items, hand off to the scrum-master for the same `PROJECT_SUMMARY.md` Overview/Architecture & Structure drafting-and-confirmation flow described under **Updating PROJECT_SUMMARY.md from onboard evidence** below (Steps A-D) — substituting the conversational pass's converged understanding for the legacy pipeline's `kept` array as the evidence source. Do not skip the stub-vs-real-content check in Step B just because the evidence came from a conversation instead of a script.
|
|
650
|
+
|
|
651
|
+
#### Legacy mode (`--legacy`)
|
|
652
|
+
|
|
653
|
+
Everything below is `E40_S04`'s original, fully-automated, zero-prompt pipeline — **unchanged** by this rework, reachable only via the explicit `--legacy` flag.
|
|
654
|
+
|
|
479
655
|
#### The subsystem cap
|
|
480
656
|
|
|
481
657
|
Live as of `E40_S04_T03`. `onboard` does **not** create one epic per discovered subsystem — a
|
|
@@ -592,7 +768,9 @@ outside `project/` ended up tracked by git.
|
|
|
592
768
|
|
|
593
769
|
#### Updating PROJECT_SUMMARY.md from onboard evidence
|
|
594
770
|
|
|
595
|
-
Live as of `E40_S04_T05
|
|
771
|
+
Live as of `E40_S04_T05`, and reused by both `onboard` modes since `E20_S08_T03` — the conversational default's own Step 5 above hands off here too, substituting its converged understanding for the `kept` array as the evidence source described below. Nested under Legacy mode structurally only because it was written before the conversational path existed; treat "Steps A-D" as shared, not legacy-only.
|
|
772
|
+
|
|
773
|
+
This is the last step of an `onboard` run, after the backfilled epics (legacy) or `[ARCH]` items (conversational)
|
|
596
774
|
above have been written (or confirmed skipped). Its job is to close the gap the epic itself
|
|
597
775
|
describes: a project adopting Jenga should not end up with a populated board sitting underneath a
|
|
598
776
|
`PROJECT_SUMMARY.md` that still reads as if the codebase were empty.
|
|
@@ -684,10 +862,11 @@ hard constraint** above, since that file was never a required write, only a perm
|
|
|
684
862
|
These hold across every mode:
|
|
685
863
|
|
|
686
864
|
- **Read-only against application code**, with one exception: `import` mode writes an acquired source to a location the user has explicitly confirmed. Nothing else in this skill modifies, moves, or restructures a consumer's code.
|
|
687
|
-
- **Confirm before writing to the board.** Proposals are presented for approval first, consistent with the brainstorm-before-commit convention.
|
|
688
|
-
- **No new rapport type.** Understanding documents use the existing `analysis` type and land in `project/rapports/analysis/`.
|
|
689
|
-
- **Never invent findings.** If the evidence does not support a conclusion, say so under Open Questions.
|
|
865
|
+
- **Confirm before writing to the board.** Proposals are presented for approval first, consistent with the brainstorm-before-commit convention. Since `E20_S08_T03`, this also covers Directory Triage's confirmation gate and each converged node's graph/board write in the conversational elicitation flow — same convention, applied to a new output type (graph nodes), not a new exception to it.
|
|
866
|
+
- **No new rapport type.** Understanding documents use the existing `analysis` type and land in `project/rapports/analysis/`. Conversational elicitation's output (graph nodes + `[ARCH]` board items) is not an understanding document and does not need one — see Conversational Elicitation above for what it produces instead.
|
|
867
|
+
- **Never invent findings.** If the evidence does not support a conclusion, say so under Open Questions (delivery-shaped/legacy paths) or state the uncertainty plainly in the node itself (conversational paths — see Human-Oracle-Availability Limitation above).
|
|
690
868
|
- **Deterministic work belongs in `skills/uncharted/scripts/`**, agent judgement belongs here.
|
|
869
|
+
- **Existing zero-prompt behavior is never silently retired.** `onboard --legacy` and `segment --mode delivery` reproduce exactly what `/uncharted` did before `E20_S08_T03`, with no behavior drift for a caller who keeps using them.
|
|
691
870
|
|
|
692
871
|
---
|
|
693
872
|
|
|
@@ -695,8 +874,8 @@ These hold across every mode:
|
|
|
695
874
|
|
|
696
875
|
`/uncharted` is invocable directly, and is also offered automatically at the two moments it is most needed (E40_S05):
|
|
697
876
|
|
|
698
|
-
- **`/init`** — when scaffolding detects a non-empty, non-Jenga-scaffolded directory, it offers `onboard` instead of proceeding as though the project were empty.
|
|
699
|
-
- **`/reconcile`** — when its normal sync pass finds code with no board linkage, it offers `segment` for the affected paths.
|
|
877
|
+
- **`/init`** — when scaffolding detects a non-empty, non-Jenga-scaffolded directory, it offers `onboard` instead of proceeding as though the project were empty. Since `E20_S08_T03`, this offer surfaces both `onboard` modes — the user picks conversational (default) or `--legacy` at that point, per the mode choice this section describes; `/init` itself makes no decision about which is appropriate.
|
|
878
|
+
- **`/reconcile`** — when its normal sync pass finds code with no board linkage, it offers `segment` for the affected paths. Since `E20_S08_T03`, this offer surfaces both `segment` modes the same way.
|
|
700
879
|
|
|
701
880
|
Both are offers, never automatic execution.
|
|
702
881
|
|
|
@@ -0,0 +1,342 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# ---------------------------------------------------------------------------
|
|
3
|
+
# skills/uncharted/scripts/directory-triage.sh
|
|
4
|
+
#
|
|
5
|
+
# Deterministic front half of the directory-triage step (E20_S08_T03) that
|
|
6
|
+
# sits between discovery and any developer/tester Investigative Mode
|
|
7
|
+
# dispatch, for both `/uncharted onboard`'s conversational default and
|
|
8
|
+
# `/uncharted segment --mode investigate`.
|
|
9
|
+
#
|
|
10
|
+
# The problem: an investigative pass that walks every directory in a
|
|
11
|
+
# candidate set wastes turns (and, worse, a user's attention at the
|
|
12
|
+
# confirmation gate) on vendor/generated noise nobody wants a coarse graph
|
|
13
|
+
# node for. This script answers the MECHANICAL half of that question —
|
|
14
|
+
# "is this directory already excluded by .gitignore, or does it match a
|
|
15
|
+
# well-known vendor/generated pattern?" — so the agent only has to spend
|
|
16
|
+
# judgement on what's left: the content-based "generalize" cases a fixed
|
|
17
|
+
# pattern list can never catch (e.g. "this directory is SOAP request
|
|
18
|
+
# mock-ups for a test suite").
|
|
19
|
+
#
|
|
20
|
+
# This script performs NO judgement of its own and writes nothing. It
|
|
21
|
+
# classifies; the agent decides what to do with `remaining`.
|
|
22
|
+
#
|
|
23
|
+
# ---------------------------------------------------------------------------
|
|
24
|
+
# INPUT
|
|
25
|
+
# ---------------------------------------------------------------------------
|
|
26
|
+
# A root (the repository, or the target `onboard`/`segment` is investigating)
|
|
27
|
+
# and a set of candidate directory paths to triage, relative to that root.
|
|
28
|
+
# Candidates come from positional arguments, or newline-delimited from stdin
|
|
29
|
+
# when none are given — composable with discover-subsystems.sh the same way
|
|
30
|
+
# apply-subsystem-cap.sh is:
|
|
31
|
+
#
|
|
32
|
+
# discover-subsystems.sh <root> | jq -r '.candidates[].path' \
|
|
33
|
+
# | directory-triage.sh <root>
|
|
34
|
+
#
|
|
35
|
+
# A candidate that does not exist under <root>, or is not a directory, is
|
|
36
|
+
# reported in `notices` and otherwise skipped — it is an input problem, not a
|
|
37
|
+
# triage verdict.
|
|
38
|
+
#
|
|
39
|
+
# ---------------------------------------------------------------------------
|
|
40
|
+
# THE DETERMINISTIC IGNORE-LIST
|
|
41
|
+
# ---------------------------------------------------------------------------
|
|
42
|
+
# Path-COMPONENT match against a fixed, known-noise pattern list — never a
|
|
43
|
+
# substring match, so a real subsystem directory named e.g. "targets/" is not
|
|
44
|
+
# caught by the "target" pattern. The list is intentionally short and named
|
|
45
|
+
# in the task's own acceptance criteria; it is expected to need occasional
|
|
46
|
+
# extension as new ecosystems are triaged, and is kept as a single array
|
|
47
|
+
# below for exactly that reason — this is a revisitable list, not a closed
|
|
48
|
+
# one:
|
|
49
|
+
#
|
|
50
|
+
# node_modules vendor dist build .venv
|
|
51
|
+
# venv __pycache__ .next coverage target
|
|
52
|
+
# .git .tox .mypy_cache .pytest_cache .cache
|
|
53
|
+
# *.egg-info (suffix match on the last path component)
|
|
54
|
+
#
|
|
55
|
+
# A candidate already excluded by .gitignore (via `git check-ignore`) is
|
|
56
|
+
# reported separately (`reason: "gitignore"`) even if it ALSO matches a
|
|
57
|
+
# pattern — gitignore is checked first and wins, since "why is this
|
|
58
|
+
# ignored" should point at the actual mechanism in effect.
|
|
59
|
+
#
|
|
60
|
+
# ---------------------------------------------------------------------------
|
|
61
|
+
# OPTIONS
|
|
62
|
+
# ---------------------------------------------------------------------------
|
|
63
|
+
# --root DIR Root the candidates are resolved against. Required as
|
|
64
|
+
# the first positional argument OR via this flag.
|
|
65
|
+
# --json-out FILE Also write the JSON report to a file. stdout gets it
|
|
66
|
+
# regardless.
|
|
67
|
+
# -h, --help Show this help and exit 0.
|
|
68
|
+
#
|
|
69
|
+
# ---------------------------------------------------------------------------
|
|
70
|
+
# OUTPUT (stdout, JSON)
|
|
71
|
+
# ---------------------------------------------------------------------------
|
|
72
|
+
# {
|
|
73
|
+
# "script": "directory-triage.sh",
|
|
74
|
+
# "version": 1,
|
|
75
|
+
# "root": "<as given>",
|
|
76
|
+
# "root_absolute": "<canonicalised>",
|
|
77
|
+
# "candidate_count": <int>,
|
|
78
|
+
# "ignored_count": <int>,
|
|
79
|
+
# "remaining_count": <int>,
|
|
80
|
+
# "ignored": [ { "path", "absolute_path", "reason" }, ... ],
|
|
81
|
+
# "remaining": [ { "path", "absolute_path" }, ... ],
|
|
82
|
+
# "notices": [ "<non-fatal diagnostic>", ... ]
|
|
83
|
+
# }
|
|
84
|
+
#
|
|
85
|
+
# `reason` on an ignored entry is either "gitignore" or "pattern:<name>"
|
|
86
|
+
# (e.g. "pattern:node_modules") — never a bare boolean, so a consumer never
|
|
87
|
+
# has to re-derive why something was excluded.
|
|
88
|
+
#
|
|
89
|
+
# `remaining` is the generalize-list INPUT, not its output — this script
|
|
90
|
+
# does not attempt content-based judgement (reading file contents to guess
|
|
91
|
+
# "this looks like SOAP mock-ups") at all. That pass is agent judgement, per
|
|
92
|
+
# the Skill Implementation Principle in CLAUDE.md, and lives in
|
|
93
|
+
# skills/uncharted/SKILL.md's Directory Triage subsection, not here.
|
|
94
|
+
#
|
|
95
|
+
# ---------------------------------------------------------------------------
|
|
96
|
+
# EXIT CODES
|
|
97
|
+
# ---------------------------------------------------------------------------
|
|
98
|
+
# 0 — success (including zero candidates, or everything remaining)
|
|
99
|
+
# 1 — usage error: unknown flag, missing root
|
|
100
|
+
# 2 — input error: root does not exist or is not a directory
|
|
101
|
+
#
|
|
102
|
+
# Examples:
|
|
103
|
+
# directory-triage.sh . src lib vendor node_modules
|
|
104
|
+
# discover-subsystems.sh . | jq -r '.candidates[].path' | directory-triage.sh .
|
|
105
|
+
#
|
|
106
|
+
# Requires: bash, python3, git (optional — gitignore check is skipped with a
|
|
107
|
+
# notice when the root is not inside a git working tree).
|
|
108
|
+
|
|
109
|
+
set -euo pipefail
|
|
110
|
+
|
|
111
|
+
SCRIPT_NAME="directory-triage.sh"
|
|
112
|
+
ROOT=""
|
|
113
|
+
JSON_OUT=""
|
|
114
|
+
CANDIDATES=()
|
|
115
|
+
|
|
116
|
+
print_help() {
|
|
117
|
+
cat <<'EOF'
|
|
118
|
+
Usage: directory-triage.sh [--root DIR] [--json-out FILE] [<root>] [candidate...]
|
|
119
|
+
discover-subsystems.sh <root> | jq -r '.candidates[].path' | directory-triage.sh <root>
|
|
120
|
+
|
|
121
|
+
Classifies candidate directories under <root> into:
|
|
122
|
+
- ignored: already excluded by .gitignore, or matching a known
|
|
123
|
+
vendor/generated pattern (node_modules, vendor, dist, build,
|
|
124
|
+
.venv, venv, __pycache__, .next, coverage, target, .git,
|
|
125
|
+
.tox, .mypy_cache, .pytest_cache, .cache, *.egg-info)
|
|
126
|
+
- remaining: everything else — the input to the agent-judgement
|
|
127
|
+
generalize-list pass documented in skills/uncharted/SKILL.md
|
|
128
|
+
|
|
129
|
+
Options:
|
|
130
|
+
--root DIR Root the candidates are resolved against (or give it as
|
|
131
|
+
the first positional argument).
|
|
132
|
+
--json-out FILE Also write the JSON report to a file. stdout gets it
|
|
133
|
+
regardless.
|
|
134
|
+
-h, --help Show this help and exit 0.
|
|
135
|
+
|
|
136
|
+
Exit codes: 0 success, 1 usage error, 2 root missing/not a directory.
|
|
137
|
+
EOF
|
|
138
|
+
}
|
|
139
|
+
|
|
140
|
+
usage_error() {
|
|
141
|
+
echo "Usage: $SCRIPT_NAME [--root DIR] [--json-out FILE] [<root>] [candidate...]" >&2
|
|
142
|
+
echo " (candidates may also be piped newline-delimited on stdin)" >&2
|
|
143
|
+
exit 1
|
|
144
|
+
}
|
|
145
|
+
|
|
146
|
+
while [ "$#" -gt 0 ]; do
|
|
147
|
+
case "$1" in
|
|
148
|
+
-h|--help)
|
|
149
|
+
print_help
|
|
150
|
+
exit 0
|
|
151
|
+
;;
|
|
152
|
+
--root)
|
|
153
|
+
[ "$#" -ge 2 ] || usage_error
|
|
154
|
+
ROOT="$2"
|
|
155
|
+
shift 2
|
|
156
|
+
;;
|
|
157
|
+
--json-out)
|
|
158
|
+
[ "$#" -ge 2 ] || usage_error
|
|
159
|
+
JSON_OUT="$2"
|
|
160
|
+
shift 2
|
|
161
|
+
;;
|
|
162
|
+
--)
|
|
163
|
+
shift
|
|
164
|
+
break
|
|
165
|
+
;;
|
|
166
|
+
-*)
|
|
167
|
+
usage_error
|
|
168
|
+
;;
|
|
169
|
+
*)
|
|
170
|
+
if [ -z "$ROOT" ]; then
|
|
171
|
+
ROOT="$1"
|
|
172
|
+
else
|
|
173
|
+
CANDIDATES+=("$1")
|
|
174
|
+
fi
|
|
175
|
+
shift
|
|
176
|
+
;;
|
|
177
|
+
esac
|
|
178
|
+
done
|
|
179
|
+
|
|
180
|
+
# Remaining positionals after `--` are also candidates.
|
|
181
|
+
for arg in "$@"; do
|
|
182
|
+
CANDIDATES+=("$arg")
|
|
183
|
+
done
|
|
184
|
+
|
|
185
|
+
[ -n "$ROOT" ] || usage_error
|
|
186
|
+
|
|
187
|
+
if [ ! -d "$ROOT" ]; then
|
|
188
|
+
echo "$SCRIPT_NAME: root '$ROOT' does not exist or is not a directory" >&2
|
|
189
|
+
exit 2
|
|
190
|
+
fi
|
|
191
|
+
|
|
192
|
+
ROOT_ABS=$(cd -- "$ROOT" && pwd -P)
|
|
193
|
+
|
|
194
|
+
# If no candidates were given on the command line, read newline-delimited
|
|
195
|
+
# candidates from stdin (only when stdin is not a terminal, so an
|
|
196
|
+
# interactive invocation with no candidates doesn't hang waiting for input).
|
|
197
|
+
if [ "${#CANDIDATES[@]}" -eq 0 ] && [ ! -t 0 ]; then
|
|
198
|
+
while IFS= read -r line; do
|
|
199
|
+
[ -n "$line" ] && CANDIDATES+=("$line")
|
|
200
|
+
done
|
|
201
|
+
fi
|
|
202
|
+
|
|
203
|
+
# Determine gitignore availability once, up front.
|
|
204
|
+
GITIGNORE_AVAILABLE=0
|
|
205
|
+
if git -C "$ROOT_ABS" rev-parse --is-inside-work-tree >/dev/null 2>&1; then
|
|
206
|
+
GITIGNORE_AVAILABLE=1
|
|
207
|
+
fi
|
|
208
|
+
|
|
209
|
+
# Fixed, path-component ignore-list. Kept as a single array so it stays easy
|
|
210
|
+
# to extend — see the header note above.
|
|
211
|
+
IGNORE_PATTERNS=(
|
|
212
|
+
node_modules vendor dist build .venv venv __pycache__ .next coverage
|
|
213
|
+
target .git .tox .mypy_cache .pytest_cache .cache
|
|
214
|
+
)
|
|
215
|
+
|
|
216
|
+
# ---------------------------------------------------------------------------
|
|
217
|
+
# Build the JSON input for the python3 classifier: one record per candidate,
|
|
218
|
+
# already resolved to an absolute path plus its git-ignore verdict (computed
|
|
219
|
+
# in bash, since `git check-ignore` is a subprocess call best kept out of the
|
|
220
|
+
# python3 half). Pattern matching itself happens in python3 for portable,
|
|
221
|
+
# component-wise path matching (os.path based) rather than a second,
|
|
222
|
+
# divergent bash implementation.
|
|
223
|
+
# ---------------------------------------------------------------------------
|
|
224
|
+
|
|
225
|
+
NOTICES=()
|
|
226
|
+
RECORDS_JSON="[]"
|
|
227
|
+
|
|
228
|
+
if [ "${#CANDIDATES[@]}" -gt 0 ]; then
|
|
229
|
+
RECORDS_TMP=$(mktemp)
|
|
230
|
+
trap 'rm -f "$RECORDS_TMP"' EXIT
|
|
231
|
+
|
|
232
|
+
{
|
|
233
|
+
for candidate in "${CANDIDATES[@]}"; do
|
|
234
|
+
CAND_PATH="$ROOT_ABS/$candidate"
|
|
235
|
+
if [ ! -e "$CAND_PATH" ]; then
|
|
236
|
+
NOTICES+=("candidate '$candidate' does not exist under '$ROOT_ABS' — skipped")
|
|
237
|
+
continue
|
|
238
|
+
fi
|
|
239
|
+
if [ ! -d "$CAND_PATH" ]; then
|
|
240
|
+
NOTICES+=("candidate '$candidate' is not a directory — skipped")
|
|
241
|
+
continue
|
|
242
|
+
fi
|
|
243
|
+
CAND_ABS=$(cd -- "$CAND_PATH" && pwd -P)
|
|
244
|
+
|
|
245
|
+
GITIGNORED="false"
|
|
246
|
+
if [ "$GITIGNORE_AVAILABLE" -eq 1 ]; then
|
|
247
|
+
if git -C "$ROOT_ABS" check-ignore -q -- "$CAND_ABS" 2>/dev/null; then
|
|
248
|
+
GITIGNORED="true"
|
|
249
|
+
fi
|
|
250
|
+
fi
|
|
251
|
+
|
|
252
|
+
printf '%s\t%s\t%s\n' "$candidate" "$CAND_ABS" "$GITIGNORED"
|
|
253
|
+
done
|
|
254
|
+
} > "$RECORDS_TMP"
|
|
255
|
+
|
|
256
|
+
if [ "$GITIGNORE_AVAILABLE" -eq 0 ]; then
|
|
257
|
+
NOTICES+=("root is not inside a git working tree — gitignore check skipped for all candidates")
|
|
258
|
+
fi
|
|
259
|
+
|
|
260
|
+
RECORDS_JSON=$(python3 -c '
|
|
261
|
+
import json, sys
|
|
262
|
+
records = []
|
|
263
|
+
with open(sys.argv[1], "r") as f:
|
|
264
|
+
for line in f:
|
|
265
|
+
line = line.rstrip("\n")
|
|
266
|
+
if not line:
|
|
267
|
+
continue
|
|
268
|
+
parts = line.split("\t")
|
|
269
|
+
if len(parts) != 3:
|
|
270
|
+
continue
|
|
271
|
+
path, abs_path, gitignored = parts
|
|
272
|
+
records.append({"path": path, "absolute_path": abs_path, "gitignored": gitignored == "true"})
|
|
273
|
+
print(json.dumps(records))
|
|
274
|
+
' "$RECORDS_TMP")
|
|
275
|
+
fi
|
|
276
|
+
|
|
277
|
+
NOTICES_JSON=$(python3 -c '
|
|
278
|
+
import json, sys
|
|
279
|
+
print(json.dumps(sys.argv[1:]))
|
|
280
|
+
' "${NOTICES[@]:-}")
|
|
281
|
+
# printf with an empty array above can leave one empty-string arg; strip it.
|
|
282
|
+
if [ "${#NOTICES[@]}" -eq 0 ]; then
|
|
283
|
+
NOTICES_JSON="[]"
|
|
284
|
+
fi
|
|
285
|
+
|
|
286
|
+
python3 -c '
|
|
287
|
+
import json, sys, os
|
|
288
|
+
|
|
289
|
+
root = sys.argv[1]
|
|
290
|
+
root_absolute = sys.argv[2]
|
|
291
|
+
records = json.loads(sys.argv[3])
|
|
292
|
+
notices = json.loads(sys.argv[4])
|
|
293
|
+
patterns = sys.argv[5:]
|
|
294
|
+
|
|
295
|
+
def pattern_reason(path):
|
|
296
|
+
# Path-COMPONENT match, never substring: split on os.sep and compare
|
|
297
|
+
# each component (and a suffix-glob case for *.egg-info) against the
|
|
298
|
+
# fixed pattern list.
|
|
299
|
+
components = [c for c in path.split(os.sep) if c]
|
|
300
|
+
for comp in components:
|
|
301
|
+
for pat in patterns:
|
|
302
|
+
if pat.startswith("*."):
|
|
303
|
+
suffix = pat[1:]
|
|
304
|
+
if comp.endswith(suffix):
|
|
305
|
+
return "pattern:" + pat
|
|
306
|
+
elif comp == pat:
|
|
307
|
+
return "pattern:" + pat
|
|
308
|
+
return None
|
|
309
|
+
|
|
310
|
+
ignored = []
|
|
311
|
+
remaining = []
|
|
312
|
+
|
|
313
|
+
for rec in records:
|
|
314
|
+
path = rec["path"]
|
|
315
|
+
abs_path = rec["absolute_path"]
|
|
316
|
+
if rec.get("gitignored"):
|
|
317
|
+
ignored.append({"path": path, "absolute_path": abs_path, "reason": "gitignore"})
|
|
318
|
+
continue
|
|
319
|
+
reason = pattern_reason(path)
|
|
320
|
+
if reason:
|
|
321
|
+
ignored.append({"path": path, "absolute_path": abs_path, "reason": reason})
|
|
322
|
+
else:
|
|
323
|
+
remaining.append({"path": path, "absolute_path": abs_path})
|
|
324
|
+
|
|
325
|
+
report = {
|
|
326
|
+
"script": "directory-triage.sh",
|
|
327
|
+
"version": 1,
|
|
328
|
+
"root": root,
|
|
329
|
+
"root_absolute": root_absolute,
|
|
330
|
+
"candidate_count": len(records),
|
|
331
|
+
"ignored_count": len(ignored),
|
|
332
|
+
"remaining_count": len(remaining),
|
|
333
|
+
"ignored": ignored,
|
|
334
|
+
"remaining": remaining,
|
|
335
|
+
"notices": notices,
|
|
336
|
+
}
|
|
337
|
+
print(json.dumps(report, indent=2))
|
|
338
|
+
' "$ROOT" "$ROOT_ABS" "$RECORDS_JSON" "$NOTICES_JSON" "${IGNORE_PATTERNS[@]}" > "${JSON_OUT:-/dev/stdout}"
|
|
339
|
+
|
|
340
|
+
if [ -n "$JSON_OUT" ]; then
|
|
341
|
+
cat "$JSON_OUT"
|
|
342
|
+
fi
|